A sequencing coverage calculator answers a question that decides both the quality of your results and the size of your bill: how many reads does this experiment actually need. Ask for too few and the data cannot support the analysis. Ask for too many and you pay for depth that changes nothing. The second error is the expensive one, and it is nearly invisible, because a run that produced more data than necessary looks exactly like a run that worked. Nobody investigates a success.
The calculator below converts three numbers you already know, the size of your target, the depth your application needs, and how many samples you are running, into the reads required, the data to sequence, and an estimated cost. The sections that follow explain each input so you can trust the output and adapt it to your own work.
|
|
|---|
|
Key Takeaways
|
How Coverage Is Calculated
The relationship between reads and coverage is one equation, long known as the Lander-Waterman relation: mean coverage equals the number of reads multiplied by read length, divided by the size of the target. Rearranged to tell you what to order, the reads you need equal your target depth multiplied by your target size, divided by your read length.
An example makes it concrete. A human genome is about 3.1 billion bases. To cover it at 30x with paired-end 150 base reads, which sequence 300 bases per read pair, you need roughly 310 million read pairs per sample. That is the ideal figure. In practice you must sequence more than that, because not every read ends up usable, which is what the usable-data fraction accounts for.
Choosing Target Depth by Application
Depth is the input people most often get wrong, usually by carrying over a number from a different application. The right depth is a property of the question you are asking, not a universal setting. The table below gives typical starting points; always confirm against your assay and any relevant clinical or platform guidance, since requirements vary with sample type, expected variant frequency, and analysis method.
|
Application |
Typical Depth |
Why |
|
Human WGS, germline variants |
30x |
Reliable diploid genotype calls across most of the genome |
|
Human WGS, somatic / tumor |
80-100x+ |
Low-frequency variants in mixed tumor and normal cell populations |
|
Whole exome |
100x |
Deeper target compensates for uneven capture efficiency |
|
Targeted panel, clinical |
500-1,000x+ |
Detecting rare variants confidently in a small, high-value region |
|
Bacterial WGS |
30-50x |
Assembly and variant calling on a small genome |
|
RNA-seq, expression |
20-30M reads |
Read count, not fold coverage, is the right unit for expression |
Table 1. Typical target depths by application, as starting points for planning. Confirm requirements against your assay, sample type, and any clinical or platform guidance before committing a run.
Reads to Flow Cell
Once you know the reads you need across all samples, the next question is which flow cell to run, and this is where cost per sample is quietly won or lost. Flow cells come in output tiers, and the cost per gigabase falls sharply as the tier gets larger. A run that fills a large flow cell efficiently costs far less per sample than the same work split across several small ones or spread thinly across a large one that is mostly empty.
The practical discipline is to match total data needed to the smallest flow cell that fits it comfortably when full, then batch samples until that flow cell is full before running. Running a large flow cell half empty pays the large-flow-cell price for half the output. The real per-genome economics of flow cell selection, with figures drawn from a published rate card, are worked through in the breakdown of what NGS actually costs.
Cost per Sample at Different Depths
This is where over-sequencing shows its price. Because cost scales with the data you sequence, and data scales with depth, asking for more depth than the application needs raises the bill almost in lockstep. The calculator makes this visible instantly: change the depth and watch the total move.
|
|
|---|
|
The Cost of Depth You Do Not Need Take a 24-sample germline whole-genome project at 30x. At a representative $2.50 per gigabase for sequencing reagents, the calculator puts the total sequencing cost near $6,975. Raise the depth to 60x, which germline variant calling does not require, and the total roughly doubles to about $13,950, for data that will not meaningfully improve a germline call set. That $7,000 difference bought nothing, and because both runs succeed, nothing in the results will ever flag it. The only place the waste is visible is the budget, and only if someone is looking. |
The reverse error, under-sequencing, is less common and usually caught quickly, because the analysis fails or the results are visibly noisy and the run gets repeated. That makes under-sequencing self-correcting in a way over-sequencing is not. This asymmetry is exactly why over-sequencing persists: one error announces itself and the other hides in a successful run. Setting depth deliberately, to the application rather than to habit, is the single cheapest efficiency available to a sequencing lab.
Using the Calculator
Start with an application preset to load typical values, then adjust for your specifics. Enter your target size in megabases, the depth your application requires, your read length as total bases per read or read pair, and your sample count. Set the usable-data fraction to reflect how much of your sequencing typically survives duplicate removal and quality filtering; 80% is a reasonable planning default for many whole-genome workflows, lower for capture-based methods. Enter your own cost per gigabase if you have it, since that figure varies by flow cell and negotiated pricing.
The output gives you reads needed per sample, data to sequence, and estimated cost, both per sample and across the whole run. Treat the cost as a reagent-level planning estimate: it covers sequencing itself, not library preparation, QC, storage, labor, or instrument capital, all of which are covered in the full cost picture in the guide to building an NGS program. For the wider operational context of how coverage planning fits into running a sequencing program, see the manager’s guide to next-generation sequencing in the lab.
This article was produced under Lab Manager's AI Editorial Guidelines.
















