Sequencing Coverage and Cost Calculator: How Much Data Do You Actually Need?

Over-sequencing is the most common and least visible waste in a genomics budget, because nobody audits a run that worked.

Written byTrevor J Henderson
| 4 min read
A lab manager works through sequencing coverage calculations at a workstation beside a benchtop sequencer, illustrating how to size a sequencing run to the data it actually needs.
Register for free to listen to this article
Listen with Speechify
0:00
4:00

A sequencing coverage calculator answers a question that decides both the quality of your results and the size of your bill: how many reads does this experiment actually need. Ask for too few and the data cannot support the analysis. Ask for too many and you pay for depth that changes nothing. The second error is the expensive one, and it is nearly invisible, because a run that produced more data than necessary looks exactly like a run that worked. Nobody investigates a success.

The calculator below converts three numbers you already know, the size of your target, the depth your application needs, and how many samples you are running, into the reads required, the data to sequence, and an estimated cost. The sections that follow explain each input so you can trust the output and adapt it to your own work.

 


Key Takeaways

  • Coverage follows a simple relation: reads needed equals target depth times target size divided by read length. The calculator handles the arithmetic; you supply the three inputs.
  • Target depth is set by the application, not by preference. Germline variant calling is well served near 30x; somatic, panel, and single-cell work need much more.
  • Over-sequencing raises cost roughly in proportion to the extra depth. Doubling depth on a whole genome roughly doubles the sequencing bill.
  • A usable-data fraction below 100% is realistic. Duplicates, off-target reads, and QC loss mean you must sequence more than the ideal figure to net your target depth.

 

How Coverage Is Calculated

The relationship between reads and coverage is one equation, long known as the Lander-Waterman relation: mean coverage equals the number of reads multiplied by read length, divided by the size of the target. Rearranged to tell you what to order, the reads you need equal your target depth multiplied by your target size, divided by your read length.

An example makes it concrete. A human genome is about 3.1 billion bases. To cover it at 30x with paired-end 150 base reads, which sequence 300 bases per read pair, you need roughly 310 million read pairs per sample. That is the ideal figure. In practice you must sequence more than that, because not every read ends up usable, which is what the usable-data fraction accounts for.

Choosing Target Depth by Application

Depth is the input people most often get wrong, usually by carrying over a number from a different application. The right depth is a property of the question you are asking, not a universal setting. The table below gives typical starting points; always confirm against your assay and any relevant clinical or platform guidance, since requirements vary with sample type, expected variant frequency, and analysis method.

Application

Typical Depth

Why

Human WGS, germline variants

30x

Reliable diploid genotype calls across most of the genome

Human WGS, somatic / tumor

80-100x+

Low-frequency variants in mixed tumor and normal cell populations

Whole exome

100x

Deeper target compensates for uneven capture efficiency

Targeted panel, clinical

500-1,000x+

Detecting rare variants confidently in a small, high-value region

Bacterial WGS

30-50x

Assembly and variant calling on a small genome

RNA-seq, expression

20-30M reads

Read count, not fold coverage, is the right unit for expression

Table 1. Typical target depths by application, as starting points for planning. Confirm requirements against your assay, sample type, and any clinical or platform guidance before committing a run.

Reads to Flow Cell

Once you know the reads you need across all samples, the next question is which flow cell to run, and this is where cost per sample is quietly won or lost. Flow cells come in output tiers, and the cost per gigabase falls sharply as the tier gets larger. A run that fills a large flow cell efficiently costs far less per sample than the same work split across several small ones or spread thinly across a large one that is mostly empty.

The practical discipline is to match total data needed to the smallest flow cell that fits it comfortably when full, then batch samples until that flow cell is full before running. Running a large flow cell half empty pays the large-flow-cell price for half the output. The real per-genome economics of flow cell selection, with figures drawn from a published rate card, are worked through in the breakdown of what NGS actually costs.

Lab manager academy logo

Advanced Lab Management Certificate

The Advanced Lab Management certificate is more than training—it’s a professional advantage.

Gain critical skills and IACET-approved CEUs that make a measurable difference.

Cost per Sample at Different Depths

This is where over-sequencing shows its price. Because cost scales with the data you sequence, and data scales with depth, asking for more depth than the application needs raises the bill almost in lockstep. The calculator makes this visible instantly: change the depth and watch the total move.


The Cost of Depth You Do Not Need

Take a 24-sample germline whole-genome project at 30x. At a representative $2.50 per gigabase for sequencing reagents, the calculator puts the total sequencing cost near $6,975. Raise the depth to 60x, which germline variant calling does not require, and the total roughly doubles to about $13,950, for data that will not meaningfully improve a germline call set. That $7,000 difference bought nothing, and because both runs succeed, nothing in the results will ever flag it. The only place the waste is visible is the budget, and only if someone is looking.

 

The reverse error, under-sequencing, is less common and usually caught quickly, because the analysis fails or the results are visibly noisy and the run gets repeated. That makes under-sequencing self-correcting in a way over-sequencing is not. This asymmetry is exactly why over-sequencing persists: one error announces itself and the other hides in a successful run. Setting depth deliberately, to the application rather than to habit, is the single cheapest efficiency available to a sequencing lab.

Using the Calculator

Start with an application preset to load typical values, then adjust for your specifics. Enter your target size in megabases, the depth your application requires, your read length as total bases per read or read pair, and your sample count. Set the usable-data fraction to reflect how much of your sequencing typically survives duplicate removal and quality filtering; 80% is a reasonable planning default for many whole-genome workflows, lower for capture-based methods. Enter your own cost per gigabase if you have it, since that figure varies by flow cell and negotiated pricing.

The output gives you reads needed per sample, data to sequence, and estimated cost, both per sample and across the whole run. Treat the cost as a reagent-level planning estimate: it covers sequencing itself, not library preparation, QC, storage, labor, or instrument capital, all of which are covered in the full cost picture in the guide to building an NGS program. For the wider operational context of how coverage planning fits into running a sequencing program, see the manager’s guide to next-generation sequencing in the lab

 

This article was produced under Lab Manager's AI Editorial Guidelines.

Add Lab Manager as a preferred source on Google

Add Lab Manager as a preferred Google source to see more of our trusted coverage.

Frequently Asked Questions (FAQs)

  • How do I calculate sequencing coverage?

    Coverage equals the number of reads multiplied by read length, divided by the size of your target. To find the reads you need, rearrange it: reads equal target depth times target size divided by read length. For a human genome at 30x with paired-end 150 base reads, that is roughly 310 million read pairs per sample before accounting for reads lost to duplicates and filtering. The calculator on this page does the arithmetic and adds the usable-data adjustment.

  • How many reads for whole genome sequencing?

    For a human genome at 30x, the standard germline depth, you need roughly 310 million paired-end 150 base read pairs per sample as an ideal figure, and somewhat more in practice to offset reads lost to duplication and off-target signal. Somatic or tumor sequencing at 80x to 100x or higher needs proportionally more. Enter your exact depth and read length into the calculator for a precise figure.

  • What depth do I need for variant calling?

    It depends on the variant type. Germline variant calling on a whole human genome is reliably served near 30x. Somatic variant calling, which must detect low-frequency variants in mixed cell populations, typically needs 80x to 100x or more. Clinical targeted panels looking for rare variants in a small region often run 500x to 1,000x or higher. Higher depth beyond what the application requires raises cost without improving germline calls.

About the Author

  • Trevor Henderson headshot

    Trevor Henderson BSc (HK), MSc, PhD (c), has more than two decades of experience in the fields of scientific and technical writing, editing, and creative content creation. With academic training in the areas of human biology, physical anthropology, and community health, he has a broad skill set of both laboratory and analytical skills. Since 2013, he has been working with LabX Media Group developing content solutions that engage and inform scientists and laboratorians. He can be reached at thenderson@labmanager.com.

    View Full Profile

Related Topics

Related Articles

Current Magazine Issue Background Image

CURRENT ISSUE - September/2026

Are You Asking the Right Questions?

How Question Framing Shapes Better Lab Decisions

Lab Manager September 2026 Cover Image