A functioning NGS workflow is easy to draw and hard to run. Sample in, library prep, sequencing run, alignment, variant calling, result out. Six boxes on a slide. What the slide does not show is that most labs never decided to build that workflow in the first place. They inherited it, one grant-funded instrument at a time, and discovered afterward that the sequencer was the cheapest and least troublesome part of the arrangement.
This guide is written for the person who has to make that inheritance work: cost it honestly, staff it, store its output, keep it running, and defend it at budget time. It deliberately does not explain sequencing chemistry. That ground is covered thoroughly in this explainer on Next-generation Sequencing: How it Works, Platforms, and Applications, and the older comparison of NGS vs. Sanger Sequencing remains the clearest starting point if you are still deciding whether NGS is the right method at all. What follows starts one step later, at the point where the science is settled and the operations are not.
Key Takeaways
|
What an NGS Operation Actually Costs to Run
Sequencing has one of the most quoted cost curves in science, and one of the most misleading. The National Human Genome Research Institute has tracked cost per genome at its funded sequencing centers since 2001, and that dataset is the source of nearly every chart you have seen on the subject. It is real data, carefully collected, and it does not answer the question a lab manager is asking. It reflects production costs at large, high-throughput, specialist centers, and it stops short of the analysis, interpretation, and long-term storage that make a sequencing result usable.
The practical consequence is that the headline number your director has read is a floor, not an estimate. Build your own number instead. The model below is deliberately transparent so you can substitute your own figures for every line.
Assume a 30x human whole genome. Reagents and flow cell allocation at $600 per sample, library prep consumables at $80, 1.5 hours of hands-on technician time at $45 an hour, $40 of compute for primary and secondary analysis, and $60 of first-year storage. Assume an 8% repeat rate, so 108 samples must be run for every 100 delivered. On the fixed side, assume a $985,000 instrument amortized over five years and a $60,000 annual service contract, giving $257,000 of annual fixed cost.
Delivered Genomes per Year | Fixed Cost per Genome | Fully Loaded Cost per Genome | Multiple of the Reagent Figure |
250 | $1,028 | $1,939 | 3.2x |
500 | $514 | $1,425 | 2.4x |
1,000 | $257 | $1,168 | 1.9x |
2,000 | $129 | $1,039 | 1.7x |
Table 1. Fully loaded cost per delivered genome against annual throughput, on the stated assumption set. Substitute your own reagent price, labor rate, and capital figures; the shape of the curve does not change.
Why the Reagent Number Never Becomes the Real NumberAn eightfold increase in throughput, from 250 to 2,000 delivered genomes a year, cuts fully loaded cost per genome by about 46%. That is a genuine and worthwhile saving, and it is entirely a fixed-cost effect: capital and service spread across more samples. What it does not do is close the gap to the reagent price. Even at 2,000 genomes a year, the loaded figure is still roughly 1.7 times the reagent quote. This matters for how you argue. A business case built on reaching the quoted reagent cost will fail on contact with the first annual review. A business case built on utilization, turnaround, and capability will survive, because those are the things that actually move. |
Two cautions on using a model like this. First, the fixed-cost column assumes the instrument is dedicated to this workload; if it is shared across programs, allocate proportionally or you will overstate the cost of your own work and understate everyone else’s. Second, the repeat rate is the assumption most likely to be wrong in a new program, and it is the one nobody measures. Record the outcome of every run from day one, not just that the run happened, because a repeat rate cannot be reconstructed later from a log that only counts runs. Full cost-stack detail, business case structure, and the in-house versus outsourced break-even sit in Building an NGS Program: Strategy, Budget, and ROI.
How Should You Choose and Qualify a Platform?
Start from the samples, not the specification sheet. The single most common procurement error in sequencing is choosing an instrument sized for the throughput a lab hopes to reach rather than the throughput it can currently fill. A high-output instrument running at a quarter of capacity is more expensive per sample than a benchtop instrument running near full, and it is harder to schedule, because larger flow cells force you to wait for enough samples to batch economically.
Work through four questions in order, and resist the temptation to answer them out of sequence.
- What applications will this instrument actually serve, and what read length and depth do they require? An instrument that is excellent for whole-genome work may be poorly matched to small targeted panels, where per-run overhead dominates.
- How many samples per year, and how are they distributed? A hundred samples arriving evenly is a completely different operational problem from a hundred arriving in two bursts.
- What turnaround does the work require? Turnaround, not output, is what determines whether you can batch, and batching is what determines your cost per sample.
- What can the site support? Floor loading, power, uninterruptible supply, temperature stability, network egress, and physical separation between pre-amplification and post-amplification areas. Site problems are the ones that delay installation, and they are almost always discovered late.
Qualification is the step most research labs skip and later wish they had not. Even outside a regulated setting, running a documented installation, operational, and performance qualification gives you something valuable: a recorded baseline of how the instrument performed when it was new and correct. Without that baseline, every future argument with a vendor about drift or degradation is an argument about impressions. With it, the conversation is about data. Detailed vendor evaluation questions, service contract terms, and site requirements are covered in Choosing an NGS Platform: A Lab Manager’s Selection and Procurement Guide. For instrument pricing, platform families, and what to check when buying used, see Buying a DNA Sequencer: The Complete Buyer’s Guide.
One decision worth making explicitly rather than by default: whether to own at all. Sequencing services have become genuinely competitive, and for a lab running low or irregular volume, outsourcing frequently wins on cost while removing the service contract, the qualification burden, and the staffing problem in a single stroke. The reason to own is rarely price. It is control of turnaround, protection of sensitive samples, method flexibility, or the strategic value of holding the capability in-house. Those are all legitimate reasons. They are just not the same reason as saving money, and conflating the two produces business cases that do not survive scrutiny.
Workflow, Throughput, and Automation
The sequencing run is almost never the bottleneck. In most labs the constraint sits upstream, in nucleic acid extraction, quantification, and library preparation, or downstream, in analysis and interpretation. The instrument is idle while the queue builds at both ends. This is why buying a faster sequencer so often fails to increase delivered output: the new instrument arrives, and the same library prep team is still the limiting step.
Diagnose before you invest. Time each stage of the workflow for a month, including queue time between stages, not just hands-on time. The stage with the longest total elapsed time is your constraint, and it is frequently not the stage that feels most burdensome. Manual library prep feels like the worst job in the lab because it is repetitive and exacting, but if libraries are waiting three days for a scheduled run slot, automating prep will simply produce libraries that wait three days.
When library prep genuinely is the constraint, automation is a strong answer, and the operational case is well established: consistent liquid handling, lower contamination risk, reduced hands-on time, and a digital chain of custody when the platform is integrated with laboratory informatics. The practical considerations, from deck layout to LIMS integration, are set out in this guide to implementing automated liquid handling in your lab, and the specific application to genomics workflows is covered in this overview of automation in genomics workflows.
Three scheduling disciplines are worth more than most equipment purchases:
- Batch to the flow cell, not to the calendar. Partial runs are the quietest form of waste in a sequencing lab, because the cost appears in the reagent budget rather than anywhere labeled inefficiency.
- Publish a run schedule and hold to it. Predictable slots let submitting groups plan backward from them, which improves batching without anyone having to negotiate.
- Set a rerun policy in advance. Decide who pays, under what circumstances, and with what evidence, before the first failed run rather than during the resulting conversation.
Throughput planning, multiplexing strategy, run failure analysis, and the automation decision are worked through in Running NGS at Scale: Throughput, Scheduling, and Automation.
How Much Data Infrastructure Will You Need?
More than the vendor quote implies, and for longer. A study of CRAM compression across sequencing platforms reports that the full set of key files generated for a single 30x whole-genome sample, covering raw output, FASTQ, BAM, and gVCF, runs to roughly 200 to 900 GB depending on the platform, at a cloud storage cost the authors put at roughly $50 to $250 per sample per year. The same work found CRAM achieving 40% to 70% compression depending on platform, without altering variant calls.
Take the conservative end of that range, 250 GB per sample, and the arithmetic becomes uncomfortable quickly.
Genomes per Year | New Data per Year | Cumulative After 5 Years | Cumulative After 5 Years, CRAM Applied |
200 | 50 TB | 250 TB | 125 TB |
500 | 125 TB | 625 TB | 313 TB |
1,000 | 250 TB | 1,250 TB | 625 TB |
Table 2. Storage accumulation at 250 GB per 30x whole-genome sample with full retention and no deletion. CRAM column applies a 50% reduction, mid-range for short-read data.
The Budget Line That Grows When Nothing Else DoesEvery other line in an NGS budget is proportional to how much sequencing you do this year. Storage is not. A lab running a steady 500 genomes a year with full retention pays for one cohort in year one and five cohorts in year five. At the low end of the published cloud figure, that is roughly $25,000 rising to roughly $125,000 annually, with sequencing volume completely unchanged. Which means retention policy is a budget decision disguised as a technical one. Decide what you keep, in what format, for how long, and on what storage tier, before the first run rather than after the first invoice that makes someone ask. |
A workable default for a research lab: keep variant calls and processed results indefinitely on accessible storage, keep aligned data in a compressed format on a lower tier, and retain raw output only for a defined window unless a specific requirement says otherwise. Funder and journal data availability requirements, institutional records policy, and consent terms all constrain this and should be checked before the policy is written rather than after. Storage sizing, cloud against on-premise, retention policy, and analysis software selection are covered in Managing NGS Data: Storage, Compute, Retention, and Staffing.
Compute deserves a separate thought from storage. The two are usually procured together and behave nothing alike. Storage demand is steady and cumulative. Compute demand is spiky, concentrated in the hours after a run completes, which is precisely the pattern that suits burst capacity and suits an on-premise cluster sized for average load very badly. And sample tracking has to hold the whole thing together: every sample needs an unbroken record from receipt through library, run, and result. Whether that lives in a LIMS or in an increasingly elaborate spreadsheet is a question most growing labs answer too late, and the tradeoffs are laid out in this comparison of paper-based records and laboratory information management systems.
Quality, Compliance, and Staffing
Quality expectations in sequencing scale with what the result is used for, and they scale sharply. A research-use workflow needs defined quality thresholds, documented procedures, tracked run metrics, and honest failure investigation. A clinical workflow needs all of that plus formal analytical validation, personnel competency records, proficiency testing, and version control extending into the bioinformatics pipeline.
That last point is where sequencing differs most from other laboratory methods, and where inspection findings concentrate. The analytical process does not end at the instrument. Alignment, variant calling, annotation, and filtering are all part of the test, which means the software versions, reference genome build, and database releases used to produce a result have to be recorded and reproducible years later. CAP’s next-generation sequencing worksheets set out expectations across both the wet-lab and bioinformatics halves of the workflow, and the CDC NGS Quality Initiative’s pathway to quality-focused testing walks through method validation planning in a form that is useful even to labs with no intention of seeking accreditation.
If regulated testing is anywhere in your three-year plan, design toward it from the start. Retrofitting version control, competency documentation, and validation records onto an established research workflow is materially harder than building them in, and the difference is usually measured in months of the laboratory director’s time. Accreditation requirements, assay validation, proficiency testing, and contamination control are covered in Quality and Compliance in NGS Labs: From Research Use to Regulated Testing.
Staffing is where NGS programs most often fail quietly. The wet-lab side is straightforward to plan for and reasonably easy to hire. The analysis side is neither. Programs that scale sequencing capacity without scaling analysis capacity do not stop producing data; they accumulate unanalyzed runs, which is worse, because the money has been spent and the result has not been delivered. The honest planning question is not how many technicians you need to run the instrument. It is who converts a completed run into an answer somebody can act on, how long that takes them, and what happens to the queue when that person is on leave.
A minimum viable staffing picture for a modest program has four functions, which may be fewer than four people: someone accountable for sample handling and library preparation, someone accountable for instrument operation and maintenance, someone accountable for analysis and data delivery, and someone accountable for quality and documentation. Programs that collapse the last function into the other three tend to discover the gap during an audit or after an unexplained run failure, and both are expensive ways to learn it.
Building a Three-Year NGS Roadmap
Sequencing programs are usually funded in one-year increments and only make sense on a three-year view. The mismatch is structural, and the way through it is to hold a three-year plan that can be defended one year at a time.
Period | Focus | Decisions to Make | What to Measure |
Months 1–6 | Stand up and stabilize | Retention policy, rerun policy, sample tracking system, quality thresholds | Run success rate, turnaround, hands-on hours per sample |
Months 7–18 | Utilization and cost control | Whether to automate, and which step; batching and scheduling rules | Instrument utilization, cost per delivered sample, repeat rate, queue length |
Months 19–36 | Scale or specialize | Second instrument, added modality, regulated testing, service model | Demand trend, groups served, analysis backlog, storage growth |
Table 3. A three-year framing for an NGS program. The measurement column matters more than the focus column, because it determines what evidence you will have when the next decision arrives.
Two habits make this work. Set the review dates in advance and write them into the original business case, so the review happens because it was scheduled rather than because someone chose to raise it. And instrument the metrics before go-live, not after, because metrics chosen retrospectively get selected to justify whatever happened.
It also helps to keep a modest sense of proportion about the technology itself. Sequencing has been through several complete generational shifts in twenty years, a history this illustrated look at the evolution of next-generation sequencing captures well, and there is no reason to assume the current arrangement is the final one. A three-year plan that assumes today’s platform is permanent will be wrong. A three-year plan built around utilization, staff capability, data discipline, and the ability to absorb a new method is much harder to invalidate, because none of those things depend on which instrument is on the bench.
The most useful thing a lab manager can do for a sequencing program is unglamorous: make its real cost visible, keep its data under a policy rather than under an assumption, and make sure somebody is accountable for turning runs into answers. The instrument will be replaced. Those three things will still be the job.
This article was produced under Lab Manager's AI Editorial Guidelines.

















