Handling an NGS run failure well is less about technical troubleshooting than most labs assume and more about having decided, in advance, how the decisions around a failure will be made. A failed or marginal run raises three questions in quick succession: what went wrong, whose fault it was, and whether to repeat it, and the last two are where the real cost and the real friction live. In the absence of a policy, those questions get answered in the heat of the moment by whoever is most frustrated, most senior, or most inconvenienced, which is exactly the wrong basis for a decision that consumes reagents, instrument time, and sample material. This guide is about making those decisions by policy rather than by pressure.
It covers how to triage a failed run quickly, how to separate a library problem from an instrument problem, how to read run metrics at the level needed to assign the cause, and, most importantly, how to set a rerun policy and a cost-recovery structure before you need them. It does not go deep into the technical interpretation of quality-control metrics, which is a subject in its own right and is covered separately, linked below. The focus here is the management decision: who decides, on what basis, and who pays.
Key Takeaways
|
Triaging a Failed Run
The first move when a run fails or comes back marginal is to triage, not to fix, because rushing to a fix before understanding the cause is how a lab reruns a failed library on a healthy instrument and fails again. Triage means answering, quickly and provisionally, a small set of questions: did the run fail outright or produce usable-but-degraded data, did it fail across the whole flow cell or only for some samples, and does the pattern point upstream to the libraries or to the instrument and its run. These questions do not require deep analysis; they require looking at the shape of the failure before reacting to it.
The single most useful discriminator is whether the problem affected everything or only some samples. A failure that hits the entire run uniformly points toward the instrument, the reagents, or the run setup, since those are common to every sample. A failure confined to particular samples while others in the same run performed well points toward those specific libraries, since whatever differed was upstream of the shared sequencing step. That one observation resolves a large fraction of failures into the right domain before any detailed metric is examined, and it is the foundation of the responsibility question that follows.
Library Problems vs. Instrument Problems
Almost every run failure resolves into one of two domains, and which one it is determines who owns the failure, what the fix is, and who bears the cost. A library problem originates upstream of sequencing, in sample quality, extraction, or library preparation, and it typically shows up as specific samples underperforming while the run itself was healthy. An instrument or run problem originates at the sequencing step, in the reagents, the flow cell, the run setup, or the instrument itself, and it typically shows up as the whole run underperforming regardless of which samples were on it. Sorting a failure into the correct domain is the pivotal act of run-failure management, because everything downstream, the fix, the responsibility, the cost, follows from it.
Symptom Pattern | Likely Domain | Who Owns the Decision and Cost |
Specific samples fail, others fine | Library or sample: upstream of sequencing | The group whose samples or prep failed; often the submitter |
Whole run underperforms uniformly | Instrument, reagent, flow cell, or run setup | The sequencing operation; the lab or a warranty or service claim |
Marginal data, all samples | Run setup or loading, or a borderline reagent lot | Shared: often the lab, pending root cause |
Repeated same-mode failures | Systemic: a recurring process or lot problem | The lab, escalated to root-cause and prevention |
Table 1. Triaging a run failure by symptom pattern to the likely domain and, from that, to who owns the rerun decision and its cost. The domain determines the responsibility, which is why triage precedes any argument about who pays.
Reading Run Metrics for Cause
Run metrics are how a provisional triage becomes a confident cause assignment, and for the management decision they need to be read only at the level that distinguishes the two domains, not exhaustively. The metrics a sequencing run reports, such as the density of usable data-generating clusters, the fraction of reads passing quality filters, the quality scores across the run, and the yield against expectation, together indicate whether the sequencing chemistry and instrument performed as intended or not. Low usable-cluster density across the whole run, for instance, points at loading or library quantification or a run-level issue rather than at any one sample.
The important discipline for a manager is to read these metrics to assign the domain and then stop, rather than attempting a deep technical diagnosis that belongs to a specialist. Whether a given metric pattern indicates over-loading, under-loading, a library quantification error, a reagent problem, or an instrument fault is a genuinely technical question. For the management decision, the metrics need to answer one question: does the evidence place this failure in the library domain or the instrument domain, so the right party owns what comes next.
When to Rerun and Who Decides
The decision to rerun is a cost-benefit judgment that should be made against criteria set in advance, not against the mood of the moment. A rerun consumes a fresh flow cell, another reagent kit, more instrument time, and, critically, more of the sample, which may be limited or irreplaceable, and it should happen only when the value of the result justifies that renewed cost and when there is reason to believe the rerun will succeed where the first attempt failed. Rerunning a failed library the same way that produced the failure, without addressing the cause, simply spends the cost twice.
A written rerun policy answers the questions that otherwise get answered by whoever pushes hardest: who has the authority to approve a rerun, what criteria must be met, whether the cause has to be identified before a rerun is approved, and what happens when sample material is too limited to repeat. Having these answers in advance does two things. It makes the decision consistent across failures and across the people involved, and it removes the emotional charge from a moment that is, by nature, frustrating for whoever owns the failed samples. The policy is not bureaucracy; it is the thing that lets a stressful decision be made calmly and fairly.
The Policy Has to Exist Before the Failure Once a run has failed, everyone in the room is an interested party. The submitter wants their result and did not cause the problem, or believes they did not. The sequencing team defends its process. Someone is paying either way. A policy written in that moment will be shaped by whoever has the most standing or the most frustration, which is precisely the decision basis to avoid. A policy written in advance, when no particular failure is at stake and everyone can reason about the general case fairly, is the only kind that holds up when a real failure puts it to the test. Write it before you need it, because you cannot write it fairly once you do. |
Cost Recovery for Reruns
The most contentious question a failure raises is who pays for the rerun, and it is the one most improved by a policy set in advance. A rerun has a real cost, the flow cell, the reagents, the instrument time, worked through in How Much Does NGS Cost? Budgeting Instruments, Reagents, and Sequencing, and someone absorbs it. The fair answer usually follows the cause: a failure originating in a submitter’s sample or library is reasonably the submitter’s cost, while a failure originating in the sequencing operation’s reagents, instrument, or process is reasonably the lab’s. Tying cost to cause, decided in advance, is what makes the outcome feel fair rather than arbitrary.
A workable cost-recovery policy states the default owner of the cost for each failure domain, an escalation path for cases where the cause is ambiguous or shared, and a provision for genuinely no-fault failures, such as a bad reagent lot that no one could have anticipated, which many labs choose to absorb centrally rather than assign. The specific structure matters less than the fact of having one, because a cost-recovery rule that everyone agreed to before any particular failure is one they will accept when a failure occurs, whereas a cost decision improvised after the fact leaves someone feeling they were charged unfairly regardless of the merits. Decide the principle once, apply it consistently, and the recurring argument disappears.
Building a Failure Log
The final discipline turns individual failures into intelligence, and it is the one most likely to be skipped because each failure feels like a one-off to be resolved and forgotten. A failure log, a simple, consistent record of every failed or marginal run, its triaged domain, its identified cause, and its resolution, converts a scatter of isolated frustrations into a dataset that reveals patterns no single failure could. A run that fails once is an incident; the same failure mode appearing five times in three months is a systemic problem the log makes visible and memory alone would miss.
That pattern visibility is what lets a lab move from repeatedly reacting to failures to preventing them, by revealing which causes recur and warrant a process change, a supplier conversation, or additional training. For labs doing or moving toward regulated work, the log is also the foundation of the formal corrective-and-preventive-action process quality systems require, as part of the framework covered in Quality and Compliance in NGS Labs: From Research Use to Regulated Testing. Even outside a regulated setting it is among the highest-return, lowest-cost habits a sequencing lab can adopt, and it connects to the throughput discipline, since failures that consume capacity are a throughput problem as much as a quality one, covered in Running NGS at Scale: Throughput, Scheduling, and Automation. The wider operational picture failure management sits within is in Next-Generation Sequencing in the Lab: A Manager’s Guide to Building, Budgeting, and Scaling NGS Capacity.
This article was produced under Lab Manager's AI Editorial Guidelines.

















