Managing Run Failures: Root Cause Analysis and Rerun Policy

Without a written rerun policy, the decision to repeat a run gets made by whoever is most upset about the result.

Written byTrevor J Henderson
| 7 min read
A supervisor and technician review a run-metrics dashboard to diagnose a failed sequencing run, illustrating root cause analysis and the rerun decision.
Register for free to listen to this article
Listen with Speechify
0:00
7:00

Handling an NGS run failure well is less about technical troubleshooting than most labs assume and more about having decided, in advance, how the decisions around a failure will be made. A failed or marginal run raises three questions in quick succession: what went wrong, whose fault it was, and whether to repeat it, and the last two are where the real cost and the real friction live. In the absence of a policy, those questions get answered in the heat of the moment by whoever is most frustrated, most senior, or most inconvenienced, which is exactly the wrong basis for a decision that consumes reagents, instrument time, and sample material. This guide is about making those decisions by policy rather than by pressure.

It covers how to triage a failed run quickly, how to separate a library problem from an instrument problem, how to read run metrics at the level needed to assign the cause, and, most importantly, how to set a rerun policy and a cost-recovery structure before you need them. It does not go deep into the technical interpretation of quality-control metrics, which is a subject in its own right and is covered separately, linked below. The focus here is the management decision: who decides, on what basis, and who pays.


Key Takeaways

  • The hardest part of a failed run is not diagnosis; it is the rerun and cost decisions, which go badly when made in the heat of the moment.
  • Triage first to one question: is this a library problem or an instrument problem? That split determines who owns the failure and who pays.
  • Set a rerun policy in advance, defining who decides and on what criteria, so the decision is not driven by whoever is most upset.
  • Set a cost-recovery policy in advance too, deciding who bears the cost under each cause, before the first failure rather than during the argument.
  • Log every failure and its cause. A failure log turns individual frustrations into the pattern that reveals a systemic problem.

 

Triaging a Failed Run

The first move when a run fails or comes back marginal is to triage, not to fix, because rushing to a fix before understanding the cause is how a lab reruns a failed library on a healthy instrument and fails again. Triage means answering, quickly and provisionally, a small set of questions: did the run fail outright or produce usable-but-degraded data, did it fail across the whole flow cell or only for some samples, and does the pattern point upstream to the libraries or to the instrument and its run. These questions do not require deep analysis; they require looking at the shape of the failure before reacting to it.

The single most useful discriminator is whether the problem affected everything or only some samples. A failure that hits the entire run uniformly points toward the instrument, the reagents, or the run setup, since those are common to every sample. A failure confined to particular samples while others in the same run performed well points toward those specific libraries, since whatever differed was upstream of the shared sequencing step. That one observation resolves a large fraction of failures into the right domain before any detailed metric is examined, and it is the foundation of the responsibility question that follows.

Library Problems vs. Instrument Problems

Almost every run failure resolves into one of two domains, and which one it is determines who owns the failure, what the fix is, and who bears the cost. A library problem originates upstream of sequencing, in sample quality, extraction, or library preparation, and it typically shows up as specific samples underperforming while the run itself was healthy. An instrument or run problem originates at the sequencing step, in the reagents, the flow cell, the run setup, or the instrument itself, and it typically shows up as the whole run underperforming regardless of which samples were on it. Sorting a failure into the correct domain is the pivotal act of run-failure management, because everything downstream, the fix, the responsibility, the cost, follows from it.

Symptom Pattern

Likely Domain

Who Owns the Decision and Cost

Specific samples fail, others fine

Library or sample: upstream of sequencing

The group whose samples or prep failed; often the submitter

Whole run underperforms uniformly

Instrument, reagent, flow cell, or run setup

The sequencing operation; the lab or a warranty or service claim

Marginal data, all samples

Run setup or loading, or a borderline reagent lot

Shared: often the lab, pending root cause

Repeated same-mode failures

Systemic: a recurring process or lot problem

The lab, escalated to root-cause and prevention

Table 1. Triaging a run failure by symptom pattern to the likely domain and, from that, to who owns the rerun decision and its cost. The domain determines the responsibility, which is why triage precedes any argument about who pays.

Lab manager academy logo

Lab Management Certificate

The Lab Management certificate is more than training—it’s a professional advantage.

Gain critical skills and IACET-approved CEUs that make a measurable difference.

Reading Run Metrics for Cause

Run metrics are how a provisional triage becomes a confident cause assignment, and for the management decision they need to be read only at the level that distinguishes the two domains, not exhaustively. The metrics a sequencing run reports, such as the density of usable data-generating clusters, the fraction of reads passing quality filters, the quality scores across the run, and the yield against expectation, together indicate whether the sequencing chemistry and instrument performed as intended or not. Low usable-cluster density across the whole run, for instance, points at loading or library quantification or a run-level issue rather than at any one sample.

The important discipline for a manager is to read these metrics to assign the domain and then stop, rather than attempting a deep technical diagnosis that belongs to a specialist. Whether a given metric pattern indicates over-loading, under-loading, a library quantification error, a reagent problem, or an instrument fault is a genuinely technical question. For the management decision, the metrics need to answer one question: does the evidence place this failure in the library domain or the instrument domain, so the right party owns what comes next.

Interested in lab leadership?

Register for a FREE Lab Manager account to subscribe to our Lab Leadership Digest Newsletter.
Subscribe for Free

When to Rerun and Who Decides

The decision to rerun is a cost-benefit judgment that should be made against criteria set in advance, not against the mood of the moment. A rerun consumes a fresh flow cell, another reagent kit, more instrument time, and, critically, more of the sample, which may be limited or irreplaceable, and it should happen only when the value of the result justifies that renewed cost and when there is reason to believe the rerun will succeed where the first attempt failed. Rerunning a failed library the same way that produced the failure, without addressing the cause, simply spends the cost twice.

A written rerun policy answers the questions that otherwise get answered by whoever pushes hardest: who has the authority to approve a rerun, what criteria must be met, whether the cause has to be identified before a rerun is approved, and what happens when sample material is too limited to repeat. Having these answers in advance does two things. It makes the decision consistent across failures and across the people involved, and it removes the emotional charge from a moment that is, by nature, frustrating for whoever owns the failed samples. The policy is not bureaucracy; it is the thing that lets a stressful decision be made calmly and fairly.


The Policy Has to Exist Before the Failure

Once a run has failed, everyone in the room is an interested party. The submitter wants their result and did not cause the problem, or believes they did not. The sequencing team defends its process. Someone is paying either way. A policy written in that moment will be shaped by whoever has the most standing or the most frustration, which is precisely the decision basis to avoid. A policy written in advance, when no particular failure is at stake and everyone can reason about the general case fairly, is the only kind that holds up when a real failure puts it to the test. Write it before you need it, because you cannot write it fairly once you do.

 

Cost Recovery for Reruns

The most contentious question a failure raises is who pays for the rerun, and it is the one most improved by a policy set in advance. A rerun has a real cost, the flow cell, the reagents, the instrument time, worked through in How Much Does NGS Cost? Budgeting Instruments, Reagents, and Sequencing, and someone absorbs it. The fair answer usually follows the cause: a failure originating in a submitter’s sample or library is reasonably the submitter’s cost, while a failure originating in the sequencing operation’s reagents, instrument, or process is reasonably the lab’s. Tying cost to cause, decided in advance, is what makes the outcome feel fair rather than arbitrary.

A workable cost-recovery policy states the default owner of the cost for each failure domain, an escalation path for cases where the cause is ambiguous or shared, and a provision for genuinely no-fault failures, such as a bad reagent lot that no one could have anticipated, which many labs choose to absorb centrally rather than assign. The specific structure matters less than the fact of having one, because a cost-recovery rule that everyone agreed to before any particular failure is one they will accept when a failure occurs, whereas a cost decision improvised after the fact leaves someone feeling they were charged unfairly regardless of the merits. Decide the principle once, apply it consistently, and the recurring argument disappears.

Building a Failure Log

The final discipline turns individual failures into intelligence, and it is the one most likely to be skipped because each failure feels like a one-off to be resolved and forgotten. A failure log, a simple, consistent record of every failed or marginal run, its triaged domain, its identified cause, and its resolution, converts a scatter of isolated frustrations into a dataset that reveals patterns no single failure could. A run that fails once is an incident; the same failure mode appearing five times in three months is a systemic problem the log makes visible and memory alone would miss.

That pattern visibility is what lets a lab move from repeatedly reacting to failures to preventing them, by revealing which causes recur and warrant a process change, a supplier conversation, or additional training. For labs doing or moving toward regulated work, the log is also the foundation of the formal corrective-and-preventive-action process quality systems require, as part of the framework covered in Quality and Compliance in NGS Labs: From Research Use to Regulated Testing. Even outside a regulated setting it is among the highest-return, lowest-cost habits a sequencing lab can adopt, and it connects to the throughput discipline, since failures that consume capacity are a throughput problem as much as a quality one, covered in Running NGS at Scale: Throughput, Scheduling, and Automation. The wider operational picture failure management sits within is in Next-Generation Sequencing in the Lab: A Manager’s Guide to Building, Budgeting, and Scaling NGS Capacity.

 

This article was produced under Lab Manager's AI Editorial Guidelines.

Add Lab Manager as a preferred source on Google

Add Lab Manager as a preferred Google source to see more of our trusted coverage.

Frequently Asked Questions (FAQs)

  • Why did my sequencing run fail?

    Start by triaging into one of two domains rather than jumping to a specific cause. If only certain samples failed while others in the same run performed well, the problem is most likely upstream, in those samples, their extraction, or their library preparation. If the whole run underperformed uniformly, the problem is more likely at the sequencing step, in the reagents, flow cell, run setup, or instrument, since those are common to every sample. That library-versus-instrument split resolves most failures into the right domain and determines who owns the fix, before any detailed metric interpretation.

  • What causes low cluster density?

    Low density of usable, data-generating clusters across a whole run generally points to a loading or library quantification issue or a run-level problem rather than to any individual sample, because it affects everything uniformly. Common contributors include loading too little material, an error in quantifying the library pool before loading, or a reagent or flow cell issue. Because the detailed interpretation of this and other run metrics is genuinely technical and platform-specific, read the metric far enough to place the failure in the instrument-or-run domain, then consult the detailed quality-metric guidance for the specific diagnosis.

  • Who pays for a failed sequencing run?

    That should be decided by a cost-recovery policy set in advance, not argued after each failure. The fair default usually follows the cause: a failure originating in a submitter’s sample or library is reasonably the submitter’s cost, while a failure originating in the sequencing operation’s reagents, instrument, or process is reasonably the lab’s, with an escalation path for ambiguous cases and a provision for no-fault failures like a bad reagent lot. A policy everyone agreed to before any particular failure is accepted when a failure occurs; a cost decision improvised afterward leaves someone feeling charged unfairly.

About the Author

  • Trevor Henderson headshot

    Trevor Henderson BSc (HK), MSc, PhD (c), has more than two decades of experience in the fields of scientific and technical writing, editing, and creative content creation. With academic training in the areas of human biology, physical anthropology, and community health, he has a broad skill set of both laboratory and analytical skills. Since 2013, he has been working with LabX Media Group developing content solutions that engage and inform scientists and laboratorians. He can be reached at thenderson@labmanager.com.

    View Full Profile

Related Topics

Related Articles

Current Magazine Issue Background Image

CURRENT ISSUE - September/2026

Are You Asking the Right Questions?

How Question Framing Shapes Better Lab Decisions

Lab Manager September 2026 Cover Image