Reliable sample tracking in sequencing is the quiet discipline that separates a result you can trust from one you cannot, and its failures are uniquely hard to catch. A sample swap in a sequencing workflow does not announce itself. The run succeeds, the data looks normal, and the error surfaces only later, when a result does not match what was expected of a sample, often discovered by an analyst who was nowhere near the bench when the swap happened. By then the sample has passed through accessioning, extraction, library preparation, pooling, and sequencing, and reconstructing where identity was lost is difficult or impossible. The cost of a tracking failure is not just the wasted run; it is every decision made on a result attributed to the wrong sample.
This guide follows sample identity through the workflow, shows where it most often breaks, and covers the controls that hold it together, from barcoding and accessioning through the particular risks of pooling to the sequencing-specific safety net that catches swaps nothing else does: confirming identity from the data itself. It ends with what regulated work additionally requires. The theme throughout is that identity has to be maintained deliberately at every handoff, because the workflow gives no natural warning when it is lost.
Key Takeaways
|
Where Identity Breaks in an NGS Workflow
Sample identity is not lost at random. It is lost at specific, predictable points, and knowing where they are is the first step to controlling them. Every handoff between people, every transfer between vessels, and every transformation that changes what the sample physically is presents an opportunity for identity to detach from the material. The workflow stages themselves are safe; the moments between them are where swaps happen.
Workflow Stage | How Identity Is Lost Here | The Control That Holds It |
Accessioning | Transcription error linking sample to record | Machine-readable barcode assigned at receipt |
Extraction | Tube-to-tube transfer into an unlabeled or mislabeled vessel | Barcoded tubes; scan at each transfer |
Plate layout | A sample placed in the wrong well | Documented, verified plate maps; scan-to-well |
Pooling and indexing | Wrong index assigned, or index collision in the pool | Verified index-to-sample map; unique index checks |
Demultiplexing | Reads assigned to the wrong sample by a bad index map | Index map validated against the pooling record |
Result delivery | Result attributed to the wrong record downstream | Genotype-based identity confirmation |
Table 1. Where sample identity breaks in a sequencing workflow, and the control that prevents each failure. The stages are safe; the handoffs and transformations between them are where swaps occur.
Barcoding and Accessioning
Identity control begins at the door. The moment a sample is received, it should be assigned a unique, machine-readable identifier, a barcode, that travels with the physical sample and links it to its record for the rest of its life in the lab. Assigning that identifier at accessioning, before any handling, is what prevents the single most common tracking failure: a transcription error made while copying a handwritten or externally supplied identifier onto a tube. A barcode scanned is a barcode not mistyped.
The barcode is only as good as the discipline of scanning it at every subsequent step, which is where a laboratory information management system earns its place: it records each scan, building an unbroken electronic trail of where each sample has been and what was done to it. The transition from informal tracking to a system that enforces scanning at each handoff is one most growing labs make later than they should, and choosing a system suited to the genomics workflow is covered in LIMS and Informatics for Genomics Labs: What to Look For. The essential point is that the identifier must be machine-readable and scanned, not read and retyped, because every manual transcription is a chance to break the link.
Tracking Through Pooling and Indexing
Pooling is the most dangerous moment in the entire workflow for sample identity, and it deserves disproportionate attention because of what it does: it deliberately combines many individually labeled samples into a single tube, after which they are physically indistinguishable and can only be told apart by the molecular index assigned to each. From the moment of pooling, identity is no longer carried by the label on the vessel. It is carried entirely by the correctness of the index-to-sample assignment, and if that assignment is wrong, no physical inspection can reveal it.
Two failure modes matter here. The first is a wrong index-to-sample mapping, where the record of which index went on which sample is incorrect, so demultiplexing confidently assigns each sample’s reads to the wrong identity. The second is an index collision, where two samples in the same pool share or have too-similar indices, so their reads cannot be cleanly separated. Both are prevented by the same discipline: a verified, documented index-to-sample map created and checked before pooling, and a check that all indices in a pool are compatible and distinct. Because this is exactly the kind of step automation performs more reliably than manual assignment, it connects to the case for automating library prep and pooling in Running NGS at Scale: Throughput, Scheduling, and Automation.
Genotype-Based Identity Checks
Here is the safety net that exists only in sequencing, and it is the reason a well-run sequencing lab can catch swaps that would be invisible in almost any other setting. Because sequencing reads the sample’s own genome, the data carries an intrinsic identity signature: the specific pattern of variants a sample possesses is, in effect, a fingerprint of that sample. That fingerprint can be compared against a known reference for the sample, or checked for internal consistency, to confirm that the data at the end of the workflow actually came from the sample that entered it.
This turns identity from something that must be perfectly maintained into something that can also be verified after the fact. If a sample has a known prior genotype, whether from an earlier test or an orthogonal method, the sequencing result can be checked against it, and a mismatch flags a swap that every barcode and every scan missed. Even without a prior reference, samples expected to be related or unrelated can be checked for the expected relationship, catching swaps that violate it. Building a genotype-based identity check into the standard analysis is one of the highest-value controls a sequencing lab can adopt, precisely because it catches the failures that occur despite good barcoding, at the one point where they can still be caught before a wrong result is delivered.
The Data Can Confirm What the Labels Cannot Barcodes and scans prevent swaps going forward; a genotype-based identity check detects them looking back. The two are complementary, not alternatives. Prevention keeps identity intact through the workflow, but no prevention is perfect, and the genotype check is the last line that catches what slipped through, using the one thing a sequencing lab uniquely has: the sample’s own genome as proof of who it is. A lab that barcodes rigorously and never checks genotype identity is trusting that prevention never fails. A lab that does both has a way to know. |
Documentation for Regulated Work
For regulated or clinical sequencing, accurate tracking is necessary but not sufficient, because the standard is not only that identity was maintained but that it can be proven to have been maintained. Chain of custody in that setting means a documented, auditable record of every person who handled a sample, every transfer it underwent, and every check performed on its identity, retained in a form an inspection can examine. The tracking that a research lab does for its own confidence becomes, in a regulated lab, a formal record that is itself part of what is audited.
This is one more area where the habits are far easier to build from the start than to retrofit, since reconstructing a chain of custody that was not documented as it happened is generally impossible. A lab moving toward regulated work should treat its tracking system as the foundation of a future chain-of-custody record, and the broader quality and compliance requirements this sits within are covered in Quality and Compliance in NGS Labs: From Research Use to Regulated Testing. The principle holds across research and regulated settings alike: maintain identity deliberately at every handoff, verify it from the data where you can, and, where it must be proven, document it as it happens. How sample tracking fits the whole operational picture is covered in Next-Generation Sequencing in the Lab: A Manager’s Guide to Building, Budgeting, and Scaling NGS Capacity.
This article was produced under Lab Manager's AI Editorial Guidelines.

















