Writing a genomic data retention policy is a task most sequencing labs postpone until storage cost forces the question, and the postponement is itself a decision with a real price. A lab without a written retention policy does not thereby retain nothing; it retains everything, indefinitely, because in the absence of a rule that authorizes deletion, deleting anything is a risk no one wants to take. The result is the most expensive possible policy, unlimited accumulation, adopted by default rather than by choice. A written policy is what replaces that default with a deliberate decision: keep what funders, regulators, and consent require, keep what is scientifically worth keeping, and delete the rest on a defensible schedule.
This guide covers what goes into such a policy: what funders and regulators actually require, how the consent under which samples were collected constrains retention, how to decide which files to keep, the difference between backup and archive that labs routinely conflate, the deletion procedures that make deletion defensible, and a template to structure your own policy. A caution applies throughout: specific retention periods vary by funder, institution, and jurisdiction, so the figures here are reference points to confirm against your own obligations, not universal rules.
Key Takeaways
|
What Funders and Regulators Require
Retention policy starts from the external requirements, because these set a floor below which the lab cannot go regardless of what it would prefer on cost grounds. Funders impose data-retention and data-availability requirements as a condition of funding, and these are the first constraint to establish. In the United States, the National Institutes of Health, under its data management and sharing policy, requires that scientific data generated by funded research be made available for at least three years following the closeout of the grant, and its Genomic Data Sharing policy applies specific expectations to research generating large-scale human or non-human genomic data. Institutions frequently impose their own requirements on top, commonly retaining research data for several years beyond publication, and some require six years after study closure for research records.
These figures are floors, not targets, and they are jurisdiction- and funder-specific: a lab funded by a different agency, or operating in another country such as Canada under its national funding agencies and provincial rules, faces its own requirements that may be longer or structured differently. Regulated clinical testing adds a further layer, with its own record-retention obligations distinct from research funder rules, as part of the compliance picture covered in Quality and Compliance in NGS Labs: From Research Use to Regulated Testing. The practical first step is to gather the specific retention requirements of every funder, institution, and regulator the lab answers to, and take the longest applicable period as the floor. Confirm your own, because the numbers above are reference points, not a substitute for the requirements that actually bind you.
Consent Terms as a Retention Constraint
Consent is the constraint labs most often overlook, and it can both require and forbid retention in ways that override a purely operational decision. The consent under which a sample was collected defines what may be done with the data derived from it, and that can include limits on how long the data may be kept, what it may be used for, and whether it must be destroyed after a defined purpose is fulfilled. A retention policy that ignores consent can put the lab in the position of holding data it no longer has permission to hold, which is a compliance and ethical problem, not merely a cost one.
Consent can also work in the other direction, permitting or requiring long-term retention and future use where participants agreed to it, which is often what makes valuable data reusable for later research. The point is that consent terms are an input to the retention policy as real as any funder rule, and a policy has to reconcile them: where consent is more restrictive than the funder floor, consent governs.
Deciding What to Keep: Raw, Aligned, or Called
Within the boundaries the external requirements set, the lab decides which of its files to keep, and this is where retention policy meets the file-type economics directly, because the three main classes of sequencing data differ enormously in both size and reconstructability. The decision turns on a simple principle: keep what is expensive or impossible to regenerate, and let go of what can be recreated or is superseded, subject always to the external floor. The sizes of these file types, which drive the cost side of this decision, are detailed in How Big Is NGS Data? Sizing Storage for Your Sequencing Volume.
File Class | Default Retention Stance | Reasoning |
Raw (FASTQ) | Retain for a defined window, then reconsider | Largest and most expensive; regenerable from aligned data in many cases; the first candidate to age out |
Aligned (BAM/CRAM) | Retain long term, compressed as CRAM | Preserves full information for reanalysis; CRAM roughly halves the cost of keeping it |
Called (VCF) | Retain indefinitely | Small, interpreted result; cheap to keep and expensive to regenerate; often the scientifically essential record |
Table 1. A default keep-decision by file class, subject always to funder, regulator, and consent requirements. Keep what is expensive or impossible to regenerate; let the largest, most regenerable class age out first.
This default, keep the small interpreted results indefinitely, keep the aligned data compressed for reanalysis, and retain raw data only for a defined window unless a requirement says otherwise, balances scientific value against cost for many research labs. It is a starting point to adapt, not a rule: a lab whose work depends on periodic reanalysis from raw data will weigh raw retention differently, and any external requirement to keep a given class overrides the cost logic entirely.
Backup vs. Archive
Backup and archive are routinely spoken of as if they were the same thing, and conflating them produces policies that do neither well, so a retention policy has to treat them as the distinct functions they are. Backup is protection against loss: a second copy of data you are actively using, kept so that a hardware failure, accidental deletion, or disaster does not destroy the only copy. Its purpose is recovery, its horizon is short, and it is judged by how quickly and completely you can restore after a failure. Archive is long-term retention: data you have deliberately decided to keep, moved to durable, lower-cost storage because you will rarely access it but must or want to retain it. Its purpose is preservation, its horizon is years, and it is judged by durability and cost.
The distinction matters operationally because the two have different requirements. A backup that cannot be restored quickly fails at its job even if the data is intact; an archive optimized for cheap long-term storage may be slow to retrieve, and that is acceptable. A common and costly error is treating an archive as a backup, assuming that data pushed to cheap cold storage is also protected against loss, when a single archived copy with no second copy is not backed up at all. A sound policy specifies both: what is backed up and how it is restored, and what is archived and on what terms. The storage tiers that archive uses, and their retrieval-cost tradeoffs, are a central part of the data-management picture.
Deletion Procedures and Documentation
Deletion is the part of a retention policy that makes the rest of it real, because a policy that never actually deletes anything is just accumulation with extra steps, and the reason labs avoid deletion is that undocumented deletion feels risky. The resolution is to make deletion procedural and documented: the policy defines when data becomes eligible for deletion, what approval deletion requires, and how the deletion is recorded, so that deleting data at the end of its retention period is a routine, defensible, logged action rather than a nervous judgment call. Documenting that data was deleted according to policy, by whom and when, is what turns deletion from a liability into a demonstration of good governance.
Two safeguards make a deletion procedure trustworthy. First, deletion should require a check that no active requirement, funder, regulator, consent, or ongoing scientific need, still applies to the data in question, so that nothing is deleted that must be kept. Second, the deletion should be recorded in a durable log, creating the evidence that the lab followed its own policy. With those in place, deletion becomes the ordinary mechanism that keeps storage cost bounded rather than the frightening act it is in a policy-free lab. This is the direct answer to the hook: the policy is precisely what gives the lab a defensible basis to delete, which is what makes keeping everything forever unnecessary.
A Policy Template
The elements below are the sections a workable genomic data retention policy needs. Adapt each to your own funders, institution, jurisdiction, consent terms, and volume, and the result is a document that replaces default accumulation with deliberate, defensible retention.
- Scope. The data types the policy covers (raw, aligned, called, and any metadata), and the work it applies to.
- External requirements. The funder, institutional, regulatory, and jurisdictional retention floors that apply, with the longest applicable period identified as the governing minimum.
- Consent constraints. How consent terms limit or extend retention, and the rule that the more restrictive of consent and funder requirements governs.
- Retention schedule by file class. How long each file class is kept and on what storage tier, consistent with the external floor.
- Backup. What is backed up, how often, and the tested procedure for restoring it.
- Archive. What is archived, to what storage, and under what access and durability terms.
- Deletion. When data becomes eligible for deletion, the pre-deletion check for active requirements, the approval needed, and how deletion is logged.
- Review. How often the policy is reviewed and who owns it, since requirements and volumes change.
A policy built from these sections is short, specific, and enforceable, and it is the document that turns the whole data-management effort from a growing liability into a controlled system. How retention fits the full storage, compute, and staffing picture is in Managing NGS Data: Storage, Compute, Retention, and Staffing, and the full operational context of running a sequencing program is in Next-Generation Sequencing in the Lab: A Manager’s Guide to Building, Budgeting, and Scaling NGS Capacity.
This article was produced under Lab Manager's AI Editorial Guidelines.















