Data Retention, Backup, and Archiving Policy for Sequencing Labs

Most labs have no retention policy, which is itself a policy: keep everything, forever, at increasing cost, with no defensible basis for deletion.

Written byTrevor J Henderson
| 7 min read
A data manager reviews a printed retention policy document beside a genomics lab, illustrating the written policy that governs what sequencing data to keep, back up, archive, and delete.
Register for free to listen to this article
Listen with Speechify
0:00
7:00

Writing a genomic data retention policy is a task most sequencing labs postpone until storage cost forces the question, and the postponement is itself a decision with a real price. A lab without a written retention policy does not thereby retain nothing; it retains everything, indefinitely, because in the absence of a rule that authorizes deletion, deleting anything is a risk no one wants to take. The result is the most expensive possible policy, unlimited accumulation, adopted by default rather than by choice. A written policy is what replaces that default with a deliberate decision: keep what funders, regulators, and consent require, keep what is scientifically worth keeping, and delete the rest on a defensible schedule.

This guide covers what goes into such a policy: what funders and regulators actually require, how the consent under which samples were collected constrains retention, how to decide which files to keep, the difference between backup and archive that labs routinely conflate, the deletion procedures that make deletion defensible, and a template to structure your own policy. A caution applies throughout: specific retention periods vary by funder, institution, and jurisdiction, so the figures here are reference points to confirm against your own obligations, not universal rules.


Key Takeaways

  • Having no retention policy is a policy: keep everything forever, the most expensive option, chosen by default.
  • Funder and institutional rules set a retention floor. NIH requires shared scientific data to remain available at least 3 years after grant closeout; institutions often require longer.
  • The consent under which samples were collected can constrain how long data may be kept and for what purpose, independently of funder rules.
  • Backup and archive are different things: backup protects against loss, archive is long-term retention of data you are keeping deliberately.
  • Deletion has to be documented to be defensible. A policy is what lets you delete without risk, and what makes keeping everything unnecessary.

 

What Funders and Regulators Require

Retention policy starts from the external requirements, because these set a floor below which the lab cannot go regardless of what it would prefer on cost grounds. Funders impose data-retention and data-availability requirements as a condition of funding, and these are the first constraint to establish. In the United States, the National Institutes of Health, under its data management and sharing policy, requires that scientific data generated by funded research be made available for at least three years following the closeout of the grant, and its Genomic Data Sharing policy applies specific expectations to research generating large-scale human or non-human genomic data. Institutions frequently impose their own requirements on top, commonly retaining research data for several years beyond publication, and some require six years after study closure for research records.

These figures are floors, not targets, and they are jurisdiction- and funder-specific: a lab funded by a different agency, or operating in another country such as Canada under its national funding agencies and provincial rules, faces its own requirements that may be longer or structured differently. Regulated clinical testing adds a further layer, with its own record-retention obligations distinct from research funder rules, as part of the compliance picture covered in Quality and Compliance in NGS Labs: From Research Use to Regulated Testing. The practical first step is to gather the specific retention requirements of every funder, institution, and regulator the lab answers to, and take the longest applicable period as the floor. Confirm your own, because the numbers above are reference points, not a substitute for the requirements that actually bind you.

Lab manager academy logo

Advanced Lab Management Certificate

The Advanced Lab Management certificate is more than training—it’s a professional advantage.

Gain critical skills and IACET-approved CEUs that make a measurable difference.

Consent Terms as a Retention Constraint

Consent is the constraint labs most often overlook, and it can both require and forbid retention in ways that override a purely operational decision. The consent under which a sample was collected defines what may be done with the data derived from it, and that can include limits on how long the data may be kept, what it may be used for, and whether it must be destroyed after a defined purpose is fulfilled. A retention policy that ignores consent can put the lab in the position of holding data it no longer has permission to hold, which is a compliance and ethical problem, not merely a cost one.

Consent can also work in the other direction, permitting or requiring long-term retention and future use where participants agreed to it, which is often what makes valuable data reusable for later research. The point is that consent terms are an input to the retention policy as real as any funder rule, and a policy has to reconcile them: where consent is more restrictive than the funder floor, consent governs. 

Interested in life sciences?

Register for a FREE Lab Manager account to subscribe to our Life Sciences Newsletter.
Subscribe for Free

Deciding What to Keep: Raw, Aligned, or Called

Within the boundaries the external requirements set, the lab decides which of its files to keep, and this is where retention policy meets the file-type economics directly, because the three main classes of sequencing data differ enormously in both size and reconstructability. The decision turns on a simple principle: keep what is expensive or impossible to regenerate, and let go of what can be recreated or is superseded, subject always to the external floor. The sizes of these file types, which drive the cost side of this decision, are detailed in How Big Is NGS Data? Sizing Storage for Your Sequencing Volume.

File Class

Default Retention Stance

Reasoning

Raw (FASTQ)

Retain for a defined window, then reconsider

Largest and most expensive; regenerable from aligned data in many cases; the first candidate to age out

Aligned (BAM/CRAM)

Retain long term, compressed as CRAM

Preserves full information for reanalysis; CRAM roughly halves the cost of keeping it

Called (VCF)

Retain indefinitely

Small, interpreted result; cheap to keep and expensive to regenerate; often the scientifically essential record

Table 1. A default keep-decision by file class, subject always to funder, regulator, and consent requirements. Keep what is expensive or impossible to regenerate; let the largest, most regenerable class age out first.

This default, keep the small interpreted results indefinitely, keep the aligned data compressed for reanalysis, and retain raw data only for a defined window unless a requirement says otherwise, balances scientific value against cost for many research labs. It is a starting point to adapt, not a rule: a lab whose work depends on periodic reanalysis from raw data will weigh raw retention differently, and any external requirement to keep a given class overrides the cost logic entirely.

Backup vs. Archive

Backup and archive are routinely spoken of as if they were the same thing, and conflating them produces policies that do neither well, so a retention policy has to treat them as the distinct functions they are. Backup is protection against loss: a second copy of data you are actively using, kept so that a hardware failure, accidental deletion, or disaster does not destroy the only copy. Its purpose is recovery, its horizon is short, and it is judged by how quickly and completely you can restore after a failure. Archive is long-term retention: data you have deliberately decided to keep, moved to durable, lower-cost storage because you will rarely access it but must or want to retain it. Its purpose is preservation, its horizon is years, and it is judged by durability and cost.

The distinction matters operationally because the two have different requirements. A backup that cannot be restored quickly fails at its job even if the data is intact; an archive optimized for cheap long-term storage may be slow to retrieve, and that is acceptable. A common and costly error is treating an archive as a backup, assuming that data pushed to cheap cold storage is also protected against loss, when a single archived copy with no second copy is not backed up at all. A sound policy specifies both: what is backed up and how it is restored, and what is archived and on what terms. The storage tiers that archive uses, and their retrieval-cost tradeoffs, are a central part of the data-management picture.

Deletion Procedures and Documentation

Deletion is the part of a retention policy that makes the rest of it real, because a policy that never actually deletes anything is just accumulation with extra steps, and the reason labs avoid deletion is that undocumented deletion feels risky. The resolution is to make deletion procedural and documented: the policy defines when data becomes eligible for deletion, what approval deletion requires, and how the deletion is recorded, so that deleting data at the end of its retention period is a routine, defensible, logged action rather than a nervous judgment call. Documenting that data was deleted according to policy, by whom and when, is what turns deletion from a liability into a demonstration of good governance.

Two safeguards make a deletion procedure trustworthy. First, deletion should require a check that no active requirement, funder, regulator, consent, or ongoing scientific need, still applies to the data in question, so that nothing is deleted that must be kept. Second, the deletion should be recorded in a durable log, creating the evidence that the lab followed its own policy. With those in place, deletion becomes the ordinary mechanism that keeps storage cost bounded rather than the frightening act it is in a policy-free lab. This is the direct answer to the hook: the policy is precisely what gives the lab a defensible basis to delete, which is what makes keeping everything forever unnecessary.

A Policy Template

The elements below are the sections a workable genomic data retention policy needs. Adapt each to your own funders, institution, jurisdiction, consent terms, and volume, and the result is a document that replaces default accumulation with deliberate, defensible retention.

  1. Scope. The data types the policy covers (raw, aligned, called, and any metadata), and the work it applies to.
  2. External requirements. The funder, institutional, regulatory, and jurisdictional retention floors that apply, with the longest applicable period identified as the governing minimum.
  3. Consent constraints. How consent terms limit or extend retention, and the rule that the more restrictive of consent and funder requirements governs.
  4. Retention schedule by file class. How long each file class is kept and on what storage tier, consistent with the external floor.
  5. Backup. What is backed up, how often, and the tested procedure for restoring it.
  6. Archive. What is archived, to what storage, and under what access and durability terms.
  7. Deletion. When data becomes eligible for deletion, the pre-deletion check for active requirements, the approval needed, and how deletion is logged.
  8. Review. How often the policy is reviewed and who owns it, since requirements and volumes change.

A policy built from these sections is short, specific, and enforceable, and it is the document that turns the whole data-management effort from a growing liability into a controlled system. How retention fits the full storage, compute, and staffing picture is in Managing NGS Data: Storage, Compute, Retention, and Staffing, and the full operational context of running a sequencing program is in Next-Generation Sequencing in the Lab: A Manager’s Guide to Building, Budgeting, and Scaling NGS Capacity.

 

This article was produced under Lab Manager's AI Editorial Guidelines.

Add Lab Manager as a preferred source on Google

Add Lab Manager as a preferred Google source to see more of our trusted coverage.

Frequently Asked Questions (FAQs)

  • How long must sequencing data be retained?

    It depends on your funders, your institution, your regulators, and the consent under which the samples were collected, so there is no single universal period. As a reference point, the US National Institutes of Health requires scientific data from funded research to be made available for at least three years following grant closeout, and institutions frequently require longer, commonly several years beyond publication or six years after study closure for research records. Regulated clinical testing carries its own separate obligations. Gather every requirement that applies to your lab and take the longest as your floor, and confirm your own figures rather than relying on a general number.

  • Can I delete FASTQ files after alignment?

    Sometimes, but only after checking the constraints. Raw FASTQ is the largest and most expensive file class, and because aligned data can regenerate it in many cases, raw data is often the first candidate to age out under a retention policy. However, deletion is only permissible if no funder, regulator, or consent requirement obliges you to keep the raw data, and if your workflow does not depend on periodic reanalysis directly from raw. The sound approach is to retain raw data for a defined window and delete it on a documented schedule once you have confirmed no active requirement applies, rather than deleting ad hoc.

  • What should a data retention policy include?

    A workable policy defines its scope, the external retention floors from funders, institution, regulators, and jurisdiction, the consent constraints that can override them, a retention schedule for each file class on each storage tier, the backup arrangements and their tested restore procedure, the archive arrangements, the deletion procedure including a pre-deletion check and a deletion log, and a review cycle with a named owner. Built from these sections, the policy replaces default indefinite accumulation with deliberate, documented, defensible retention that satisfies external requirements while keeping storage cost bounded.

About the Author

  • Trevor Henderson headshot

    Trevor Henderson BSc (HK), MSc, PhD (c), has more than two decades of experience in the fields of scientific and technical writing, editing, and creative content creation. With academic training in the areas of human biology, physical anthropology, and community health, he has a broad skill set of both laboratory and analytical skills. Since 2013, he has been working with LabX Media Group developing content solutions that engage and inform scientists and laboratorians. He can be reached at thenderson@labmanager.com.

    View Full Profile

Related Topics

Related Articles

Current Magazine Issue Background Image

CURRENT ISSUE - September/2026

Are You Asking the Right Questions?

How Question Framing Shapes Better Lab Decisions

Lab Manager September 2026 Cover Image