Findable, accessible, interoperable, and reusable (FAIR) data principles have become even more critical now that AI is becoming the norm. AI is generating mountains of data via in silico analysis, labs are increasing throughput using AI agents, and researchers are still generating structured, unstructured, and semi structured data with wet bench experiments. All this data must be unified within a FAIR infrastructure so it can be used and reused without causing data havoc.
FAIR isn’t just a fad that will be important to labs now and replaced with something else down the line. By operating with FAIR principles, labs will accelerate throughput, streamline compliance, and create a single source of truth across the organization, providing labs and researchers with unified data access, whether generated by humans, instruments, or AI.
The philosophy of FAIR: A data continuum
“Living” by the FAIR principles isn’t a call for rigidity or a checklist of rules and regulations. It’s a continuum of data management across a variety of lab environments. If someone spends time and effort generating data, it should be used as much as possible. But if you can’t find the data, you can’t reuse it. Some of the best examples: rerunning an experiment because you can’t find the results of the previous experiment, or you can find the results but can’t reuse them because there isn’t sufficient information about how the experiment was run.
Operating with FAIR data principles across the organization minimizes barriers to data access. Its implementation speeds the research timeline by ensuring scientists stop wasting time repeating work that’s already been done.
Transforming your data for AI within the FAIR context
Pharma, life sciences, and academic research teams must support both local work and global collaboration, making organization-wide adherence to FAIR principles critical. The data, no matter the source, must be findable, accessible, and reusable across the organization; accessibility does not mean less security. Data access management rules still apply.
Data may be in traditional databases, within individual Excel files, or across SharePoint environments. It may be trapped within a variety of silos, based on vendor storage formats, department, experiment type, equipment type, etc. The priority is to get the data in order.
A variety of AI tools can be used to standardize data formats, transforming them from proprietary, semi structured, or unstructured data to more easily accessible formats. While AI can extract structure from unstructured data, the transformation may not be as accurate. Scientists will need to train the AI tools and models so they “understand” what needs to be done. Consistency is key. Ideally, of course, scientists will generate data in such ways as to ensure data is structured from day one.
Managing your AI- and human-generated lab data
While AI is the “new” way of doing things, procedures and standards based on learned knowledge and experience should still apply. Lab managers should designate a common structure or formal framework defining relationships, properties, and concepts across all experiments to streamline knowledge sharing within the FAIR environment.
It really takes a team to implement FAIR processes across the lab, including developers, UX design, data product teams, scientific experts, and team members with expertise in data governance. The team should include members from within the lab as well as a management sponsor who can build organizational buy-in.
Work closely with the IT team when creating a data management process. They fully understand the existing infrastructure, while you understand your specific requirements. Together, you can learn from their best practices, while they can help you custom craft the exact data flows and storage necessary for your lab’s success.
Using a purpose-built software system to manage your lab data can help with FAIRness. Many laboratory information management systems and electronic lab notebooks automatically provide connectivity to a variety of instruments, from scales to microscopes, to standard lab equipment, liquid handlers and inventory freezers. Lab managers will need to assess compatibility, data mappings, security, validation needs, and vendor support to ensure they’re choosing the right connectivity platforms.
Attributing the data source is critical; metadata is key. Metadata includes instrument maintenance and quality control records, the instrument settings, and the experimental protocols, with clear timestamps and provenance. In addition to creating a clear audit trail for FDA and EMA regulatory compliance, the process demonstrates the provenance of data, preventing redundant experiments. Further complementing FAIR in the AI age is ALCOA+, ensuring data is attributable, legible, contemporaneous, original, and accurate—whether human- or AI-generated. Meanwhile, the C in ALCOA is achieved by recording the data in real time, not doing an experiment and then writing up the results days after. Even in the age of AI, human behavior drives what’s possible within a FAIR environment.
Furthermore, AI data management is slowly but surely becoming standardized, with new avenues of connectivity, such as the model context protocol (MCP), which streamlines communications between AI and lab software platforms, allowing AI to “see” and use data tools securely, integrating AI- and human-generated data in a single source of truth. Agent-to-agent connectivity via the Agent2Agent (A2A) protocol is alternative to MCP. They both streamline the flow of data within and between organizations, from lab to lab and between a contract research organization (CRO) and its pharma clients. These massive data volumes can be stored on the cloud or within data warehouses or lakes, making them easily available for holistic analysis.
Building a FAIR environment across the organization
Lab R&D and lab management are already full-time jobs. While adding in the responsibility of data cleaning and FAIR principle implementation may add significant responsibilities in the short term, it benefits all labs and researchers across the entire organization for the long term.
Experimental data is not “owned” by a single individual; it’s an asset to all researchers across the organization. Establishing governance of a data management strategy enables everyone to benefit from these advances. Providing cross-organizational access to data could mean that a team working on drug A in lab Z can accelerate the discovery process by finding, accessing, and reusing data from the team working on drug B in lab Y.
Becoming a FAIR lab is another step in the evolution of lab research. It’s reminiscent of the transition from lab scientists holding all their research materials in individual cupboards and freezers to having a central sample management function.
Ensuring AI data stays FAIR
Managing and integrating AI into laboratory environments itself is changing the way data is handled, especially as AI generates high volumes of data when doing in silico experimentation. Of course, AI should be trained so the data it generates complies with the FAIR principles.
Making sure data and new data fall within the FAIR principles isn’t the end of the process. Researchers will be able to make faster, more informed decisions and increase operational efficiency with organization-wide access to the most valuable IP–the data. FAIR is the base upon which new discoveries will be built.
Implementing FAIR in your lab
It may not be easy to get all the relevant data within the FAIR “environment” for the first project. Starting small may be best. Establish a proof of concept by beginning with a specific protocol, project, or assay. Bioinformaticians have significant experience with data management in lab environments, so they can be asked to champion the projects if available. Ultimately, the key is making FAIR data principles a priority and adopting processes, procedures, tools, and protocols that make it easier for the individual scientist to work within a FAIR environment, ensuring the data is FAIR for everyone else.
It’s easy to get started. Select one high-value data workflow, document where the data originates and travels, identify the minimum required metadata, assign ownership, and define one measurable result. Track every step to ensure that it’s being developed appropriately. Once the guidelines are in place, it’s easier to scale the process within the lab, across labs, and into the greater organization.










