Artificial intelligence is finding its way into pathology and laboratory workflows, but identifying a useful application is only the first step. Labs also need to understand where automation can save expert time, how data quality shapes model performance, and what evidence is needed before a tool becomes part of routine decision-making.
Michael Staup and Aleksandra Zuraw, both with Charles River Laboratories, discuss where AI earns its keep in laboratory operations, and what has to be true before you rely on it.
Q: Where do you see AI providing the most practical value in pathology and laboratory operations today?
MS: Slide quality control. It may not be the most exciting application, but it is the most practical. Every case starts with a set of physical questions: is the tissue there, is there enough of it, is it the right tissue, was it cut and stained well, and did it scan cleanly? Today, that check is usually done by senior technicians one slide at a time, first at the microscope, and then again on the computer, often spending more time per slide than the pathologist will. An algorithm can run those same checks on the digital image automatically. That lets you design the workflow so the slide goes straight from the stainer to the scanner, problems get caught before they reach anyone's queue, and skilled technicians go back to skilled work.
AZ: And each of those checks is for a separate failure mode. A tool that reliably flags a missing section is not automatically a tool that can judge stain quality. They require separate performance testing.
Q: You've described AI as a way to reduce routine burden and help experts focus on meaningful signals. What does that look like in practice?
MS: Digital image analysis. Routine morphometry in developmental and neurotoxicity work or grading inflammation in colitis models used to sit on a pathologist's desk. Automated, it takes a fraction of the time and is more consistent run to run.
AZ: The computer does the counting; the expert does the interpreting. Instead of visually estimating how much of a section is inflamed, you get a measured number. Ordinal scoring is still the foundation of how findings get reported; image analysis adds continuous data underneath it. What changes is where expert attention goes: the scientist is no longer generating the number; they're deciding what it means biologically. That is the work their training is for.
Q: What should lab leaders understand about the relationship between an AI model's performance and the quality of the data and annotations used to develop it?
AZ: Garbage in, garbage out, and annotations are the part people underestimate. A model cannot be more consistent than the expert judgment encoded in its training labels. If they are rushed, inconsistent, or drawn by someone without the domain knowledge to recognize an edge case, the performance ceiling is set before a line of code runs.
Q: What types of validation and performance testing should take place before a lab relies on an AI tool within an operational workflow?
MS: It runs in a sequence. First, write down what the tool is meant to do and where, which is the context of use. Then ask the harder question: how can this fail, and what is the highest risk of automating the analysis within the context of the study design? Those two things drive the shape of your credibility assessment plan, which documents the performance testing strategy arrived at by your subject matter experts.
AZ: The metrics have to match what the tool is for. If its job is to catch bad slides, the number that matters is how many of the bad slides it caught, not how often it was right overall. Most slides are fine, so "right overall" mostly measures how well it handles the easy cases. Then, execute the plan: compare the output against a trained expert's assessment, or an orthogonal method, using held-out data. That means slides you set aside before training and never showed the model, drawn from the same kind of material it was built for. Testing on data it already learned from measures its memory, not performance. The FDA's January 2025 draft guidance on using AI in drug development lays out this whole sequence in seven steps. It applies when an AI output supports a regulatory decision about a drug, but the logic is sound for any lab.
Q: How should labs evaluate whether an AI system performs consistently across different sample types, disease states, instruments, sites, or patient populations?
AZ: Start from the assumption that it doesn't. An AI model developed on one type of data should not be expected to perform on data it never saw: a different stainer, a different scanner, a different species, a different population. This is domain shift, and the test is direct. Run the tool on the new data and compare against its performance on the development data. If it's worse, bring that data into training and revalidate.
MS: Anything that deviates from the parameters you defined in your context of use is outside what your credibility assessment covered. And separate from domain shift, there is biological drift. Even on the same kind of study material, performance can deteriorate over time.
Q: From a quality and regulatory perspective, what documentation should labs maintain to demonstrate how an AI tool was evaluated, validated, and used in decision-making?
MS: In GLP nonclinical work, three documents defined by the FDA guidance carry the weight: the context of use, the risk assessment, and the credibility assessment plan, with results captured in a report and signed by the subject matter expert who defined the context of use and approved the plan. We keep all three in one controlled document. Worth saying plainly, though: labs doing this well were doing it before the guidance existed. We had validated our tools, and the guidance largely confirmed our method. It gave the field shared vocabulary for evidence we were already generating.
AZ: The framework determines the form, not the substance. GLP, CLIA, and GMP labs each have their own document structures, and the authority overseeing yours tells you how AI fits inside it. Underneath, it's the same scientific method. Work out what would convince you the tool performs, then map that to whatever your framework calls those documents.
Q: For a lab leader who sees potential for AI but is unsure where to begin, what would you recommend evaluating before selecting a use case or technology?
MS: Ask what your team spends the most time on where their talents contribute the least. What is preventing you from getting the most value out of your people? Offload that. It saves time and money and gives staff a more rewarding job doing what they are trained to do.
AZ: Also pick something small enough that failure doesn't break the lab. A first win teaches your team how these tools behave and brings people along with you. And find someone who has already done what you're trying to do, then ask what went wrong. Publications tell you what worked; a conversation tells you what didn't. Don't adopt AI because it's expected of you. The use case has to be useful to your lab.












