In pharmaceutical R&D, the promise of artificial intelligence (AI) has often been framed in terms of speed: faster design cycles, faster screening, faster optimization, and faster decisions. But faster experimental execution does not automatically produce quicker or better development decisions.
The challenge may not be technology alone. As more AI-enabled tools enter the laboratory, lab leaders should ask whether traditional metrics used to evaluate automation truly capture what matters in exploratory research.
Lab automation inherited incomplete measures of success
For decades, laboratory productivity has been assessed using measures borrowed from manufacturing, including throughput, utilization rate, experiments per day, and cost per run. In a production setting, these can be appropriate measures because the process and output are defined in advance. For example, a manufacturing process that produces more units at a lower cost is genuinely more productive. The process repeats reliably, and the goal is to eliminate variability and waste. In that context, overall equipment effectiveness,1 a canonical manufacturing benchmark that combines availability, performance, and quality, can be useful because the intended process and output are already defined.
Exploratory R&D is different. Manufacturing metrics are not irrelevant, but they are an incomplete north star for discovery.
In science, goals change as evidence accumulates. A negative result, one that efficiently eliminates a hypothesis, can be as valuable as a positive one. The highest-value moment in a discovery program is often not when the best candidate is identified, but when the team realizes they have been asking the wrong question. A lab optimized only to minimize variability may overlook the uncertainty, iteration, and course correction that discovery depends on.
The first wave of automation made it easier to run more experiments. The next wave should make it easier to run the right experiments.
The throughput trap
High-throughput screening defined the early era of automated drug discovery. The underlying premise was that testing more compounds would increase the likelihood of identifying a promising candidate. Enormous resources followed this logic, resulting in larger screening platforms and rapidly expanding datasets.
Yet, over the same period in which these technologies became widespread, pharmaceutical R&D productivity fell sharply. Eroom's Law, named as the deliberate inverse of Moore's Law, describes an approximately 80-fold decline in the number of new drugs approved per billion dollars of inflation-adjusted R&D investment between the early 1950s and 2010. Over that same period, DNA sequencing became exponentially cheaper, computational power advanced rapidly, and high-throughput screening automated experiments that had previously taken weeks. The tools improved while productivity continued to move in the opposite direction.
Better tools, larger budgets, and faster execution did not automatically translate into better development outcomes. Today, this pattern persists. At BIO 2026, pharmaceutical executives broadly reported that AI tools had yet to deliver anticipated gains, citing data quality, integration challenges, and the disconnect between laboratory outputs and downstream development decisions as persistent barriers. Together, these challenges suggest that the field needs metrics that capture not only how quickly experiments are running, but how much each one contributes to the decisions that matter.
A better question: how much did you learn?
One of the most useful KPIs for AI-driven R&D is rarely featured prominently in vendor pitches: how many experiments did your system need to reach a decision, and what is it compared against?
This reframes the productivity question in a fundamental way. The goal is not to run as many experiments as possible. It is to make each experiment carry as much decision-relevant information as it can, reducing uncertainty about which direction to pursue next. AI-driven optimization approaches can support this by selecting subsequent experiments based on what the available data reveal. Rather than executing a fixed plan designed before any results exist, the system updates its model as results become available and uses that information to recommend what should be tested next.
The practical difference can be significant. Systems designed to maximize information gained per experiment have reached their targets in a fraction of the experimental budget required by conventional screening approaches.
These results can be translated into terms that matter operationally: uncertainty removed per experiment, per gram of active pharmaceutical ingredient (API) consumed, or per week of development time. A system that reduces the number of experiments needed to reach a defensible decision does not simply cut costs, it compresses the time between questions and answers.
What this means for formulation scientists
The distinction between throughput and learning is particularly consequential in pharmaceutical formulation. Formulation turns molecules into medicines and is increasingly recognized as a major value driver in drug development.
Formulation work operates under constraints that make experimental efficiency critical. APIs are often scarce and expensive during early development. Timelines between chemistry milestones and clinical entry may be constrained by formulation development, while bench-scale measurements are imperfect proxies for downstream development. In vitro release measurements, for example, do not always translate directly to how a drug product will behave in vivo.
In this environment, every formulation experiment needs to do more than generate a result. It needs to help the team decide what to test, change, or eliminate next. A formulation experiment that does not change what the team already knows consumes time and material that early-stage programs cannot spare. That is why sample efficiency is not a secondary consideration in formulation development—it is one of the central constraints work is planned around.
Questions worth asking
This shift from throughput to learning changes how lab managers should evaluate AI-driven automation platforms. The buying criteria should extend beyond traditional equipment specifications. Alongside throughput capacity, reproducibility, and instrument integration, the conversation should include questions about learning efficiency and decision quality, including:
- How many experiments does this system require to reach a decision, and what benchmark is that compared against?
- Does the system update its experiment recommendations as results come in, or does it execute a fixed experimental design at the outset?
- How quickly does a result translate into a revised experimental strategy, not simply a new data point, but an updated recommendation for what to run next?
- Are the targets the system optimizes for connected to what matters downstream, or only to what bench-scale instruments can conveniently measure?
These questions are not just theoretical. A 2024 analysis of self-driving laboratory studies in the field found that 65 percent of published implementations included no form of algorithmic comparison. Without relevant benchmarking against a reference strategy, it is effectively impossible to evaluate whether those systems were learning efficiently or simply running experiments. Asking vendors these questions directly is both reasonable and necessary.
The standard worth setting
The next productivity breakthrough in R&D will not come from running every experiment faster. It will come from needing fewer experiments because each one teaches more.
That does not mean laboratories should stop measuring essential operational metrics. It means those metrics should be paired with measures that reflect how effectively experiments reduce uncertainty and support decisions. In manufacturing, the ideal process minimizes unwanted variability around a defined output. In discovery, the ideal process converts uncertainty into evidence and evidence into better-informed hypotheses.
The labs that close the gap between those two standards, by treating each experimental result not just as a data point but as a training signal for what to do next, are building a compounding advantage. Laboratories that measure only how many experiments they ran may be missing the more important question: how much did those experiments help them learn?
References:
1. Nakajima S. Introduction to Total Productive Maintenance (TPM). Productivity Press; 1988.











