AI Thinks Fast.Biology Experiments Must Catch Up.

AI can now reason quickly. In mathematics, a capable model, a good reasoning method, and a verifierVerifier. A tool that tells the AI if an answer is correct. In mathematics, it can be a program that checks a proof. In biology, it is the experiment. An experiment is a weaker verifier than a proof checker. It tests a claim only in the cells and conditions used, so biology needs feedback that is strong enough, not just fast enough, to separate competing explanations. have already solved problems that human experts had not. In the physical sciences, that verifier is usually an experiment, and most experiments are slow. Particularly in biology, our experimental tools are nowhere near the speed of AI or the needs of medicine, turning ‘lab-in-the-loop’ into ‘lab-blocking-in-the-loop’.

This is an efficiency problem, so we first need a clear objective. We measure AI by how fast and cheaply it reaches a correct answer, and experiments in the AI era should be measured the same way. Science is built on causal claims that experiments support. An LLM, like a scientist, takes causal claims and produces new claims to test and plans for new experiments. The objective should be how many causal claims experiments settle (support or rule out) per hour, at a given cost.Counting claims. The count is a rough guide, since one result can be written as one claim or a thousand. What matters is how fast the questions that decide the next step get answered.

A causal claim looks like this: “p53 stops the cell cycle at G1 through p21.”p53 and p21. Two proteins in the cell. p53 stops damaged cells from dividing, and p21 is one of the proteins it uses to do this. In p21-deficient mice, G1 arrest was impaired while other p53 functions remained, so p21 mediates this arrest, not everything p53 does.Cell cycle. The series of steps by which a cell copies its DNA and divides into two cells. G1 is the gap before the cell starts copying its DNA. To test it, activate p53 and check that cells stop at G1 and p21 rises. Then remove p21 and activate p53 again. If cells no longer stop, and adding p21 back restores the stop, the claim is supported. Each experiment was chosen after seeing the last. That sequence is the loop this essay is about.

But claims are not all equal. If our goal is to improve human health, a claim’s value is how much it changes what we know about a disease or can do about it. A claim about a famous cancer gene can be worth little; one showing that an unknown protein is needed for a disease can be worth a lot. Drug targetsDrug target. The molecule in the body, usually a protein, that a drug acts on. with strong evidence in cells, mice, and sometimes humans usually come from claims like these.

Some of the earliest such claims came from the cell cycle. Labs must grow cells to study them, and growth was easy to measure, even with the simple microscopes of a century ago. Cancer is uncontrolled cell growth, so the chain was short: cell cycle, cell growth, cancer. Today the cell cycle is among the best-understood processes in the cell, and what remains is harder. Recent advances often needed longer chains. Immunotherapy, which helps immune cells attack cancer, came from studying how cells interact.

Many causal claims behind today’s therapies came from studying one gene at a time, before omics. Around 2000, biologists began to see the limits: genes act in networks, and the same gene can act differently in different cell types and conditions. Meanwhile, the first draft of the human genome appeared, and microarrays could measure thousands of genes at once. Together, these gave rise to systems biologySystems biology. A field that measures many parts of the cell at the same time to study how they work together. Systems biology was never purely observational. Early work such as Ideker et al. (2001) already cycled between perturbing, measuring broadly, and updating a model. Many later large datasets, however, were observational., which measures many signals in the cell simultaneously. Some of these studies perturb cells; others observe them. But observation is often an experiment too. Comparing patients with healthy donors, or following tissue across age or development, varies one thing (disease, age or stage) and measures everything else.

Run at scale for 30 years, with cheap, broad measurements on samples from across the body, this approach gave us many atlasesAtlas. A large reference map of which genes and molecules are active in each cell type and tissue of the body.. They describe cell types and molecular states, provide pretraining data for biological AI models, and have already pointed to early-stage molecular targets for many diseases. Therapeutic development, though, often needs tight causality across scales: how a molecule leads to disease through every level in between. Between a protein and a human lie groups of proteins, cell parts, cells, tissues, organs, and the full body.Levels. Climbing levels is not the goal in itself. An organ-level readout is not automatically closer to health than a molecular one. What matters is whether the model and readout are relevant to the disease. Comparisons such as disease versus healthy, old versus young, or female versus male are usually not enough to give the causal evidence therapeutic development needs.

Show How long biological events take, at every scale
How long biological events take
These selected examples use different endpoints. There is no universal link between how big a thing is and how long it takes.
Hover a point for the numbers, the source, and what it has to do with disease.

I think the goal of experiments is not more samples or more measurements per sample, but more health-linked causal claims settled per hour. Two numbers matter, and they are easy to confuse. Throughput is how many results come out each hour. Loop time is how long it takes from choosing an experiment to getting a result that can change the next choice. A lab that starts a 72-hour experiment every hour gets a result every hour, but any experiment chosen in response to one still waits 72 hours. Throughput is enough when experiments are independent. Loop time matters when each result changes what to try next, and that is where a fast-reasoning AI helps most.

Show What happens to the belt when the first result arrives
What happens when the first result arrives
Six culture slots, each running a 24-hour experiment: 8 hours to prepare the cells, then 16 hours of biology and readout. New experiments start every 2 hours. The first answer lands at hour 24 and changes the question. What should the rest of the belt do?
QUESTIONS ANSWERED BY 72 H
0
SLOT-HOURS SPENT ON A QUESTION ALREADY ANSWERED
0h
BATCHES OF CELLS THROWN AWAY
0
Pick what the AI decides when the first answer lands, and watch what the rest of the belt does.

Some biological processes are slow by nature, which sets a limit. But we often do not need to measure everything at once, especially since breadth often trades off against loop speed. Instead, change one variable quickly: one gene, a group of related genes, or one drug. Then measure, just as quickly, only what the claim needs: a few genes, a thousand, or, if necessary, the full transcriptome or epigenomeOmics, transcriptome, epigenome. Omics methods measure all molecules of one kind at the same time. The transcriptome is the activity of all the genes in a cell. The epigenome is the set of chemical marks on the DNA that control which genes are active.. Early evidence can save us a week on the wrong experiment.

Show How fast each tool acts, and how long until the cell recovers
How fast each tool acts, and how long until the cell recovers
Hover a row for what the tool is; select it for the full detail. The bars measure different things — a release, a halving, or a chosen observation window — so read each row’s own definition before comparing them.
The “can you undo it?” column is a judgement, not a measurement: it says whether the tool can be switched back, and roughly how fast the cell follows. Any dose big enough to kill or permanently reprogram a cell is not reversible, whatever the tool. These are single steps, not whole experiments. “Half-time” means half of the change has happened, not all of it. Building the cell line, delivering the reagent, waiting for the cell to respond and measuring it all take their own time on top.
Hover a row for what the tool is; select it for the full detail.

Broad and focused data play different roles for AI. Broad datasets are used for pretraining. Fast change-and-measure loops could be the reinforcement learning (RL) environment, where an AI chooses what to test and sees the result.Reinforcement learning. RL is one option, not a requirement; an agent can revise its plans from results without updating its weights. If used for RL, the reward should be for settling questions, including negative results. Rewarding positive claims would favor easy or overstated conclusions. Together, they let AI test causal claims and use the results to choose the next experiment.

I used Codex and Claude to review technologies that could shorten these loops. The figures group them by the scale they observe. I see four practical changes worth considering.

Show Fourteen ways to measure a cell, and where each one’s time goes
Fourteen ways to measure a cell, and where each one’s time goes
Select a row for what the method does, how often you can look, and its sources. Each bar uses the slowest end of the estimated times, with measuring and analysis counted together. These are planning estimates, not measured breakdowns.
get cells ready wait for the biology measure and work out the answer
Waiting for a busy instrument and redoing failed runs are not counted here. Neither is building a new cell line, growing immune cells, or repeating an experiment because the cells turned out to have an unexpected history. Add those when they apply.
Select a row for the details and sources.

First, reduce repeated setup. Measure live cells several times, then divide the perturbed culture among compatible endpoint assays, such as sequencing and microscopy.

Second, keep cultures growing and stagger the experiments so cells are ready when needed. This reduces gaps between experiments, although it does not shorten a 72-hour response. Early readouts can help if they tell us to change the plan before the experiment ends. Keeping this running takes considerable culture capacity. At 80% yield, one valid result per hour requires starting 1.25 batches per hour. If each batch needs a day of growth, all those batches must be prepared in advance.

Third, use barcodes to trace outcomes to their perturbations, allowing experiments to share a pooled readout. This can test many claims in one pass. It raises throughput even when the time to a result stays the same.

Fourth, prepare the tools needed for the next experiment. A fast readout helps little if building the next reporter or cell line takes two weeks. Keep common reagents available, install reusable switches in cells, and maintain cultures with the treatment histories the next test may need.

These changes require keeping track of what is growing, what each well has been through, and which experiments can start with the cells and tools available. Every result may change that schedule. An agent could help track this information and choose the next experiment, provided it knows the lab’s actual constraints.

Here is how I would combine these ideas in a machine.

Show A possible machine for fast loops
A possible machine for fast loops
Station times are design targets for a chosen workload, not measured.
The same machine runs three speeds of measurement. Fast: about 20–60 min, for early signals. Slower: hours to days, to see a cell divide, die, kill or recover. Deep: on request, for RNA panels or electron microscopy. Starred times assume the cells are already growing and nothing is queued; the electron microscopy figure is its full published workflow. Cells are only grown from scratch once, not on every pass.
Live cells go round the loop keeping their history; the one-way exit leads to deeper measurements that use the cells up. Select a station to see what limits it.

Science has always moved by making a guess and letting an experiment correct it, and AI has so far sped up only the first half of that exchange. A model can now propose more careful claims in an afternoon than a lab could test in a year, which leaves the cell, growing at its own unhurried pace, as the slowest participant in the conversation. I find something hopeful in that, because cells have answered our questions faithfully for as long as we have known how to ask them, and what we need now is a way to ask more often and hear the answer sooner. I like to picture the plate in this machine turning beneath its stations through the night, each well carrying its small history forward until the next question arrives. Most of the answers that come back will rule a claim out, and each of those spares some future scientist a month spent on a path that leads nowhere. Every so often an answer will hold up through every level between a single protein and a patient, and that kind of claim can become a target worth building a medicine around. The walk from molecule to person has always been long, and the hope behind this machine is that each step of it could take hours where it now takes weeks. That will not come from faster instruments alone. The experiment and the reasoning that drives it will have to be designed together, each built around what the other needs: assays chosen because an agent can act on what they return, and agents built around what a living cell can actually be asked.

—

Sources

Primary studies, protocols, and documentation cited in the figures.