AI Thinks Fast.Biology Experiments Must Catch Up.AI 运算如飞,生物实验亟待破局
AI can now reason quickly. In mathematics, a capable model, a good reasoning method, and a verifierVerifier. A tool that tells the AI if an answer is correct. In mathematics, it can be a program that checks a proof. In biology, it is the experiment. An experiment is a weaker verifier than a proof checker. It tests a claim only in the cells and conditions used, so biology needs feedback that is strong enough, not just fast enough, to separate competing explanations. have already solved problems that human experts had not. In the physical sciences, that verifier is usually an experiment, and most experiments are slow. Particularly in biology, our experimental tools are nowhere near the speed of AI or the needs of medicine, turning ‘lab-in-the-loop’ into ‘lab-blocking-in-the-loop’.
如今,AI 已具备强大的快速推理能力。在数学领域,凭借强大的模型、出色的推理方法以及高效的验证器,AI 已攻克了人类专家都曾束手无策的难题。而在自然科学领域,所谓的“验证器”通常就是实验,但大多数实验的进程却十分缓慢。尤其是在生物学领域,现有实验工具的运行速度远远跟不上 AI 的推演速度,也满足不了医学发展的迫切需求。这就导致原本理想的“实验在环”(lab-in-the-loop)模式,沦为了“实验在环中受阻”(lab-blocking-in-the-loop)的瓶颈。
This is an efficiency problem, so we first need a clear objective. We measure AI by how fast and cheaply it reaches a correct answer, and experiments in the AI era should be measured the same way. Science is built on causal claims that experiments support. An LLM, like a scientist, takes causal claims and produces new claims to test and plans for new experiments. The objective should be how many causal claims experiments settle (support or rule out) per hour, at a given cost.Counting claims. The count is a rough guide, since one result can be written as one claim or a thousand. What matters is how fast the questions that decide the next step get answered.
归根结底,这是一个效率问题,因此我们首先需要确立一个清晰的目标。我们衡量 AI 的标准是它得出正确答案的速度与成本,而在 AI 时代,我们也应以同样的标准来衡量实验。科学的基石在于实验所支撑的“因果论断”(causal claims)。大语言模型(LLM)就像科学家一样,接收已有的因果论断,推演出需要检验的新论断,并为新实验制定计划。因此,我们的核心目标应定为:在给定成本下,实验每小时能够敲定(证实或证伪)多少个因果论断。
A causal claim looks like this: “p53 stops the cell cycle at G1 through p21.”p53 and p21. Two proteins in the cell. p53 stops damaged cells from dividing, and p21 is one of the proteins it uses to do this. In p21-deficient mice, G1 arrest was impaired while other p53 functions remained, so p21 mediates this arrest, not everything p53 does.Cell cycle. The series of steps by which a cell copies its DNA and divides into two cells. G1 is the gap before the cell starts copying its DNA. To test it, activate p53 and check that cells stop at G1 and p21 rises. Then remove p21 and activate p53 again. If cells no longer stop, and adding p21 back restores the stop, the claim is supported. Each experiment was chosen after seeing the last. That sequence is the loop this essay is about.
一个典型的因果论断大致如此:“p53 蛋白通过 p21 蛋白在 G1 期阻断细胞周期。” 为了检验这一点,我们需要激活 p53,并观察细胞是否停滞在 G1 期且 p21 水平是否升高。随后,移除 p21 并再次激活 p53。如果此时细胞不再停滞,而重新引入 p21 后细胞又恢复停滞,那么该论断便得到了证实。这里的每一次实验,都是在观察到上一次实验结果后才做出的抉择。这般环环相扣的迭代过程,正是本文所要探讨的“闭环”(loop)。
But claims are not all equal. If our goal is to improve human health, a claim’s value is how much it changes what we know about a disease or can do about it. A claim about a famous cancer gene can be worth little; one showing that an unknown protein is needed for a disease can be worth a lot. Drug targetsDrug target. The molecule in the body, usually a protein, that a drug acts on. with strong evidence in cells, mice, and sometimes humans usually come from claims like these.
然而,这些论断的价值并非等量齐观。如果我们的目标是改善人类健康,那么一个论断的价值就在于它在多大程度上颠覆了我们对某种疾病的认知,或为治疗该疾病提供了多少新手段。一个关于知名癌症基因的常规论断可能价值寥寥;但一个证明某种未知蛋白是致病关键的论断,则可能价值连城。在细胞、小鼠甚至人类体内具备确凿证据的药物靶点,往往正是脱胎于此类极具洞见的论断。
Some of the earliest such claims came from the cell cycle. Labs must grow cells to study them, and growth was easy to measure, even with the simple microscopes of a century ago. Cancer is uncontrolled cell growth, so the chain was short: cell cycle, cell growth, cancer. Today the cell cycle is among the best-understood processes in the cell, and what remains is harder. Recent advances often needed longer chains. Immunotherapy, which helps immune cells attack cancer, came from studying how cells interact.
最早的一批此类论断,部分源于对细胞周期的研究。实验室必须通过培养细胞来进行观察,而“细胞生长”这一现象哪怕是用一个世纪前简陋的显微镜也很容易进行测量。癌症本质上就是失控的细胞生长,因此这其中的逻辑链条很短:细胞周期——细胞生长——癌症。如今,细胞周期已成为细胞内被研究得最透彻的机制之一,而留给我们的则是更难啃的骨头。近年来的突破往往依赖于更长的逻辑链条。例如,助力免疫细胞攻击癌细胞的免疫疗法,便源自于对细胞间相互作用的深度探索。
Many causal claims behind today’s therapies came from studying one gene at a time, before omics. Around 2000, biologists began to see the limits: genes act in networks, and the same gene can act differently in different cell types and conditions. Meanwhile, the first draft of the human genome appeared, and microarrays could measure thousands of genes at once. Together, these gave rise to systems biologySystems biology. A field that measures many parts of the cell at the same time to study how they work together. Systems biology was never purely observational. Early work such as Ideker et al. (2001) already cycled between perturbing, measuring broadly, and updating a model. Many later large datasets, however, were observational., which measures many signals in the cell simultaneously. Some of these studies perturb cells; others observe them. But observation is often an experiment too. Comparing patients with healthy donors, or following tissue across age or development, varies one thing (disease, age or stage) and measures everything else.
在“组学”(omics)时代到来之前,当今众多疗法背后的因果论断,多是通过“一次研究一个基因”的方式得出的。到了 2000 年左右,生物学家们开始意识到这种方法的局限性:基因是在网络中协同运作的,而同一个基因在不同的细胞类型和环境条件下可能表现出截然不同的行为。与此同时,人类基因组草图问世,微阵列技术(microarrays)使得同时测量成千上万个基因成为可能。这些因素共同催生了系统生物学(systems biology),它能够同时捕捉细胞内的多种信号。在这些研究中,有些旨在对细胞施加扰动,有些则侧重于观察。但实际上,观察本身往往也是一种实验。无论是比较患者与健康捐献者的样本,还是追踪组织在不同年龄或发育阶段的变化,其本质都是改变一个变量(如疾病有无、年龄大小或发育阶段),进而测量其余所有的系统性变化。
Run at scale for 30 years, with cheap, broad measurements on samples from across the body, this approach gave us many atlasesAtlas. A large reference map of which genes and molecules are active in each cell type and tissue of the body.. They describe cell types and molecular states, provide pretraining data for biological AI models, and have already pointed to early-stage molecular targets for many diseases. Therapeutic development, though, often needs tight causality across scales: how a molecule leads to disease through every level in between. Between a protein and a human lie groups of proteins, cell parts, cells, tissues, organs, and the full body.Levels. Climbing levels is not the goal in itself. An organ-level readout is not automatically closer to health than a molecular one. What matters is whether the model and readout are relevant to the disease. Comparisons such as disease versus healthy, old versus young, or female versus male are usually not enough to give the causal evidence therapeutic development needs.
这种在全身样本上进行廉价、广泛测量的方法,在经过 30 年的规模化应用后,为我们编纂了许多生物学“图谱”(atlases)。它们描绘了各种细胞类型和分子状态,为生物学 AI 模型提供了预训练数据,并已为许多疾病指明了早期的分子靶点。然而,疗法的开发通常需要在跨尺度层面上建立严密的因果关系:即一个分子是如何穿透层层生物学结构最终导致疾病的。在一个蛋白质和一个活生生的人之间,横亘着蛋白群、细胞器、细胞、组织、器官乃至整个身体等多重层级。仅仅进行诸如疾病与健康、年老与年轻、或雌性与雄性之间的横向比较,通常不足以提供药物研发所必需的严谨因果证据。
Show How long biological events take, at every scale
I think the goal of experiments is not more samples or more measurements per sample, but more health-linked causal claims settled per hour. Two numbers matter, and they are easy to confuse. Throughput is how many results come out each hour. Loop time is how long it takes from choosing an experiment to getting a result that can change the next choice. A lab that starts a 72-hour experiment every hour gets a result every hour, but any experiment chosen in response to one still waits 72 hours. Throughput is enough when experiments are independent. Loop time matters when each result changes what to try next, and that is where a fast-reasoning AI helps most.
在我看来,实验的终极目标并非获取更多的样本或对单个样本进行更多的测量,而是提高每小时能够敲定的、与健康相关的因果论断的数量。这里有两个至关重要却极易混淆的指标:吞吐量(Throughput)和闭环时间(Loop time)。吞吐量是指每小时能产出多少结果;而闭环时间是指从决定做某项实验开始,直到获得能影响下一步决策的结果所需的时间。一个实验室如果每小时启动一个耗时 72 小时的实验,固然也能做到每小时都有结果产出,但任何基于前一个结果而启动的新实验,依然需要苦等 72 小时。当实验彼此独立时,追求吞吐量便已足够。但当每一个结果都决定着下一步该探索什么时,闭环时间才是关键所在——这恰恰是具备快速推理能力的 AI 最能大显身手的地方。
Show What happens to the belt when the first result arrives
- QUESTIONS ANSWERED BY 72 H
- 0
- SLOT-HOURS SPENT ON A QUESTION ALREADY ANSWERED
- 0h
- BATCHES OF CELLS THROWN AWAY
- 0
Some biological processes are slow by nature, which sets a limit. But we often do not need to measure everything at once, especially since breadth often trades off against loop speed. Instead, change one variable quickly: one gene, a group of related genes, or one drug. Then measure, just as quickly, only what the claim needs: a few genes, a thousand, or, if necessary, the full transcriptome or epigenomeOmics, transcriptome, epigenome. Omics methods measure all molecules of one kind at the same time. The transcriptome is the activity of all the genes in a cell. The epigenome is the set of chemical marks on the DNA that control which genes are active.. Early evidence can save us a week on the wrong experiment.
某些生物学过程天生缓慢,这设定了不可逾越的物理极限。但我们往往不需要一次性测量所有指标,尤其是考虑到测量的“广度”常常以牺牲“闭环速度”为代价。相反,我们应当迅速改变单一变量——无论是单个基因、一组相关基因还是某种药物,随后以同样迅捷的速度,仅测量验证该论断所必需的指标:也许是几个基因、一千个基因,或者在必要时,测量完整的转录组或表观基因组。哪怕是早期的初步证据,也能让我们免于在一个错误的实验方向上白白浪费一周光阴。
Show How fast each tool acts, and how long until the cell recovers
Broad and focused data play different roles for AI. Broad datasets are used for pretraining. Fast change-and-measure loops could be the reinforcement learning (RL) environment, where an AI chooses what to test and sees the result.Reinforcement learning. RL is one option, not a requirement; an agent can revise its plans from results without updating its weights. If used for RL, the reward should be for settling questions, including negative results. Rewarding positive claims would favor easy or overstated conclusions. Together, they let AI test causal claims and use the results to choose the next experiment.
泛化数据与聚焦数据对 AI 而言扮演着截然不同的角色。广泛的数据集主要用于模型的预训练;而快速的“改变-测量”(change-and-measure)闭环则可以充当强化学习(RL)的环境,在其中,AI 可以自主选择测试内容并实时观察结果。两者相辅相成,赋予了 AI 检验因果论断并依据结果规划下一次实验的能力。
I used Codex and Claude to review technologies that could shorten these loops. The figures group them by the scale they observe. I see four practical changes worth considering.
我曾利用 Codex 和 Claude 来梳理那些有望缩短这些实验闭环的技术。我认为有四项切实可行的变革值得考量:
Show Fourteen ways to measure a cell, and where each one’s time goes
First, reduce repeated setup. Measure live cells several times, then divide the perturbed culture among compatible endpoint assays, such as sequencing and microscopy.
第一,减少重复性的实验准备工作。对活细胞进行多次动态测量,随后将受扰动的细胞培养物分发到互相兼容的终点检测(如测序和显微成像)中。
Second, keep cultures growing and stagger the experiments so cells are ready when needed. This reduces gaps between experiments, although it does not shorten a 72-hour response. Early readouts can help if they tell us to change the plan before the experiment ends. Keeping this running takes considerable culture capacity. At 80% yield, one valid result per hour requires starting 1.25 batches per hour. If each batch needs a day of growth, all those batches must be prepared in advance.
第二,保持细胞持续培养并错开实验时间,确保在需要时随时有细胞可用。这能有效缩减实验之间的空窗期,尽管它无法从根本上缩短 72 小时的响应时间。如果早期的数据读取能提示我们在实验结束前及时更改计划,那将大有裨益。维持这种运转模式需要极大的细胞培养能力。假设成功率为 80%,要实现每小时获得一个有效结果,就必须每小时启动 1.25 批实验。如果每批细胞需要生长一天,那么所有的批次都必须提前准备就绪。
Third, use barcodes to trace outcomes to their perturbations, allowing experiments to share a pooled readout. This can test many claims in one pass. It raises throughput even when the time to a result stays the same.
第三,利用条形码(barcodes)技术将结果与特定的扰动一一对应,从而允许在混合体系中进行集中读取。这样一次操作就能同时验证许多论断。即便获得单个结果的时间没有改变,它也能显著提升系统的吞吐量。
Fourth, prepare the tools needed for the next experiment. A fast readout helps little if building the next reporter or cell line takes two weeks. Keep common reagents available, install reusable switches in cells, and maintain cultures with the treatment histories the next test may need.
第四,未雨绸缪,为下一步实验备好工具。如果构建下一个报告基因或细胞系需要两周时间,那么再快的读取速度也无济于事。应当保持常用试剂储备充足,在细胞中植入可重复使用的“开关”,并维持带有特定处理背景的细胞培养,以备下一次测试的不时之需。
These changes require keeping track of what is growing, what each well has been through, and which experiments can start with the cells and tools available. Every result may change that schedule. An agent could help track this information and choose the next experiment, provided it knows the lab’s actual constraints.
实施这些变革,意味着需要精确追踪哪些细胞正在生长、培养板上的每个孔经历了怎样的处理,以及利用现有的细胞和工具可以立即启动哪些实验。任何一个新结果都可能打乱原有的时间表。只要一个智能体(agent)了解实验室的实际条件约束,它就能协助追踪这些错综复杂的信息,并妥善决策下一步的最佳实验。
Here is how I would combine these ideas in a machine.
综上所述,以下是我设想中如何将这些理念融入到一台机器中的方案。
Show A possible machine for fast loops
Science has always moved by making a guess and letting an experiment correct it, and AI has so far sped up only the first half of that exchange. A model can now propose more careful claims in an afternoon than a lab could test in a year, which leaves the cell, growing at its own unhurried pace, as the slowest participant in the conversation. I find something hopeful in that, because cells have answered our questions faithfully for as long as we have known how to ask them, and what we need now is a way to ask more often and hear the answer sooner. I like to picture the plate in this machine turning beneath its stations through the night, each well carrying its small history forward until the next question arrives. Most of the answers that come back will rule a claim out, and each of those spares some future scientist a month spent on a path that leads nowhere. Every so often an answer will hold up through every level between a single protein and a patient, and that kind of claim can become a target worth building a medicine around. The walk from molecule to person has always been long, and the hope behind this machine is that each step of it could take hours where it now takes weeks. That will not come from faster instruments alone. The experiment and the reasoning that drives it will have to be designed together, each built around what the other needs: assays chosen because an agent can act on what they return, and agents built around what a living cell can actually be asked.
科学的进步,历来是建立在“提出假设,而后以实验纠偏”的循环之上。迄今为止,AI 仅仅加速了这一交互过程的前半部分。如今,一个模型在一下午提出的严谨论断,甚至比一个实验室一年能测试的还要多。这使得按照自身节奏慢条斯理生长的细胞,成了这场对话中最迟缓的参与者。
但我却从中看到了一线希望。因为自从人类掌握了向细胞提问的方法以来,它们便一直忠实地给予解答。我们现在亟需的,是一种能够更频繁发问并更早聆听答案的机制。我喜欢脑海中的这幅画面:在这台机器里,培养板在各个工作站下彻夜流转,每一个孔位都承载着其微小的历史轨迹,静静等待着下一个问题的降临。
传回的大多数答案都会否定某个假设,而每一次否定,都能将未来的科学家从徒劳无功的死胡同里拯救出数月的光阴。偶尔,也会有一个答案在从单一蛋白质到患者个体的重重验证中屹立不倒,这样的论断便能化作值得围绕其研发药物的珍贵靶点。从分子到人体的探索之路历来漫长,而这台机器所承载的愿景,便是将这漫漫长路上的每一步从“数周”缩短为“数小时”。
要实现这一跨越,单靠更快的仪器是远远不够的。实验本身及其背后的推理逻辑必须被统筹设计,彼此契合互补:选择何种检测方法,取决于智能体能否基于其反馈采取行动;而设计何种智能体,则必须顾及一个活细胞真正能够承受并回答的提问边界。
Sources
Primary studies, protocols, and documentation cited in the figures.