Closing the data loop in AI-driven drug discovery

Closing the data loop in AI-driven drug discovery

闭环数据:人工智能驱动的药物研发

Drug discovery is a high-cost, high-risk endeavor that is under growing pressure from a market increasingly defined by first-mover advantage. Since the 1950s, the cost of developing new pharmaceuticals has roughly doubled every nine years—a phenomenon known as Eroom’s Law. Today, bringing a new drug to market takes an average of 10-15 years and costs anywhere from $1 billion to $2.5 billion, with failure rates upward of 90%.

药物研发是一项高成本、高风险的事业,正面临着日益增长的市场压力,而该市场越来越看重“先发优势”。自20世纪50年代以来,开发新药的成本大约每九年翻一番——这一现象被称为“反摩尔定律”(Eroom’s Law)。如今,将一种新药推向市场平均需要10到15年,成本在10亿至25亿美元之间,且失败率高达90%以上。

AI has become the pharmaceutical industry’s biggest bet on bringing success rates up and timelines down. The faster drug companies can identify, test, and optimize new chemical compounds, the lower the risk of costly failures later in development. “The main cost in drug discovery is still the clinical phase, so trying to reduce risk and increase your success rates there is obviously hugely beneficial,” says Paul Belcher, director of protein research strategy at global life sciences company Cytiva. “AI is one approach that drug companies hope will not only save time and compress timelines, but enable better quality candidates to reach the clinic.”

人工智能已成为制药行业提高成功率、缩短研发周期的最大赌注。制药公司识别、测试和优化新化合物的速度越快,后期开发中出现昂贵失败的风险就越低。全球生命科学公司Cytiva的蛋白质研究策略总监Paul Belcher表示:“药物研发的主要成本仍在于临床阶段,因此试图降低风险并提高该阶段的成功率显然非常有益。人工智能是制药公司寄予厚望的一种方法,它不仅能节省时间、压缩周期,还能让更高质量的候选药物进入临床试验。”

Early use of AI in drug discovery shows potential, but also highlights the need for robust and authentic data, as well as integration in lab systems.

人工智能在药物研发中的早期应用展现了潜力,但也凸显了对稳健、真实数据以及实验室系统集成的需求。

AI brings efficiency to the lab

人工智能为实验室带来效率

One of the most promising early-stage applications of AI in drug discovery is in hit identification. This involves screening libraries of molecular entities against a disease-related target, such as a protein, to find molecules that bind to it. A successful hit gives researchers a starting point for further testing and refinement, with the aim of eventually developing a viable drug.

人工智能在药物研发早期阶段最有前景的应用之一是“苗头化合物”(hit)的识别。这涉及针对疾病相关靶点(如蛋白质)筛选分子实体库,以寻找能与之结合的分子。一个成功的“苗头”为研究人员提供了进一步测试和优化的起点,目标是最终开发出可行的药物。

Belcher has seen a shift from empirical screening to predictive design: Instead of physically screening libraries, drug companies are now using AI to design drug candidates from scratch and predict how they will interact with disease targets before committing anything to research and development (R&D). This means companies are no longer limited by how much they can physically screen to identify starting points. “AI does away with that,” says Belcher. “And it can help eliminate low-quality candidates before you have to physically test them, saving time and resources.”

Belcher观察到,行业正从经验性筛选转向预测性设计:制药公司不再仅仅依靠物理筛选分子库,而是利用人工智能从零开始设计候选药物,并在投入研发(R&D)之前预测它们与疾病靶点的相互作用。这意味着公司不再受限于物理筛选的能力来寻找研发起点。Belcher说:“人工智能消除了这种限制,它还能在进行物理测试之前帮助剔除低质量的候选药物,从而节省时间和资源。”

What AI can’t do yet is reliably predict kinetics or developability of new compounds, says Belcher. This means every AI-generated candidate still needs to be validated in the lab. Traditional screening workflows were built to identify hits at scale, not to profile large numbers of complex candidates in detail. This is placing more pressure on lab teams, who now have to test, characterize, and purify a growing volume of more diverse, AI-generated compounds.

Belcher指出,人工智能目前还无法可靠地预测新化合物的动力学或可开发性。这意味着每一个由人工智能生成的候选药物仍需在实验室中进行验证。传统的筛选工作流程旨在进行大规模的“苗头”识别,而非对大量复杂的候选药物进行详细分析。这给实验室团队带来了更大的压力,他们现在必须测试、表征和纯化数量日益增加且种类更多样的人工智能生成化合物。

“The current techniques used in hit identification can screen hundreds of thousands, sometimes millions of compounds, using binary or threshold-based techniques producing low-fidelity data—yes-or-no responses,” Belcher explains. “AI can increase the number of hits you get and potentially give you better quality hits as well. That increases demand for higher-throughput, information-rich technologies to then validate and characterize those hits.”

“目前用于识别苗头化合物的技术可以筛选数十万甚至数百万种化合物,但使用的是基于二元或阈值的技术,产生的是低保真度的数据——即‘是’或‘否’的反馈,”Belcher解释道,“人工智能可以增加你获得的苗头数量,并可能提供质量更好的苗头。这就增加了对高通量、信息丰富技术的需求,以便随后对这些苗头进行验证和表征。”

Models need complete, quality data

模型需要完整、高质量的数据

As AI has accelerated demand for data-rich lab systems, it has also highlighted a fundamental need for better, more complete data. Many earlier AI models were trained on publicly available datasets and are now hitting what Belcher calls a data wall. Because models have access to the same data, they all reach similar conclusions, with diminishing returns over time. Additionally, the datasets weren’t built with AI in mind, meaning they lack the structure, labeling, and diversity needed to keep models accurate and free of bias.

随着人工智能加速了对数据丰富型实验室系统的需求,它也凸显了对更好、更完整数据的根本需求。许多早期的人工智能模型是在公开数据集上训练的,现在正撞上Belcher所说的“数据墙”。由于模型访问的是相同的数据,它们得出的结论往往相似,且随着时间的推移,边际收益递减。此外,这些数据集在构建时并未考虑人工智能的需求,这意味着它们缺乏保持模型准确性和无偏见所需的结构、标签和多样性。

Publication bias reinforces the problem. “Most publicly available datasets and scientific publications focus exclusively on positive results,” says Belcher. “No one wants to share their failures. This bias is almost like having one hand tied behind your back. AI models can identify patterns associated with success, but they lack the comprehensive understanding of failures that would make predictions more reliable.”

发表偏倚加剧了这一问题。“大多数公开数据集和科学出版物只关注阳性结果,”Belcher说,“没人愿意分享他们的失败。这种偏见就像被绑住了一只手。人工智能模型可以识别与成功相关的模式,但它们缺乏对失败的全面理解,而这种理解本可以让预测变得更可靠。”

The data Belcher believes would markedly improve models—the failed experiments, the compounds that don’t bind—remains frustratingly difficult to come by. “We often joke that there should be a journal of negative data,” he says. “It’s often buried in lab notebooks, and it’s never used to inform or guide future research.” This lack of negative data creates a fundamental problem: Without access to a broad range of data, models can’t be adequately trained to avoid bias. “In all machine learning applications, the model’s performance relies heavily on the quality and scope of the training data,” notes Belcher.

Belcher认为能显著改善模型的数据——即失败的实验、不结合的化合物——仍然难以获得,令人沮丧。“我们常开玩笑说应该有一本‘负面数据期刊’,”他说,“这些数据通常被埋在实验室笔记本里,从未被用来为未来的研究提供信息或指导。”这种负面数据的缺失造成了一个根本性问题:如果没有广泛的数据支持,模型就无法得到充分训练以避免偏见。Belcher指出:“在所有机器学习应用中,模型的性能在很大程度上依赖于训练数据的质量和范围。”

Fabrication has also become much easier with AI, compounding concerns around data integrity. Take Western blots, for example. These are part of a standard technique for identifying proteins in blood or tissue samples, and they are among the most common targets for manipulation in biomedical research. Belcher cites research by Dutch microbiologist Elisabeth Bik, who found that almost 4% of biomedical papers contained duplicated or manipulated images. This was back in 2016, before generative AI made fabrication trivial. “Manipulated or faked data has always been a problem in science, but in the AI world, especially when used to train models, it could have potentially disastrous consequences,” says Belcher. “There needs to be more tools to verify that data is not manipulated.”

随着人工智能的出现,数据造假变得更加容易,这加剧了人们对数据完整性的担忧。以蛋白质印迹(Western blot)为例,这是识别血液或组织样本中蛋白质的标准技术之一,也是生物医学研究中最常被篡改的目标。Belcher引用了荷兰微生物学家Elisabeth Bik的研究,她发现近4%的生物医学论文包含重复或被篡改的图像。这还是在2016年,即生成式人工智能让造假变得轻而易举之前。“篡改或伪造数据在科学界一直是个问题,但在人工智能时代,特别是当这些数据被用于训练模型时,可能会产生灾难性的后果,”Belcher说,“需要更多的工具来验证数据是否被篡改。”

Some vendors are starting to tackle this challenge. Belcher points to solutions like Cytiva’s Image Integrity Checker, for instance, which uses secure hash algorithms—the same technology used in blockchain—to detect whether scientific images have been tampered with. “We’re starting to see a lot of interest from publishing houses that want to adopt this as standard because it’s a quick way to ensure that what gets published in the literature is…”

一些供应商已开始应对这一挑战。Belcher举例提到了Cytiva的“图像完整性检查器”(Image Integrity Checker),它使用安全哈希算法(与区块链使用的技术相同)来检测科学图像是否被篡改。“我们开始看到许多出版机构对此表现出浓厚兴趣,他们希望将其作为标准,因为这是一种确保文献发表内容真实可靠的快捷方式……”