How AI helps scientists design the next generation of medicines
How AI helps scientists design the next generation of medicines
人工智能如何帮助科学家设计下一代药物
Designing and developing a new medicine is an expensive, failure-prone scientific challenge. A new drug can take many years to develop, at the cost of a significant investment. And even then, most possible candidates never reach the patient. 设计和开发一种新药是一项昂贵且容易失败的科学挑战。一种新药的研发可能需要多年时间,并耗费巨额投资。即便如此,大多数候选药物最终也无法到达患者手中。
For biologic medicines, therapies made from engineered proteins rather than synthetic chemistry (which are often used to treat conditions across most major acute and chronic diseases), the complexity is even greater. Scientists explore vast quantities of possible molecules, looking for the rare few that will bind to the right target, remain stable in the human body, and be manufacturable at scale. 对于生物药物(由工程蛋白质而非合成化学制成,常用于治疗大多数主要的急性和慢性疾病)而言,其复杂性更高。科学家需要探索海量的潜在分子,寻找那些能够与正确靶点结合、在人体内保持稳定并能大规模生产的极少数分子。
Today, AI is speeding up these processes and has quickly become a core part of the infrastructure in pharmaceutical R&D. AI-assisted design is a growing part of how biologic drug candidates are developed, and companies like AstraZeneca are actively building its engineering teams to push this further. 如今,人工智能正在加速这些进程,并迅速成为制药研发基础设施的核心部分。人工智能辅助设计在生物药物候选物的开发中占比日益增加,像阿斯利康(AstraZeneca)这样的公司正在积极组建工程团队,以进一步推动这一进程。
“Everything we do, whether it’s design, make, test, or analyze, is now computationally enhanced,” says Puja Sapra, senior vice president and head of R&D biologics engineering and oncology targeted discovery at AstraZeneca. “The cycle times are getting shorter while productivity and innovation increase.” “我们所做的一切,无论是设计、制造、测试还是分析,现在都得到了计算能力的增强,”阿斯利康研发生物制剂工程与肿瘤靶向发现高级副总裁兼负责人 Puja Sapra 表示,“周期正在缩短,而生产力和创新能力却在提高。”
Sapra explains that AstraZeneca’s approach follows a build-measure-learn loop. AI generates or prioritizes candidate molecules computationally, predicting which designs are most likely to succeed. Scientists then focus lab resources only on the top-ranked candidates. This leads to a tighter feedback cycle with fewer dead ends, faster iteration, and the ability to go after disease targets that were previously considered untreatable by medicine. Sapra 解释说,阿斯利康的方法遵循“构建-测量-学习”循环。人工智能通过计算生成或优先筛选候选分子,预测哪些设计最有可能成功。随后,科学家仅将实验室资源集中在排名最高的候选物上。这带来了更紧密的反馈循环,减少了死胡同,加快了迭代速度,并使我们能够攻克以往被认为无法通过药物治疗的疾病靶点。
Because the number of possible molecular combinations far exceeds what any human team can systematically explore, using AI to narrow and refine the options for testing has become a major focus in biologics drug design. 由于可能的分子组合数量远远超过了任何人类团队所能系统探索的范围,利用人工智能来缩小和优化测试选项已成为生物药物设计的核心重点。
Navigating complex drug design problems
应对复杂的药物设计难题
Beyond accelerating timelines, AI is also being applied to the discovery of entirely new classes of medicines. Traditional biologics typically target one disease pathway. The next generation of drugs can hit multiple targets simultaneously or precisely deliver therapeutic payloads to specific cells. Achieving this requires optimization across many variables at once. 除了缩短时间表外,人工智能还被应用于发现全新的药物类别。传统的生物制剂通常针对单一疾病通路。而下一代药物可以同时打击多个靶点,或将治疗载荷精确递送至特定细胞。实现这一目标需要同时对多个变量进行优化。
Looking ahead
展望未来
AI-driven models could help design these increasingly complex, multi-specific biologics, explains Puja Sapra. “For example,” she continues, “such models could help identify which two or three targets to prioritize based on the underlying biology, then optimize across multiple parameters to balance a molecule’s potency, stability, manufacturability, and safety.” Puja Sapra 解释说,人工智能驱动的模型可以帮助设计这些日益复杂的多特异性生物制剂。“例如,”她继续说道,“此类模型可以帮助根据基础生物学确定优先考虑哪两到三个靶点,然后在多个参数之间进行优化,以平衡分子的效力、稳定性、可制造性和安全性。”
“Drugging the undruggable is becoming a reality,” Sapra says. “These technologies will eventually enable us to develop medicines against targets once thought impossible to reach. The potential for benefit to patients is remarkable.” “‘攻克不可成药靶点’正在成为现实,”Sapra 说,“这些技术最终将使我们能够开发出针对曾经被认为无法触及的靶点的药物。其对患者的潜在益处是巨大的。”
The data moat
数据护城河
McKinsey estimates that generative AI, combined with other computational tools, could cut drug discovery timelines by as much as 50%. But every AI model is only as good as its training data. In drug discovery, that means ample quantities of high-quality biological data. Experiments can provide a rich source of such data. Whether they succeed or fail, each experiment generates a signal about what does and does not work. 麦肯锡估计,生成式人工智能结合其他计算工具,可以将药物发现的时间表缩短多达 50%。但任何人工智能模型的好坏都取决于其训练数据。在药物发现中,这意味着需要海量的高质量生物数据。实验可以提供此类数据的丰富来源。无论成功还是失败,每一次实验都会产生关于什么有效、什么无效的信号。
“Data is our differentiator,” says Sapra, explaining how the company’s datasets are proprietary and multimodal and include molecular structures, binding measurements, safety profiles, and manufacturing outcomes. “We’ve built an intentionally diverse portfolio across multiple disease areas and drug types. All of that data empowers us to fine-tune frontier AI models with richer, more representative training sets.” “数据是我们的差异化优势,”Sapra 说。她解释了公司的数据集是如何专有且多模态的,包括分子结构、结合测量值、安全性概况和制造结果。“我们在多个疾病领域和药物类型中构建了刻意多元化的产品组合。所有这些数据使我们能够利用更丰富、更具代表性的训练集来微调前沿人工智能模型。”
She continues, “Further, we have invested in deep screening technologies to generate additional datasets required in volume to constantly refine and validate our models.” 她继续说道:“此外,我们还投资了深度筛选技术,以生成不断优化和验证模型所需的大量额外数据集。”
Building an autonomous discovery engine
构建自主发现引擎
To bring all of that data together in one place, AstraZeneca is building what it calls a “lab of the future” facility in Kendall Square, Cambridge, Massachusetts where AI and robotic automation will be able to form a continuous, closed-loop discovery system. 为了将所有这些数据汇集到一处,阿斯利康正在马萨诸塞州剑桥市的肯德尔广场(Kendall Square)建设所谓的“未来实验室”,在那里,人工智能和机器人自动化将能够形成一个连续的、闭环的发现系统。
“Where a self-driving car uses sensors and models to navigate its environment, this system uses AI to make predictions, robotic systems to execute experiments, and instruments to generate data,” explains Sapra. That data feeds directly back into the models, accelerating each subsequent cycle. “Throughout, scientists will remain central to the process, providing the oversight, judgement, and strategic direction that ensure outputs are explainable, tolerable, and directed toward potential patient benefit,” she adds. “就像自动驾驶汽车使用传感器和模型来导航环境一样,该系统使用人工智能进行预测,使用机器人系统执行实验,并使用仪器生成数据,”Sapra 解释道。这些数据直接反馈回模型,加速随后的每一个周期。“在此过程中,科学家始终处于核心地位,提供监督、判断和战略指导,确保产出是可解释的、可耐受的,并指向潜在的患者获益,”她补充道。
Eventually, automated high-throughput systems will be able to make and evaluate thousands of molecular interactions on a weekly basis. “This will generate AI-ready data at a scale that traditional workflows cannot match,” Sapra says. “Robotic sample handling, automated quality checks, and integrated data pipelines also have the potential to help accelerate early drug development timelines significantly.” 最终,自动化高通量系统将能够每周制造并评估数千种分子相互作用。“这将以传统工作流程无法比拟的规模生成人工智能就绪的数据,”Sapra 说,“机器人样本处理、自动化质量检查和集成数据管道也有潜力帮助显著加快早期药物开发的时间表。”
The next frontier: Generating medicines from scratch
下一个前沿:从零开始生成药物
Ultimately, Sapra says, the end-state vision for AI in biologic drug discovery is what the field calls “de novo” design. For this, the goal is for AI to generate entirely new protein sequences that precisely fit the desired drug properties. This includes designing the structure, predicting safety, how it will behave in the body and how to make it manufacturable. Sapra 表示,最终,人工智能在生物药物发现中的终极愿景是该领域所谓的“从头(de novo)”设计。其目标是让人工智能生成完全符合所需药物特性的全新蛋白质序列。这包括设计结构、预测安全性、预测其在体内的表现以及如何使其具备可制造性。
“The field is making great progress toward a completely AI-generated biologic, designed from scratch all the way to a clinical candidate,” Sapra says. “As we continue to leverage frontier models and fine-tune them with the right datasets, we bring ourselves closer to this reality. I believe it will come. It’s a matter of time.” “该领域在实现完全由人工智能生成的生物制剂方面取得了巨大进展,从零开始一直到临床候选药物,”Sapra 说,“随着我们继续利用前沿模型并用正确的数据集对其进行微调,我们正越来越接近这一现实。我相信这一天会到来。这只是时间问题。”