Team uses AlphaFold AI to redesign gene-editing proteins to make them safer
Team uses AlphaFold AI to redesign gene-editing proteins to make them safer
研究团队利用 AlphaFold AI 重新设计基因编辑蛋白,以提高其安全性
A couple of decades after the discovery of systems that could selectively target DNA, we’re starting to see the first therapies based on gene editing. One challenge these developments have faced is safety. While we can make them pretty specific to the gene we want edited, the human genome is very large, and even rare DNA sequences can appear a couple of times by chance. As a result, all the original gene-editing systems had known rates of what are called off-target effects, in which they simply edit the wrong sequence. This may be a low-probability event, but edit enough cells—and therapies generally have to edit many—and errors become inevitable. A lot of effort has gone into finding ways to minimize or eliminate off-target edits. In a recent issue of Nature, researchers described modifying the AI protein-folding software AlphaFold to help identify key areas of gene-editing proteins responsible for off-target effects. Those areas were then modified to reduce the problems.
在发现能够选择性靶向 DNA 的系统二十年后,我们开始看到首批基于基因编辑的疗法。这些发展面临的一个挑战是安全性。虽然我们可以使它们对想要编辑的基因具有很高的特异性,但人类基因组非常庞大,即使是罕见的 DNA 序列也可能偶然出现多次。因此,所有原始的基因编辑系统都存在已知的“脱靶效应”发生率,即它们会错误地编辑了错误的序列。这可能是一个低概率事件,但如果编辑的细胞数量足够多(疗法通常需要编辑大量细胞),错误就变得不可避免。目前,人们投入了大量精力来寻找最小化或消除脱靶编辑的方法。在最近一期的《自然》杂志上,研究人员描述了对 AI 蛋白质折叠软件 AlphaFold 进行改进,以帮助识别基因编辑蛋白中导致脱靶效应的关键区域。随后,这些区域被修改以减少此类问题。
Gene editing and off-target effects
基因编辑与脱靶效应
Gene-editing systems have three key components. The first is guide RNA, which can base-pair with the targeted genome sequence. It’s possible to design many guide RNAs that all target the same gene, so one approach to making the system safer is to pick sequences that don’t share much similarity with any other genomic locations. This is now widely used as part of the basic design process for a gene-editing system.
基因编辑系统有三个关键组成部分。第一是引导 RNA(guide RNA),它可以与目标基因组序列进行碱基配对。设计多个靶向同一基因的引导 RNA 是可能的,因此使系统更安全的一种方法是选择与其他基因组位置相似度较低的序列。这目前已被广泛用作基因编辑系统基本设计过程的一部分。
The second component is a Cas protein named after the CRISPR system’s Cas9 protein. It interacts with both the guide RNA and genomic RNA and helps enforce the specificity of the interactions. If there are too many mismatches in the base pairing between the RNA and DNA, Cas9 (or other Cas family members) won’t stick there. Various approaches have led to improved Cas family members that have reduced the tendency to enable off-target edits.
第二个组成部分是 Cas 蛋白,它以 CRISPR 系统的 Cas9 蛋白命名。它与引导 RNA 和基因组 DNA 相互作用,并有助于强化相互作用的特异性。如果 RNA 和 DNA 之间的碱基配对存在过多错配,Cas9(或其他 Cas 家族成员)就不会结合在那里。各种方法已经促成了改进型 Cas 家族成员的诞生,这些成员降低了产生脱靶编辑的倾向。
The final key component of the system is the protein that interacts with Cas9 when it’s bound to DNA and modifies the DNA. In the original CRISPR system, this cut both strands of the double helix, producing damage that’s difficult to control. Researchers have since modified other proteins to interact with Cas9 but catalyze more subtle changes to DNA, such as lopping off a single base or making chemical modifications that alter how it base-pairs.
系统的最后一个关键组成部分是当 Cas9 与 DNA 结合时与之相互作用并修饰 DNA 的蛋白质。在最初的 CRISPR 系统中,这会切断双螺旋的两条链,产生难以控制的损伤。此后,研究人员修改了其他蛋白质以与 Cas9 相互作用,但催化更微妙的 DNA 变化,例如切除单个碱基或进行改变其碱基配对方式的化学修饰。
Overall, the length of the base pairing between the guide RNAs and the genome is on the order of 18 bases long. That should show up in random DNA sequences only about once in 70 billion bases, and our genome is only about 3 billion bases. By that measure, we should be good. But it turns out that Cas9 can tolerate a small number of mispaired bases without losing its ability to stick to DNA. The exact number and location of the bases where variations are tolerated can vary somewhat, making it difficult to identify in advance which guide RNAs might pose a danger. One route to improving safety is to better understand how these off-target interactions occur.
总体而言,引导 RNA 与基因组之间的碱基配对长度约为 18 个碱基。在随机 DNA 序列中,这种情况大约每 700 亿个碱基才会出现一次,而我们的人类基因组只有约 30 亿个碱基。按此衡量,我们应该很安全。但事实证明,Cas9 可以容忍少量错配碱基而不失去其结合 DNA 的能力。容忍变异的碱基的具体数量和位置可能会有所不同,这使得提前识别哪些引导 RNA 可能构成危险变得困难。提高安全性的一条途径是更好地了解这些脱靶相互作用是如何发生的。
Making contact
建立接触
The team behind the new work, based at a variety of institutions in China, reasoned out their approach in advance. A perfectly matched DNA-RNA hybrid will have one structure, while one with one or more mispaired bases will have a slightly different structure. Evolution has optimized the structure of the Cas9 to stick to the former. But it apparently hasn’t prevented Cas9 from adopting slightly different conformations that can interact with mispaired structures. If we can identify the portions of Cas9 that mediate these problematic interactions, we can modify and potentially block them.
这项新工作的团队来自中国多家机构,他们提前推导出了自己的研究方法。完美匹配的 DNA-RNA 杂合体具有一种结构,而具有一个或多个错配碱基的杂合体则具有略微不同的结构。进化已经优化了 Cas9 的结构以结合前者。但显然,这并没有阻止 Cas9 采用可能与错配结构相互作用的略微不同的构象。如果我们能识别出介导这些问题相互作用的 Cas9 部分,我们就可以对其进行修改并可能阻断它们。
The team’s first step was to build a large library of off-target editing sites. They did this by using a modified CRISPR system that converts the DNA base adenine to a related chemical, inosine, and then isolating any DNA fragments that contain it. They repeated this process with 10 different guide RNAs and analyzed a large number of modified DNA fragments from each to get a broad picture of the types of off-target sequences present.
团队的第一步是建立一个大型脱靶编辑位点库。他们通过使用一种改进的 CRISPR 系统来实现这一点,该系统将 DNA 碱基腺嘌呤转化为相关的化学物质肌苷,然后分离出任何包含该物质的 DNA 片段。他们用 10 种不同的引导 RNA 重复了这一过程,并分析了每种引导 RNA 产生的大量修饰 DNA 片段,从而获得了现有脱靶序列类型的广泛图景。
The next step was to examine how the CRISPR complex interacted with them, using the AlphaFold AI-based protein-folding software. Updated versions were designed to handle interactions between proteins and nucleic acids, as well as complexes of multiple proteins. So the team fed AlphaFold versions of a target DNA sequence, along with a guide RNA, the Cas9 sequence, and an enzyme that chemically modifies bases and can stick to Cas9. Unfortunately, it choked, placing one of the proteins in what was clearly the wrong location. Undeterred, the team simplified things and fed AlphaFold only the DNA, RNA, and Cas9 protein, since the latter is the primary factor determining its sequence specificity. This worked much better, producing a structure that agreed with ones determined by experiments with actual nucleic acids and proteins.
下一步是使用基于 AI 的蛋白质折叠软件 AlphaFold 来检查 CRISPR 复合物如何与它们相互作用。更新后的版本旨在处理蛋白质与核酸之间的相互作用,以及多种蛋白质的复合物。因此,团队向 AlphaFold 输入了目标 DNA 序列、引导 RNA、Cas9 序列以及一种能化学修饰碱基并能结合 Cas9 的酶。不幸的是,它“卡壳”了,将其中一种蛋白质放置在明显错误的位置。团队并未气馁,简化了流程,只向 AlphaFold 输入了 DNA、RNA 和 Cas9 蛋白,因为后者是决定其序列特异性的主要因素。这种方法效果好得多,产生的结构与通过实际核酸和蛋白质实验确定的结构相吻合。
By comparing the structures AlphaFold generated when fed different on- and off-target sites, the researchers found a general pattern. Many (about two-thirds) of the off-target sites caused the Cas9 protein to adopt a slightly different structure. But nearly all (over 95 percent) of them altered which amino acids contacted the RNA. So there are clearly some cases where Cas9 maintains its normal structure but amino acids within it flex around in ways that accommodate the mispaired bases of off-target sites. Conveniently, AlphaFold was already set up to identify what is termed the “contact probability,” namely, the chance that any two items, such as amino acids or nucleotides, are within a very small distance (eight Angstroms). The researchers could take the output of the contact probability analysis for on- and off-target sites and compare them, identifying exactly which amino acids in Cas9 have altered contacts when there’s a mismatch between the guide RNA and the DNA.
通过比较 AlphaFold 在输入不同靶向和脱靶位点时生成的结构,研究人员发现了一个普遍规律。许多(约三分之二)脱靶位点导致 Cas9 蛋白采用了略微不同的结构。但几乎所有(超过 95%)的脱靶位点都改变了与 RNA 接触的氨基酸。因此,显然在某些情况下,Cas9 保持其正常结构,但其内部的氨基酸会以某种方式弯曲,以适应脱靶位点的错配碱基。方便的是,AlphaFold 已经具备识别所谓“接触概率”的功能,即任何两个项目(如氨基酸或核苷酸)处于极小距离(8 埃)内的几率。研究人员可以获取靶向和脱靶位点的接触概率分析输出并进行比较,从而准确识别出当引导 RNA 和 DNA 之间存在错配时,Cas9 中哪些氨基酸的接触发生了改变。