Not just Neanderthals: Ghost lineage in Africa left its mark on our DNA
Not just Neanderthals: Ghost lineage in Africa left its mark on our DNA
不仅仅是尼安德特人:非洲的“幽灵血统”在我们的DNA中留下了印记
When the history of our ancestry is written, the fact that we’ve interbred with some of our closest relatives, the Neanderthals and Denisovans, will have a central role. And it will be tempting to write it as a very tidy story: Once we got genomes from these other groups, it was possible to identify the sequences in our genomes that we shared with them. But in reality, as scientists started poking large enough collections of data, there were regular hints of some strange ancestry in our genomes. It was hard to pin down, though, at least in part because the 2 percent on average of Neanderthal DNA found in many populations does not guarantee that any two individuals will have the same 2 percent. So having the genomes of those two groups made sense of some things researchers had already been seeing.
当我们书写人类祖先的历史时,我们曾与尼安德特人和丹尼索瓦人这些近亲杂交的事实将占据核心地位。人们很容易将其描述为一个非常整洁的故事:一旦我们获得了这些群体的基因组,就有可能识别出我们基因组中与他们共享的序列。但现实情况是,随着科学家开始深入研究足够庞大的数据集,我们的基因组中不断出现一些奇怪血统的迹象。然而,这很难确定,部分原因是许多人群中平均 2% 的尼安德特人 DNA 并不保证任何两个个体拥有相同的 2%。因此,拥有这两个群体的基因组确实解释了研究人员此前观察到的一些现象。
But knowing what we do about Neanderthal and Denisovan DNA is now allowing researchers to answer a somewhat different question: Is there anything else? Using recently developed analytical techniques, they find evidence of a third lineage that we apparently interbred with before any modern humans left Africa. Again, there were hints of this earlier, but so far, there’s been no genome from a modern human relative to help us understand the details—the source of this DNA remains a “ghost lineage.”
但现在,我们对尼安德特人和丹尼索瓦人 DNA 的了解,使研究人员能够回答一个截然不同的问题:还有其他发现吗?利用最近开发的分析技术,他们发现了第三个血统的证据,显然我们在现代人类离开非洲之前就曾与该血统杂交。同样,此前也有过这方面的迹象,但到目前为止,还没有任何现代人类近亲的基因组能帮助我们了解细节——这种 DNA 的来源仍然是一个“幽灵血统”。
Old, and yet young
既古老,又年轻
The new work, done by a group largely based at Berkeley, relies on developments from elsewhere in the field of genomic analysis. Any site in a given person’s genome is the product of a mixture of common descent and random mutations, and its relationship to its neighbors can be mixed up by recombination, when pairs of chromosomes swap segments of DNA. With enough genomic data, computers can be used to reconstruct what are called ancestral recombination graphs that try to reconstruct this history. For each base in the genome, ancestral recombination graphs estimate its history: How many generations back that particular base first appeared in the genome and when it has been involved with recombinations.
这项由主要位于伯克利的研究小组完成的新工作,依赖于基因组分析领域其他地方的发展。一个人基因组中的任何位点都是共同祖先和随机突变的产物,而它与邻近位点的关系可能会因重组(即染色体对交换 DNA 片段)而被打乱。有了足够的基因组数据,计算机就可以用来重建所谓的“祖先重组图谱”,试图还原这段历史。对于基因组中的每一个碱基,祖先重组图谱都会估算其历史:该特定碱基在基因组中首次出现是在多少代之前,以及它何时参与了重组。
Because of the randomness of some of these things and complexities like deletions, many of the individual inferences about history will be wrong. But those are likely to be the exceptions, and the average picture across the genome’s three billion bases should be informative. The Berkeley team made a few inferences about what these ancestral recombination graphs should look like in cases where a separate lineage contributed DNA to modern humans (a process called “introgression”). One is that there should be a cluster of sequences that look consistently old, since they shouldn’t have as many of the same variants that the modern human genomes have picked up while the lineages were separate.
由于其中一些因素的随机性以及缺失等复杂情况,许多关于历史的个体推断可能会出错。但这些很可能只是例外,基因组 30 亿个碱基的平均图景应该是具有参考价值的。伯克利团队对在独立血统向现代人类贡献 DNA(这一过程称为“基因渗入”)的情况下,祖先重组图谱应该是什么样子做出了一些推断。其中之一是,应该存在一组看起来始终“古老”的序列,因为它们不应该拥有现代人类基因组在血统分离期间所积累的那么多相同变异。
Normally, sequences that have been around for a while have more chances to be involved in a recombination. But these sequences were reintroduced to the human genome later in our history, so recombination should be far less frequent relative to a genome that’s been in the modern human lineage the whole time. So any part of the genome that introgressed from a separate lineage should have two properties: Many of its bases should look “old” in the sense of how far back their common ancestry can be traced, yet they should look “young” in terms of how much recombination has taken place. So the researchers developed a software tool they call TRACE to look for these sequences.
通常情况下,存在时间较长的序列有更多机会参与重组。但这些序列是在我们历史的后期重新引入人类基因组的,因此相对于一直存在于现代人类血统中的基因组,其重组频率应该低得多。因此,任何从独立血统渗入的基因组部分都应具备两个特性:从共同祖先可追溯的时间跨度来看,其许多碱基应该看起来很“古老”;但从发生的重组次数来看,它们又应该看起来很“年轻”。因此,研究人员开发了一种名为 TRACE 的软件工具来寻找这些序列。
Like a ghost
如同幽灵
To test the tool, they had convenient examples: the Neanderthal and Denisovan sequences we’ve already identified in the human genome. If TRACE couldn’t pick those out, it wouldn’t find anything else useful. The development of ancestral recombination graphs is still a work in progress, so the researchers used two different tools to generate them and stuck with the one that produced the best results. The combination produced a very low false discovery rate (less than a quarter of a percent) while having an accuracy of over 90 percent.
为了测试该工具,他们有现成的例子:我们已经在人类基因组中识别出的尼安德特人和丹尼索瓦人序列。如果 TRACE 无法识别出这些序列,它也就无法找到其他有用的东西。祖先重组图谱的开发仍处于进行中,因此研究人员使用了两种不同的工具来生成图谱,并坚持使用效果最好的一种。这种组合产生了极低的错误发现率(不到 0.25%),同时准确率超过 90%。
In another demonstration, the researchers performed the analysis on African populations, which only received Neanderthal and Denisovan DNA when individuals from Eurasia migrated back. TRACE found only 0.1 percent of ancestry from these archaic lineages in these African genomes, consistent with its low false error rate. Given about 500 modern human genomes, TRACE pulled out the expected Neanderthal and Denisovan segments. But it also revealed a lot of DNA from a ghost lineage. The percentage was small, at about 0.5 to 1.1 percent of current human genomes. But collectively, they covered nearly 1.5 billion bases. For context, the full human genome is three billion bases.
在另一次演示中,研究人员对非洲人群进行了分析,这些人群只有在欧亚大陆个体迁回时才获得了尼安德特人和丹尼索瓦人的 DNA。TRACE 在这些非洲基因组中仅发现了 0.1% 来自这些古老血统的祖先成分,这与其较低的错误率相符。在分析了约 500 个现代人类基因组后,TRACE 提取出了预期的尼安德特人和丹尼索瓦人片段。但它也揭示了大量来自“幽灵血统”的 DNA。其比例很小,约占当前人类基因组的 0.5% 到 1.1%。但总计覆盖了近 15 亿个碱基。作为参考,完整的人类基因组包含 30 亿个碱基。
It was present in all modern human populations, indicating that the interbreeding occurred before the out-of-Africa expansion. But African populations carry more diverse segments and some that are distinct, a consequence of some of the diversity that was lost when only a fraction of the total population expanded into Eurasia. In fact, the researchers could identify about 100 areas of the genome that lack ghost lineage DNA entirely in non-African populations. Based on the last common ancestor of the sequences, the researchers estimate that the lineage last shared a common ancestor with modern humans over 800,000 years ago, making it about the same age as the split with the ancestor of Neanderthals and Denisovans.
它存在于所有现代人类群体中,表明这种杂交发生在“走出非洲”扩张之前。但非洲人群携带的片段更多样化,其中一些是独特的,这是因为只有一小部分总人口扩张到欧亚大陆时,部分多样性随之丧失。事实上,研究人员可以在非非洲人群的基因组中识别出约 100 个完全缺乏幽灵血统 DNA 的区域。根据这些序列的最后共同祖先,研究人员估计该血统与现代人类最后一次共享共同祖先是在 80 万年前,这使其与尼安德特人和丹尼索瓦人祖先的分离时间大致相同。
The segments in the human genome are, on average, shorter than Neanderthal and Denisovan segments. This means they’ve been around in the modern human genome longer, which is consistent with them having arrived before modern humans left Africa. The modern human genome has areas, called “deserts,” that lack Neanderthal and Denisovan DNA entirely. That has led to the suggestion that the human genome can’t tolerate dramatically different variants in these areas. But the ghost lineage DNA shows up.
人类基因组中的这些片段平均比尼安德特人和丹尼索瓦人的片段更短。这意味着它们在现代人类基因组中存在的时间更长,这与它们在现代人类离开非洲之前就已进入的推论相一致。现代人类基因组中存在被称为“沙漠”的区域,这些区域完全缺乏尼安德特人和丹尼索瓦人的 DNA。这导致了一种观点,即人类基因组无法容忍这些区域中存在差异巨大的变异。但幽灵血统的 DNA 却出现在了这些地方。