An AI “mind-reading” tool can reconstruct what you’re looking at from a brain scan
An AI “mind-reading” tool can reconstruct what you’re looking at from a brain scan
一种人工智能“读心”工具可以通过脑部扫描重建你所看到的画面
EXECUTIVE SUMMARY A new AI tool can guess what you’re looking at just by analyzing your brain scans—and re-create that image with remarkable precision. It can go the other way, too, and predict a person’s brain activity based on what they’re looking at. In the image above, for example, the left-hand image of each pair is what the user actually saw—and its right-hand counterpart is what the model reconstructed from the brain scan. 执行摘要 一种新的人工智能工具仅通过分析脑部扫描结果,就能猜出你正在看什么,并以惊人的精度重建该图像。它还可以反向操作,根据一个人所看到的画面来预测其大脑活动。例如,在上图中,每一对图像的左侧是用户实际看到的画面,而右侧则是模型根据脑部扫描重建出的图像。
Michal Irani, who developed the tool with her colleagues at the Weizmann Institute of Science in Rehovot, Israel, hopes her “mind-reading” tool will ultimately reveal more about how the brain works, and could perhaps be used to help locked-in people communicate, or allow scientists to re-create the content of dreams. Michal Irani 与她在以色列雷霍沃特魏茨曼科学研究所的同事共同开发了该工具。她希望这种“读心”工具最终能揭示更多关于大脑运作机制的奥秘,并可能用于帮助“闭锁综合征”患者进行交流,或者让科学家能够重建梦境的内容。
Judy Illes, a neuroethicist and professor of neurology at the University of British Columbia in Canada, who was not involved in the research, describes the work as “magnificent.” “The idea [of using this approach] to help people with neurologic conditions … therapeutically is tremendously exciting,” she says. 加拿大不列颠哥伦比亚大学的神经伦理学家兼神经学教授 Judy Illes(未参与此项研究)将这项工作描述为“宏伟的”。她说:“利用这种方法在治疗上帮助神经系统疾病患者的想法……令人极其兴奋。”
But other scientists warn that a similar approach could be used to reveal people’s inner thoughts and mental imagery, potentially without their consent. “The results seem very impressive,” says Tommy Sprague, a neuroscientist at the University of California, Santa Barbara. “But if there’s a way to surreptitiously extract information about what you’re thinking about, then …150 years of sci-fi can come true anytime, and that’s worrisome in a lot of ways.” 但其他科学家警告称,类似的方法可能会被用于揭示人们的内心想法和心理意象,且可能是在未经本人同意的情况下进行的。加州大学圣塔芭芭拉分校的神经科学家 Tommy Sprague 表示:“研究结果看起来非常令人印象深刻。但如果有一种方法可以秘密提取你正在思考的内容,那么……150年来的科幻小说随时可能成为现实,这在很多方面都令人担忧。”
Peeking into the brain 窥探大脑
Neuroscientists have been working for years on ways to use functional magnetic resonance imaging (fMRI) to reconstruct what people see and what’s going on in their minds. The first attempts produced images that were blurry and hard to make sense of. Advances in technology—both in the fMRI scans themselves and in the tools used to make sense of the results—have led to improvements over the years. 多年来,神经科学家一直致力于研究如何利用功能性磁共振成像(fMRI)来重建人们所看到的画面以及大脑中的思维活动。最初的尝试所生成的图像模糊不清,难以辨认。随着技术的进步——包括 fMRI 扫描本身以及用于解读结果的工具——这些年来已经取得了显著改善。
Irani and her colleagues started by analyzing publicly available brain-scan data. Other researchers had already collected scans from volunteers who were shown hundreds of images while they lay in fMRI scanners. fMRI uses a giant magnet to track the flow of oxygenated blood through the brain. Brain areas that “light up” on the scans are thought to be those that are particularly active at any given moment. Irani 和她的同事首先分析了公开的脑部扫描数据。其他研究人员此前已经收集了志愿者在躺进 fMRI 扫描仪时观看数百张图像时的扫描结果。fMRI 利用巨大的磁铁来追踪大脑中含氧血液的流动。扫描中“亮起”的大脑区域被认为是在特定时刻特别活跃的区域。
The results are not especially specific—in typical fMRI scanners, each highlighted “voxel” of activity covers around three cubic millimeters, containing around 16,000 neurons. But Irani and her colleagues used newer datasets collected using scanners with a higher resolution—each voxel covered around one cubic millimeter of neurons, she says. Those datasets showed what the brain activity of volunteers looked like when they viewed various images. 这些结果并不十分精确——在典型的 fMRI 扫描仪中,每个高亮的活动“体素”(voxel)覆盖约三立方毫米,包含约 16,000 个神经元。但 Irani 和她的同事使用了通过更高分辨率扫描仪收集的较新数据集——她说,每个体素覆盖约一立方毫米的神经元。这些数据集展示了志愿者在观看各种图像时大脑活动的样子。
Other teams have done this too, and several other tools have been used to re-create images from brain-scan data. But they’re not good enough, says Irani. Say a person saw a banana. These models can generate an image of a banana, but it would look different, she says. “It wouldn’t have the same structure, the same position.” 其他团队也做过类似的研究,并且已经有几种其他工具被用于从脑部扫描数据中重建图像。但 Irani 表示,它们还不够好。她说,假设一个人看到了一根香蕉,这些模型可以生成一张香蕉的图像,但它看起来会不一样。“它不会有相同的结构,也不会有相同的位置。”
A better decoder 更好的解码器
The researchers wanted to more closely re-create the images that had been seen. The first step was to train an AI model on already available data from eight people who each had been shown around 9,000 images while in a high-resolution fMRI scanner. Crucially, their “brain decoder” has two branches—one to predict the structure of an image (where the colors are, for instance) and a second to predict its content (for example, a bunch of bananas on a plate). 研究人员希望更精确地重建所看到的图像。第一步是在现有数据上训练一个人工智能模型,这些数据来自八个人,他们在高分辨率 fMRI 扫描仪中每人观看了约 9,000 张图像。至关重要的是,他们的“大脑解码器”有两个分支——一个用于预测图像的结构(例如颜色分布的位置),另一个用于预测其内容(例如盘子里的一串香蕉)。
The predictions allow a diffusion model, a type of AI best known for creating video and images by gradually cleaning up a noisy mess of pixels, to produce a much more accurate representation of what the person saw. But to improve the models they needed more data—far more than was actually available. 这些预测结果使得扩散模型(一种以通过逐渐清理杂乱像素来生成视频和图像而闻名的人工智能)能够生成更准确的视觉还原。但为了改进模型,他们需要更多的数据——远超目前实际可用的数据量。
To get around this problem, Irani and her colleagues trained another model—an encoder that can predict brain activity from an image. The team then used the encoder and decoder together to improve both tools. It works like this: Start with a new image of, say, a leopard. Then use the encoder to predict what someone’s fMRI brain scan would look like when the person saw that picture. The decoder is then used to reconstruct the image. 为了解决这个问题,Irani 和她的同事训练了另一个模型——一个可以根据图像预测大脑活动的编码器。随后,团队将编码器和解码器结合起来,以改进这两种工具。其工作原理如下:从一张新图像开始,比如一只豹子。然后使用编码器预测当某人看到这张图片时,其 fMRI 脑部扫描会是什么样子。接着,使用解码器来重建该图像。
At first, the reconstruction probably won’t look much like a leopard, says Irani. But repeatedly training the models this way eventually leads to dramatic improvements. This approach also allows the team to train their models on as many images as they want, even images that have never been shown to a person in an fMRI scanner. Irani says that around 70% of the training data is from images that were not originally paired with fMRI scans. Irani 说,起初,重建的图像可能看起来不太像豹子。但通过这种方式反复训练模型,最终会带来显著的改进。这种方法还允许团队使用任意数量的图像来训练模型,甚至是那些从未在 fMRI 扫描仪中展示给受试者看过的图像。Irani 表示,约 70% 的训练数据来自最初并未与 fMRI 扫描配对的图像。
By combining data from multiple studies, they were also able to identify brain regions that seem to share functions across all individuals. One region seemed to respond to images of food, for example, while another responded to images of sports. Irani, a computer scientist, says she is now working with neuroscientists “to see if we can actually use these tools that we’ve developed to really find out new things about the brain.” 通过结合多项研究的数据,他们还能够识别出在所有个体中功能似乎共享的大脑区域。例如,一个区域似乎对食物图像有反应,而另一个区域则对体育图像有反应。身为计算机科学家的 Irani 表示,她目前正在与神经科学家合作,“看看我们是否真的能利用这些开发的工具,去发现关于大脑的新事物。”
The resulting “universal brain encoder” can work on a scan from a new person with minimal calibration. In other attempts, a tool has typically required about 40 hours of fMRI data on anyone new before it can be used to predict what that person is seeing. Irani’s decoder only needs one hour of data, she says. 由此产生的“通用大脑编码器”只需极少的校准即可处理新人的扫描结果。在其他尝试中,工具通常需要约 40 小时的 fMRI 数据才能预测某人正在看什么。Irani 说,她的解码器只需要一小时的数据。
The finding was presented at the Cognitive Computational Neuroscience conference in New York last month. That could make it valuable for neuroscientists studying the brain, says Sprague. “None of us can afford 40 hours of imaging for a new subject,” he says. “It’s something like $600 to $1,000 an hour.” Tools like this one could speed up research, he says. 这一发现于上个月在纽约举行的认知计算神经科学会议上发表。Sprague 表示,这可能对研究大脑的神经科学家具有重要价值。“我们没有人能负担得起为新受试者进行 40 小时的成像,”他说,“每小时的费用大约在 600 到 1,000 美元之间。”他说,像这样的工具可以加快研究速度。
State of the art 最先进水平
The encoder and decoder aren’t perfect. “Of course we have failures,” says Irani. Over a Zoom call, she… 该编码器和解码器并不完美。“当然,我们也会有失败的时候,”Irani 说。在一次 Zoom 通话中,她……