Encoding EEG Signals to Examine Human-Like Next-Word Prediction Behaviour in Language Models

Encoding EEG Signals to Examine Human-Like Next-Word Prediction Behaviour in Language Models

通过编码脑电图(EEG)信号研究语言模型中类人的下一词预测行为

Abstract: Language models (LMs) are trained to excel at predicting the next word in the sequence given prior context, and humans also share this predictability in reading comprehension. Neuroscience research reveals that next-word predictability influences brain response, as recorded at millisecond resolution using electroencephalography (EEG).

摘要: 语言模型(LM)经过训练,能够根据先前的上下文出色地预测序列中的下一个词,而人类在阅读理解中也具有这种可预测性。神经科学研究表明,下一词的可预测性会影响大脑反应,这可以通过脑电图(EEG)以毫秒级的分辨率记录下来。

While our evidence indicates that advanced LMs achieve accuracies closely aligned with human performance at the next-word prediction task, this raises the question: Does higher prediction accuracy necessarily mean that these models adequately capture the cognitive signals associated with human reading comprehension?

尽管我们的证据表明,先进的语言模型在下一词预测任务上的准确率与人类表现非常接近,但这引发了一个问题:更高的预测准确率是否一定意味着这些模型能够充分捕捉到与人类阅读理解相关的认知信号?

Here, we generate regressors for both humans and LMs based on two information measures, including top-1 prediction and surprisal, to predict event-related potential (ERP) elicited from EEG recordings which reflect different stages of cognitive processing during reading. We argue that modelling ERP patterns offers fine-grained analysis of the cognitive plausibility of various LMs during reading.

在此,我们基于两种信息度量(包括 Top-1 预测和惊奇度/Surprisal)为人类和语言模型生成回归量,以预测从 EEG 记录中提取的事件相关电位(ERP),这些电位反映了阅读过程中不同阶段的认知处理。我们认为,对 ERP 模式进行建模,可以对各种语言模型在阅读过程中的认知合理性进行细粒度的分析。

Our results indicate that only surprisal potentially correlates with language-processing ERPs, especially for open-class words with high semantic content. Moreover, our findings challenge the assumption that scaling LMs with increased parameters and computational budgets will consistently lead to improved convergence with human-like linguistic处理.

我们的研究结果表明,只有惊奇度(Surprisal)可能与语言处理的 ERP 相关,特别是对于具有高语义内容的开放类词汇。此外,我们的发现挑战了一种假设,即通过增加参数和计算预算来扩展语言模型,将始终能够改善其与人类语言处理过程的趋同性。