LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization

LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization

LongNovel:用于长文本小说摘要幻觉检测的多尺度基准

Although context windows have expanded significantly in recent years, hallucinations in long-context summarization remain a challenge. Long novels are better suited than news or papers for researching these hallucinations, due to their intrinsic information and detailed descriptions of events and dialogues. 尽管近年来上下文窗口已显著扩大,但长文本摘要中的幻觉问题仍然是一项挑战。与新闻或论文相比,长篇小说因其内在的信息量以及对事件和对话的详细描述,更适合用于研究此类幻觉。

However, current research lacks a multi-scale benchmark for hallucination detection in long-context novel summarization and does not fully explore how hallucinations change as the context grows longer. In this study, we propose LongNovel, a multi-scale long-context bilingual (Chinese and English) novel benchmark for hallucination detection. 然而,目前的研究缺乏用于长文本小说摘要幻觉检测的多尺度基准,且未能充分探讨幻觉如何随着上下文长度的增加而变化。在本研究中,我们提出了 LongNovel,这是一个用于幻觉检测的多尺度长文本双语(中英文)小说基准。

This benchmark is constructed from 29 Chinese novels (ranging from 16k to 100k tokens) and chapter-level data from the BookSum dataset. We design 8 hallucination types and employ a combination of Multi-Model Arbitration and Entity-Referenced Hallucination Generation to ensure both data authenticity and a balanced distribution of hallucination categories. 该基准由 29 部中文小说(长度从 1.6 万到 10 万 token 不等)以及 BookSum 数据集中的章节级数据构建而成。我们设计了 8 种幻觉类型,并结合了多模型仲裁(Multi-Model Arbitration)和基于实体的幻觉生成(Entity-Referenced Hallucination Generation)方法,以确保数据的真实性以及幻觉类别的分布平衡。

Furthermore, we manually revise the content in the test set to guarantee data reliability. Extensive experimental results demonstrate that LongNovel is a challenging benchmark. We release LongNovel for future research. 此外,我们对测试集的内容进行了人工修订,以保证数据的可靠性。大量的实验结果表明,LongNovel 是一个具有挑战性的基准。我们现已发布 LongNovel 以供后续研究使用。