LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs

LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs

学术工作流中的大语言模型:基于短长上下文窗口生成文献综述的评估

Abstract: Our research focuses on evaluating literature reviews generated in short and long context settings of large language models (LLMs) to investigate the impact of context window on the quality of AI-generated literature reviews and the role of AI in supporting literature review writing.

摘要: 我们的研究重点是评估大语言模型(LLM)在短上下文和长上下文设置下生成的文献综述,旨在探讨上下文窗口对人工智能生成文献综述质量的影响,以及人工智能在辅助文献综述写作中的作用。

Twenty AI-generated literature reviews based on research sources from Semantic Scholar and Arxiv were evaluated by two researchers across 15 dimensions.

由两名研究人员从 15 个维度对 20 篇基于 Semantic Scholar 和 Arxiv 研究来源的人工智能生成文献综述进行了评估。

Our findings reveal that AI-generated literature reviews require human oversight to meet academic publishing standards. As context windows increase, LLMs can incorporate broader information and maintain coherence across longer inputs, but they also exacerbate issues such as content repetition, omission of critical work, and a tendency towards descriptiveness over synthesis.

研究结果表明,人工智能生成的文献综述需要人工监督才能达到学术出版标准。随着上下文窗口的增加,大语言模型能够整合更广泛的信息并在更长的输入中保持连贯性,但同时也加剧了内容重复、遗漏关键研究以及倾向于描述而非综合等问题。

Our work shows that AI-generated reviews can provide foundational overviews, but their output must be critically evaluated and refined by domain experts. Future research should consider integrating other LLMs and fine-tuned models in different domains with hybrid approaches that combine human expertise with AI capabilities to address the limitations identified in this study.

我们的研究表明,人工智能生成的综述可以提供基础性概述,但其输出结果必须经过领域专家的批判性评估和润色。未来的研究应考虑整合其他大语言模型以及针对不同领域微调的模型,采用结合人类专业知识与人工智能能力的混合方法,以解决本研究中发现的局限性。


Paper Details:

  • Authors: Muhammad Ali Chaudhry, Xinyuan Hao, Haifa Alwahaby
  • Submission Date: 28 Jun 2026
  • Subjects: Artificial Intelligence (cs.AI); Human-Computer Interaction (cs.HC); Information Retrieval (cs.IR)
  • DOI: 10.48550/arXiv.2608.26145

论文详情:

  • 作者: Muhammad Ali Chaudhry, Xinyuan Hao, Haifa Alwahaby
  • 提交日期: 2026 年 6 月 28 日
  • 学科分类: 人工智能 (cs.AI);人机交互 (cs.HC);信息检索 (cs.IR)
  • 数字对象唯一标识符 (DOI): 10.48550/arXiv.2608.26145