GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents

GraphEcho: Structural Redundancy and Evidence Provenance in LLM Graph Agents

GraphEcho:大语言模型图智能体中的结构冗余与证据溯源

Abstract: A large language model (LLM) agent can follow more graph paths without acquiring more independent evidence. GraphEcho tests whether agents mistake these repeated encounters for additional corroboration. The benchmark varies path counts and evidential origins while holding evidence content fixed, and evaluates both judgments and active exploration.

摘要: 大语言模型(LLM)智能体可以在不获取更多独立证据的情况下,追踪更多的图路径。GraphEcho 旨在测试智能体是否会将这些重复的路径探索误认为是额外的证据支持。该基准测试在保持证据内容不变的前提下,改变路径数量和证据来源,并对智能体的判断能力和主动探索行为进行评估。

Controlled synthetic experiments reveal model-dependent judgment shifts, but redundant supporting paths increase the share of repeated walks across all evaluated frozen agents. Provenance-aware post-training (PAPT) reduces revisits and improves synthetic accuracy, yet covers fewer distinct sources.

受控合成实验表明,判断偏差因模型而异,但冗余的支撑路径增加了所有被评估的冻结模型(frozen agents)中重复路径探索的比例。溯源感知后训练(PAPT)减少了重复访问并提高了合成任务的准确性,但其覆盖的独立信息源却有所减少。

On scientific claims, it continues to reduce repetition while accuracy declines. These findings expose a gap between efficient exploration and effective evidence use: an agent can learn to stop repeating itself while overlooking information it needs. GraphEcho provides a controlled way to evaluate both what graph agents conclude and whether their exploration reaches distinct evidential sources.

在科学声明任务中,该方法虽然持续减少了重复,但准确性却有所下降。这些发现揭示了高效探索与有效证据利用之间的差距:智能体可能学会了停止自我重复,却忽略了其真正需要的信息。GraphEcho 提供了一种受控方法,既能评估图智能体的结论,也能评估其探索过程是否触及了不同的证据来源。