Tree Navigation Without LLM Summaries: A Matched-Cost Study of Hierarchical Retrieval for Long-Document QA

本文为原文前 6,000 字符的节选翻译,完整内容请查看原文。

Retrieval-augmented generation grounds language models in external context, but for long documents flat top-$k$ retrieval can cluster on a single region and miss complementary evidence. RAPTOR-style summary trees address this by recursively clustering chunks and using a language model to summarize each cluster at indexing time, then ranking summary nodes alongside raw chunks at query time.

检索增强生成(RAG)将语言模型建立在外部语境之上,但对于长文档而言,扁平化的 top-$k$ 检索可能会集中在单一区域,从而遗漏互补的证据。RAPTOR 式的摘要树通过递归聚类文本块,并在索引时利用语言模型对每个聚类进行总结,随后在查询时将摘要节点与原始文本块一同排序,以此解决上述问题。

We show the main benefit of summary trees in long-document QA can come from navigation rather than the generated summary content. We introduce NavTree, a leaves-only retriever that builds a deterministic balanced segment tree over chunks (zero language-model calls at indexing) and uses the tree purely as a navigation scaffold: a hybrid lexical-and-dense frontier walk, anchored on top retrieved leaves, descends from the root and emits only leaf chunks to the reader.

我们证明了长文档问答中摘要树的主要优势可能源于导航,而非生成的摘要内容本身。我们引入了 NavTree,这是一种仅检索叶子节点的检索器,它在文本块之上构建确定性的平衡线段树(索引时无需调用语言模型),并将该树纯粹用作导航框架:通过一种混合词法与稠密向量的前沿遍历方式,以检索到的顶级叶子节点为锚点,从根节点向下搜索,并仅向阅读器输出叶子文本块。

On a matched-cost evaluation against flat retrievers and an extractive re-implementation of RAPTOR, NavTree is the strongest matched-cost hierarchical retriever in our evaluated grid and ties the strongest flat baseline. On long-document multi-hop QA, it is the only hierarchical method that significantly beats BM25 on a class-vs-class basis, corroborated by a reader-free retrieval-recall check.

在与扁平化检索器及 RAPTOR 的抽取式重新实现进行的成本匹配评估中,NavTree 是我们评估体系中最强的成本匹配分层检索器,并与最强的扁平化基准持平。在长文档多跳问答任务中,它是唯一一种在类对类基础上显著优于 BM25 的分层方法,这一点已通过无需阅读器的检索召回检查得到证实。

A matched-reader replication of the published abstractive RAPTOR variant, given strong cluster summaries, still loses to NavTree at every multi-chunk budget, at zero indexing cost. The ranking carries across stronger and open-weight readers, a stronger encoder, and a full factorial that isolates leaves-only emission as the structural lever.

即便在提供高质量聚类摘要的情况下,对已发表的抽象式 RAPTOR 变体进行匹配阅读器复现,其表现仍逊于 NavTree,且 NavTree 在索引时无需任何成本。这种排名优势在更强的阅读器、开源权重阅读器、更强的编码器以及通过全因子实验分离出“仅输出叶子节点”作为结构性杠杆的测试中均保持一致。