Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents

Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents

不仅要事实正确,更要来源准确:面向 MCP 智能体的来源感知验证

Tool-using LLM agents no longer read from a single retrieved passage. Through the Model Context Protocol (MCP), an agent can call a search tool, inspect a structured patient or account record, query a database, and pull metadata, then weave all of it into one answer. That makes the usual question of factuality more subtle than it looks. 使用工具的 LLM 智能体不再仅仅从单一的检索段落中读取信息。通过模型上下文协议(MCP),智能体可以调用搜索工具、检查结构化的患者或账户记录、查询数据库并提取元数据,然后将所有这些信息编织成一个答案。这使得通常的事实性问题变得比看起来更微妙。

Most of the systems built to check LLM answers, from RAGAS faithfulness to fine-grained checkers like MiniCheck, AlignScore, and SummaC, ask whether a claim is supported by the available evidence once that evidence has been pooled together. In their usual form, they do not tell us which MCP tool output supports each claim, or whether that is the source the answer names. 大多数用于检查 LLM 答案的系统(从 RAGAS 的忠实度评估到 MiniCheck、AlignScore 和 SummaC 等细粒度检查器)都在询问:一旦证据被汇总在一起,某个声明是否得到了现有证据的支持。在常规形式下,它们无法告诉我们每个声明是由哪个 MCP 工具输出所支持的,也无法判断这是否就是答案中所指明的来源。

Our latest paper, ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents (read it on Hugging Face, or on arXiv in the meantime), targets that gap. The failure mode we care about is one we call cross-source conflation: a claim that is true somewhere in the evidence, but attributed to the wrong source. A source-blind verifier may pass it, because the fact does exist in the pool. A source-aware verifier should not. 我们最新的论文《ProvenanceGuard:面向基于 MCP 的 LLM 智能体的来源感知事实性验证》(可在 Hugging Face 或 arXiv 上阅读)旨在填补这一空白。我们关注的故障模式被称为“跨来源混淆”:即某个声明在证据中的某处是真实的,但被归因于错误的来源。来源盲目的验证器可能会通过它,因为该事实确实存在于证据池中;但来源感知的验证器则不应通过。

The problem: supported somewhere is not the same as supported by the right source

问题所在:在某处得到支持并不等同于由正确的来源支持

Consider a customer support agent that answers, “According to the account record, this plan includes a 30-day refund window.” The refund window may be perfectly real, but stated in a policy document, not in the account record the answer points to. Pool the two together and the claim looks supported. Keep them separate and the attribution is wrong, and in a data-sensitive setting a wrong attribution can be as damaging as a wrong fact. 考虑一个客户支持智能体回答道:“根据账户记录,该计划包含 30 天的退款窗口。”这个退款窗口可能是完全真实的,但它是在政策文件中说明的,而不是在答案所指向的账户记录中。如果将两者汇总在一起,该声明看起来是得到支持的;但如果将它们分开,归因就是错误的。在数据敏感的环境中,错误的归因可能与错误的事实一样具有破坏性。

The same pattern shows up in a clinical agent, where a patient-specific medication detail taken from a patient-history tool becomes misleading the moment the answer presents it as a finding from the medical literature. A claim can be supported by one MCP source while the answer attributes it to another. Source-blind scoring sees support in the pooled evidence and passes it; ProvenanceGuard separately checks whether the supporting source matches the one the answer states or implies. 同样的模式也出现在临床智能体中:当答案将从患者病史工具中获取的特定药物细节呈现为医学文献的研究发现时,它就变得具有误导性。一个声明可能由某个 MCP 来源支持,但答案却将其归因于另一个来源。来源盲目的评分系统会在汇总的证据中看到支持并予以通过;而 ProvenanceGuard 会单独检查支持来源是否与答案陈述或暗示的来源相匹配。

What ProvenanceGuard does

ProvenanceGuard 的工作原理

ProvenanceGuard is a post-generation verification layer that sits on top of a black-box MCP agent. It runs after an agent produces an answer, and never collapses the evidence into one anonymous context. Instead it carries the source identity all the way through the pipeline. It reads the captured MCP trace, including the tool outputs and their source IDs, without retraining the agent. ProvenanceGuard 是一个位于黑盒 MCP 智能体之上的生成后验证层。它在智能体生成答案后运行,且从不将证据合并为一个匿名的上下文。相反,它在整个流程中始终携带来源标识。它读取捕获的 MCP 追踪记录(包括工具输出及其来源 ID),而无需重新训练智能体。

Then it does five things in sequence: it breaks the answer into specific claims, finds the source most relevant to each one, checks whether that source actually supports it, compares the source with the one the answer names or implies, and finally emits both a per-claim source verdict and a global, answer-level allow or block decision. 随后,它按顺序执行五项操作:将答案拆解为具体声明,找到与每个声明最相关的来源,检查该来源是否确实支持该声明,将该来源与答案中指明或暗示的来源进行比较,最后输出针对每个声明的来源判定以及全局性的答案级允许或拦截决策。

Results

结果

We tested ProvenanceGuard on answers from a medical agent that had used patient records, research articles, and other tools. This gave us 281 real traces to study. Medicine is a useful test because a fact from a patient’s record and a fact from general research cannot be treated as the same source. 我们测试了 ProvenanceGuard 在医疗智能体答案上的表现,该智能体使用了患者记录、研究文章和其他工具。这为我们提供了 281 条真实追踪记录进行研究。医学是一个有用的测试领域,因为来自患者记录的事实和来自一般研究的事实不能被视为同一来源。

For the main test, human experts checked 361 claims from 40 answers set aside from the data used to develop the system. The most direct result is this: experts said 139 claims should not pass, and ProvenanceGuard caught 138 of them. It let one through. It also held 67 claims that the experts considered supported, sending them for review or repair. This reflects the cautious setting we tested: it favors a second look at some supported claims over letting unsupported ones through. 在主要测试中,人类专家检查了从开发系统所用数据中预留出的 40 个答案中的 361 条声明。最直接的结果是:专家认为 139 条声明不应通过,而 ProvenanceGuard 捕获了其中的 138 条,仅漏过了一条。它还拦截了 67 条专家认为已得到支持的声明,将其送交审查或修复。这反映了我们测试中所采用的谨慎设置:它宁愿对某些已支持的声明进行二次检查,也不愿让不支持的声明通过。