Cartograph: Federated Tool Discovery with Operator-Attested Retrieval for AI Agents
Cartograph: Federated Tool Discovery with Operator-Attested Retrieval for AI Agents
Cartograph:面向 AI 智能体的联邦式工具发现与操作员认证检索系统
Abstract: The Model Context Protocol (MCP) enables AI agents to discover and call tools, but loading every definition becomes expensive as connected catalogs grow. We present Cartograph, a federated MCP proxy that changes agent-visible tool discovery from $O(n)$ catalog traversal to $O(k)$ progressive disclosure.
摘要: 模型上下文协议(MCP)使 AI 智能体能够发现并调用工具,但随着连接目录的增长,加载所有定义变得成本高昂。我们提出了 Cartograph,这是一个联邦式 MCP 代理,它将智能体可见的工具发现方式从 $O(n)$ 的目录遍历转变为 $O(k)$ 的渐进式披露。
Cartograph combines three mechanisms: (1) operator-attested capability cards, Ed25519-signed descriptions generated under the deploying operator’s control rather than ranked publisher copy; (2) Rift, a three-layer confusable-cluster analysis comprising density clustering, query-margin analysis, and token diagnosis; and (3) two-stage retrieval, which ranks servers before tools.
Cartograph 结合了三种机制:(1) 操作员认证的能力卡(Capability Cards),即由部署操作员控制生成的 Ed25519 签名描述,而非由发布者排名的副本;(2) Rift,一种包含密度聚类、查询边界分析和 Token 诊断的三层混淆聚类分析;以及 (3) 两阶段检索,即先对服务器进行排序,再对工具进行排序。
On a 22-server, 374-tool deployment, Cartograph exposes three proxy tools instead of 374 definitions. A 49-query author-constructed benchmark yields R@5 of 0.816, compared with 0.592 for a Jaccard keyword baseline, while a measured top-5 discovery exchange uses 475 tokens rather than 42,450 under the stated full-catalog accounting.
在包含 22 台服务器和 374 个工具的部署环境中,Cartograph 仅暴露了 3 个代理工具,而非 374 个定义。在作者构建的 49 个查询基准测试中,其 R@5 指标达到 0.816,而 Jaccard 关键词基准仅为 0.592;同时,在测量的前 5 名发现交换中,仅使用了 475 个 Token,远低于全目录统计下的 42,450 个 Token。
Rift identifies 49 confusable clusters, including four HIGH-risk clusters in bootstrap-generated cards. An exploratory comparison of 119 LLM-generated descriptions removes the observed zero-distance cluster but shows that mixing card-generation regimes can reduce R@5. Gateway measurements over ten trials add 5ms mean latency (0.8%) relative to direct stdio MCP calls.
Rift 识别出了 49 个混淆聚类,其中包括引导生成卡片中的 4 个高风险聚类。对 119 个由 LLM 生成的描述进行的探索性比较表明,虽然可以消除观察到的零距离聚类,但混合使用不同的卡片生成机制可能会降低 R@5。经过十次试验的网关测量显示,相对于直接的 stdio MCP 调用,其平均延迟仅增加了 5 毫秒(0.8%)。
Cartograph is complementary to code-execution approaches: it controls which tool descriptions are surfaced and records the provenance of the descriptions used for ranking for each query.
Cartograph 与代码执行方法互补:它控制哪些工具描述被呈现,并记录每次查询中用于排序的描述的来源。