One MCP Gateway for All Your Internal APIs: The Aggregation Pattern
One MCP Gateway for All Your Internal APIs: The Aggregation Pattern
为所有内部 API 提供统一的 MCP 网关:聚合模式
A single MCP server in front of one service is a solved problem: generate tools from the OpenAPI spec, run it over stdio or HTTP, done. At company scale the problem changes shape. A midsize platform team has forty services, each with its own spec, its own auth, its own staging and production hosts. 在单个服务前部署一个 MCP 服务器是一个已解决的问题:根据 OpenAPI 规范生成工具,通过 stdio 或 HTTP 运行,搞定。但在公司规模下,问题就变了样。一个中型平台团队可能有四十个服务,每个服务都有自己的规范、自己的认证方式以及自己的测试和生产环境主机。
Let every team publish an MCP endpoint and you quickly get: Agents configured against dozens of URLs, each with its own OAuth consent. Tool catalogs in the hundreds, past the limit most clients expose to the model, so tools silently disappear. Name collisions: three list_users, two create_order, no way to tell them apart. No central point for rate limiting, audit logs, or revocation.
如果让每个团队都发布一个 MCP 端点,你很快会遇到这些问题:代理(Agents)需要配置几十个 URL,每个都需要单独的 OAuth 授权;工具目录达到数百个,超过了大多数客户端能向模型展示的上限,导致工具悄无声息地消失;名称冲突:三个 list_users,两个 create_order,无法区分;没有统一的限流、审计日志或权限撤销中心。
The answer that keeps working as services multiply is an MCP gateway: one endpoint an agent authenticates against, which aggregates many backend APIs into one namespaced, governed tool catalog. 随着服务数量的增加,唯一行之有效的方案是 MCP 网关:代理只需向一个端点进行身份验证,该网关将多个后端 API 聚合为一个带有命名空间、受治理的工具目录。
The shape of the pattern
模式架构
AI clients (Claude Desktop, Cursor, VS Code, CI agents)
│ OAuth 2.1, one consent, one token
▼
┌──────────────── MCP gateway ────────────────┐
│ auth & scopes rate limits audit log │
│ catalog aggregation name normalization │
└───────┬───────────┬───────────┬─────────────┘
▼ ▼ ▼
orders API billing API support API
(OpenAPI) (OpenAPI) (OpenAPI)
│ │ │
▼ ▼ ▼
services + databases (private network)
The gateway is not a new API implementation. It is a composition layer: each backend keeps owning its spec and its service; the gateway compiles those specs into one MCP catalog and proxies calls. 网关并不是一个新的 API 实现,而是一个组合层:每个后端继续维护自己的规范和服务;网关将这些规范编译成一个 MCP 目录并代理调用请求。
Catalog aggregation: namespace, do not flatten
目录聚合:使用命名空间,而非扁平化
The core design decision is how services appear in tools/list. Flattening produces collisions and ambiguity. Prefixing every tool with a service namespace produces a predictable, greppable catalog:
核心设计决策在于服务如何在 tools/list 中呈现。扁平化会导致冲突和歧义。为每个工具添加服务命名空间前缀,可以生成一个可预测且易于搜索的目录:
{
"tools": [
{
"name": "orders__create_order",
"description": "[orders] Create a pending order and reserve inventory for 15 minutes.",
"inputSchema": { "$ref": "#/$defs/orders.CreateOrderRequest" }
},
{
"name": "billing__create_invoice",
"description": "[billing] Issue an invoice for a fulfilled order.",
"inputSchema": { "$ref": "#/$defs/billing.CreateInvoiceRequest" }
}
]
}
Three rules keep the catalog usable: 三条规则确保目录保持可用性:
- Namespaces are stable and short (orders, billing), taken from a registry rather than guessed from repo names.
- 命名空间必须稳定且简短(如 orders, billing),应从注册表中获取,而不是从仓库名称中猜测。
- Schema components are namespaced too (orders.Order, billing.Invoice) so
$refresolution never merges two teams’ Error models. - Schema 组件也必须命名空间化(如 orders.Order, billing.Invoice),这样
$ref解析就不会合并两个团队的错误模型。 - Descriptions carry the namespace as a prefix, because models read descriptions more reliably than tool-name conventions.
- 描述应以命名空间作为前缀,因为模型读取描述比读取工具命名约定更可靠。
The gateway should also deduplicate genuinely shared schemas by reference rather than copying them, so the model sees one Money concept, not twelve. 网关还应通过引用对真正共享的 Schema 进行去重,而不是复制它们,这样模型看到的只是一个“货币”概念,而不是十二个。
Tool budget: more services does not mean more exposed tools
工具预算:服务增多不代表工具暴露增多
Clients and models have practical limits on tool count; hundreds of tools degrade selection accuracy even where the client technically accepts them. A gateway earns its keep by managing that budget: 客户端和模型对工具数量有实际限制;即使客户端在技术上能接受,数百个工具也会降低选择准确性。网关通过管理这一预算来体现其价值:
- Capability scopes filter the catalog. A token with
orders:read,billing:readsees only those namespaces;tools/listitself is authorization-aware. - 能力范围(Scopes)过滤目录。 拥有
orders:read,billing:read权限的令牌只能看到这些命名空间;tools/list本身就是具备授权意识的。 - Task-scoped views expose a curated subset for common workflows (“incident triage”, “order-to-cash”) instead of the entire platform.
- 任务范围视图为常见工作流(如“事件分类”、“订单到现金”)提供精选子集,而不是暴露整个平台。
- Read and write tools split by scope, so broad onboarding can start read-only; the same principle as MCP least-privilege design.
- 读写工具按范围拆分,因此广泛的入职培训可以从只读开始;这与 MCP 的最小权限设计原则一致。
- Deep query operations stay resources. Search endpoints and document lookups fit MCP resources better than dozens of getter tools.
- 深度查询操作应保留为资源(Resources)。 搜索端点和文档查找比几十个 getter 工具更适合作为 MCP 资源。
Auth: one token in, service credentials out
认证:令牌入,服务凭证出
The agent authenticates once against the gateway using OAuth 2.1 with PKCE. The gateway then holds service-to-service credentials downstream and maps the caller’s identity and scopes onto each request: 代理使用带有 PKCE 的 OAuth 2.1 向网关进行一次身份验证。随后,网关在下游持有服务间凭证,并将调用者的身份和范围映射到每个请求中:
agent token (scopes: orders:read, billing:write)
│
▼
gateway authorizes orders__get_order → allowed, signs request as gateway, forwards user identity
gateway authorizes billing__create_invoice → allowed, mints scoped service token
gateway authorizes admin__delete_account → denied at the edge with a structured MCP error
Forward the authenticated user’s identity to backends in a signed header or token exchange rather than making every call anonymous; otherwise audit trails end at the gateway. 应通过签名头或令牌交换将已认证的用户身份转发给后端,而不是让每个调用都匿名;否则审计追踪将在网关处中断。
The operational features that belong at the edge
属于边缘层的运维功能
Because every tool call crosses the gateway, concerns that would be duplicated forty times are implemented once: 由于每个工具调用都会经过网关,原本需要重复实现四十次的功能现在只需实现一次:
- Rate limiting and quotas, per user and per namespace, with
Retry-Afteron throttles. - 限流与配额,按用户和命名空间划分,并在受限时返回
Retry-After。 - Audit logging in one schema, including the calling client, scope used, and arguments with PII redacted.
- 审计日志,采用统一 Schema,包含调用客户端、使用的范围以及脱敏后的参数。
- Versioning and deprecation: the gateway can serve
orders__create_orderandorders__create_order_v2side by side while backends migrate. - 版本控制与弃用:网关可以在后端迁移期间同时提供
orders__create_order和orders__create_order_v2。 - Observability: latency and error-rate metrics per namespace give teams feedback on how agents actually use their APIs.
- 可观测性:按命名空间统计的延迟和错误率指标,能让团队了解代理实际使用其 API 的情况。
- Fail containment: a backend outage degrades one namespace and returns a clean, retryable MCP error instead of failing the whole catalog.
- 故障隔离:后端中断只会导致一个命名空间降级,并返回一个清晰、可重试的 MCP 错误,而不会导致整个目录失效。
Build vs buy, and what to do on day one
构建还是购买,以及第一天该做什么
Do not start by building a platform. The aggregation pattern can be introduced incrementally: 不要一上来就构建平台。聚合模式可以循序渐进地引入:
- Week one: publish one generated MCP server for the highest-value service, remotely, with OAuth and logs.
- 第一周:为最有价值的服务发布一个远程生成的 MCP 服务器,并配置 OAuth 和日志。
- Week two: put a thin gateway in front of it that does nothing but auth and pass-through, so clients already point at the stable URL.
- 第二周:在其前方放置一个轻量级网关,仅处理认证和透传,以便客户端指向稳定的 URL。
- As demand grows: register additional services by adding their OpenAPI specs to the gateway’s catalog build, each in its namespace, without clients changing their config.
- 随着需求增长:通过将其他服务的 OpenAPI 规范添加到网关的目录构建中来注册新服务,每个服务使用各自的命名空间,客户端无需更改配置。
- Only when needed: add scoped catalog views, quotas, and curated task bundles.
- 仅在需要时:添加范围化的目录视图、配额和精选任务包。
Because the tool catalog is generated from specs, adding a service is a configuration and build step — register the spec, declare the namespace, choose the scopes — rather than an integration project. 由于工具目录是根据规范生成的,添加服务只是一个配置和构建步骤(注册规范、声明命名空间、选择范围),而不是一个集成项目。