LLM Structured Outputs and JSON Schema: Tool Calling That Never Drifts

LLM Structured Outputs and JSON Schema: Tool Calling That Never Drifts

LLM 结构化输出与 JSON Schema:永不偏移的工具调用

The gap between a demo agent and a reliable one is usually not model intelligence; it is output discipline. In a demo, the model calls get_weather({"city": "SF"}) and looks magical. In production it calls refund_payment({"payment_id": 7712, "amount": "49.00", "currency": "usd", "reason": "user asked nicely"}) — integer where you expect a string, string where you expect cents, an undeclared reason, and an extra field your handler silently ignores. Every one of those mismatches is a bug report that says “the AI is unreliable” when the real cause is an unconstrained output contract.

演示级 Agent 与可靠生产级 Agent 之间的差距,通常不在于模型智能,而在于输出的规范性。在演示中,模型调用 get_weather({"city": "SF"}) 看起来非常神奇。但在生产环境中,它可能会调用 refund_payment({"payment_id": 7712, "amount": "49.00", "currency": "usd", "reason": "user asked nicely"}) —— 本该是字符串的地方变成了整数,本该是分的地方变成了字符串,还包含了一个未定义的理由字段,以及一个被你的处理器静默忽略的额外字段。每一个此类不匹配都会导致一份“AI 不可靠”的错误报告,但其根本原因其实是输出契约缺乏约束。

Structured outputs fix this at the model layer, and JSON Schema is the language every provider converged on. If you maintain an OpenAPI document, you already own most of the schemas involved.

结构化输出在模型层解决了这个问题,而 JSON Schema 是所有服务商达成共识的通用语言。如果你维护着 OpenAPI 文档,那么你其实已经拥有了大部分所需的 Schema。

What structured outputs actually guarantees

结构化输出究竟保证了什么

Providers implement this under different names — OpenAI’s structured outputs and function calling strict mode, Anthropic’s tool-use input schemas, Gemini’s responseSchema — but the guarantee is the same shape: given a schema the provider can enforce, the response is guaranteed to validate against it. Invalid JSON, missing required fields, and types that do not match are eliminated at decoding time rather than surfacing in your application.

各服务商以不同的名称实现了这一功能——OpenAI 的结构化输出与函数调用严格模式、Anthropic 的工具使用输入 Schema、Gemini 的 responseSchema——但其保证的核心是一致的:给定一个服务商可强制执行的 Schema,响应保证符合该 Schema。无效的 JSON、缺失的必填字段以及类型不匹配的问题,会在解码阶段被直接消除,而不会在你的应用程序中暴露出来。

That guarantee is deliberately narrow. It does not promise the values are correct (the model can still refund the wrong payment), only that they are valid. Validation is exactly the layer you should never have had to hand-write in natural language; correctness still requires good tool design, confirmation gates, and tests.

这一保证是有意为之的窄范围保证。它不承诺数值是正确的(模型仍然可能退错款),只承诺数值是有效的。验证层正是你不应该通过自然语言手写的部分;而正确性仍然依赖于良好的工具设计、确认机制和测试。

The strict subset you must design within

你必须遵循的严格子集

Providers enforce structured outputs by constraining the JSON Schema dialect they accept. The rules are consistent across the major platforms and easy to internalize: Every object must list "additionalProperties": false and enumerate all properties. Every property must appear in required; optionality is expressed with a union including null, not by omission. Types are explicit; use type: ["string", "null"] for nullable fields. $ref is supported against $defs (or definitions), which keeps schemas DRY. Avoid unsupported keywords (format is often advisory, not enforced; avoid patternProperties, conditional schemas, and tuple-heavy arrays unless your provider documents them).

服务商通过限制其接受的 JSON Schema 方言来强制执行结构化输出。各大平台遵循的规则是一致的且易于掌握:每个对象必须列出 "additionalProperties": false 并枚举所有属性。每个属性都必须出现在 required 列表中;可选性通过包含 null 的联合类型来表达,而不是通过省略字段。类型必须明确;对于可为空的字段,请使用 type: ["string", "null"]。支持针对 $defs(或 definitions)的 $ref 引用,这有助于保持 Schema 的 DRY(不重复)原则。避免使用不支持的关键字(format 通常仅供参考,不强制执行;除非服务商有明确文档说明,否则应避免使用 patternProperties、条件 Schema 和复杂的元组数组)。

A strict refund schema looks like this: 一个严格的退款 Schema 如下所示:

{
  "type": "object",
  "additionalProperties": false,
  "properties": {
    "payment_id": { "type": "string" },
    "amount_cents": { "type": ["integer", "null"], "minimum": 1 },
    "currency": { "type": "string", "enum": ["USD", "EUR", "GBP"] },
    "reason": { "type": "string", "enum": ["customer_request", "duplicate", "fraud"] },
    "notify_customer": { "type": "boolean" }
  },
  "required": ["payment_id", "amount_cents", "currency", "reason", "notify_customer"]
}

There is nowhere for an invented field to land, no ambiguity about whether the amount is a string, and a closed enum for the reason instead of a prose judgment the support team later has to parse.

这里没有多余字段的容身之地,关于金额是否为字符串也没有歧义,且“理由”字段使用了封闭的枚举值,而不是需要支持团队后续解析的散文式判断。

Enums and unions carry the semantics

枚举与联合类型承载语义

Models are remarkably good at mapping messy human intent onto closed value sets, and remarkably bad at inventing consistent open-ended codes. Push decisions into schemas wherever a value set is finite: statuses, categories, sort orders, time windows. When a concept is genuinely open (a customer-facing note), keep it a string and constrain the rest of the shape. For nullable optionality, prefer the union form over omitting the key. A consistent object shape simplifies both the model’s job and your handler: there is no difference between “key absent” and “key present with null” to test.

模型非常擅长将混乱的人类意图映射到封闭的值集合上,但非常不擅长发明一致的开放式代码。只要值集合是有限的,就应将其决策推入 Schema 中:例如状态、类别、排序方式、时间窗口等。当概念确实是开放的(例如面向客户的备注)时,将其保留为字符串,并约束其余的结构。对于可为空的可选性,优先使用联合类型形式,而不是省略键。一致的对象结构简化了模型的工作和你的处理器逻辑:无需测试“键缺失”与“键存在但为 null”之间的区别。

Reuse the schemas you already publish

复用你已发布的 Schema

If the tool wraps an HTTP API, its arguments are the API’s request body, path parameters, and query parameters. Those are already modeled in OpenAPI under components.schemas. Duplicating them into tool definitions creates two contracts that drift the first time someone edits one.

如果工具封装了 HTTP API,那么它的参数就是 API 的请求体、路径参数和查询参数。这些已经在 OpenAPI 的 components.schemas 中建模了。将它们复制到工具定义中会创建两份契约,一旦有人修改其中一份,它们就会产生偏移。

A build step can combine the path parameters and the request schema into one strict tool input object, namespacing $defs to avoid collisions across services. The same component schema then drives:

  • Server-side validation (the API itself).
  • Generated typed clients in TypeScript and other languages.
  • MCP tool input schemas for agents.
  • Documentation examples.

构建步骤可以将路径参数和请求 Schema 合并为一个严格的工具输入对象,并通过命名空间 $defs 来避免跨服务冲突。同一个组件 Schema 随后可驱动:

  • 服务端验证(API 本身)。
  • TypeScript 及其他语言的生成式类型化客户端。
  • Agent 的 MCP 工具输入 Schema。
  • 文档示例。

This is the same single-source-of-truth argument behind generating a TypeScript client from OpenAPI, applied to the model interface.

这与从 OpenAPI 生成 TypeScript 客户端背后的“单一事实来源”逻辑相同,只是将其应用到了模型接口上。

Validate anyway, and make errors teachable

无论如何都要验证,并让错误具有可教性

Provider guarantees remove malformed output; they do not remove business-rule violations (refund over the captured amount, payment already refunded). Keep a server-side validation pass using the same schema, and return field-level, machine-readable errors so the model can self-correct in one turn.

服务商的保证消除了格式错误的输出,但无法消除业务规则违规(例如退款金额超过已捕获金额、支付已退款)。请使用相同的 Schema 保留服务端验证环节,并返回字段级、机器可读的错误,以便模型可以在一轮对话中自我修正。

A model receiving that response retries with a corrected value. A model receiving “refund failed” guesses. Error design for this audience is covered in “designing APIs for AI agents” and the error format itself in RFC 9457 Problem Details.

收到此类响应的模型会使用修正后的值重试。而收到“退款失败”这种模糊提示的模型只能靠猜。针对此类受众的错误设计可参考“为 AI Agent 设计 API”相关内容,错误格式本身可参考 RFC 9457 Problem Details。

Test the schema boundary like any other contract

像对待其他契约一样测试 Schema 边界

  • Schema meta-validation: every tool input schema validates against the strict subset; reject PRs that introduce unsupported keywords, which providers either reject at registration or silently downgrade.

  • Representative prompts: run a fixed set of natural-language requests through the model in CI against recorded responses, asserting they validate. This catches description problems (the model cannot tell which field is the amount) without testing the model’s logic.

  • Schema 元验证: 每个工具输入 Schema 都必须通过严格子集的验证;拒绝引入不支持关键字的 PR,因为这些关键字要么会被服务商在注册时拒绝,要么会被静默降级。

  • 代表性提示词: 在 CI 中运行一组固定的自然语言请求,对比记录的响应,并断言它们通过验证。这可以在不测试模型逻辑的情况下,捕获描述性问题(例如模型无法区分哪个字段是金额)。