从智能体获取结构化输出
用 JSON Schema、Zod 或 Pydantic 从智能体工作流返回经过验证的 JSON:outputFormat 配置、类型安全的 schema、TODO 跟踪示例、错误处理(error_max_structured_output_retries)与避免错误的建议。
结构化输出让你定义想从智能体得到的数据的确切形状。智能体可以用它需要的任何工具完成任务,最后你仍能得到与你的 schema 匹配的、经过验证的 JSON。为你需要的结构定义一个 JSON Schema,SDK 会按它验证输出,不匹配时重新提示。如果在重试限制内验证没有成功,结果就是错误而不是结构化数据(见「错误处理」)。要获得完整的类型安全,用 Zod(TypeScript)或 Pydantic(Python)定义你的 schema,并得到强类型对象。
为什么用结构化输出?
智能体默认返回自由格式的文本,这适合聊天,但当你需要以编程方式使用输出时就不行了。结构化输出给你类型化的数据,可以直接传给你的应用逻辑、数据库或 UI 组件。
设想一个食谱应用,智能体在网上搜索并带回食谱。没有结构化输出时,你得到的是需要自己解析的自由文本:
Here's a classic chocolate chip cookie recipe!
**Chocolate Chip Cookies**
Prep time: 15 minutes | Cook time: 10 minutes
Ingredients:
- 2 1/4 cups all-purpose flour
- 1 cup butter, softened
...要在应用里用它,你得解析出标题、把 "15 minutes" 转成数字、把配料与步骤分开,并处理各次响应里不一致的格式。有了结构化输出,你定义想要的形状,得到可以直接在应用里使用的类型化数据:
{
"name": "Chocolate Chip Cookies",
"prep_time_minutes": 15,
"cook_time_minutes": 10,
"ingredients": [
{ "item": "all-purpose flour", "amount": 2.25, "unit": "cups" },
{ "item": "butter, softened", "amount": 1, "unit": "cup" }
// ...
],
"steps": ["Preheat oven to 375°F", "Cream butter and sugar" /* ... */]
}快速开始
要使用结构化输出,定义一个描述所需数据形状的 JSON Schema,再通过 outputFormat 选项(TypeScript)或 output_format 选项(Python)把它传给 query()。智能体完成后,结果消息包含带有与你的 schema 匹配的、经过验证数据的 structured_output 字段。下面的例子让智能体研究 Anthropic,并以结构化输出返回公司名、成立年份和总部。
import { query } from "@anthropic-ai/claude-agent-sdk";
// 定义你想要返回的数据形状
const schema = {
type: "object",
properties: {
company_name: { type: "string" },
founded_year: { type: "number" },
headquarters: { type: "string" }
},
required: ["company_name"]
};
try {
for await (const message of query({
prompt: "Research Anthropic and provide key company information",
options: {
outputFormat: {
type: "json_schema",
schema: schema
}
}
})) {
// 结果消息包含带验证数据的 structured_output
if (message.type === "result" && message.subtype === "success" && message.structured_output) {
console.log(message.structured_output);
// { company_name: "Anthropic", founded_year: 2021, headquarters: "San Francisco, CA" }
}
}
} catch (error) {
// 单次 query() 在产出错误结果(如 error_max_structured_output_retries)
// 之后抛出;见「错误处理」一节。
console.error(`Session ended with an error: ${error}`);
}import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions, ResultMessage
# 定义你想要返回的数据形状
schema = {
"type": "object",
"properties": {
"company_name": {"type": "string"},
"founded_year": {"type": "number"},
"headquarters": {"type": "string"},
},
"required": ["company_name"],
}
async def main():
try:
async for message in query(
prompt="Research Anthropic and provide key company information",
options=ClaudeAgentOptions(
output_format={"type": "json_schema", "schema": schema}
),
):
# 结果消息包含带验证数据的 structured_output
if isinstance(message, ResultMessage) and message.structured_output:
print(message.structured_output)
# {'company_name': 'Anthropic', 'founded_year': 2021, 'headquarters': 'San Francisco, CA'}
except Exception as error:
# 单次 query() 在产出错误结果(如 error_max_structured_output_retries)
# 之后抛出;见「错误处理」一节。
print(f"Session ended with an error: {error}")
asyncio.run(main())用 Zod 和 Pydantic 得到类型安全的 schema
你可以用 Zod(TypeScript)或 Pydantic(Python)来定义 schema,而不是手写 JSON Schema。这些库替你生成 JSON Schema,并让你把响应解析成完整类型的对象,在整个代码库里享受自动补全和类型检查。下面的例子为功能实现计划定义 schema,含摘要、步骤列表(每步带复杂度级别)和潜在风险。智能体规划该功能并返回类型化的 FeaturePlan 对象,之后你可以访问 plan.summary 这样的属性并遍历 plan.steps,全程类型安全。
SDK 用 JSON Schema draft-07 验证 schema,所以声明更新版本的 schema 会被拒绝。Zod 默认面向 draft 2020-12,所以转换 schema 时要传 target: "draft-7"。
import { z } from "zod";
import { query } from "@anthropic-ai/claude-agent-sdk";
// 用 Zod 定义 schema
const FeaturePlan = z.object({
feature_name: z.string(),
summary: z.string(),
steps: z.array(
z.object({
step_number: z.number(),
description: z.string(),
estimated_complexity: z.enum(["low", "medium", "high"])
})
),
risks: z.array(z.string())
});
type FeaturePlan = z.infer<typeof FeaturePlan>;
// 用 SDK 期望的 draft-07 目标转换成 JSON Schema
const schema = z.toJSONSchema(FeaturePlan, { target: "draft-7" });
// 在查询里使用
try {
for await (const message of query({
prompt:
"Plan how to add dark mode support to a React app. Break it into implementation steps.",
options: {
outputFormat: {
type: "json_schema",
schema: schema
}
}
})) {
if (message.type === "result" && message.subtype === "success" && message.structured_output) {
// 验证并得到完整类型的结果
const parsed = FeaturePlan.safeParse(message.structured_output);
if (parsed.success) {
const plan: FeaturePlan = parsed.data;
console.log(`Feature: ${plan.feature_name}`);
console.log(`Summary: ${plan.summary}`);
plan.steps.forEach((step) => {
console.log(`${step.step_number}. [${step.estimated_complexity}] ${step.description}`);
});
}
}
}
} catch (error) {
// 单次 query() 在产出错误结果(如 error_max_structured_output_retries)
// 之后抛出;见「错误处理」一节。
console.error(`Session ended with an error: ${error}`);
}import asyncio
from pydantic import BaseModel
from claude_agent_sdk import query, ClaudeAgentOptions, ResultMessage
class Step(BaseModel):
step_number: int
description: str
estimated_complexity: str # 'low'、'medium'、'high'
class FeaturePlan(BaseModel):
feature_name: str
summary: str
steps: list[Step]
risks: list[str]
async def main():
try:
async for message in query(
prompt="Plan how to add dark mode support to a React app. Break it into implementation steps.",
options=ClaudeAgentOptions(
output_format={
"type": "json_schema",
"schema": FeaturePlan.model_json_schema(),
}
),
):
if isinstance(message, ResultMessage) and message.structured_output:
# 验证并得到完整类型的结果
plan = FeaturePlan.model_validate(message.structured_output)
print(f"Feature: {plan.feature_name}")
print(f"Summary: {plan.summary}")
for step in plan.steps:
print(
f"{step.step_number}. [{step.estimated_complexity}] {step.description}"
)
except Exception as error:
# 单次 query() 在产出错误结果(如 error_max_structured_output_retries)
# 之后抛出;见「错误处理」一节。
print(f"Session ended with an error: {error}")
asyncio.run(main())输出格式配置
outputFormat(TypeScript)或 output_format(Python)选项接受一个对象,含:type——设为 "json_schema" 表示结构化输出;schema——定义输出结构的 JSON Schema 对象,可以用 z.toJSONSchema(schema, { target: "draft-7" }) 从 Zod schema 生成,或用 .model_json_schema() 从 Pydantic 模型生成。SDK 支持标准的 JSON Schema 特性,包括所有基本类型(object、array、string、number、boolean、null)、enum、const、required、嵌套对象和 $ref 定义(支持的特性和限制的完整列表见「JSON Schema 限制」)。
不是有效 JSON Schema 的 schema 会在启动时让运行失败,并给出点名问题的错误;v2.1.205 之前,无效的 schema 被静默忽略,智能体返回非结构化文本。format 关键字(如 "format": "email")被作为注解接受,SDK 的验证器不强制它;v2.1.205 之前,任何含 format 的 schema 都被当作无效。
示例:TODO 跟踪智能体
这个例子演示结构化输出如何配合多步骤工具使用。智能体需要在代码库里找出 TODO 注释,再为每一个查询 git blame 信息。它自主决定使用哪些工具(用 Grep 搜索、用 Bash 运行 git 命令),并把结果合并成一个结构化响应。schema 包含可选字段(author 和 date),因为并非所有文件都有 git blame 信息;智能体填入它能找到的,省略其余的。
import { query } from "@anthropic-ai/claude-agent-sdk";
// 定义 TODO 提取的结构
const todoSchema = {
type: "object",
properties: {
todos: {
type: "array",
items: {
type: "object",
properties: {
text: { type: "string" },
file: { type: "string" },
line: { type: "number" },
author: { type: "string" },
date: { type: "string" }
},
required: ["text", "file", "line"]
}
},
total_count: { type: "number" }
},
required: ["todos", "total_count"]
};
// 智能体用 Grep 找 TODO,用 Bash 获取 git blame 信息
try {
for await (const message of query({
prompt: "Find all TODO comments in this codebase and identify who added them",
options: {
outputFormat: {
type: "json_schema",
schema: todoSchema
}
}
})) {
if (message.type === "result" && message.subtype === "success" && message.structured_output) {
const data = message.structured_output as { total_count: number; todos: Array<{ file: string; line: number; text: string; author?: string; date?: string }> };
console.log(`Found ${data.total_count} TODOs`);
data.todos.forEach((todo) => {
console.log(`${todo.file}:${todo.line} - ${todo.text}`);
if (todo.author) {
console.log(` Added by ${todo.author} on ${todo.date}`);
}
});
}
}
} catch (error) {
// 单次 query() 在产出错误结果(如 error_max_structured_output_retries)
// 之后抛出;见「错误处理」一节。
console.error(`Session ended with an error: ${error}`);
}import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions, ResultMessage
# 定义 TODO 提取的结构
todo_schema = {
"type": "object",
"properties": {
"todos": {
"type": "array",
"items": {
"type": "object",
"properties": {
"text": {"type": "string"},
"file": {"type": "string"},
"line": {"type": "number"},
"author": {"type": "string"},
"date": {"type": "string"},
},
"required": ["text", "file", "line"],
},
},
"total_count": {"type": "number"},
},
"required": ["todos", "total_count"],
}
async def main():
# 智能体用 Grep 找 TODO,用 Bash 获取 git blame 信息
try:
async for message in query(
prompt="Find all TODO comments in this codebase and identify who added them",
options=ClaudeAgentOptions(
output_format={"type": "json_schema", "schema": todo_schema}
),
):
if isinstance(message, ResultMessage) and message.structured_output:
data = message.structured_output
print(f"Found {data['total_count']} TODOs")
for todo in data["todos"]:
print(f"{todo['file']}:{todo['line']} - {todo['text']}")
if "author" in todo:
print(f" Added by {todo['author']} on {todo['date']}")
except Exception as error:
# 单次 query() 在产出错误结果(如 error_max_structured_output_retries)
# 之后抛出;见「错误处理」一节。
print(f"Session ended with an error: {error}")
asyncio.run(main())错误处理
当智能体无法产生与你的 schema 匹配的有效 JSON 时,结构化输出的生成可能失败。这通常发生在 schema 对任务太复杂、任务本身含糊,或智能体在试图修复验证错误时达到重试限制。它也可能在没有任何验证失败的情况下发生:模型回退可能在流中途撤回已完成的输出,如果没有重试来替换它,运行就以同样的错误结束。调试 schema 之前,先检查结果消息上的 errors 列表来区分这两种原因。出错时,结果消息的 subtype 指示出了什么问题:
| Subtype | 含义 |
|---|---|
success | 输出已生成并验证成功 |
error_max_structured_output_retries | 多次尝试后没有剩余的有效输出(验证失败,或模型回退撤回后没有成功的重试) |
结果也可能以 subtype success 结束却没有 structured_output 值,例如运行完成了但智能体没有产生结构化输出;这种情况也要当作失败处理(排障条目「structured_output 是 None 但结果显示 success」涵盖它)。下面的例子只在 subtype 是 success 且 structured_output 存在时才把结果当作成功,并把其他每种结果都当作失败处理:
import { query } from "@anthropic-ai/claude-agent-sdk";
const contactSchema = {
type: "object",
properties: {
name: { type: "string" },
email: { type: "string" }
},
required: ["name"]
};
try {
for await (const msg of query({
prompt: "Extract contact info from the document",
options: {
outputFormat: {
type: "json_schema",
schema: contactSchema
}
}
})) {
if (msg.type === "result") {
if (msg.subtype === "success" && msg.structured_output) {
// 使用验证过的输出
console.log(msg.structured_output);
} else if (msg.subtype === "error_max_structured_output_retries") {
console.error("Could not produce valid output");
} else {
console.error("Run ended without a structured output");
}
}
}
} catch (error) {
// 单次 query() 在产出错误结果之后抛出。如果失败是错误结果,
// 上面的错误 subtype 分支已经运行过了;连接或进程失败不产生结果消息。
console.log(`Session ended with an error: ${error}`);
}import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions, ResultMessage
contact_schema = {
"type": "object",
"properties": {
"name": {"type": "string"},
"email": {"type": "string"},
},
"required": ["name"],
}
async def main():
try:
async for message in query(
prompt="Extract contact info from the document",
options=ClaudeAgentOptions(
output_format={"type": "json_schema", "schema": contact_schema}
),
):
if isinstance(message, ResultMessage):
if message.subtype == "success" and message.structured_output:
# 使用验证过的输出
print(message.structured_output)
elif message.subtype == "error_max_structured_output_retries":
print("Could not produce valid output")
else:
print("Run ended without a structured output")
except Exception as error:
# 单次 query() 在产出错误结果之后抛出。如果失败是错误结果,
# 上面的错误 subtype 分支已经运行过了;连接或进程失败不产生结果消息。
print(f"Session ended with an error: {error}")
asyncio.run(main())避免错误的建议:
- 让 schema 聚焦。 带许多必填字段的深层嵌套 schema 更难满足;从简单开始,按需增加复杂度。
- 让 schema 与任务匹配。 如果任务可能没有你的 schema 所要求的全部信息,就把这些字段设为可选。
- 使用清晰的提示。 含糊的提示让智能体更难知道该产生什么输出。
相关资源
- JSON Schema 文档:学习用于定义带嵌套对象、数组、枚举和验证约束的复杂 schema 的 JSON Schema 语法
- API 结构化输出:直接对 Claude API 使用结构化输出,用于不使用工具的单轮请求
- 自定义工具:给你的智能体自定义工具,供它在返回结构化输出之前的执行期间调用