Schema 驱动的决策
Schema 驱动的决策
把一份 JSON schema 或一个 pydantic 模型变成 Laya 问题,拿回带校准置信度的类型化值。正是这座桥 让 Laya 成为一个结构化输出引擎:你描述想要的形状,Laya 在一次前向传播里给出答案。
import laya
schema = {
"type": "object",
"properties": {
"department": {"type": "string", "enum": ["billing", "support", "sales"],
"description": "Which team should handle this?"},
"urgency": {"type": "integer", "minimum": 0, "maximum": 2},
"needs_human": {"type": "boolean"},
},
}
agent = laya.load("convaiinnovations/laya")
values = agent.decide("I was charged twice, refund me.", schema=schema)
# {"department": "billing", "urgency": 2, "needs_human": True}
用 pydantic(安装 laya[structured]):
from typing import Literal
from pydantic import BaseModel
class Ticket(BaseModel):
department: Literal["billing", "support", "sales"]
urgency: Literal[0, 1, 2]
needs_human: bool
ticket = agent.decide("I was charged twice, refund me.", schema=Ticket)
支持的子集
顶层必须是一个带 properties 的对象。每个属性变成一个 Laya 问题。
下面每一行都是真实的 schema:tests/test_structured_docs.py 编译第一列,并断言编译器实际产出的
问题,所以这张表不会和代码脱节。一个单元格要么是单独的属性 schema,要么是对某个入口点的调用。
| JSON schema | Laya 问题 | 返回的值 |
|---|---|---|
{"enum": ["billing", "support"]} |
choice |
被选中的值,保留它原本的类型 |
{"const": "billing"} |
choice |
那一个值 |
{"type": "boolean"} |
noul |
true / false |
{"type": "integer", "minimum": 0, "maximum": 5} |
score |
概率最高的档位,作为一个整数 |
{"type": "number", "minimum": 0, "maximum": 5} |
score |
那个档位,作为一个整数 |
{"type": "string", "enum": ["low", "high"], "description": "How urgent?"} |
choice |
How urgent? 就是问题措辞 |
{"anyOf": [{"enum": ["x", "y"]}, {"type": "null"}]} |
choice |
和普通的 enum 行一样;没有答案时这个键不出现 |
{"oneOf": [{"type": "boolean"}, {"type": "null"}]} |
noul |
和普通的 boolean 行一样 |
{"type": ["integer", "null"], "minimum": 1, "maximum": 3} |
score |
和普通的带界整数行一样 |
Literal[...] 和 Optional[...] 是 pydantic 里对应 enum 和 anyOf 两行的写法:
questions_from_pydantic 把它们渲染成那些形状,同样的行适用。
title 不会被读取。pydantic 不管你有没有要求,都会给 model_json_schema() 的每个字段放上
一个,而一个按属性的名字没法给一个问题所据以构建的那些逐选项 choice 打标签,所以措辞的杠杆是
description —— 见下面的它内部如何映射。
投影是精确的:enum: [1, 2, 3] 返回 2,不是 "2";带界整数返回 minimum 和 maximum
之间的一个档位;布尔值就是 noul >= 0.5。
拒绝
一个无法从固定选项集作答的 schema 会抛出 laya.structured.SchemaError(一个 ValueError),
并点出确切的路径。每一行也都会被执行,字段名取 name:
| 属性 schema | 消息 |
|---|---|
{"type": "string"} |
properties.name: a free string cannot be a fixed option set; use 'enum' or a boolean |
{"type": "array", "items": {"type": "string"}} |
properties.name: arrays are not supported; ask one field per element |
{"type": "object", "properties": {"inner": {"type": "boolean"}}} |
properties.name: nested objects are not supported; flatten the schema |
{"$ref": "#/definitions/node"} |
properties.name: $ref/recursion is not supported; flatten the schema |
{"enum": [1, "1"]} |
properties.name: enum values produce duplicate choice labels |
{"enum": []} |
properties.name: 'enum' must not be empty |
{"type": "number"} |
properties.name: a numeric field needs integer 'minimum' and 'maximum' to become a score |
{"type": "integer", "minimum": 5, "maximum": 2} |
properties.name: 'maximum' 2 is below 'minimum' 5 |
{"type": "integer", "minimum": 0, "maximum": 10} |
properties.name: 11 levels exceeds MAX_SCORE_LEVELS=10; narrow the range or use an enum |
{"anyOf": [{"type": "string"}, {"type": "integer"}]} |
properties.name: only 'Optional[...]' unions (one non-null branch) are supported, got 2 |
{"type": ["string", "integer"]} |
properties.name: 'type' has multiple non-null types; unions are not supported |
{"format": "date"} |
properties.name: unsupported schema {'format': 'date'} |
"boolean" |
properties.name: property must be an object, got str |
入口点自身会拒绝这些:
| 调用 | 消息 |
|---|---|
plan_from_json_schema("not a schema") |
expected a JSON schema object, got str |
plan_from_json_schema({"type": "object"}) |
the top level must be an object with 'properties' |
plan_from_json_schema({"type": "object", "properties": {}}) |
'properties' must be a non-empty object |
plan_from_json_schema({"type": "object", "properties": {"p%d" % i: {"type": "boolean"} for i in range(33)}}) |
33 properties exceeds MAX_PROPERTIES=32 |
plan_from_json_schema({"type": "object", "properties": {"name": {"enum": ["v%d" % i for i in range(33)]}}}) |
properties.name: 33 options exceeds MAX_OPTIONS=32 |
decide(None, "I was charged twice.", schema=42) |
expected a JSON schema dict or a pydantic model, got int |
限制:MAX_PROPERTIES = 32、MAX_OPTIONS = 32、MAX_SCORE_LEVELS = 10。
API
| 函数 | 用途 |
|---|---|
laya.decide(runner, state, schema=..., *, questions=..., return_details=..., min_confidence=..., **predict_kwargs) |
自由函数,对 Agent 和 Router 都适用 |
agent.decide(state, schema=..., ...) / router.decide(state, schema=..., ...) |
便捷方法 |
laya.decide_batch(runner, states, schema=..., ...) / agent.decide_batch(...) / router.decide_batch(...) |
在多个状态上做同样的事,一次批调用 |
questions_from_json_schema(schema) |
schema 转成 Laya 问题 |
questions_from_pydantic(model) |
pydantic 模型转成问题(需要 pydantic) |
answers_to_json(answers, schema) |
把原始答案投影到 schema 的值上 |
answer_to_pydantic(model, answers) |
把原始答案投影成一个 pydantic 实例 |
plan_from_json_schema(schema) |
经过校验的字段计划(进阶) |
schema 和 questions 只能传其中一个。传 questions 时,decide 返回原始答案而不是做投影。
额外的关键字参数会转发给 predict,所以钩子、model=、task= 和 token 预算都能用:
router.decide(state, schema=Ticket, model="multilingual", hooks=[Metrics()])
给多个状态打分
decide_batch 是吞吐的形式:schema 只规划一次,它的问题通过 predict_batch 在每一个状态上跑,
于是多个状态共享前向传播,而不是每次调用一次。结果按输入顺序返回,投影方式和 decide 完全一样,
而 return_details=True 会给每个状态一个 DecisionResult:
values = agent.decide_batch(ticket_texts, schema=Ticket) # values[i] matches ticket_texts[i]
results = router.decide_batch(states, schema=Ticket, return_details=True, batch_size=64)
在 Router 上,每个状态仍然各自路由,所以一次调用可以横跨多个 checkpoint。关键字参数会到达
predict_batch,所以 batch_size=、model= 和钩子都和 decide 一样能用。Agent、ONNXAgent
和 Router 都有它;没有 predict_batch 的 runner 会抛 TypeError,而不是静默回退成一个循环
—— 那种情况就逐状态调用 decide。批处理会像 predict_batch 那样让临界 argmax 发生偏移;README
记录了两种设备上实测的加速。
置信度和概率
默认 decide 只返回值。传 return_details=True 得到一个 DecisionResult,带逐字段的置信度、
概率、原始答案,以及这次调用的 usage 和路由:
result = agent.decide(state, schema=Ticket, return_details=True)
result.values["department"] # "billing"
result.answer_confidence["department"] # 0.94 max(p): the quantity min_confidence gates on
result.confidence["department"] # 0.71 normalized entropy, which depends on label count
result.probabilities["department"] # {"billing": 0.94, "support": 0.06, "sales": 0.0}
result.usage # {"input_tokens": 42, "output_tokens": 0}
result.routing # the Router decision, when a Router answered
confidence 和 answer_confidence 是两个不同的量,名字本身就说明了各自的定义。
answer_confidence 是 max(p),即报告出来的那个答案所分到的概率质量。它正是温度缩放所拟合的
东西,是本仓库里每一个校准数字所据以计算的东西,也是 min_confidence 所比较的对象 —— 这正是
该对它、而不是对 confidence 做门控的原因。confidence 是归一化熵,它取决于问题当时有多少个
选项:tests/test_confidence.py 固定了这一点 —— 一个两选项分布在一个 noul 上返回 0.90,在
一个等价的 choice 上返回 0.53 —— 所以它不能拿去和阈值比较。一个没有报告出可用
answer_confidence 的字段会映射到 None,这和报告出一个 0.0 不是一回事。
门控要和门控本身所用的量一致:
if result.answer_confidence["department"] < 0.6:
result.values["department"] = "human-review"
answer_confidence 适合用来筛选,这不等于它就是一个可信的概率。把它读作「在 c 处返回的答案里
约有 c 的比例是对的」,只有在温度已经拟合、并且针对该 checkpoint 和问题形状在留出数据上验证过
之后才成立。随包发布的 checkpoint 按发布状态看是过度自信的,而 laya-multilingual 干脆没有自带
任何已拟合的温度 —— 见 README 的
校准 和
诚实的限制 两节,以及
微调 notebook
里的拟合循环。在依赖这个级别之前先拟合;报告它,是因为门控和评估 harness 用的都是它。
它内部如何映射
- Enum 和
Literal变成choice问题,值作为字符串标签;标签在返回时映射回原来的值,所以整数 仍然是整数。 - 带界整数变成一个
score问题,每个值一个档位;返回的值是minimum + argmax。 - 布尔值变成一个
noul问题;值为noul >= 0.5。 description变成问题的指令,所以一个好的描述正是决策准确的原因。这遵循与钩子指南 相同的规则:把每个选项的含义说清楚。null分支在字段被规划之前就被丢掉,所以Optional[X]问的正是X问的问题。没有答案时,该 字段的键就直接从值里缺席,这正是能安全地把一个字段声明为可选、又不改变模型所见内容的原因。
另见
- 预测钩子:观察、塑造、缓存或门控它产生的这些决策。
- 决策原语:深入讲
choice、score和noul。 - LangChain 与 LangGraph:
LayaDecision就是这次调用作为一个 runnable 放进链里。