文档导航

Schema 驱动的决策

Schema 驱动的决策

把一份 JSON schema 或一个 pydantic 模型变成 Laya 问题,拿回带校准置信度的类型化值。正是这座桥 让 Laya 成为一个结构化输出引擎:你描述想要的形状,Laya 在一次前向传播里给出答案。

import laya

schema = {
    "type": "object",
    "properties": {
        "department": {"type": "string", "enum": ["billing", "support", "sales"],
                       "description": "Which team should handle this?"},
        "urgency": {"type": "integer", "minimum": 0, "maximum": 2},
        "needs_human": {"type": "boolean"},
    },
}

agent = laya.load("convaiinnovations/laya")
values = agent.decide("I was charged twice, refund me.", schema=schema)
# {"department": "billing", "urgency": 2, "needs_human": True}

用 pydantic(安装 laya[structured]):

from typing import Literal
from pydantic import BaseModel

class Ticket(BaseModel):
    department: Literal["billing", "support", "sales"]
    urgency: Literal[0, 1, 2]
    needs_human: bool

ticket = agent.decide("I was charged twice, refund me.", schema=Ticket)

支持的子集

顶层必须是一个带 properties 的对象。每个属性变成一个 Laya 问题。

下面每一行都是真实的 schema:tests/test_structured_docs.py 编译第一列,并断言编译器实际产出的 问题,所以这张表不会和代码脱节。一个单元格要么是单独的属性 schema,要么是对某个入口点的调用。

JSON schema Laya 问题 返回的值
{"enum": ["billing", "support"]} choice 被选中的值,保留它原本的类型
{"const": "billing"} choice 那一个值
{"type": "boolean"} noul true / false
{"type": "integer", "minimum": 0, "maximum": 5} score 概率最高的档位,作为一个整数
{"type": "number", "minimum": 0, "maximum": 5} score 那个档位,作为一个整数
{"type": "string", "enum": ["low", "high"], "description": "How urgent?"} choice How urgent? 就是问题措辞
{"anyOf": [{"enum": ["x", "y"]}, {"type": "null"}]} choice 和普通的 enum 行一样;没有答案时这个键不出现
{"oneOf": [{"type": "boolean"}, {"type": "null"}]} noul 和普通的 boolean 行一样
{"type": ["integer", "null"], "minimum": 1, "maximum": 3} score 和普通的带界整数行一样

Literal[...] 和 Optional[...] 是 pydantic 里对应 enum 和 anyOf 两行的写法: questions_from_pydantic 把它们渲染成那些形状,同样的行适用。

title 不会被读取。pydantic 不管你有没有要求,都会给 model_json_schema() 的每个字段放上 一个,而一个按属性的名字没法给一个问题所据以构建的那些逐选项 choice 打标签,所以措辞的杠杆是 description —— 见下面的它内部如何映射。

投影是精确的:enum: [1, 2, 3] 返回 2,不是 "2";带界整数返回 minimum 和 maximum 之间的一个档位;布尔值就是 noul >= 0.5。

拒绝

一个无法从固定选项集作答的 schema 会抛出 laya.structured.SchemaError(一个 ValueError), 并点出确切的路径。每一行也都会被执行,字段名取 name:

属性 schema 消息
{"type": "string"} properties.name: a free string cannot be a fixed option set; use 'enum' or a boolean
{"type": "array", "items": {"type": "string"}} properties.name: arrays are not supported; ask one field per element
{"type": "object", "properties": {"inner": {"type": "boolean"}}} properties.name: nested objects are not supported; flatten the schema
{"$ref": "#/definitions/node"} properties.name: $ref/recursion is not supported; flatten the schema
{"enum": [1, "1"]} properties.name: enum values produce duplicate choice labels
{"enum": []} properties.name: 'enum' must not be empty
{"type": "number"} properties.name: a numeric field needs integer 'minimum' and 'maximum' to become a score
{"type": "integer", "minimum": 5, "maximum": 2} properties.name: 'maximum' 2 is below 'minimum' 5
{"type": "integer", "minimum": 0, "maximum": 10} properties.name: 11 levels exceeds MAX_SCORE_LEVELS=10; narrow the range or use an enum
{"anyOf": [{"type": "string"}, {"type": "integer"}]} properties.name: only 'Optional[...]' unions (one non-null branch) are supported, got 2
{"type": ["string", "integer"]} properties.name: 'type' has multiple non-null types; unions are not supported
{"format": "date"} properties.name: unsupported schema {'format': 'date'}
"boolean" properties.name: property must be an object, got str

入口点自身会拒绝这些:

调用 消息
plan_from_json_schema("not a schema") expected a JSON schema object, got str
plan_from_json_schema({"type": "object"}) the top level must be an object with 'properties'
plan_from_json_schema({"type": "object", "properties": {}}) 'properties' must be a non-empty object
plan_from_json_schema({"type": "object", "properties": {"p%d" % i: {"type": "boolean"} for i in range(33)}}) 33 properties exceeds MAX_PROPERTIES=32
plan_from_json_schema({"type": "object", "properties": {"name": {"enum": ["v%d" % i for i in range(33)]}}}) properties.name: 33 options exceeds MAX_OPTIONS=32
decide(None, "I was charged twice.", schema=42) expected a JSON schema dict or a pydantic model, got int

限制:MAX_PROPERTIES = 32、MAX_OPTIONS = 32、MAX_SCORE_LEVELS = 10。

API

函数 用途
laya.decide(runner, state, schema=..., *, questions=..., return_details=..., min_confidence=..., **predict_kwargs) 自由函数,对 Agent 和 Router 都适用
agent.decide(state, schema=..., ...) / router.decide(state, schema=..., ...) 便捷方法
laya.decide_batch(runner, states, schema=..., ...) / agent.decide_batch(...) / router.decide_batch(...) 在多个状态上做同样的事,一次批调用
questions_from_json_schema(schema) schema 转成 Laya 问题
questions_from_pydantic(model) pydantic 模型转成问题(需要 pydantic)
answers_to_json(answers, schema) 把原始答案投影到 schema 的值上
answer_to_pydantic(model, answers) 把原始答案投影成一个 pydantic 实例
plan_from_json_schema(schema) 经过校验的字段计划(进阶)

schema 和 questions 只能传其中一个。传 questions 时,decide 返回原始答案而不是做投影。 额外的关键字参数会转发给 predict,所以钩子、model=、task= 和 token 预算都能用:

router.decide(state, schema=Ticket, model="multilingual", hooks=[Metrics()])

给多个状态打分

decide_batch 是吞吐的形式:schema 只规划一次,它的问题通过 predict_batch 在每一个状态上跑, 于是多个状态共享前向传播,而不是每次调用一次。结果按输入顺序返回,投影方式和 decide 完全一样, 而 return_details=True 会给每个状态一个 DecisionResult:

values = agent.decide_batch(ticket_texts, schema=Ticket)          # values[i] matches ticket_texts[i]
results = router.decide_batch(states, schema=Ticket, return_details=True, batch_size=64)

在 Router 上,每个状态仍然各自路由,所以一次调用可以横跨多个 checkpoint。关键字参数会到达 predict_batch,所以 batch_size=、model= 和钩子都和 decide 一样能用。Agent、ONNXAgent 和 Router 都有它;没有 predict_batch 的 runner 会抛 TypeError,而不是静默回退成一个循环 —— 那种情况就逐状态调用 decide。批处理会像 predict_batch 那样让临界 argmax 发生偏移;README 记录了两种设备上实测的加速。

置信度和概率

默认 decide 只返回值。传 return_details=True 得到一个 DecisionResult,带逐字段的置信度、 概率、原始答案,以及这次调用的 usage 和路由:

result = agent.decide(state, schema=Ticket, return_details=True)
result.values["department"]            # "billing"
result.answer_confidence["department"] # 0.94  max(p): the quantity min_confidence gates on
result.confidence["department"]        # 0.71  normalized entropy, which depends on label count
result.probabilities["department"]     # {"billing": 0.94, "support": 0.06, "sales": 0.0}
result.usage                           # {"input_tokens": 42, "output_tokens": 0}
result.routing                         # the Router decision, when a Router answered

confidence 和 answer_confidence 是两个不同的量,名字本身就说明了各自的定义。 answer_confidence 是 max(p),即报告出来的那个答案所分到的概率质量。它正是温度缩放所拟合的 东西,是本仓库里每一个校准数字所据以计算的东西,也是 min_confidence 所比较的对象 —— 这正是 该对它、而不是对 confidence 做门控的原因。confidence 是归一化熵,它取决于问题当时有多少个 选项:tests/test_confidence.py 固定了这一点 —— 一个两选项分布在一个 noul 上返回 0.90,在 一个等价的 choice 上返回 0.53 —— 所以它不能拿去和阈值比较。一个没有报告出可用 answer_confidence 的字段会映射到 None,这和报告出一个 0.0 不是一回事。

门控要和门控本身所用的量一致:

if result.answer_confidence["department"] < 0.6:
    result.values["department"] = "human-review"

answer_confidence 适合用来筛选,这不等于它就是一个可信的概率。把它读作「在 c 处返回的答案里 约有 c 的比例是对的」,只有在温度已经拟合、并且针对该 checkpoint 和问题形状在留出数据上验证过 之后才成立。随包发布的 checkpoint 按发布状态看是过度自信的,而 laya-multilingual 干脆没有自带 任何已拟合的温度 —— 见 README 的 校准 和 诚实的限制 两节,以及 微调 notebook 里的拟合循环。在依赖这个级别之前先拟合;报告它,是因为门控和评估 harness 用的都是它。

它内部如何映射

  • Enum 和 Literal 变成 choice 问题,值作为字符串标签;标签在返回时映射回原来的值,所以整数 仍然是整数。
  • 带界整数变成一个 score 问题,每个值一个档位;返回的值是 minimum + argmax。
  • 布尔值变成一个 noul 问题;值为 noul >= 0.5。
  • description 变成问题的指令,所以一个好的描述正是决策准确的原因。这遵循与钩子指南 相同的规则:把每个选项的含义说清楚。
  • null 分支在字段被规划之前就被丢掉,所以 Optional[X] 问的正是 X 问的问题。没有答案时,该 字段的键就直接从值里缺席,这正是能安全地把一个字段声明为可选、又不改变模型所见内容的原因。

另见