问题与答案
问题与答案
Laya 的一个问题就是一次类型化决策,类型同时决定了你问什么、拿回什么。共有三种,一个 state 可以 在一次前向传播里带上全部三种:
import laya
agent = laya.load("convaiinnovations/laya")
questions = {
"dept": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "money and invoices",
"technical": "bugs and outages",
"sales": "pricing and contracts"}},
"urgent": {"type": "noul", "instructions": "Is this urgent?"},
"severity": {"type": "score", "instructions": "How severe is this?",
"criteria": ["trivial", "minor", "moderate", "serious", "critical"]},
}
result = agent.system_one({"text": "I was charged twice and nobody has replied for a week. "
"Please refund me."}, questions)
result["answers"]["dept"]["choice"] # 'billing'
result["answers"]["urgent"]["noul"] # 0.8727
result["answers"]["severity"]["score"] # 2.9046
一次前向传播回答全部三种。这正是类型化接口的意义:一个 noul 问题不是碰巧带「是」和「否」两个
选项的 choice —— 它是不同的 head,输出形状也不同,类型告诉 Laya 该用哪个。
三种类型
choice —— 从一组里选一个
{"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {"billing": "money and invoices", "technical": "bugs and outages"}}
criteria 是一个从标签到描述的有序映射。顺序是有位置意义的:渲染出来的问题里,两组标签
相同但顺序不同的问题就是不同的问题。描述是可选的;{"billing": None} 只渲染标签本身。标签会按
你写的原样返回,所以非字符串标签在 choice 里还是它本身,在 probabilities 里则作为键。
描述值得写。它不是装饰:渲染出来的问题文本就是模型读到的内容,所以一个光秃秃的
{"a": None, "b": None} 给不了它任何区分选项的信息。
{"type": "choice", "choice": "billing",
"probabilities": {"billing": 0.9881, "technical": 0.0057, "sales": 0.0062},
"confidence": 0.9339, "answer_confidence": 0.9881,
"action": {"act_probability": 1.0}}
noul —— 是或否
{"type": "noul", "instructions": "Is this urgent?"}
noul 是这个项目给一个二值决策起的
名字,它的答案是为真的概率,不是一个砍过阈值的布尔值:
{"type": "noul", "noul": 0.8727, "confidence": 0.8727, "answer_confidence": 0.8727,
"action": {"act_probability": 1.0}}
阈值由你自己选,因为它取决于一次误报会付出什么代价。这里没有一个 bool 字段会让你误当成决策。
当决策本身不是天然的是/否题时,你可以给两个选项换标签 —— labels 只接受 false 和 true
这两个键,极性不变:noul 仍然是 P(true)。
{"type": "noul", "instructions": "Does this need a human?",
"labels": {"false": "automatic", "true": "escalate"}}
换标签不是修正一个模型做错的问题的办法。noul 会跟随它自己的选项标签,在 English checkpoint 上
尤其明显,所以一对读起来像个决策的标签(「approve」/「reject」)可能把答案拉向标签,而不是拉向
state。任何换标签都要先在你自己的数据上验证,再依赖它。
score —— 一个有序的档位
{"type": "score", "instructions": "How severe is this?",
"criteria": ["trivial", "minor", "moderate", "serious", "critical"]}
criteria 是一个有序列表,必须按升序 —— 位置就是刻度。
{"type": "score", "score": 2.9046,
"legend": {"0": "trivial", "1": "minor", "2": "moderate", "3": "serious", "4": "critical"},
"probabilities": {"0": 0.0134, "1": 0.05, "2": 0.0518, "3": 0.7881, "4": 0.0967},
"confidence": 0.5187, "answer_confidence": 0.7881,
"action": {"act_probability": 1.0}}
score 是期望值,不是最可能的档位。 上面那个例子里,score 是 2.90,而单独最可能的档位
是 0.788 处的 serious(3)。两者都有用,回答的是不同的问题:期望值在整条刻度上最小化平方误差,
argmax 最小化与模型的分歧。如果你想要标签,就对 probabilities 取 argmax,或者从档位说明里读
answer_confidence 对应的那一项 —— 不要把 score 四舍五入,然后以为它就是标签。legend 的
存在就是为了让你永远不必猜哪个下标是什么意思。
读取置信度
每个答案都带两个置信度数字,它们衡量的东西不同。
| 字段 | 是什么 | 用它做门控? |
|---|---|---|
answer_confidence |
max(p) —— 所报告答案的概率 |
是的,但要在拟合之后 |
confidence |
1 - H(p) / log(k) —— 整个分布有多集中 |
否 |
probabilities |
完整的分布(choice、score) |
— |
answer_confidence 正是温度缩放所拟合的量,也是仓库里每张校准图都据以计算的量 —— 这才是拿它做
门控的原因。它按出厂状态并未校准:通常归到它头上的那条性质 —— 以置信度 c 返回的那些答案里,
约 c 的比例是对的 —— 只有在温度拟合完毕、并在你的 checkpoint、你的选项数下的留出数据上
验证过之后才成立。出厂的 checkpoint 是过度自信的,过度多少取决于选项数,所以一个没调过的阈值
可能把低于模型自身准确率的答案选进来(#394)。
# THRESHOLD is a number you measured on your own held-out data, not one the model ships.
# Fit and validate the temperatures first — the fine-tuning notebook has the loop:
# notebooks/laya_finetune_typed_decisions_2xT4_kaggle.ipynb
ans = result["answers"]["dept"]
if ans["answer_confidence"] >= THRESHOLD:
...
confidence 是归一化熵:分布尖锐时高,分散时低,与排第一的答案对不对无关。它是个有用的信号,
但它不在同一量纲上,所以两者不能用同一个数字来门控:
# the same three answers, and the two numbers are not the same
dept confidence 0.9339 answer_confidence 0.9881
urgent confidence 0.8727 answer_confidence 0.8727
severity confidence 0.5187 answer_confidence 0.7881
对 noul 而言,两者按构造就是相等的 —— 两个选项下,max(p, 1-p) 就是 max(p) —— 所以一个
noul 答案本身分不出你读的是哪一个。每个类型上两个键都存在,这样这个选择是显式的,而不是隐含的。
阈值是一项策略,不是模型的属性。 两个 checkpoint 出厂都是过度自信的,过度多少取决于选项数,
所以在一个 3 选项问题上量出的数字不能搬到 20 选项问题上。在你自己的数据上量它;BENCHMARKS.md
的校准一节里有拟合循环和拟合出的值。
action 与 act_probability
action.act_probability 是另一个 head 的分数,回答「agent 到底该不该据此行动」,与答案自身的
置信度不同。每种问题类型都会报告它。库里没有任何东西替你给它设阈值。
预设
三套现成的问题集,常见情形不必手写判定标准:
from laya import triage_questions, guard_questions, moderation_questions
agent.system_one(ticket, triage_questions())
把它们当起点,而不是当契约 —— 用 render_options 读出它们产生的问题,并在发布之前检查标签是否
适合你的领域。
把选项读回来
因为选项顺序有位置意义,而选项文本正是模型读到的内容,所以能确切看到发出去的是什么,是值得的:
from laya import render_options
render_options({"t": "choice", "crit": {"billing": None, "sales": "pricing"}})
# ['billing', 'sales: pricing']
render_options({"t": "score", "crit": ["low", "high"]})
# ['level 0: low', 'level 1: high']
同样的标签换个顺序,就按那个顺序渲染,这正是「顺序是问题身份的一部分」的原因:
render_options({"t": "choice", "crit": {"x": "first", "y": "second"}})
# ['x: first', 'y: second']
render_options({"t": "choice", "crit": {"y": "second", "x": "first"}})
# ['y: second', 'x: first']
注意键名。 render_options 接受的是内部的短键形式 {"t": ..., "crit": ...},不是你在
问题里写的 {"type": ..., "criteria": ...} 形式 —— 传公开的形状会抛 KeyError: 't'。
这个转换是一段很短、很稳定的映射,你可以直接内联,从而避免去碰一个私有辅助函数 ——
Agent._to_internal 是内部的,可能会变:
def as_internal(q):
"""The short-key shape `render_options` reads, from a question as you wrote it."""
crit = q.get("criteria")
if q["type"] == "choice" and isinstance(crit, list):
crit = {c: None for c in crit}
return {"t": q["type"], "ins": q["instructions"], "crit": crit}
render_options(as_internal(question))
这复刻了库对一个写成标签列表的 choice 问题所做的处理;一个 criteria dict 和一个 score
列表则原样通过。