文档导航

问题与答案

问题与答案

Laya 的一个问题就是一次类型化决策,类型同时决定了你问什么、拿回什么。共有三种,一个 state 可以 在一次前向传播里带上全部三种:

import laya

agent = laya.load("convaiinnovations/laya")

questions = {
    "dept": {"type": "choice", "instructions": "Which team should handle this?",
             "criteria": {"billing": "money and invoices",
                          "technical": "bugs and outages",
                          "sales": "pricing and contracts"}},
    "urgent": {"type": "noul", "instructions": "Is this urgent?"},
    "severity": {"type": "score", "instructions": "How severe is this?",
                 "criteria": ["trivial", "minor", "moderate", "serious", "critical"]},
}

result = agent.system_one({"text": "I was charged twice and nobody has replied for a week. "
                                   "Please refund me."}, questions)
result["answers"]["dept"]["choice"]        # 'billing'
result["answers"]["urgent"]["noul"]        # 0.8727
result["answers"]["severity"]["score"]     # 2.9046

一次前向传播回答全部三种。这正是类型化接口的意义:一个 noul 问题不是碰巧带「是」和「否」两个 选项的 choice —— 它是不同的 head,输出形状也不同,类型告诉 Laya 该用哪个。

三种类型

choice —— 从一组里选一个

{"type": "choice",
 "instructions": "Which team should handle this?",
 "criteria": {"billing": "money and invoices", "technical": "bugs and outages"}}

criteria 是一个从标签到描述的有序映射。顺序是有位置意义的:渲染出来的问题里,两组标签 相同但顺序不同的问题就是不同的问题。描述是可选的;{"billing": None} 只渲染标签本身。标签会按 你写的原样返回,所以非字符串标签在 choice 里还是它本身,在 probabilities 里则作为键。

描述值得写。它不是装饰:渲染出来的问题文本就是模型读到的内容,所以一个光秃秃的 {"a": None, "b": None} 给不了它任何区分选项的信息。

{"type": "choice", "choice": "billing",
 "probabilities": {"billing": 0.9881, "technical": 0.0057, "sales": 0.0062},
 "confidence": 0.9339, "answer_confidence": 0.9881,
 "action": {"act_probability": 1.0}}

noul —— 是或否

{"type": "noul", "instructions": "Is this urgent?"}

noul 是这个项目给一个二值决策起的 名字,它的答案是为真的概率,不是一个砍过阈值的布尔值:

{"type": "noul", "noul": 0.8727, "confidence": 0.8727, "answer_confidence": 0.8727,
 "action": {"act_probability": 1.0}}

阈值由你自己选,因为它取决于一次误报会付出什么代价。这里没有一个 bool 字段会让你误当成决策。

当决策本身不是天然的是/否题时,你可以给两个选项换标签 —— labels 只接受 false 和 true 这两个键,极性不变:noul 仍然是 P(true)。

{"type": "noul", "instructions": "Does this need a human?",
 "labels": {"false": "automatic", "true": "escalate"}}

换标签不是修正一个模型做错的问题的办法。noul 会跟随它自己的选项标签,在 English checkpoint 上 尤其明显,所以一对读起来像个决策的标签(「approve」/「reject」)可能把答案拉向标签,而不是拉向 state。任何换标签都要先在你自己的数据上验证,再依赖它。

score —— 一个有序的档位

{"type": "score", "instructions": "How severe is this?",
 "criteria": ["trivial", "minor", "moderate", "serious", "critical"]}

criteria 是一个有序列表,必须按升序 —— 位置就是刻度。

{"type": "score", "score": 2.9046,
 "legend": {"0": "trivial", "1": "minor", "2": "moderate", "3": "serious", "4": "critical"},
 "probabilities": {"0": 0.0134, "1": 0.05, "2": 0.0518, "3": 0.7881, "4": 0.0967},
 "confidence": 0.5187, "answer_confidence": 0.7881,
 "action": {"act_probability": 1.0}}

score 是期望值,不是最可能的档位。 上面那个例子里,score 是 2.90,而单独最可能的档位 是 0.788 处的 serious(3)。两者都有用,回答的是不同的问题:期望值在整条刻度上最小化平方误差, argmax 最小化与模型的分歧。如果你想要标签,就对 probabilities 取 argmax,或者从档位说明里读 answer_confidence 对应的那一项 —— 不要把 score 四舍五入,然后以为它就是标签。legend 的 存在就是为了让你永远不必猜哪个下标是什么意思。

读取置信度

每个答案都带两个置信度数字,它们衡量的东西不同。

字段 是什么 用它做门控?
answer_confidence max(p) —— 所报告答案的概率 是的,但要在拟合之后
confidence 1 - H(p) / log(k) —— 整个分布有多集中 否
probabilities 完整的分布(choice、score) —

answer_confidence 正是温度缩放所拟合的量,也是仓库里每张校准图都据以计算的量 —— 这才是拿它做 门控的原因。它按出厂状态并未校准:通常归到它头上的那条性质 —— 以置信度 c 返回的那些答案里, 约 c 的比例是对的 —— 只有在温度拟合完毕、并在你的 checkpoint、你的选项数下的留出数据上 验证过之后才成立。出厂的 checkpoint 是过度自信的,过度多少取决于选项数,所以一个没调过的阈值 可能把低于模型自身准确率的答案选进来(#394)。

# THRESHOLD is a number you measured on your own held-out data, not one the model ships.
# Fit and validate the temperatures first — the fine-tuning notebook has the loop:
#   notebooks/laya_finetune_typed_decisions_2xT4_kaggle.ipynb
ans = result["answers"]["dept"]
if ans["answer_confidence"] >= THRESHOLD:
    ...

confidence 是归一化熵:分布尖锐时高,分散时低,与排第一的答案对不对无关。它是个有用的信号, 但它不在同一量纲上,所以两者不能用同一个数字来门控:

# the same three answers, and the two numbers are not the same
dept      confidence 0.9339   answer_confidence 0.9881
urgent    confidence 0.8727   answer_confidence 0.8727
severity  confidence 0.5187   answer_confidence 0.7881

对 noul 而言,两者按构造就是相等的 —— 两个选项下,max(p, 1-p) 就是 max(p) —— 所以一个 noul 答案本身分不出你读的是哪一个。每个类型上两个键都存在,这样这个选择是显式的,而不是隐含的。

阈值是一项策略,不是模型的属性。 两个 checkpoint 出厂都是过度自信的,过度多少取决于选项数, 所以在一个 3 选项问题上量出的数字不能搬到 20 选项问题上。在你自己的数据上量它;BENCHMARKS.md 的校准一节里有拟合循环和拟合出的值。

action 与 act_probability

action.act_probability 是另一个 head 的分数,回答「agent 到底该不该据此行动」,与答案自身的 置信度不同。每种问题类型都会报告它。库里没有任何东西替你给它设阈值。

预设

三套现成的问题集,常见情形不必手写判定标准:

from laya import triage_questions, guard_questions, moderation_questions

agent.system_one(ticket, triage_questions())

把它们当起点,而不是当契约 —— 用 render_options 读出它们产生的问题,并在发布之前检查标签是否 适合你的领域。

把选项读回来

因为选项顺序有位置意义,而选项文本正是模型读到的内容,所以能确切看到发出去的是什么,是值得的:

from laya import render_options

render_options({"t": "choice", "crit": {"billing": None, "sales": "pricing"}})
# ['billing', 'sales: pricing']

render_options({"t": "score", "crit": ["low", "high"]})
# ['level 0: low', 'level 1: high']

同样的标签换个顺序,就按那个顺序渲染,这正是「顺序是问题身份的一部分」的原因:

render_options({"t": "choice", "crit": {"x": "first", "y": "second"}})
# ['x: first', 'y: second']
render_options({"t": "choice", "crit": {"y": "second", "x": "first"}})
# ['y: second', 'x: first']

注意键名。 render_options 接受的是内部的短键形式 {"t": ..., "crit": ...},不是你在 问题里写的 {"type": ..., "criteria": ...} 形式 —— 传公开的形状会抛 KeyError: 't'。

这个转换是一段很短、很稳定的映射,你可以直接内联,从而避免去碰一个私有辅助函数 —— Agent._to_internal 是内部的,可能会变:

def as_internal(q):
    """The short-key shape `render_options` reads, from a question as you wrote it."""
    crit = q.get("criteria")
    if q["type"] == "choice" and isinstance(crit, list):
        crit = {c: None for c in crit}
    return {"t": q["type"], "ins": q["instructions"], "crit": crit}

render_options(as_internal(question))

这复刻了库对一个写成标签列表的 choice 问题所做的处理;一个 criteria dict 和一个 score 列表则原样通过。

在设计方案之前值得知道的限制

  • 选项多的时候,置信度不是正确性的保证。 在一个 20 选项问题上,正确答案和错误答案的分布 大量重叠,阈值可能最终选出低于模型自身准确率的结果。见 #394。
  • 强制选择式问题里的否定不能可靠处理。 一个关于「取消」的问题,可能对一个写着不要取消的 state 返回取消标签,而且置信度很高,两个 checkpoint 都是如此。见 #377。
  • noul 可能跟随它的标签,而不是 state,所以任何换标签都要验证。
  • 选项多不是免费的。 超过大约 20 个之后模型会很快退化;用 shortlist 辅助函数在提问之前把 大的标签空间缩小,或者把它拆成一个粗问题和一个细问题。