問題與答案
Laya 的一個問題就是一次型別化決策,型別同時決定了你問什麼、拿回什麼。共有三種,一個 state 可以 在一次前向傳播裡帶上全部三種:
import laya
agent = laya.load("convaiinnovations/laya")
questions = {
"dept": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "money and invoices",
"technical": "bugs and outages",
"sales": "pricing and contracts"}},
"urgent": {"type": "noul", "instructions": "Is this urgent?"},
"severity": {"type": "score", "instructions": "How severe is this?",
"criteria": ["trivial", "minor", "moderate", "serious", "critical"]},
}
result = agent.system_one({"text": "I was charged twice and nobody has replied for a week. "
"Please refund me."}, questions)
result["answers"]["dept"]["choice"] # 'billing'
result["answers"]["urgent"]["noul"] # 0.8727
result["answers"]["severity"]["score"] # 2.9046
一次前向傳播回答全部三種。這正是型別化介面的意義:一個 noul 問題不是碰巧帶「是」和「否」兩個
選項的 choice —— 它是不同的 head,輸出形狀也不同,型別告訴 Laya 該用哪個。
三種類型
choice —— 從一組裡選一個
{"type": "choice",
"instructions": "Which team should handle this?",
"criteria": {"billing": "money and invoices", "technical": "bugs and outages"}}
criteria 是一個從標籤到描述的有序對映。順序是有位置意義的:渲染出來的問題裡,兩組標籤
相同但順序不同的問題就是不同的問題。描述是可選的;{"billing": None} 只渲染標籤本身。標籤會按
你寫的原樣返回,所以非字串標籤在 choice 裡還是它本身,在 probabilities 裡則作為鍵。
描述值得寫。它不是裝飾:渲染出來的問題文本就是模型讀到的內容,所以一個光禿禿的
{"a": None, "b": None} 給不了它任何區分選項的資訊。
{"type": "choice", "choice": "billing",
"probabilities": {"billing": 0.9881, "technical": 0.0057, "sales": 0.0062},
"confidence": 0.9339, "answer_confidence": 0.9881,
"action": {"act_probability": 1.0}}
noul —— 是或否
{"type": "noul", "instructions": "Is this urgent?"}
noul 是這個專案給一個二值決策起的
名字,它的答案是為真的機率,不是一個砍過閾值的布林值:
{"type": "noul", "noul": 0.8727, "confidence": 0.8727, "answer_confidence": 0.8727,
"action": {"act_probability": 1.0}}
閾值由你自己選,因為它取決於一次誤報會付出什麼代價。這裡沒有一個 bool 欄位會讓你誤當成決策。
當決策本身不是天然的是/否題時,你可以給兩個選項換標籤 —— labels 只接受 false 和 true
這兩個鍵,極性不變:noul 仍然是 P(true)。
{"type": "noul", "instructions": "Does this need a human?",
"labels": {"false": "automatic", "true": "escalate"}}
換標籤不是修正一個模型做錯的問題的辦法。noul 會跟隨它自己的選項標籤,在 English checkpoint 上
尤其明顯,所以一對讀起來像個決策的標籤(「approve」/「reject」)可能把答案拉向標籤,而不是拉向
state。任何換標籤都要先在你自己的資料上驗證,再依賴它。
score —— 一個有序的檔位
{"type": "score", "instructions": "How severe is this?",
"criteria": ["trivial", "minor", "moderate", "serious", "critical"]}
criteria 是一個有序列表,必須按升序 —— 位置就是刻度。
{"type": "score", "score": 2.9046,
"legend": {"0": "trivial", "1": "minor", "2": "moderate", "3": "serious", "4": "critical"},
"probabilities": {"0": 0.0134, "1": 0.05, "2": 0.0518, "3": 0.7881, "4": 0.0967},
"confidence": 0.5187, "answer_confidence": 0.7881,
"action": {"act_probability": 1.0}}
score 是期望值,不是最可能的檔位。 上面那個例子裡,score 是 2.90,而單獨最可能的檔位
是 0.788 處的 serious(3)。兩者都有用,回答的是不同的問題:期望值在整條刻度上最小化平方誤差,
argmax 最小化與模型的分歧。如果你想要標籤,就對 probabilities 取 argmax,或者從檔位說明裡讀
answer_confidence 對應的那一項 —— 不要把 score 四捨五入,然後以為它就是標籤。legend 的
存在就是為了讓你永遠不必猜哪個下標是什麼意思。
讀取置信度
每個答案都帶兩個置信度數字,它們衡量的東西不同。
| 欄位 | 是什麼 | 用它做門控? |
|---|---|---|
answer_confidence |
max(p) —— 所報告答案的機率 |
是的,但要在擬合之後 |
confidence |
1 - H(p) / log(k) —— 整個分佈有多集中 |
否 |
probabilities |
完整的分佈(choice、score) |
— |
answer_confidence 正是溫度縮放所擬合的量,也是倉庫裡每張校準圖都據以計算的量 —— 這才是拿它做
門控的原因。它按出廠狀態並未校準:通常歸到它頭上的那條性質 —— 以置信度 c 返回的那些答案裡,
約 c 的比例是對的 —— 只有在溫度擬合完畢、並在你的 checkpoint、你的選項數下的留出資料上
驗證過之後才成立。出廠的 checkpoint 是過度自信的,過度多少取決於選項數,所以一個沒調過的閾值
可能把低於模型自身準確率的答案選進來(#394)。
# THRESHOLD is a number you measured on your own held-out data, not one the model ships.
# Fit and validate the temperatures first — the fine-tuning notebook has the loop:
# notebooks/laya_finetune_typed_decisions_2xT4_kaggle.ipynb
ans = result["answers"]["dept"]
if ans["answer_confidence"] >= THRESHOLD:
...
confidence 是歸一化熵:分佈尖銳時高,分散時低,與排第一的答案對不對無關。它是個有用的訊號,
但它不在同一量綱上,所以兩者不能用同一個數字來門控:
# the same three answers, and the two numbers are not the same
dept confidence 0.9339 answer_confidence 0.9881
urgent confidence 0.8727 answer_confidence 0.8727
severity confidence 0.5187 answer_confidence 0.7881
對 noul 而言,兩者按構造就是相等的 —— 兩個選項下,max(p, 1-p) 就是 max(p) —— 所以一個
noul 答案本身分不出你讀的是哪一個。每個型別上兩個鍵都存在,這樣這個選擇是顯式的,而不是隱含的。
閾值是一項策略,不是模型的屬性。 兩個 checkpoint 出廠都是過度自信的,過度多少取決於選項數,
所以在一個 3 選項問題上量出的數字不能搬到 20 選項問題上。在你自己的資料上量它;BENCHMARKS.md
的校準一節裡有擬合迴圈和擬合出的值。
action 與 act_probability
action.act_probability 是另一個 head 的分數,回答「agent 到底該不該據此行動」,與答案自身的
置信度不同。每種問題型別都會報告它。庫裡沒有任何東西替你給它設閾值。
預設
三套現成的問題集,常見情形不必手寫判定標準:
from laya import triage_questions, guard_questions, moderation_questions
agent.system_one(ticket, triage_questions())
把它們當起點,而不是當契約 —— 用 render_options 讀出它們產生的問題,並在釋出之前檢查標籤是否
適合你的領域。
把選項讀回來
因為選項順序有位置意義,而選項文本正是模型讀到的內容,所以能確切看到發出去的是什麼,是值得的:
from laya import render_options
render_options({"t": "choice", "crit": {"billing": None, "sales": "pricing"}})
# ['billing', 'sales: pricing']
render_options({"t": "score", "crit": ["low", "high"]})
# ['level 0: low', 'level 1: high']
同樣的標籤換個順序,就按那個順序渲染,這正是「順序是問題身份的一部分」的原因:
render_options({"t": "choice", "crit": {"x": "first", "y": "second"}})
# ['x: first', 'y: second']
render_options({"t": "choice", "crit": {"y": "second", "x": "first"}})
# ['y: second', 'x: first']
注意鍵名。 render_options 接受的是內部的短鍵形式 {"t": ..., "crit": ...},不是你在
問題裡寫的 {"type": ..., "criteria": ...} 形式 —— 傳公開的形狀會拋 KeyError: 't'。
這個轉換是一段很短、很穩定的對映,你可以直接內聯,從而避免去碰一個私有輔助函式 ——
Agent._to_internal 是內部的,可能會變:
def as_internal(q):
"""The short-key shape `render_options` reads, from a question as you wrote it."""
crit = q.get("criteria")
if q["type"] == "choice" and isinstance(crit, list):
crit = {c: None for c in crit}
return {"t": q["type"], "ins": q["instructions"], "crit": crit}
render_options(as_internal(question))
這復刻了庫對一個寫成標籤列表的 choice 問題所做的處理;一個 criteria dict 和一個 score
列表則原樣通過。