文件導航

問題與答案

Laya 的一個問題就是一次型別化決策,型別同時決定了你問什麼、拿回什麼。共有三種,一個 state 可以 在一次前向傳播裡帶上全部三種:

import laya

agent = laya.load("convaiinnovations/laya")

questions = {
    "dept": {"type": "choice", "instructions": "Which team should handle this?",
             "criteria": {"billing": "money and invoices",
                          "technical": "bugs and outages",
                          "sales": "pricing and contracts"}},
    "urgent": {"type": "noul", "instructions": "Is this urgent?"},
    "severity": {"type": "score", "instructions": "How severe is this?",
                 "criteria": ["trivial", "minor", "moderate", "serious", "critical"]},
}

result = agent.system_one({"text": "I was charged twice and nobody has replied for a week. "
                                   "Please refund me."}, questions)
result["answers"]["dept"]["choice"]        # 'billing'
result["answers"]["urgent"]["noul"]        # 0.8727
result["answers"]["severity"]["score"]     # 2.9046

一次前向傳播回答全部三種。這正是型別化介面的意義:一個 noul 問題不是碰巧帶「是」和「否」兩個 選項的 choice —— 它是不同的 head,輸出形狀也不同,型別告訴 Laya 該用哪個。

三種類型

choice —— 從一組裡選一個

{"type": "choice",
 "instructions": "Which team should handle this?",
 "criteria": {"billing": "money and invoices", "technical": "bugs and outages"}}

criteria 是一個從標籤到描述的有序對映。順序是有位置意義的:渲染出來的問題裡,兩組標籤 相同但順序不同的問題就是不同的問題。描述是可選的;{"billing": None} 只渲染標籤本身。標籤會按 你寫的原樣返回,所以非字串標籤在 choice 裡還是它本身,在 probabilities 裡則作為鍵。

描述值得寫。它不是裝飾:渲染出來的問題文本就是模型讀到的內容,所以一個光禿禿的 {"a": None, "b": None} 給不了它任何區分選項的資訊。

{"type": "choice", "choice": "billing",
 "probabilities": {"billing": 0.9881, "technical": 0.0057, "sales": 0.0062},
 "confidence": 0.9339, "answer_confidence": 0.9881,
 "action": {"act_probability": 1.0}}

noul —— 是或否

{"type": "noul", "instructions": "Is this urgent?"}

noul 是這個專案給一個二值決策起的 名字,它的答案是為真的機率,不是一個砍過閾值的布林值:

{"type": "noul", "noul": 0.8727, "confidence": 0.8727, "answer_confidence": 0.8727,
 "action": {"act_probability": 1.0}}

閾值由你自己選,因為它取決於一次誤報會付出什麼代價。這裡沒有一個 bool 欄位會讓你誤當成決策。

當決策本身不是天然的是/否題時,你可以給兩個選項換標籤 —— labels 只接受 false 和 true 這兩個鍵,極性不變:noul 仍然是 P(true)。

{"type": "noul", "instructions": "Does this need a human?",
 "labels": {"false": "automatic", "true": "escalate"}}

換標籤不是修正一個模型做錯的問題的辦法。noul 會跟隨它自己的選項標籤,在 English checkpoint 上 尤其明顯,所以一對讀起來像個決策的標籤(「approve」/「reject」)可能把答案拉向標籤,而不是拉向 state。任何換標籤都要先在你自己的資料上驗證,再依賴它。

score —— 一個有序的檔位

{"type": "score", "instructions": "How severe is this?",
 "criteria": ["trivial", "minor", "moderate", "serious", "critical"]}

criteria 是一個有序列表,必須按升序 —— 位置就是刻度。

{"type": "score", "score": 2.9046,
 "legend": {"0": "trivial", "1": "minor", "2": "moderate", "3": "serious", "4": "critical"},
 "probabilities": {"0": 0.0134, "1": 0.05, "2": 0.0518, "3": 0.7881, "4": 0.0967},
 "confidence": 0.5187, "answer_confidence": 0.7881,
 "action": {"act_probability": 1.0}}

score 是期望值,不是最可能的檔位。 上面那個例子裡,score 是 2.90,而單獨最可能的檔位 是 0.788 處的 serious(3)。兩者都有用,回答的是不同的問題:期望值在整條刻度上最小化平方誤差, argmax 最小化與模型的分歧。如果你想要標籤,就對 probabilities 取 argmax,或者從檔位說明裡讀 answer_confidence 對應的那一項 —— 不要把 score 四捨五入,然後以為它就是標籤。legend 的 存在就是為了讓你永遠不必猜哪個下標是什麼意思。

讀取置信度

每個答案都帶兩個置信度數字,它們衡量的東西不同。

欄位 是什麼 用它做門控?
answer_confidence max(p) —— 所報告答案的機率 是的,但要在擬合之後
confidence 1 - H(p) / log(k) —— 整個分佈有多集中 否
probabilities 完整的分佈(choice、score) —

answer_confidence 正是溫度縮放所擬合的量,也是倉庫裡每張校準圖都據以計算的量 —— 這才是拿它做 門控的原因。它按出廠狀態並未校準:通常歸到它頭上的那條性質 —— 以置信度 c 返回的那些答案裡, 約 c 的比例是對的 —— 只有在溫度擬合完畢、並在你的 checkpoint、你的選項數下的留出資料上 驗證過之後才成立。出廠的 checkpoint 是過度自信的,過度多少取決於選項數,所以一個沒調過的閾值 可能把低於模型自身準確率的答案選進來(#394)。

# THRESHOLD is a number you measured on your own held-out data, not one the model ships.
# Fit and validate the temperatures first — the fine-tuning notebook has the loop:
#   notebooks/laya_finetune_typed_decisions_2xT4_kaggle.ipynb
ans = result["answers"]["dept"]
if ans["answer_confidence"] >= THRESHOLD:
    ...

confidence 是歸一化熵:分佈尖銳時高,分散時低,與排第一的答案對不對無關。它是個有用的訊號, 但它不在同一量綱上,所以兩者不能用同一個數字來門控:

# the same three answers, and the two numbers are not the same
dept      confidence 0.9339   answer_confidence 0.9881
urgent    confidence 0.8727   answer_confidence 0.8727
severity  confidence 0.5187   answer_confidence 0.7881

對 noul 而言,兩者按構造就是相等的 —— 兩個選項下,max(p, 1-p) 就是 max(p) —— 所以一個 noul 答案本身分不出你讀的是哪一個。每個型別上兩個鍵都存在,這樣這個選擇是顯式的,而不是隱含的。

閾值是一項策略,不是模型的屬性。 兩個 checkpoint 出廠都是過度自信的,過度多少取決於選項數, 所以在一個 3 選項問題上量出的數字不能搬到 20 選項問題上。在你自己的資料上量它;BENCHMARKS.md 的校準一節裡有擬合迴圈和擬合出的值。

action 與 act_probability

action.act_probability 是另一個 head 的分數,回答「agent 到底該不該據此行動」,與答案自身的 置信度不同。每種問題型別都會報告它。庫裡沒有任何東西替你給它設閾值。

預設

三套現成的問題集,常見情形不必手寫判定標準:

from laya import triage_questions, guard_questions, moderation_questions

agent.system_one(ticket, triage_questions())

把它們當起點,而不是當契約 —— 用 render_options 讀出它們產生的問題,並在釋出之前檢查標籤是否 適合你的領域。

把選項讀回來

因為選項順序有位置意義,而選項文本正是模型讀到的內容,所以能確切看到發出去的是什麼,是值得的:

from laya import render_options

render_options({"t": "choice", "crit": {"billing": None, "sales": "pricing"}})
# ['billing', 'sales: pricing']

render_options({"t": "score", "crit": ["low", "high"]})
# ['level 0: low', 'level 1: high']

同樣的標籤換個順序,就按那個順序渲染,這正是「順序是問題身份的一部分」的原因:

render_options({"t": "choice", "crit": {"x": "first", "y": "second"}})
# ['x: first', 'y: second']
render_options({"t": "choice", "crit": {"y": "second", "x": "first"}})
# ['y: second', 'x: first']

注意鍵名。 render_options 接受的是內部的短鍵形式 {"t": ..., "crit": ...},不是你在 問題裡寫的 {"type": ..., "criteria": ...} 形式 —— 傳公開的形狀會拋 KeyError: 't'。

這個轉換是一段很短、很穩定的對映,你可以直接內聯,從而避免去碰一個私有輔助函式 —— Agent._to_internal 是內部的,可能會變:

def as_internal(q):
    """The short-key shape `render_options` reads, from a question as you wrote it."""
    crit = q.get("criteria")
    if q["type"] == "choice" and isinstance(crit, list):
        crit = {c: None for c in crit}
    return {"t": q["type"], "ins": q["instructions"], "crit": crit}

render_options(as_internal(question))

這復刻了庫對一個寫成標籤列表的 choice 問題所做的處理;一個 criteria dict 和一個 score 列表則原樣通過。

在設計方案之前值得知道的限制

  • 選項多的時候,置信度不是正確性的保證。 在一個 20 選項問題上,正確答案和錯誤答案的分佈 大量重疊,閾值可能最終選出低於模型自身準確率的結果。見 #394。
  • 強制選擇式問題裡的否定不能可靠處理。 一個關於「取消」的問題,可能對一個寫著不要取消的 state 返回取消標籤,而且置信度很高,兩個 checkpoint 都是如此。見 #377。
  • noul 可能跟隨它的標籤,而不是 state,所以任何換標籤都要驗證。
  • 選項多不是免費的。 超過大約 20 個之後模型會很快退化;用 shortlist 輔助函式在提問之前把 大的標籤空間縮小,或者把它拆成一個粗問題和一個細問題。