Score
Score 是一種 System One 問題型別,用於按有序的、帶描述的檔位給內容打分。答案包含一個 score、每個檔位的機率和置信度。
當答案是你可以分步驟描述的譜系上的一個位置時,就用 Score。例如,一個 bug 有多嚴重、客戶有多滿意、候選人有多年 Python 經驗。如果答案是固定選項集合中的一個、彼此之間沒有順序,就用 Choice。如果是是/否,就用 Noul。選擇問題型別對比了這三種。
Score 的答案是 score 中沿你設定的檔位的一個位置,可以落在兩個檔位之間。模型還會在 probabilities 中返回每個檔位的機率,併為答案返回一個 confidence 值。
示例評分問題
How severe is the reported issue?
狀態(待評估的內容)
The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari.
答案
各檔位的機率
置信度
得分: 1.43
得分與置信度是怎麼算出來的
得分:
把每個檔位的編號乘以它的機率,再相加:
0 × 0 + 1 × 0.57 + 2 × 0.43 ≈ 1.43
置信度
TypeSafe 根據機率在各檔位上的分散程度算出它。全部集中在一檔是 1.0;分佈越平均,置信度越低。
How formal is this outfit based on the description?
狀態(待評估的內容)
A navy blazer over a plain white T-shirt, dark jeans, and clean leather loafers. No tie.
答案
各檔位的機率
置信度
得分: 1.86
得分與置信度是怎麼算出來的
得分:
把每個檔位的編號乘以它的機率,再相加:
0 × 0 + 1 × 0.14 + 2 × 0.86 + 3 × 0 + 4 × 0 ≈ 1.86
置信度
TypeSafe 根據機率在各檔位上的分散程度算出它。全部集中在一檔是 1.0;分佈越平均,置信度越低。
How relevant is this candidate's experience to the job posting?
狀態(待評估的內容)
Job posting: Senior backend engineer building Python APIs and PostgreSQL services. Candidate: Three years building Django REST APIs with PostgreSQL, preceded by two years in frontend JavaScript. Has owned small services but has not led a backend team.
答案
各檔位的機率
置信度
得分: 2.52
得分與置信度是怎麼算出來的
得分:
把每個檔位的編號乘以它的機率,再相加:
0 × 0 + 1 × 0 + 2 × 0.48 + 3 × 0.52 ≈ 2.52
置信度
TypeSafe 根據機率在各檔位上的分散程度算出它。全部集中在一檔是 1.0;分佈越平均,置信度越低。
How frustrated is the customer?
狀態(待評估的內容)
Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. This is the third time I'm writing in and honestly I'm done. Steps: open any report, click Export, choose PDF. Chrome 128 on macOS.
答案
各檔位的機率
置信度
得分: 1.26
得分與置信度是怎麼算出來的
得分:
把每個檔位的編號乘以它的機率,再相加:
0 × 0 + 1 × 0.74 + 2 × 0.26 ≈ 1.26
置信度
TypeSafe 根據機率在各檔位上的分散程度算出它。全部集中在一檔是 1.0;分佈越平均,置信度越低。
How much does the report give an engineer to work with?
狀態(待評估的內容)
Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. This is the third time I'm writing in and honestly I'm done. Steps: open any report, click Export, choose PDF. Chrome 128 on macOS.
答案
各檔位的機率
置信度
得分: 3.00
得分與置信度是怎麼算出來的
得分:
把每個檔位的編號乘以它的機率,再相加:
0 × 0 + 1 × 0 + 2 × 0 + 3 × 1 ≈ 3.00
置信度
TypeSafe 根據機率在各檔位上的分散程度算出它。全部集中在一檔是 1.0;分佈越平均,置信度越低。
每一步前面的數字是位置,詳見檔位。
請求結構
發往 TypeSafe API 的 POST 請求體,和任何其它問題型別一樣有三個頂層欄位:state,即要評估的內容;model;以及 questions。每個 Score 問題有以下欄位:
type:始終為"score"。instructions:模型要回答的問題。即它在給什麼打分。criteria:按順序排列的檔位描述陣列,從量表低端到高階。至少應有兩個檔位;API 最多接受 10 個。
下面這個請求,state 是一份 bug 報告,問題是這個 bug 有多嚴重:
{
"state": "The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari.",
"questions": {
"bug_severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists"
]
}
}
}問題 id 由你選,這裡是 bug_severity。這個 id 不會發送給模型。答案會以同一個 id 返回。
檔位
criteria 中的每一項是一個檔位:可能答案譜系上的一個點,用文字描述。檔位的編號就是它在 criteria 陣列中的位置,從 0 開始,所以上面三項分別是檔位 0、1、2。陣列的順序就是編號。
模型只拿到這些描述,別的什麼都沒有,而且每個檔位都單獨對照 state 來判斷。
響應中的 score 是檔位譜系上的一個位置。三檔量表上,它的範圍是 0 到 2,並且可以落在兩個檔位之間。
我們的客戶端 SDK提供型別化的問題。在 Python 裡,同一個問題就是一個 Score:
from typesafe_sdk import Score, TypeSafeClient
with TypeSafeClient() as client:
response = client.system_one(
state="The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari.",
questions={
"bug_severity": Score(
instructions="How severe is the reported issue?",
criteria=[
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists",
],
),
},
)
print(response.answers["bug_severity"].score)
用 system_one 方法或 https://api.typesafe.ai/v1/systemone 端點來呼叫 System One 模型。model 欄位選擇由哪個模型處理請求。如何用 TypeSafe 構建講了在程式碼裡的什麼位置呼叫它。
使用我們的某個客戶端 SDK,或直接呼叫 TypeSafe API。如果由編碼 agent 替你寫整合,先安裝 TypeSafe agent skill,這樣它就知道請求和響應的結構。
響應結構
響應在 answers 中每個問題一條,放在請求裡的 id 之下。這是上面示例請求的響應:
{
"model": "jev-1.13.0",
"answers": {
"bug_severity": {
"type": "score",
"score": 1.43,
"confidence": 0.35,
"legend": {
"0": "Cosmetic; no impact to functionality",
"1": "Broken or degraded feature, but workaround exists",
"2": "Blocking issue; no workaround exists"
},
"probabilities": {
"0": 0.0,
"1": 0.57,
"2": 0.43
}
}
},
"usage": {
"input_tokens": 332,
"output_tokens": 18
}
}
每個 Score 答案有五個值:
type:TypeSafe 問題的型別。probabilities:每個檔位的機率,以檔位編號的字串為鍵。所有值之和為 1。score:檔位編號數軸上的位置,從 0 到最高檔位編號(這裡是 2)。它是每個檔位編號乘以其機率再相加:0 x 0.0 + 1 x 0.57 + 2 x 0.43 = 1.43。legend:每個檔位編號映射回它的描述。confidence:一個 0 到 1 的數字,由probabilities的分散程度算出。機率集中在單個檔位意味著高置信度。機率分散在多個檔位意味著低置信度。
1.43 的 score 意味著模型在檔位 1 和 2 之間搖擺,略偏向檔位 1。這和報告吻合:匯出壞了,換到 Chrome 對大多數客戶算是變通辦法,但對只用 Safari 的那部分客戶不是。模型把 0.57 放在“有變通辦法”上,0.43 放在“沒有變通辦法”上,置信度是 0.35,因為它是分裂的。
用 Python SDK 時,ScoreAnswer 有 score、confidence、probabilities 和 legend 這幾個型別化欄位。SDK 用整數檔位而不是字串給 probabilities 和 legend 作鍵。
閱讀 Score
我們來看看不同輸入下 score 如何變化。例如,用上面請求中的問題及其檔位:
"How severe is the reported issue?"
→ 0: Cosmetic; no impact to functionality
→ 1: Broken or degraded feature, but workaround exists
→ 2: Blocking issue; no workaround exists
可以看到不同的 bug 報告如何改變 score:
probabilities | |||||
|---|---|---|---|---|---|
| 狀態 | score | confidence | 檔位 0 | 檔位 1 | 檔位 2 |
| The export button is misaligned by a few pixels on the settings page. | 0.0 | 1.0 | 1.0 | 0.0 | 0.0 |
| The PDF export button does nothing when clicked. I can still export to CSV and convert it myself, but that takes ages. | 1.0 | 1.0 | 0.0 | 1.0 | 0.0 |
| Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. | 1.11 | 0.84 | 0.0 | 0.89 | 0.11 |
| The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari. | 1.43 | 0.35 | 0.0 | 0.57 | 0.43 |
| Nobody on our team can log in since this morning. We get a 500 error on every attempt. | 2.0 | 1.0 | 0.0 | 0.0 | 1.0 |
在這些例子裡,置信度 1.0 意味著返回的分佈把所有機率都放在一個檔位上。這描述的是模型的答案,並不保證答案正確。
score 是檔位編號的機率加權平均。在第三和第四個例子裡,機率分裂在檔位 1 和 2 之間。檔位 2 權重越大,score 越高。它衡量的不是沒有變通辦法的客戶佔比。
不同的分佈可以產生相同的 score。1.0 的 score 可能意味著全部機率都在檔位 1,也可能意味著檔位 0 和 2 各佔一半。把 probabilities 和 confidence 與 score 一起讀,才能區分這些情況。
分數形式的 score 是一個位置。你可以用它按嚴重程度給報告排序,或在程式碼需要一個確定結果時四捨五入到最近的檔位。我們的實體對齊 cookbook有一個四捨五入到最近檔位以做決策的例子。
Score 上偏低的置信度通常意味著三種情況之一。這些檔位對這個 state 有重疊,問題衡量了不止一件事,或者 state 說得不夠多、無法定位。我們的置信度文件講了如何在程式碼裡使用它。
寫出好的檔位
描述情境,而不是程度。“Broken or degraded feature, but workaround exists”給了模型可以拿 state 去匹配的東西。“Moderately severe”則沒有。具體的描述能幫模型區分檔位。把答案和已知例子對照檢查;單憑置信度更高,並不能說明描述更好。
每個檔位都是單獨評估的。模型看不到檔位的編號,也看不到相鄰檔位,所以“比上一檔更嚴重”對它沒有任何意義,描述或 instructions 裡的數字也幫不上忙。下面是在上面表格裡那條“按鈕錯位”報告上,檔位只有數字時會發生什麼:
instructions: "Rate severity from 0 to 2, where 2 is worst"
criteria: ["0", "1", "2"]
→ score 0.55, confidence 0.33, probabilities 0: 0.45, 1: 0.55, 2: 0.0
同一份報告,用那三個帶描述的檔位,得分 0.0,置信度 1.0。只用數字時,模型沒有可匹配的東西,就把機率分裂在 0 和 1 之間。
用盡可能多的、你能清楚區分開的檔位,最多 10 個。三個也可以。別加你沒法清楚區分的檔位。
讓每個 Score 問題只針對一個維度。如果一段描述說“punctual and smart and experienced”,那這個問題就在衡量三件事,而在一個維度高、另一個維度低的輸入就無法定位。置信度下降,score 的意義也變弱。把它拆成每個維度一個 Score 問題,在程式碼裡組合,下一節會演示。
如果你量表的頂端有一個需要區別對待的罕見極端情況,就給它單獨一個檔位。一個以“very angry”結尾的情感量表,可以加上“abusive or threatening”。沒有這個檔位,兩條訊息可能都得到接近頂端的 score。單憑 score 可能無法區分它們。
如果完全沒有中間地帶,答案只是少數幾個離散類別之一,那就改用 Choice,或把問題拆成幾個 Noul 問題。一定要用你自己的資料測試你的檔位。同一個量表的兩種措辭,在你的資料上可能表現不同。
把一個複雜判斷拆成幾個 Score 問題
一個依賴多件事的複雜判斷,最好拆成每件事一個 Score 問題。然後你可以在程式碼裡把 TypeSafe 返回的各個 Score 組合起來,做出這個判斷。有些 Score 問題可能比其它的更重要,所以給每個 Score 問題配一個表示相對重要性的權重。權重由你定。當組合結果和你的團隊會做出的決定不符時,在程式碼裡改權重再跑一次。把這些 Score 問題放在一個請求裡傳送。它們會並行評估。增加問題幾乎不改變響應時間,只多花幾個問題 token;見一次提多個問題。
下面的請求是上面表格裡那條轉圈的工單,上下文更多一些。它提三個 Score 問題:bug 有多嚴重、客戶有多沮喪、這份報告給工程師多少可用的資訊。
{
"state": "Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. This is the third time I'm writing in and honestly I'm done. Steps: open any report, click Export, choose PDF. Chrome 128 on macOS.",
"questions": {
"severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists"
]
},
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer?",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language or threatening to leave"
]
},
"report_quality": {
"type": "score",
"instructions": "How much does the report give an engineer to work with?",
"criteria": [
"No detail; just says something is broken",
"Names the feature but no steps or environment",
"Steps to reproduce or environment, but not both",
"Steps to reproduce and environment"
]
}
}
}TypeSafe 的響應:
{
"model": "jev-1.13.0",
"answers": {
"severity": {
"type": "score",
"score": 1.24,
"confidence": 0.64,
"legend": {
"0": "Cosmetic; no impact to functionality",
"1": "Broken or degraded feature, but workaround exists",
"2": "Blocking issue; no workaround exists"
},
"probabilities": {
"0": 0.0,
"1": 0.76,
"2": 0.24
}
},
"frustration": {
"type": "score",
"score": 1.28,
"confidence": 0.58,
"legend": {
"0": "Calm, just stating facts",
"1": "Frustrated but civil",
"2": "Very angry, strong language or threatening to leave"
},
"probabilities": {
"0": 0.0,
"1": 0.72,
"2": 0.28
}
},
"report_quality": {
"type": "score",
"score": 3.0,
"confidence": 1.0,
"legend": {
"0": "No detail; just says something is broken",
"1": "Names the feature but no steps or environment",
"2": "Steps to reproduce or environment, but not both",
"3": "Steps to reproduce and environment"
},
"probabilities": {
"0": 0.0,
"1": 0.0,
"2": 0.0,
"3": 1.0
}
}
},
"usage": {
"input_tokens": 468,
"output_tokens": 43
}
}
每個問題都單獨對照工單作答,並給出一個 score:
severity為 1.24,置信度 0.64。讀法和開頭的例子一樣:匯出壞了,有些人有變通辦法。frustration為 1.28,置信度 0.58。措辭還算客氣,但“third time”和“I’m done”把一部分分數推向最高檔,於是模型把 0.72 和 0.28 分裂在“frustrated but civil”和“very angry”之間。對這條工單來說,兩個檔位有重疊,所以置信度中等。report_quality為 3.0,置信度 1.0。復現步驟和瀏覽器版本都寫明瞭。
這三個量表的長度不同,所以在組合之前,先歸一化每個 score。四檔量表返回 0 到 3,三檔量表返回 0 到 2,所以一個的滿分比另一個的滿分大。把每個 score 除以它的最高檔位編號 len(criteria) - 1,就把所有 score 都放到 0 到 1 上。這樣權重才名副其實:severity 給 0.6、frustration 給 0.3,意味著 severity 的分量是 frustration 的兩倍。
下面的 TypeSafe Python SDK 程式碼提這三個問題,歸一化每個 score,並用一個示例的優先順序計算把它們組合起來:
from typesafe_sdk import Score, TypeSafeClient
TRIAGE_QUESTIONS = {
"severity": Score(
instructions="How severe is the reported issue?",
criteria=[
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists",
],
),
"frustration": Score(
instructions="How frustrated is the customer?",
criteria=[
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language or threatening to leave",
],
),
"report_quality": Score(
instructions="How much does the report give an engineer to work with?",
criteria=[
"No detail; just says something is broken",
"Names the feature but no steps or environment",
"Steps to reproduce or environment, but not both",
"Steps to reproduce and environment",
],
),
}
def normalized(answers, question_id: str) -> float:
"""Put a score on 0 to 1 by dividing by its top level number."""
top_level = len(TRIAGE_QUESTIONS[question_id].criteria) - 1
return answers[question_id].score / top_level
def priority(ticket: str) -> float:
with TypeSafeClient() as client:
response = client.system_one(
state=ticket,
questions=TRIAGE_QUESTIONS,
)
answers = response.answers
severity = normalized(answers, "severity")
frustration = normalized(answers, "frustration")
report_quality = normalized(answers, "report_quality")
# A detailed report helps an engineer investigate, so it raises priority a little.
return 0.6 * severity + 0.3 * frustration + 0.1 * report_quality
對於上面的示例響應,歸一化後的分數是 severity 0.62、frustration 0.64、report quality 1.0。優先順序是 0.6 × 0.62 + 0.3 × 0.64 + 0.1 × 1.0 = 0.664,四捨五入為 0.66。
權重存在於你的程式碼裡,所以你能確切看到這個數字是怎麼來的,並在排序和你的團隊會做的決定不符時改它。如果之後需要更多 Score 問題,把它們加進 TRIAGE_QUESTIONS。請求數仍然是一個。這種把複雜判斷拆成一個個獨立的 Score、再用程式碼裡的權重組合起來的手法,叫作組合評分模式。
結構化檔位描述
先給每個檔位一段基本的文本描述。當模型在你認為清晰的輸入上總是打在兩個相鄰檔位之間時,就把每個檔位從字串換成一個物件,一個欄位說明這個檔位涵蓋什麼,另一個欄位放幾個示例情境。每個檔位用相同的欄位名,這樣模型才能同類相比。
下面的請求是我們之前用過的那條轉圈工單,但每個檔位都帶了示例:
{
"state": "Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too.",
"questions": {
"bug_severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
{
"what": "Cosmetic; no impact to functionality",
"examples": [
"typo in a label",
"misaligned icon"
]
},
{
"what": "Broken or degraded feature, but workaround exists",
"examples": [
"export fails in one browser but works in another"
]
},
{
"what": "Blocking issue; no workaround exists",
"examples": [
"cannot log in",
"data loss"
]
}
]
}
}
}響應:
{
"model": "jev-1.13.0",
"answers": {
"bug_severity": {
"type": "score",
"score": 1.09,
"confidence": 0.87,
"legend": {
"0": {
"what": "Cosmetic; no impact to functionality",
"examples": [
"typo in a label",
"misaligned icon"
]
},
"1": {
"what": "Broken or degraded feature, but workaround exists",
"examples": [
"export fails in one browser but works in another"
]
},
"2": {
"what": "Blocking issue; no workaround exists",
"examples": [
"cannot log in",
"data loss"
]
}
},
"probabilities": {
"0": 0.0,
"1": 0.91,
"2": 0.09
}
}
},
"usage": {
"input_tokens": 379,
"output_tokens": 18
}
}
用純字串時,這條工單得分 1.11,置信度 0.84。帶示例時得分 1.09,置信度 0.87,變化很小,因為純字串已經把它放得很好了。當純字串讓模型搖擺時,效果會更大,如下一張表所示。
示例會引導模型,只有當它們看起來像你的真實輸入時才有幫助。下表是開頭那條 Safari 報告,配三組不同的檔位物件:
| 檔位描述 | score |
confidence |
|---|---|---|
| 純字串:沒有帶示例的物件 | 1.43 | 0.35 |
| 加入的 examples 陣列帶一個有用的示例:“export fails in one browser but works in another” | 1.03 | 0.96 |
| 加入的 examples 陣列帶的示例與瀏覽器無關:“search fails, but browsing categories still works” | 1.43 | 0.35 |
在這個對比裡,匹配的示例把幾乎全部機率集中到一個檔位上。無關的示例給出的結果和純字串一樣。更高的置信度並不能確立哪個答案是對的。選擇那些你已知預期檔位的示例,然後在一批獨立的輸入上測試修改後的描述,再決定保留。