Noul
Noul 問題讓 TypeSafe 模型評估一個是/否問題,並返回答案為「是」的機率。
當答案只有「是」或「否」時,就用 Noul。比如:這條訊息是在要求退款嗎、這份簡歷提到分散式系統了嗎、這條評論含有個人資料嗎。如果答案是一組選項之一,就用 Choice。如果答案是一個連續譜系上的位置,就用 Score。選擇題型別對三者做了比較。
Noul 的答案是一個數,表示答案為「是」的機率,其中 0 表示否,1 表示是。
請求結構
發往 TypeSafe API 的 POST 請求體,和任何其它問題型別一樣有同樣的三個頂層欄位:state,即要評估的內容;model;以及 questions。每個 Noul 問題都有以下欄位:
type:固定為"noul"。instructions:模型要回答的是/否問題,或者要它判斷的一個陳述。criteria:可選。一個物件,用true和false兩個描述說明「是」和「否」分別意味著什麼。
下面這個請求裡,狀態是一條客服訊息,兩個問題分別是:客戶是否想找真人,以及客戶之前是否聯絡過客服:
{
"state": "I have asked three times now. Can I please just talk to a real person?",
"questions": {
"is_human_escalation": {
"type": "noul",
"instructions": "Is the customer asking for a human agent?"
},
"is_repeat_contact": {
"type": "noul",
"instructions": "Has the customer contacted support about this before?",
"criteria": {
"true": "Mentions a prior attempt, ticket, or that they have asked before",
"false": "No sign of any previous contact"
}
}
}
}問題 id,也就是這裡的 is_human_escalation 和 is_repeat_contact,由你選擇。這些 id 不會發送給模型。每個答案都以相同的 id 返回。第一個問題只靠 instructions。第二個額外用 criteria 說明什麼算「是」、什麼算「否」。
用 Python SDK 時,同樣的問法是 Noul 物件:
from typesafe_sdk import Noul, NoulCriteria, TypeSafeClient
with TypeSafeClient() as client:
response = client.system_one(
model="jev-latest",
state="I have asked three times now. Can I please just talk to a real person?",
questions={
"is_human_escalation": Noul(
instructions="Is the customer asking for a human agent?",
),
"is_repeat_contact": Noul(
instructions="Has the customer contacted support about this before?",
criteria=NoulCriteria(
true="Mentions a prior attempt, ticket, or that they have asked before",
false="No sign of any previous contact",
),
),
},
)
print(response.answers["is_human_escalation"].noul)
print(response.answers["is_repeat_contact"].noul)
system_one 方法和 https://api.typesafe.ai/v1/systemone 端點,都是以 TypeSafe 的 AI 模型 System One 命名的。如何用 TypeSafe 構建介紹了該在程式碼裡的哪些地方用它。
如果你在用 coding agent,先安裝 TypeSafe agent skill,這樣它就知道請求和響應長什麼樣。
響應結構
響應裡的 answers 為每個問題各有一個條目,鍵是請求裡用的 id:
{
"model": "jev-1.13.0",
"answers": {
"is_human_escalation": {
"type": "noul",
"noul": 0.99
},
"is_repeat_contact": {
"type": "noul",
"noul": 0.93
}
},
"usage": {
"input_tokens": 360,
"output_tokens": 39
}
}
這裡兩個答案都接近 1。客戶說的是 “talk to a real person”,所以 is_human_escalation 是 0.99。“I have asked three times now” 與 is_repeat_contact 的 true 描述相符,所以是 0.93。
如何解讀 Noul
這個數同時給出了答案和確定性。接近 1 是強「是」。接近 0 是強「否」。接近 0.5 表示模型給「是」和「否」的機率差不多。
下表是 jev-1.13.0 對不同客戶訊息在 is_human_escalation 問題上記錄下來的答案:
| 狀態 | noul |
|---|---|
| Thanks, that fixed it! | 0.02 |
| How do I reset my password? | 0.07 |
| I need this sorted today, whatever it takes. | 0.26 |
| Are you a bot? | 0.40 |
| Is there any way to speak to someone about my invoice? | 0.84 |
| I have asked three times now. Can I please just talk to a real person? | 0.99 |
前兩條和後兩條都很明確。“I need this sorted today” 很急,但從未要求找真人,得到 0.26。“Are you a bot?” 暗示想要真人,卻沒有明說,模型幾乎五五開,給出 0.40。這兩類訊息,都需要靠程式碼裡的閾值來做決定。
與 Choice 或 Score 不同,Noul 沒有單獨的 confidence 值。Noul 的機率分佈只有兩個結果,是和非,所以單個 noul 值就完整描述了它。Choice 或 Score 把機率鋪在多個選項或檔位上,confidence 則是對這種鋪開的概括。
最常見的是,程式碼把 noul 取個閾值變成布林值:
wants_human = response.answers["is_human_escalation"].noul > 0.9
if wants_human:
route_to_agent(ticket)
else:
route_to_bot(ticket)
閾值設在哪,取決於判斷錯了要付出多大代價。當「是」和「否」都同樣容易處理時,用 0.5。把假「是」當真代價很大時(比如呼叫值班人員或退款),就調高。漏掉真「是」代價很大時(比如沒能標出一個安全問題),就調低。中間的值可以交給人工,而不走任何一條程式碼分支。這和在置信度頁裡為 Choice 和 Score 答案描述的三路分流是一樣的。
Noul 的取值從 0 到 1,但它不是你問的那件事本身的刻度。它是答案為「是」的機率。如果問題實際上問的是程度,這個值並不衡量程度。下面,對四位候選人問了 “Is the candidate strong in Python?”,並排給出一個含四個檔位的 Score:no experience、some familiarity、regular use in a job、deep expertise。
| 候選人 | Noul: “Is the candidate strong in Python?” | Score: “How much Python experience does the candidate have?” |
|---|---|---|
| My experience is in Java and Go. I have not used Python. | 0.03 | 0.0 (No experience) |
| I have used Python occasionally for small scripts alongside my main Java work. | 0.14 | 1.0 (Some familiarity) |
| I used Python every day for two years in my last job, mostly data pipelines. | 0.81 | 2.05 (Regular use in a job) |
| I have written Python daily for eight years, including maintaining a large Django codebase. | 0.92 | 2.89 (Deep expertise) |
Noul 只判斷一個命題——“strong”,值就是它成立的可能性有多大。你可以在程式碼裡給 0 到 1 劃分區間,比如把 0.3 到 0.7 當作 “some experience”,但模型看不到這些區間,所以答案裡沒有任何東西是照著它們判斷的。中間值可能意味著中等經驗,也可能意味著情況不明;而且候選人之間的間隔也不是你選的。Score 則逐條判斷每個檔位描述,所以每位候選人都落到了你寫的某個檔位之上或附近,返回的機率顯示了模型如何在各檔位之間分配它的判斷。如果你不認同,就改寫某個檔位再跑一次。選擇題型別解釋了其中的區別。
如何寫 Noul 問題
一個 Noul 只問一個是/否問題。如果一個問題裡有兩個條件,比如 “Is the customer angry and asking for a refund?”,模型就得同時判斷兩者,值的資訊量就變小了。應該問兩個 Noul,再在程式碼裡把它們組合起來。
把問題措辭成「值高就代表是」。“Does the message contain personal data?” 很清晰。“Is the message free of personal data?” 則把含義反了過來,之後讀它的程式碼會弄反。
用陳述句和用疑問句一樣有效。對於 “The customer is requesting a refund”,接近 1 的值表示該陳述為真。拿你自己的資料把兩種措辭都試一遍,看看哪種更好。
把「是」和「否」之間的邊界劃得毫不含糊。“Does this candidate have any Python experience?” 效果不錯,因為 “any” 不留中間地帶。當邊界很微妙時,就加上帶 true 和 false 描述的 criteria,就像上面 is_repeat_contact 那樣。對大多數 Noul 來說 instructions 就夠了,所以把你的問題分別在有、沒有 criteria 的情況下各試一遍,留下在你的文件上答案更好的那種。
好做法:一次呼叫問多個問題
對於一張條件清單,就在一次請求裡問許多 Noul 問題:每個條件一個問題,由程式碼來決定這些結果的組合意味著什麼。問題是並行評估的,所以增加 Noul 幾乎不改變響應時間。一起問多個問題有更詳細的說明。
在程式碼裡處理多個 Noul 答案
上面那個兩問題的請求,已經足夠讓程式碼把訊息分流了。下面的示例在客戶要求找真人時升級給人工,在客戶之前聯絡過時提高優先順序。任一問題取到中間值時,就交給複核人,而不走程式碼分支:
from typesafe_sdk import Noul, NoulCriteria, TypeSafeClient
SUPPORT_QUESTIONS = {
"is_human_escalation": Noul(
instructions="Is the customer asking for a human agent?",
),
"is_repeat_contact": Noul(
instructions="Has the customer contacted support about this before?",
criteria=NoulCriteria(
true="Mentions a prior attempt, ticket, or that they have asked before",
false="No sign of any previous contact",
),
),
}
YES = 0.8
NO = 0.2
def route(message: str) -> None:
with TypeSafeClient() as client:
response = client.system_one(
model="jev-latest",
state=message,
questions=SUPPORT_QUESTIONS,
)
answers = response.answers
wants_human = answers["is_human_escalation"].noul
repeat = answers["is_repeat_contact"].noul
if NO < wants_human < YES or NO < repeat < YES:
# The model isn't sure either way. Let a person decide.
send_to_review(message)
return
priority = "high" if repeat > YES else "normal"
if wants_human > YES:
route_to_agent(message, priority=priority)
else:
route_to_bot(message, priority=priority)
對上面那條訊息,is_human_escalation 的 noul 答案是 0.99,is_repeat_contact 是 0.93,所以程式碼以高優先順序把它路由給人工客服。“How do I reset my password?” 這條訊息在兩個問題上都是 0.07,被路由給機器人。
閾值在你自己的程式碼裡。如果複核人收到的訊息太多,就縮小 NO 和 YES 之間的間距。如果放過去的錯路由太多,就把它拉大。如果以後還需要知道訊息是否提到付款、是否含有個人資料,就往 SUPPORT_QUESTIONS 裡再加一個 Noul。請求次數仍然是一次。
結構化 instructions
Instructions 可以是一個物件而不是字串,問題放在一個欄位裡,補充資料放在其它欄位裡。在問題中使用結構化資料講了它什麼時候有用。這裡它用於一個用程式碼拼出來的問題:把一份剛到的簡歷,與候選人庫裡可能是同一個人的記錄逐一比對。每條記錄都原樣放進一個 potential_duplicate 欄位,question 對每條記錄都一樣,所有記錄都在一次請求裡檢查。程式碼生成的問題鍵裡含有每條記錄的資料庫 ID:
{
"state": {
"resume": {
"name": "John Smith",
"location": "Oakland, CA",
"summary": "Backend engineer with eight years of Python and Go experience.",
"experience": [
{
"employer": "Google",
"title": "Senior Backend Engineer",
"years": "2021-2025"
},
{
"employer": "Microsoft",
"title": "Software Engineer",
"years": "2017-2021"
}
]
}
},
"questions": {
"same_as_record_18": {
"type": "noul",
"instructions": {
"potential_duplicate": {
"name": "Jon Smith",
"location": "Oakland, CA",
"last_employer": "Google"
},
"question": "Is the resume for the same person as `potential_duplicate`?"
}
},
"same_as_record_42": {
"type": "noul",
"instructions": {
"potential_duplicate": {
"name": "John Smith",
"location": "Austin, TX",
"last_employer": "Lone Star Freight"
},
"question": "Is the resume for the same person as `potential_duplicate`?"
}
},
"same_as_record_77": {
"type": "noul",
"instructions": {
"potential_duplicate": {
"name": "John Smithers",
"location": "Oakland, CA",
"last_employer": "Bay Health Clinic"
},
"question": "Is the resume for the same person as `potential_duplicate`?"
}
}
}
}響應:
{
"model": "jev-1.13.0",
"answers": {
"same_as_record_18": {
"type": "noul",
"noul": 0.74
},
"same_as_record_42": {
"type": "noul",
"noul": 0.09
},
"same_as_record_77": {
"type": "noul",
"noul": 0.08
}
},
"usage": {
"input_tokens": 535,
"output_tokens": 58
}
}
每個答案都是「這份簡歷就是該記錄裡那個人」的機率。記錄 18 名字拼寫不同,但地點和僱主對得上,得到 0.74。記錄 42 同名,但城市不同、僱主也不同,得到 0.09。記錄 77 名字相近、地點相同,但僱主不同,得到 0.08。在程式碼裡給每個值取閾值,就像在程式碼裡處理多個 Noul 答案那樣,把中間值交給人工。
用 Python SDK 時,問題是根據候選人記錄拼出來的。問題文本固定,記錄會變:
from typesafe_sdk import Noul, TypeSafeClient
SAME_PERSON = "Is the resume for the same person as `potential_duplicate`?"
def duplicate_questions(candidates: list[dict]) -> dict[str, Noul]:
"""One Noul per candidate record, all asking the same question."""
return {
f"same_as_record_{candidate['id']}": Noul(
instructions={
"potential_duplicate": {
"name": candidate["name"],
"location": candidate["location"],
"last_employer": candidate["last_employer"],
},
"question": SAME_PERSON,
},
)
for candidate in candidates
}
def find_duplicates(resume: dict, candidates: list[dict]) -> list[str]:
with TypeSafeClient() as client:
response = client.system_one(
model="jev-latest",
state={"resume": resume},
questions=duplicate_questions(candidates),
)
return [
question_id
for question_id, answer in response.answers.items()
if answer.noul > 0.7
]
結構化資料抽取級聯 cookbook用結構化的 instructions 來校驗一條抽取出來的記錄。每個欄位都得到同一組問題。每個問題的 instructions 物件把問題文本放在 main_question 屬性裡。還有 field_spec 和 extracted_field 兩個屬性,會隨欄位變化。
cookbook 裡的 Noul
看看我們的 cookbooks,裡面有使用 Noul 問題的應用: