Noul
Noul
Noul 问题让 TypeSafe 模型评估一个是/否问题,并返回答案为「是」的概率。
当答案只有「是」或「否」时,就用 Noul。比如:这条消息是在要求退款吗、这份简历提到分布式系统了吗、这条评论含有个人数据吗。如果答案是一组选项之一,就用 Choice。如果答案是一个连续谱系上的位置,就用 Score。选择题类型对三者做了比较。
Noul 的答案是一个数,表示答案为「是」的概率,其中 0 表示否,1 表示是。
请求结构
发往 TypeSafe API 的 POST 请求体,和任何其它问题类型一样有同样的三个顶层字段:state,即要评估的内容;model;以及 questions。每个 Noul 问题都有以下字段:
type:固定为"noul"。instructions:模型要回答的是/否问题,或者要它判断的一个陈述。criteria:可选。一个对象,用true和false两个描述说明「是」和「否」分别意味着什么。
下面这个请求里,状态是一条客服消息,两个问题分别是:客户是否想找真人,以及客户之前是否联系过客服:
{
"state": "I have asked three times now. Can I please just talk to a real person?",
"questions": {
"is_human_escalation": {
"type": "noul",
"instructions": "Is the customer asking for a human agent?"
},
"is_repeat_contact": {
"type": "noul",
"instructions": "Has the customer contacted support about this before?",
"criteria": {
"true": "Mentions a prior attempt, ticket, or that they have asked before",
"false": "No sign of any previous contact"
}
}
}
}问题 id,也就是这里的 is_human_escalation 和 is_repeat_contact,由你选择。这些 id 不会发送给模型。每个答案都以相同的 id 返回。第一个问题只靠 instructions。第二个额外用 criteria 说明什么算「是」、什么算「否」。
用 Python SDK 时,同样的问法是 Noul 对象:
from typesafe_sdk import Noul, NoulCriteria, TypeSafeClient
with TypeSafeClient() as client:
response = client.system_one(
model="jev-latest",
state="I have asked three times now. Can I please just talk to a real person?",
questions={
"is_human_escalation": Noul(
instructions="Is the customer asking for a human agent?",
),
"is_repeat_contact": Noul(
instructions="Has the customer contacted support about this before?",
criteria=NoulCriteria(
true="Mentions a prior attempt, ticket, or that they have asked before",
false="No sign of any previous contact",
),
),
},
)
print(response.answers["is_human_escalation"].noul)
print(response.answers["is_repeat_contact"].noul)
system_one 方法和 https://api.typesafe.ai/v1/systemone 端点,都是以 TypeSafe 的 AI 模型 System One 命名的。如何用 TypeSafe 构建介绍了该在代码里的哪些地方用它。
如果你在用 coding agent,先安装 TypeSafe agent skill,这样它就知道请求和响应长什么样。
响应结构
响应里的 answers 为每个问题各有一个条目,键是请求里用的 id:
{
"model": "jev-1.13.0",
"answers": {
"is_human_escalation": {
"type": "noul",
"noul": 0.99
},
"is_repeat_contact": {
"type": "noul",
"noul": 0.93
}
},
"usage": {
"input_tokens": 360,
"output_tokens": 39
}
}
这里两个答案都接近 1。客户说的是 “talk to a real person”,所以 is_human_escalation 是 0.99。“I have asked three times now” 与 is_repeat_contact 的 true 描述相符,所以是 0.93。
如何解读 Noul
这个数同时给出了答案和确定性。接近 1 是强「是」。接近 0 是强「否」。接近 0.5 表示模型给「是」和「否」的概率差不多。
下表是 jev-1.13.0 对不同客户消息在 is_human_escalation 问题上记录下来的答案:
| 状态 | noul |
|---|---|
| Thanks, that fixed it! | 0.02 |
| How do I reset my password? | 0.07 |
| I need this sorted today, whatever it takes. | 0.26 |
| Are you a bot? | 0.40 |
| Is there any way to speak to someone about my invoice? | 0.84 |
| I have asked three times now. Can I please just talk to a real person? | 0.99 |
前两条和后两条都很明确。“I need this sorted today” 很急,但从未要求找真人,得到 0.26。“Are you a bot?” 暗示想要真人,却没有明说,模型几乎五五开,给出 0.40。这两类消息,都需要靠代码里的阈值来做决定。
与 Choice 或 Score 不同,Noul 没有单独的 confidence 值。Noul 的概率分布只有两个结果,是和非,所以单个 noul 值就完整描述了它。Choice 或 Score 把概率铺在多个选项或档位上,confidence 则是对这种铺开的概括。
最常见的是,代码把 noul 取个阈值变成布尔值:
wants_human = response.answers["is_human_escalation"].noul > 0.9
if wants_human:
route_to_agent(ticket)
else:
route_to_bot(ticket)
阈值设在哪,取决于判断错了要付出多大代价。当「是」和「否」都同样容易处理时,用 0.5。把假「是」当真代价很大时(比如呼叫值班人员或退款),就调高。漏掉真「是」代价很大时(比如没能标出一个安全问题),就调低。中间的值可以交给人工,而不走任何一条代码分支。这和在置信度页里为 Choice 和 Score 答案描述的三路分流是一样的。
Noul 的取值从 0 到 1,但它不是你问的那件事本身的刻度。它是答案为「是」的概率。如果问题实际上问的是程度,这个值并不衡量程度。下面,对四位候选人问了 “Is the candidate strong in Python?”,并排给出一个含四个档位的 Score:no experience、some familiarity、regular use in a job、deep expertise。
| 候选人 | Noul: “Is the candidate strong in Python?” | Score: “How much Python experience does the candidate have?” |
|---|---|---|
| My experience is in Java and Go. I have not used Python. | 0.03 | 0.0 (No experience) |
| I have used Python occasionally for small scripts alongside my main Java work. | 0.14 | 1.0 (Some familiarity) |
| I used Python every day for two years in my last job, mostly data pipelines. | 0.81 | 2.05 (Regular use in a job) |
| I have written Python daily for eight years, including maintaining a large Django codebase. | 0.92 | 2.89 (Deep expertise) |
Noul 只判断一个命题——“strong”,值就是它成立的可能性有多大。你可以在代码里给 0 到 1 划分区间,比如把 0.3 到 0.7 当作 “some experience”,但模型看不到这些区间,所以答案里没有任何东西是照着它们判断的。中间值可能意味着中等经验,也可能意味着情况不明;而且候选人之间的间隔也不是你选的。Score 则逐条判断每个档位描述,所以每位候选人都落到了你写的某个档位之上或附近,返回的概率显示了模型如何在各档位之间分配它的判断。如果你不认同,就改写某个档位再跑一次。选择题类型解释了其中的区别。
如何写 Noul 问题
一个 Noul 只问一个是/否问题。如果一个问题里有两个条件,比如 “Is the customer angry and asking for a refund?”,模型就得同时判断两者,值的信息量就变小了。应该问两个 Noul,再在代码里把它们组合起来。
把问题措辞成「值高就代表是」。“Does the message contain personal data?” 很清晰。“Is the message free of personal data?” 则把含义反了过来,之后读它的代码会弄反。
用陈述句和用疑问句一样有效。对于 “The customer is requesting a refund”,接近 1 的值表示该陈述为真。拿你自己的数据把两种措辞都试一遍,看看哪种更好。
把「是」和「否」之间的边界划得毫不含糊。“Does this candidate have any Python experience?” 效果不错,因为 “any” 不留中间地带。当边界很微妙时,就加上带 true 和 false 描述的 criteria,就像上面 is_repeat_contact 那样。对大多数 Noul 来说 instructions 就够了,所以把你的问题分别在有、没有 criteria 的情况下各试一遍,留下在你的文档上答案更好的那种。
好做法:一次调用问多个问题
对于一张条件清单,就在一次请求里问许多 Noul 问题:每个条件一个问题,由代码来决定这些结果的组合意味着什么。问题是并行评估的,所以增加 Noul 几乎不改变响应时间。一起问多个问题有更详细的说明。
在代码里处理多个 Noul 答案
上面那个两问题的请求,已经足够让代码把消息分流了。下面的示例在客户要求找真人时升级给人工,在客户之前联系过时提高优先级。任一问题取到中间值时,就交给复核人,而不走代码分支:
from typesafe_sdk import Noul, NoulCriteria, TypeSafeClient
SUPPORT_QUESTIONS = {
"is_human_escalation": Noul(
instructions="Is the customer asking for a human agent?",
),
"is_repeat_contact": Noul(
instructions="Has the customer contacted support about this before?",
criteria=NoulCriteria(
true="Mentions a prior attempt, ticket, or that they have asked before",
false="No sign of any previous contact",
),
),
}
YES = 0.8
NO = 0.2
def route(message: str) -> None:
with TypeSafeClient() as client:
response = client.system_one(
model="jev-latest",
state=message,
questions=SUPPORT_QUESTIONS,
)
answers = response.answers
wants_human = answers["is_human_escalation"].noul
repeat = answers["is_repeat_contact"].noul
if NO < wants_human < YES or NO < repeat < YES:
# The model isn't sure either way. Let a person decide.
send_to_review(message)
return
priority = "high" if repeat > YES else "normal"
if wants_human > YES:
route_to_agent(message, priority=priority)
else:
route_to_bot(message, priority=priority)
对上面那条消息,is_human_escalation 的 noul 答案是 0.99,is_repeat_contact 是 0.93,所以代码以高优先级把它路由给人工客服。“How do I reset my password?” 这条消息在两个问题上都是 0.07,被路由给机器人。
阈值在你自己的代码里。如果复核人收到的消息太多,就缩小 NO 和 YES 之间的间距。如果放过去的错路由太多,就把它拉大。如果以后还需要知道消息是否提到付款、是否含有个人数据,就往 SUPPORT_QUESTIONS 里再加一个 Noul。请求次数仍然是一次。
结构化 instructions
Instructions 可以是一个对象而不是字符串,问题放在一个字段里,补充数据放在其它字段里。在问题中使用结构化数据讲了它什么时候有用。这里它用于一个用代码拼出来的问题:把一份刚到的简历,与候选人库里可能是同一个人的记录逐一比对。每条记录都原样放进一个 potential_duplicate 字段,question 对每条记录都一样,所有记录都在一次请求里检查。代码生成的问题键里含有每条记录的数据库 ID:
{
"state": {
"resume": {
"name": "John Smith",
"location": "Oakland, CA",
"summary": "Backend engineer with eight years of Python and Go experience.",
"experience": [
{
"employer": "Google",
"title": "Senior Backend Engineer",
"years": "2021-2025"
},
{
"employer": "Microsoft",
"title": "Software Engineer",
"years": "2017-2021"
}
]
}
},
"questions": {
"same_as_record_18": {
"type": "noul",
"instructions": {
"potential_duplicate": {
"name": "Jon Smith",
"location": "Oakland, CA",
"last_employer": "Google"
},
"question": "Is the resume for the same person as `potential_duplicate`?"
}
},
"same_as_record_42": {
"type": "noul",
"instructions": {
"potential_duplicate": {
"name": "John Smith",
"location": "Austin, TX",
"last_employer": "Lone Star Freight"
},
"question": "Is the resume for the same person as `potential_duplicate`?"
}
},
"same_as_record_77": {
"type": "noul",
"instructions": {
"potential_duplicate": {
"name": "John Smithers",
"location": "Oakland, CA",
"last_employer": "Bay Health Clinic"
},
"question": "Is the resume for the same person as `potential_duplicate`?"
}
}
}
}响应:
{
"model": "jev-1.13.0",
"answers": {
"same_as_record_18": {
"type": "noul",
"noul": 0.74
},
"same_as_record_42": {
"type": "noul",
"noul": 0.09
},
"same_as_record_77": {
"type": "noul",
"noul": 0.08
}
},
"usage": {
"input_tokens": 535,
"output_tokens": 58
}
}
每个答案都是「这份简历就是该记录里那个人」的概率。记录 18 名字拼写不同,但地点和雇主对得上,得到 0.74。记录 42 同名,但城市不同、雇主也不同,得到 0.09。记录 77 名字相近、地点相同,但雇主不同,得到 0.08。在代码里给每个值取阈值,就像在代码里处理多个 Noul 答案那样,把中间值交给人工。
用 Python SDK 时,问题是根据候选人记录拼出来的。问题文本固定,记录会变:
from typesafe_sdk import Noul, TypeSafeClient
SAME_PERSON = "Is the resume for the same person as `potential_duplicate`?"
def duplicate_questions(candidates: list[dict]) -> dict[str, Noul]:
"""One Noul per candidate record, all asking the same question."""
return {
f"same_as_record_{candidate['id']}": Noul(
instructions={
"potential_duplicate": {
"name": candidate["name"],
"location": candidate["location"],
"last_employer": candidate["last_employer"],
},
"question": SAME_PERSON,
},
)
for candidate in candidates
}
def find_duplicates(resume: dict, candidates: list[dict]) -> list[str]:
with TypeSafeClient() as client:
response = client.system_one(
model="jev-latest",
state={"resume": resume},
questions=duplicate_questions(candidates),
)
return [
question_id
for question_id, answer in response.answers.items()
if answer.noul > 0.7
]
结构化数据抽取级联 cookbook用结构化的 instructions 来校验一条抽取出来的记录。每个字段都得到同一组问题。每个问题的 instructions 对象把问题文本放在 main_question 属性里。还有 field_spec 和 extracted_field 两个属性,会随字段变化。
cookbook 里的 Noul
看看我们的 cookbooks,里面有使用 Noul 问题的应用: