Score
Score
Score 是一种 System One 问题类型,用于按有序的、带描述的档位给内容打分。答案包含一个 score、每个档位的概率和置信度。
当答案是你可以分步骤描述的谱系上的一个位置时,就用 Score。例如,一个 bug 有多严重、客户有多满意、候选人有多年 Python 经验。如果答案是固定选项集合中的一个、彼此之间没有顺序,就用 Choice。如果是是/否,就用 Noul。选择问题类型对比了这三种。
Score 的答案是 score 中沿你设定的档位的一个位置,可以落在两个档位之间。模型还会在 probabilities 中返回每个档位的概率,并为答案返回一个 confidence 值。
示例评分问题
How severe is the reported issue?
状态(待评估的内容)
The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari.
答案
各档位的概率
置信度
得分: 1.43
得分与置信度是怎么算出来的
得分:
把每个档位的编号乘以它的概率,再相加:
0 × 0 + 1 × 0.57 + 2 × 0.43 ≈ 1.43
置信度
TypeSafe 根据概率在各档位上的分散程度算出它。全部集中在一档是 1.0;分布越平均,置信度越低。
How formal is this outfit based on the description?
状态(待评估的内容)
A navy blazer over a plain white T-shirt, dark jeans, and clean leather loafers. No tie.
答案
各档位的概率
置信度
得分: 1.86
得分与置信度是怎么算出来的
得分:
把每个档位的编号乘以它的概率,再相加:
0 × 0 + 1 × 0.14 + 2 × 0.86 + 3 × 0 + 4 × 0 ≈ 1.86
置信度
TypeSafe 根据概率在各档位上的分散程度算出它。全部集中在一档是 1.0;分布越平均,置信度越低。
How relevant is this candidate's experience to the job posting?
状态(待评估的内容)
Job posting: Senior backend engineer building Python APIs and PostgreSQL services. Candidate: Three years building Django REST APIs with PostgreSQL, preceded by two years in frontend JavaScript. Has owned small services but has not led a backend team.
答案
各档位的概率
置信度
得分: 2.52
得分与置信度是怎么算出来的
得分:
把每个档位的编号乘以它的概率,再相加:
0 × 0 + 1 × 0 + 2 × 0.48 + 3 × 0.52 ≈ 2.52
置信度
TypeSafe 根据概率在各档位上的分散程度算出它。全部集中在一档是 1.0;分布越平均,置信度越低。
How frustrated is the customer?
状态(待评估的内容)
Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. This is the third time I'm writing in and honestly I'm done. Steps: open any report, click Export, choose PDF. Chrome 128 on macOS.
答案
各档位的概率
置信度
得分: 1.26
得分与置信度是怎么算出来的
得分:
把每个档位的编号乘以它的概率,再相加:
0 × 0 + 1 × 0.74 + 2 × 0.26 ≈ 1.26
置信度
TypeSafe 根据概率在各档位上的分散程度算出它。全部集中在一档是 1.0;分布越平均,置信度越低。
How much does the report give an engineer to work with?
状态(待评估的内容)
Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. This is the third time I'm writing in and honestly I'm done. Steps: open any report, click Export, choose PDF. Chrome 128 on macOS.
答案
各档位的概率
置信度
得分: 3.00
得分与置信度是怎么算出来的
得分:
把每个档位的编号乘以它的概率,再相加:
0 × 0 + 1 × 0 + 2 × 0 + 3 × 1 ≈ 3.00
置信度
TypeSafe 根据概率在各档位上的分散程度算出它。全部集中在一档是 1.0;分布越平均,置信度越低。
每一步前面的数字是位置,详见档位。
请求结构
发往 TypeSafe API 的 POST 请求体,和任何其它问题类型一样有三个顶层字段:state,即要评估的内容;model;以及 questions。每个 Score 问题有以下字段:
type:始终为"score"。instructions:模型要回答的问题。即它在给什么打分。criteria:按顺序排列的档位描述数组,从量表低端到高端。至少应有两个档位;API 最多接受 10 个。
下面这个请求,state 是一份 bug 报告,问题是这个 bug 有多严重:
{
"state": "The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari.",
"questions": {
"bug_severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists"
]
}
}
}问题 id 由你选,这里是 bug_severity。这个 id 不会发送给模型。答案会以同一个 id 返回。
档位
criteria 中的每一项是一个档位:可能答案谱系上的一个点,用文字描述。档位的编号就是它在 criteria 数组中的位置,从 0 开始,所以上面三项分别是档位 0、1、2。数组的顺序就是编号。
模型只拿到这些描述,别的什么都没有,而且每个档位都单独对照 state 来判断。
响应中的 score 是档位谱系上的一个位置。三档量表上,它的范围是 0 到 2,并且可以落在两个档位之间。
我们的客户端 SDK提供类型化的问题。在 Python 里,同一个问题就是一个 Score:
from typesafe_sdk import Score, TypeSafeClient
with TypeSafeClient() as client:
response = client.system_one(
state="The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari.",
questions={
"bug_severity": Score(
instructions="How severe is the reported issue?",
criteria=[
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists",
],
),
},
)
print(response.answers["bug_severity"].score)
用 system_one 方法或 https://api.typesafe.ai/v1/systemone 端点来调用 System One 模型。model 字段选择由哪个模型处理请求。如何用 TypeSafe 构建讲了在代码里的什么位置调用它。
使用我们的某个客户端 SDK,或直接调用 TypeSafe API。如果由编码 agent 替你写集成,先安装 TypeSafe agent skill,这样它就知道请求和响应的结构。
响应结构
响应在 answers 中每个问题一条,放在请求里的 id 之下。这是上面示例请求的响应:
{
"model": "jev-1.13.0",
"answers": {
"bug_severity": {
"type": "score",
"score": 1.43,
"confidence": 0.35,
"legend": {
"0": "Cosmetic; no impact to functionality",
"1": "Broken or degraded feature, but workaround exists",
"2": "Blocking issue; no workaround exists"
},
"probabilities": {
"0": 0.0,
"1": 0.57,
"2": 0.43
}
}
},
"usage": {
"input_tokens": 332,
"output_tokens": 18
}
}
每个 Score 答案有五个值:
type:TypeSafe 问题的类型。probabilities:每个档位的概率,以档位编号的字符串为键。所有值之和为 1。score:档位编号数轴上的位置,从 0 到最高档位编号(这里是 2)。它是每个档位编号乘以其概率再相加:0 x 0.0 + 1 x 0.57 + 2 x 0.43 = 1.43。legend:每个档位编号映射回它的描述。confidence:一个 0 到 1 的数字,由probabilities的分散程度算出。概率集中在单个档位意味着高置信度。概率分散在多个档位意味着低置信度。
1.43 的 score 意味着模型在档位 1 和 2 之间摇摆,略偏向档位 1。这和报告吻合:导出坏了,换到 Chrome 对大多数客户算是变通办法,但对只用 Safari 的那部分客户不是。模型把 0.57 放在“有变通办法”上,0.43 放在“没有变通办法”上,置信度是 0.35,因为它是分裂的。
用 Python SDK 时,ScoreAnswer 有 score、confidence、probabilities 和 legend 这几个类型化字段。SDK 用整数档位而不是字符串给 probabilities 和 legend 作键。
阅读 Score
我们来看看不同输入下 score 如何变化。例如,用上面请求中的问题及其档位:
"How severe is the reported issue?"
→ 0: Cosmetic; no impact to functionality
→ 1: Broken or degraded feature, but workaround exists
→ 2: Blocking issue; no workaround exists
可以看到不同的 bug 报告如何改变 score:
probabilities | |||||
|---|---|---|---|---|---|
| 状态 | score | confidence | 档位 0 | 档位 1 | 档位 2 |
| The export button is misaligned by a few pixels on the settings page. | 0.0 | 1.0 | 1.0 | 0.0 | 0.0 |
| The PDF export button does nothing when clicked. I can still export to CSV and convert it myself, but that takes ages. | 1.0 | 1.0 | 0.0 | 1.0 | 0.0 |
| Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. | 1.11 | 0.84 | 0.0 | 0.89 | 0.11 |
| The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari. | 1.43 | 0.35 | 0.0 | 0.57 | 0.43 |
| Nobody on our team can log in since this morning. We get a 500 error on every attempt. | 2.0 | 1.0 | 0.0 | 0.0 | 1.0 |
在这些例子里,置信度 1.0 意味着返回的分布把所有概率都放在一个档位上。这描述的是模型的答案,并不保证答案正确。
score 是档位编号的概率加权平均。在第三和第四个例子里,概率分裂在档位 1 和 2 之间。档位 2 权重越大,score 越高。它衡量的不是没有变通办法的客户占比。
不同的分布可以产生相同的 score。1.0 的 score 可能意味着全部概率都在档位 1,也可能意味着档位 0 和 2 各占一半。把 probabilities 和 confidence 与 score 一起读,才能区分这些情况。
分数形式的 score 是一个位置。你可以用它按严重程度给报告排序,或在代码需要一个确定结果时四舍五入到最近的档位。我们的实体对齐 cookbook有一个四舍五入到最近档位以做决策的例子。
Score 上偏低的置信度通常意味着三种情况之一。这些档位对这个 state 有重叠,问题衡量了不止一件事,或者 state 说得不够多、无法定位。我们的置信度文档讲了如何在代码里使用它。
写出好的档位
描述情境,而不是程度。“Broken or degraded feature, but workaround exists”给了模型可以拿 state 去匹配的东西。“Moderately severe”则没有。具体的描述能帮模型区分档位。把答案和已知例子对照检查;单凭置信度更高,并不能说明描述更好。
每个档位都是单独评估的。模型看不到档位的编号,也看不到相邻档位,所以“比上一档更严重”对它没有任何意义,描述或 instructions 里的数字也帮不上忙。下面是在上面表格里那条“按钮错位”报告上,档位只有数字时会发生什么:
instructions: "Rate severity from 0 to 2, where 2 is worst"
criteria: ["0", "1", "2"]
→ score 0.55, confidence 0.33, probabilities 0: 0.45, 1: 0.55, 2: 0.0
同一份报告,用那三个带描述的档位,得分 0.0,置信度 1.0。只用数字时,模型没有可匹配的东西,就把概率分裂在 0 和 1 之间。
用尽可能多的、你能清楚区分开的档位,最多 10 个。三个也可以。别加你没法清楚区分的档位。
让每个 Score 问题只针对一个维度。如果一段描述说“punctual and smart and experienced”,那这个问题就在衡量三件事,而在一个维度高、另一个维度低的输入就无法定位。置信度下降,score 的意义也变弱。把它拆成每个维度一个 Score 问题,在代码里组合,下一节会演示。
如果你量表的顶端有一个需要区别对待的罕见极端情况,就给它单独一个档位。一个以“very angry”结尾的情感量表,可以加上“abusive or threatening”。没有这个档位,两条消息可能都得到接近顶端的 score。单凭 score 可能无法区分它们。
如果完全没有中间地带,答案只是少数几个离散类别之一,那就改用 Choice,或把问题拆成几个 Noul 问题。一定要用你自己的数据测试你的档位。同一个量表的两种措辞,在你的数据上可能表现不同。
把一个复杂判断拆成几个 Score 问题
一个依赖多件事的复杂判断,最好拆成每件事一个 Score 问题。然后你可以在代码里把 TypeSafe 返回的各个 Score 组合起来,做出这个判断。有些 Score 问题可能比其它的更重要,所以给每个 Score 问题配一个表示相对重要性的权重。权重由你定。当组合结果和你的团队会做出的决定不符时,在代码里改权重再跑一次。把这些 Score 问题放在一个请求里发送。它们会并行评估。增加问题几乎不改变响应时间,只多花几个问题 token;见一次提多个问题。
下面的请求是上面表格里那条转圈的工单,上下文更多一些。它提三个 Score 问题:bug 有多严重、客户有多沮丧、这份报告给工程师多少可用的信息。
{
"state": "Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. This is the third time I'm writing in and honestly I'm done. Steps: open any report, click Export, choose PDF. Chrome 128 on macOS.",
"questions": {
"severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists"
]
},
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer?",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language or threatening to leave"
]
},
"report_quality": {
"type": "score",
"instructions": "How much does the report give an engineer to work with?",
"criteria": [
"No detail; just says something is broken",
"Names the feature but no steps or environment",
"Steps to reproduce or environment, but not both",
"Steps to reproduce and environment"
]
}
}
}TypeSafe 的响应:
{
"model": "jev-1.13.0",
"answers": {
"severity": {
"type": "score",
"score": 1.24,
"confidence": 0.64,
"legend": {
"0": "Cosmetic; no impact to functionality",
"1": "Broken or degraded feature, but workaround exists",
"2": "Blocking issue; no workaround exists"
},
"probabilities": {
"0": 0.0,
"1": 0.76,
"2": 0.24
}
},
"frustration": {
"type": "score",
"score": 1.28,
"confidence": 0.58,
"legend": {
"0": "Calm, just stating facts",
"1": "Frustrated but civil",
"2": "Very angry, strong language or threatening to leave"
},
"probabilities": {
"0": 0.0,
"1": 0.72,
"2": 0.28
}
},
"report_quality": {
"type": "score",
"score": 3.0,
"confidence": 1.0,
"legend": {
"0": "No detail; just says something is broken",
"1": "Names the feature but no steps or environment",
"2": "Steps to reproduce or environment, but not both",
"3": "Steps to reproduce and environment"
},
"probabilities": {
"0": 0.0,
"1": 0.0,
"2": 0.0,
"3": 1.0
}
}
},
"usage": {
"input_tokens": 468,
"output_tokens": 43
}
}
每个问题都单独对照工单作答,并给出一个 score:
severity为 1.24,置信度 0.64。读法和开头的例子一样:导出坏了,有些人有变通办法。frustration为 1.28,置信度 0.58。措辞还算客气,但“third time”和“I’m done”把一部分分数推向最高档,于是模型把 0.72 和 0.28 分裂在“frustrated but civil”和“very angry”之间。对这条工单来说,两个档位有重叠,所以置信度中等。report_quality为 3.0,置信度 1.0。复现步骤和浏览器版本都写明了。
这三个量表的长度不同,所以在组合之前,先归一化每个 score。四档量表返回 0 到 3,三档量表返回 0 到 2,所以一个的满分比另一个的满分大。把每个 score 除以它的最高档位编号 len(criteria) - 1,就把所有 score 都放到 0 到 1 上。这样权重才名副其实:severity 给 0.6、frustration 给 0.3,意味着 severity 的分量是 frustration 的两倍。
下面的 TypeSafe Python SDK 代码提这三个问题,归一化每个 score,并用一个示例的优先级计算把它们组合起来:
from typesafe_sdk import Score, TypeSafeClient
TRIAGE_QUESTIONS = {
"severity": Score(
instructions="How severe is the reported issue?",
criteria=[
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists",
],
),
"frustration": Score(
instructions="How frustrated is the customer?",
criteria=[
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language or threatening to leave",
],
),
"report_quality": Score(
instructions="How much does the report give an engineer to work with?",
criteria=[
"No detail; just says something is broken",
"Names the feature but no steps or environment",
"Steps to reproduce or environment, but not both",
"Steps to reproduce and environment",
],
),
}
def normalized(answers, question_id: str) -> float:
"""Put a score on 0 to 1 by dividing by its top level number."""
top_level = len(TRIAGE_QUESTIONS[question_id].criteria) - 1
return answers[question_id].score / top_level
def priority(ticket: str) -> float:
with TypeSafeClient() as client:
response = client.system_one(
state=ticket,
questions=TRIAGE_QUESTIONS,
)
answers = response.answers
severity = normalized(answers, "severity")
frustration = normalized(answers, "frustration")
report_quality = normalized(answers, "report_quality")
# A detailed report helps an engineer investigate, so it raises priority a little.
return 0.6 * severity + 0.3 * frustration + 0.1 * report_quality
对于上面的示例响应,归一化后的分数是 severity 0.62、frustration 0.64、report quality 1.0。优先级是 0.6 × 0.62 + 0.3 × 0.64 + 0.1 × 1.0 = 0.664,四舍五入为 0.66。
权重存在于你的代码里,所以你能确切看到这个数字是怎么来的,并在排序和你的团队会做的决定不符时改它。如果之后需要更多 Score 问题,把它们加进 TRIAGE_QUESTIONS。请求数仍然是一个。这种把复杂判断拆成一个个独立的 Score、再用代码里的权重组合起来的手法,叫作组合评分模式。
结构化档位描述
先给每个档位一段基本的文本描述。当模型在你认为清晰的输入上总是打在两个相邻档位之间时,就把每个档位从字符串换成一个对象,一个字段说明这个档位涵盖什么,另一个字段放几个示例情境。每个档位用相同的字段名,这样模型才能同类相比。
下面的请求是我们之前用过的那条转圈工单,但每个档位都带了示例:
{
"state": "Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too.",
"questions": {
"bug_severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
{
"what": "Cosmetic; no impact to functionality",
"examples": [
"typo in a label",
"misaligned icon"
]
},
{
"what": "Broken or degraded feature, but workaround exists",
"examples": [
"export fails in one browser but works in another"
]
},
{
"what": "Blocking issue; no workaround exists",
"examples": [
"cannot log in",
"data loss"
]
}
]
}
}
}响应:
{
"model": "jev-1.13.0",
"answers": {
"bug_severity": {
"type": "score",
"score": 1.09,
"confidence": 0.87,
"legend": {
"0": {
"what": "Cosmetic; no impact to functionality",
"examples": [
"typo in a label",
"misaligned icon"
]
},
"1": {
"what": "Broken or degraded feature, but workaround exists",
"examples": [
"export fails in one browser but works in another"
]
},
"2": {
"what": "Blocking issue; no workaround exists",
"examples": [
"cannot log in",
"data loss"
]
}
},
"probabilities": {
"0": 0.0,
"1": 0.91,
"2": 0.09
}
}
},
"usage": {
"input_tokens": 379,
"output_tokens": 18
}
}
用纯字符串时,这条工单得分 1.11,置信度 0.84。带示例时得分 1.09,置信度 0.87,变化很小,因为纯字符串已经把它放得很好了。当纯字符串让模型摇摆时,效果会更大,如下一张表所示。
示例会引导模型,只有当它们看起来像你的真实输入时才有帮助。下表是开头那条 Safari 报告,配三组不同的档位对象:
| 档位描述 | score |
confidence |
|---|---|---|
| 纯字符串:没有带示例的对象 | 1.43 | 0.35 |
| 加入的 examples 数组带一个有用的示例:“export fails in one browser but works in another” | 1.03 | 0.96 |
| 加入的 examples 数组带的示例与浏览器无关:“search fails, but browsing categories still works” | 1.43 | 0.35 |
在这个对比里,匹配的示例把几乎全部概率集中到一个档位上。无关的示例给出的结果和纯字符串一样。更高的置信度并不能确立哪个答案是对的。选择那些你已知预期档位的示例,然后在一批独立的输入上测试修改后的描述,再决定保留。