Score
Score
Score は、順序付きで説明的なレベルに対してコンテンツを評価するための System One の質問タイプです。答えにはスコア、各レベルの確率、信頼度が含まれます。
答えが、段階で記述できるスペクトル上の位置であるときに Score を使います。たとえば、バグがどれほど深刻か、顧客がどれほど満足しているか、候補者がどれだけ Python の経験があるか。答えが固定の選択肢集合のいずれかで、それらの間に順序がないなら Choice を使います。はい か いいえ なら Noul を使います。質問タイプを選ぶで三つを比較しています。
Score の答えは、score にあるレベルに沿った位置で、二つのレベルの間に落ちることもあります。モデルは probabilities にすべてのレベルの確率も返し、答えの confidence 値も返します。
スコアの質問の例
How severe is the reported issue?
状態(評価する内容)
The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari.
回答
各段階の確率
信頼度
スコア: 1.43
スコアと信頼度の計算方法
スコア:
各段階の番号にその確率を掛けて足し合わせます:
0 × 0 + 1 × 0.57 + 2 × 0.43 ≈ 1.43
信頼度
TypeSafe は確率が各段階にどれだけ散っているかからこれを算出します。1 つの段階に集中していれば 1.0、均等に散るほど低くなります。
How formal is this outfit based on the description?
状態(評価する内容)
A navy blazer over a plain white T-shirt, dark jeans, and clean leather loafers. No tie.
回答
各段階の確率
信頼度
スコア: 1.86
スコアと信頼度の計算方法
スコア:
各段階の番号にその確率を掛けて足し合わせます:
0 × 0 + 1 × 0.14 + 2 × 0.86 + 3 × 0 + 4 × 0 ≈ 1.86
信頼度
TypeSafe は確率が各段階にどれだけ散っているかからこれを算出します。1 つの段階に集中していれば 1.0、均等に散るほど低くなります。
How relevant is this candidate's experience to the job posting?
状態(評価する内容)
Job posting: Senior backend engineer building Python APIs and PostgreSQL services. Candidate: Three years building Django REST APIs with PostgreSQL, preceded by two years in frontend JavaScript. Has owned small services but has not led a backend team.
回答
各段階の確率
信頼度
スコア: 2.52
スコアと信頼度の計算方法
スコア:
各段階の番号にその確率を掛けて足し合わせます:
0 × 0 + 1 × 0 + 2 × 0.48 + 3 × 0.52 ≈ 2.52
信頼度
TypeSafe は確率が各段階にどれだけ散っているかからこれを算出します。1 つの段階に集中していれば 1.0、均等に散るほど低くなります。
How frustrated is the customer?
状態(評価する内容)
Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. This is the third time I'm writing in and honestly I'm done. Steps: open any report, click Export, choose PDF. Chrome 128 on macOS.
回答
各段階の確率
信頼度
スコア: 1.26
スコアと信頼度の計算方法
スコア:
各段階の番号にその確率を掛けて足し合わせます:
0 × 0 + 1 × 0.74 + 2 × 0.26 ≈ 1.26
信頼度
TypeSafe は確率が各段階にどれだけ散っているかからこれを算出します。1 つの段階に集中していれば 1.0、均等に散るほど低くなります。
How much does the report give an engineer to work with?
状態(評価する内容)
Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. This is the third time I'm writing in and honestly I'm done. Steps: open any report, click Export, choose PDF. Chrome 128 on macOS.
回答
各段階の確率
信頼度
スコア: 3.00
スコアと信頼度の計算方法
スコア:
各段階の番号にその確率を掛けて足し合わせます:
0 × 0 + 1 × 0 + 2 × 0 + 3 × 1 ≈ 3.00
信頼度
TypeSafe は確率が各段階にどれだけ散っているかからこれを算出します。1 つの段階に集中していれば 1.0、均等に散るほど低くなります。
各段の前にある数字は位置で、レベルで説明しています。
リクエストの構造
TypeSafe API への POST リクエストボディは、他のどの質問タイプとも同じ三つのトップレベルフィールドを持ちます。評価する内容である state、model、questions です。各 Score 質問は次のフィールドを持ちます:
type:常に"score"。instructions:モデルが答える質問。何を評価しているか。criteria:スケールの低い端から高い端までの、レベルの説明の順序付き配列。少なくとも二つのレベルが必要です。API は最大 10 まで受け付けます。
以下は、状態がバグ報告で、質問がバグの深刻さであるリクエストです:
{
"state": "The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari.",
"questions": {
"bug_severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists"
]
}
}
}質問 ID は自分で選び、この場合は bug_severity です。この ID はモデルには送られません。答えは同じ ID の下に返ります。
レベル
criteria の各エントリは一つのレベルです。考えられる答えのスペクトル上の一点を、言葉で記述したものです。レベルの番号は criteria 配列内での位置で、0 から始まります。したがって上の三つのエントリはレベル 0、1、2 です。配列の順序が番号付けです。
モデルが受け取るのは説明だけで、それ以外は何もありません。各レベルは状態に対して単独で判断されます。
応答の score はレベルスペクトル上の位置です。三つのレベルのスケールでは 0 から 2 の範囲で、二つのレベルの間に着地することもあります。
クライアント SDK は型付きの質問を提供します。Python では、同じ質問は Score です:
from typesafe_sdk import Score, TypeSafeClient
with TypeSafeClient() as client:
response = client.system_one(
state="The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari.",
questions={
"bug_severity": Score(
instructions="How severe is the reported issue?",
criteria=[
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists",
],
),
},
)
print(response.answers["bug_severity"].score)
system_one メソッドまたは https://api.typesafe.ai/v1/systemone エンドポイントで System One モデルを呼び出します。model フィールドがどのモデルがリクエストを処理するかを選びます。TypeSafe での構築方法で、コードのどこで呼ぶかを説明しています。
クライアント SDK のいずれかを使うか、TypeSafe API を直接呼び出します。コーディングエージェントに連携コードを書いてもらう場合は、先に TypeSafe agent skill をインストールして、リクエストと応答の形を把握させてください。
応答の構造
応答は、リクエストの ID の下に、質問ごとに answers のエントリを一つ持ちます。これは上の例のリクエストへの応答です:
{
"model": "jev-1.13.0",
"answers": {
"bug_severity": {
"type": "score",
"score": 1.43,
"confidence": 0.35,
"legend": {
"0": "Cosmetic; no impact to functionality",
"1": "Broken or degraded feature, but workaround exists",
"2": "Blocking issue; no workaround exists"
},
"probabilities": {
"0": 0.0,
"1": 0.57,
"2": 0.43
}
}
},
"usage": {
"input_tokens": 332,
"output_tokens": 18
}
}
各 Score の答えには五つの値があります:
type:TypeSafe の質問タイプ。probabilities:各レベルの確率で、レベル番号を文字列としたキーで示されます。すべての値の合計は 1 です。score:レベル番号の数直線上の位置で、0 から一番上のレベル番号まで、ここでは 2 です。各レベル番号にその確率を掛けて足し合わせたものです。0 x 0.0 + 1 x 0.57 + 2 x 0.43 = 1.43。legend:各レベル番号をその説明に対応付けたもの。confidence:probabilitiesがどれほど広がっているかから計算される 0 から 1 の数値。一つのレベルに単一のピークがあると高い信頼度を意味します。確率が複数のレベルに広がっていると低い信頼度を意味します。
スコア 1.43 は、モデルがレベル 1 と 2 の間で割れており、レベル 1 に傾いていることを意味します。これは報告内容と一致します。エクスポートは壊れており、Chrome に切り替えるのがほとんどの顧客にとっては回避策ですが、Safari しか使わない顧客にはそうではありません。モデルは「回避策あり」に 0.57、「回避策なし」に 0.43 を置き、割れているため信頼度は 0.35 です。
Python SDK を使うと、ScoreAnswer は score、confidence、probabilities、legend を型付きのフィールドとして持ちます。SDK は probabilities と legend のキーを、文字列ではなく整数のレベルにします。
Score を読む
異なる入力でスコアがどう変わるかを見てみましょう。たとえば、上のリクエストの質問とそのレベルを使います:
"How severe is the reported issue?"
→ 0: Cosmetic; no impact to functionality
→ 1: Broken or degraded feature, but workaround exists
→ 2: Blocking issue; no workaround exists
異なるバグ報告がスコアをどう変えるかが分かります:
probabilities | |||||
|---|---|---|---|---|---|
| 状態 | score | confidence | レベル 0 | レベル 1 | レベル 2 |
| The export button is misaligned by a few pixels on the settings page. | 0.0 | 1.0 | 1.0 | 0.0 | 0.0 |
| The PDF export button does nothing when clicked. I can still export to CSV and convert it myself, but that takes ages. | 1.0 | 1.0 | 0.0 | 1.0 | 0.0 |
| Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. | 1.11 | 0.84 | 0.0 | 0.89 | 0.11 |
| The export button crashes the settings page in Safari. It works in Chrome, but a few of our customers only use Safari. | 1.43 | 0.35 | 0.0 | 0.57 | 0.43 |
| Nobody on our team can log in since this morning. We get a 500 error on every attempt. | 2.0 | 1.0 | 0.0 | 0.0 | 1.0 |
これらの例で、信頼度 1.0 は、返された分布がその確率をすべて一つのレベルに置いていることを意味します。これはモデルの答えを説明するもので、答えが正しいことの保証ではありません。
スコアはレベル番号の確率加重平均です。第三と第四の例では、確率がレベル 1 と 2 の間に分かれています。レベル 2 の重みが大きいほどスコアが上がります。回避策のない顧客の割合を測っているのではありません。
異なる分布が同じスコアを生むことがあります。スコア 1.0 は、確率がすべてレベル 1 にあることも、レベル 0 と 2 に半分ずつあることも意味しえます。これらの場合を区別するには、probabilities と confidence をスコアと一緒に読んでください。
小数のスコアは位置です。これを使って報告を深刻さでランキングしたり、コードが一つの結果を必要とするときは最も近いレベルに丸めたりできます。私たちのエンティティアラインメントクックブックに、意思決定のために最も近いレベルに丸める例があります。
Score の低い信頼度は通常、三つのうちのどれかを意味します。この状態に対してレベルが重なっているか、質問が複数のことを測っているか、状態に位置付けるだけの情報がないかです。私たちの信頼度ドキュメントで、コードでの使い方を扱っています。
良いレベルを書く
程度ではなく状況を記述します。「Broken or degraded feature, but workaround exists」は、モデルが状態を照らし合わせるものを与えます。「Moderately severe」は与えません。具体的な説明はモデルがレベルを区別する助けになります。答えを既知の例と照合してください。信頼度が高いことだけでは、説明がより良いことにはなりません。
各レベルは別々に評価されます。モデルはレベルの番号も隣のレベルも見ないので、「前のレベルより悪い」はモデルにとって何の意味もなく、説明や instructions の中の数字も助けになりません。以下は、上の表のボタンずれの報告に対して、レベルが数字だけの場合に起きることです:
instructions: "Rate severity from 0 to 2, where 2 is worst"
criteria: ["0", "1", "2"]
→ score 0.55, confidence 0.33, probabilities 0: 0.45, 1: 0.55, 2: 0.0
同じ報告を三つの説明的なレベルで評価すると、信頼度 1.0 で 0.0 になります。数字だけだと、モデルには照らし合わせるものが何もなく、確率を 0 と 1 に分けます。
区別して記述できるだけの数のレベルを使い、最大 10 までにします。三つで十分です。区別して記述できないレベルは加えないでください。
各 Score 質問は一次元に保ってください。説明が「punctual and smart and experienced」と言っていると、質問は三つのことを測っていることになり、あるものは高く別のものは低い入力は位置付けられません。信頼度が下がり、スコアの意味が薄れます。次節で示すように、事柄ごとに一つの Score 質問に分割し、コードで組み合わせます。
スケールの一番上に、別扱いが必要なまれな極端なケースがあるなら、それに独自のレベルを与えます。「very angry」で終わる感情スケールには「abusive or threatening」を加えられます。そのレベルがなければ、両方のメッセージが上位近くのスコアを受け取るかもしれません。スコアだけでは両者を区別できないことがあります。
中間が一切なく、答えがいくつかの離散的なカテゴリのいずれかなら、代わりに Choice を使うか、質問を複数の Noul 質問に分割してください。自分のデータでレベルをテストすることが重要です。同じスケールでも二つの言い回しは、あなたのデータ上で違う振る舞いをすることがあります。
複雑な判断を複数の Score 質問に分割する
複数の事柄に依存する複雑な判断は、事柄ごとに一つの Score 質問に分割するのが最善です。そうすれば、TypeSafe から返された Score をコード内で組み合わせて判断を下せます。ある Score 質問が他より重要なこともあるので、各 Score 質問に相対的な重要度の重みを与えます。重みはあなたのものです。組み合わせた結果がチームの出す判断と合わないときは、コード内で重みを変えて再実行します。Score 質問は一つのリクエストで送ります。並列に評価されます。質問を増やしても応答時間はほとんど変わらず、質問のトークンが少し増えるだけです。複数の質問をまとめて尋ねるを参照してください。
下のリクエストは、上の表のスピナーのチケットにコンテキストを少し加えたものです。三つの Score 質問を尋ねます。バグがどれほど深刻か、顧客がどれほど不満か、報告がエンジニアにどれだけ材料を与えているか。
{
"state": "Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too. This is the third time I'm writing in and honestly I'm done. Steps: open any report, click Export, choose PDF. Chrome 128 on macOS.",
"questions": {
"severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists"
]
},
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer?",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language or threatening to leave"
]
},
"report_quality": {
"type": "score",
"instructions": "How much does the report give an engineer to work with?",
"criteria": [
"No detail; just says something is broken",
"Names the feature but no steps or environment",
"Steps to reproduce or environment, but not both",
"Steps to reproduce and environment"
]
}
}
}TypeSafe の応答:
{
"model": "jev-1.13.0",
"answers": {
"severity": {
"type": "score",
"score": 1.24,
"confidence": 0.64,
"legend": {
"0": "Cosmetic; no impact to functionality",
"1": "Broken or degraded feature, but workaround exists",
"2": "Blocking issue; no workaround exists"
},
"probabilities": {
"0": 0.0,
"1": 0.76,
"2": 0.24
}
},
"frustration": {
"type": "score",
"score": 1.28,
"confidence": 0.58,
"legend": {
"0": "Calm, just stating facts",
"1": "Frustrated but civil",
"2": "Very angry, strong language or threatening to leave"
},
"probabilities": {
"0": 0.0,
"1": 0.72,
"2": 0.28
}
},
"report_quality": {
"type": "score",
"score": 3.0,
"confidence": 1.0,
"legend": {
"0": "No detail; just says something is broken",
"1": "Names the feature but no steps or environment",
"2": "Steps to reproduce or environment, but not both",
"3": "Steps to reproduce and environment"
},
"probabilities": {
"0": 0.0,
"1": 0.0,
"2": 0.0,
"3": 1.0
}
}
},
"usage": {
"input_tokens": 468,
"output_tokens": 43
}
}
各質問はそのチケットに対して単独で答えられ、スコアが与えられます:
severityは信頼度 0.64 で 1.24 です。冒頭の例と同じ読み方です。エクスポートは壊れており、回避策がある人もいます。frustrationは信頼度 0.58 で 1.28 です。文言は穏やかですが、「三度目」と「もうたくさんだ」がスコアの一部を一番上のレベルへ寄せるので、モデルは「frustrated but civil」と「very angry」に 0.72 と 0.28 を分けています。このチケットでは二つのレベルが重なっており、だから信頼度が中程度なのです。report_qualityは信頼度 1.0 で 3.0 です。手順とブラウザのバージョンが両方書かれています。
三つのスケールは長さが違うので、それらを組み合わせる前に各スコアを正規化します。四つのレベルのスケールは 0 から 3 を返し、三つのレベルのスケールは 0 から 2 を返すので、一方の最高スコアは他方の最高スコアより大きくなります。各スコアをその一番上のレベル番号 len(criteria) - 1 で割り、すべてのスコアを 0 から 1 の範囲に置きます。すると重みは言ったとおりの意味になります。severity に 0.6、frustration に 0.3 なら、severity が二倍の重みで数えられます。
下の TypeSafe Python SDK のコードは、三つの質問を尋ね、各スコアを正規化し、例の優先度計算でそれらを組み合わせます:
from typesafe_sdk import Score, TypeSafeClient
TRIAGE_QUESTIONS = {
"severity": Score(
instructions="How severe is the reported issue?",
criteria=[
"Cosmetic; no impact to functionality",
"Broken or degraded feature, but workaround exists",
"Blocking issue; no workaround exists",
],
),
"frustration": Score(
instructions="How frustrated is the customer?",
criteria=[
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language or threatening to leave",
],
),
"report_quality": Score(
instructions="How much does the report give an engineer to work with?",
criteria=[
"No detail; just says something is broken",
"Names the feature but no steps or environment",
"Steps to reproduce or environment, but not both",
"Steps to reproduce and environment",
],
),
}
def normalized(answers, question_id: str) -> float:
"""Put a score on 0 to 1 by dividing by its top level number."""
top_level = len(TRIAGE_QUESTIONS[question_id].criteria) - 1
return answers[question_id].score / top_level
def priority(ticket: str) -> float:
with TypeSafeClient() as client:
response = client.system_one(
state=ticket,
questions=TRIAGE_QUESTIONS,
)
answers = response.answers
severity = normalized(answers, "severity")
frustration = normalized(answers, "frustration")
report_quality = normalized(answers, "report_quality")
# A detailed report helps an engineer investigate, so it raises priority a little.
return 0.6 * severity + 0.3 * frustration + 0.1 * report_quality
上の例の応答に対して、正規化されたスコアは severity が 0.62、frustration が 0.64、report quality が 1.0 です。優先度は 0.6 × 0.62 + 0.3 × 0.64 + 0.1 × 1.0 = 0.664 で、丸めると 0.66 です。
重みはあなたのコードの中にあるので、その数値がどう作られているかを正確に見られ、ランキングがチームの行動と合わないときに変えられます。後でさらに Score 質問が必要になったら、TRIAGE_QUESTIONS に加えます。リクエスト数は 1 のままです。複雑な判断を別々の Score に分解し、コード内で重みを付けて組み合わせるこの手法は、複合スコアリングパターンと呼ばれます。
構造化されたレベルの説明
まず各レベルに基本的なテキストの説明から始めます。明確だと思う入力でモデルが隣り合う二つのレベルの間でスコアを付け続けるときは、各レベルに文字列ではなくオブジェクトを与え、そのレベルがカバーするもののフィールドと、いくつかの例の状況のフィールドを付けます。モデルが同じものを同じものとして比較できるよう、すべてのレベルで同じフィールド名を使ってください。
下のリクエストは、先に使ったスピナーのチケットに、各レベルに例を付けたものです:
{
"state": "Export to PDF fails with a spinner that never finishes. Some of our team say CSV export still works for them, others say it fails too.",
"questions": {
"bug_severity": {
"type": "score",
"instructions": "How severe is the reported issue?",
"criteria": [
{
"what": "Cosmetic; no impact to functionality",
"examples": [
"typo in a label",
"misaligned icon"
]
},
{
"what": "Broken or degraded feature, but workaround exists",
"examples": [
"export fails in one browser but works in another"
]
},
{
"what": "Blocking issue; no workaround exists",
"examples": [
"cannot log in",
"data loss"
]
}
]
}
}
}応答:
{
"model": "jev-1.13.0",
"answers": {
"bug_severity": {
"type": "score",
"score": 1.09,
"confidence": 0.87,
"legend": {
"0": {
"what": "Cosmetic; no impact to functionality",
"examples": [
"typo in a label",
"misaligned icon"
]
},
"1": {
"what": "Broken or degraded feature, but workaround exists",
"examples": [
"export fails in one browser but works in another"
]
},
"2": {
"what": "Blocking issue; no workaround exists",
"examples": [
"cannot log in",
"data loss"
]
}
},
"probabilities": {
"0": 0.0,
"1": 0.91,
"2": 0.09
}
}
},
"usage": {
"input_tokens": 379,
"output_tokens": 18
}
}
プレーンな文字列では、このチケットは信頼度 0.84 で 1.11 でした。例を付けると信頼度 0.87 で 1.09 になり、プレーンな文字列がすでにうまく位置付けていたので小さな変化です。効果は、プレーンな文字列がモデルを割ったままにするときの方が大きくなります。次の表が示します。
例はモデルを導き、それらは実際の入力に似ているときにだけ役立ちます。下の表は、冒頭の Safari の報告に三つの異なるレベルのオブジェクトを組み合わせたものです:
| レベルの説明 | score |
confidence |
|---|---|---|
| プレーンな文字列:例を含むオブジェクトなし | 1.43 | 0.35 |
| 有用な例を含む examples 配列を追加:「export fails in one browser but works in another」 | 1.03 | 0.96 |
| ブラウザと無関係な例を含む examples 配列を追加:「search fails, but browsing categories still works」 | 1.43 | 0.35 |
この比較では、合致する例が確率のほとんどすべてを一つのレベルに集中させます。無関係な例はプレーンな文字列と同じ結果を返します。信頼度が高いことは、どの答えが正しいかを確定しません。既知の期待レベルを持つ例を選び、改訂した説明を別の入力でテストしてから採用してください。