자기 일관성: choice
모더레이션 결정에 불확실한 결과를 추가하고, 레이블 일치율을 자동 조치 비율과 비교합니다.
이 cookbook은 경계선에 있는 사용자 게시물 하나를 가져와 모더레이션 루브릭으로 15번 실행하고, 각 답이 반복 내내 흔들리지 않는지 확인합니다. 모든 검사가 Choice이므로, 각 답은 고정된 집합에서 고른 레이블 하나입니다. 모더레이션 파이프라인에서 그 레이블은 라우팅 결정입니다. 삭제할지 그대로 둘지, 에스컬레이션할지 자동 해결할지, 위협·스팸·일반 큐 중 어디로 보낼지입니다. 레이블이 실행마다 흔들리면, 같은 게시물이 뚜렷한 이유 없이 서로 다른 곳으로 라우팅됩니다.
루브릭은 8개의 Choice 질문이며, 각 실행은 8개를 모두 답하는 호출 하나입니다. 조건마다 15회 반복하는데, 여기서 조건이란 모델 하나에 설정 하나를 더한 것입니다. 그리고 돌아온 모든 레이블을 그립니다.
조건은 다음과 같습니다.
- 비추론 LLM
claude-haiku-4-5와gpt-5.4-mini, temperature0과 API 기본값에서. - 추론 LLM
gpt-5.5와claude-opus-4-8, 이들에는 temperature 조절이 없습니다. - TypeSafe: 8개
Choice질문에 대한system_one호출 하나이며, 호출마다 새uid필드(일회성 고유 값)를 붙이며, noul cookbook 설정과 맞춥니다.
주목할 점: 선택된 레이블은 TypeSafe를 포함해 단일 조건 안에서도 뒤집힐 수 있고, 조건들끼리 서로 어긋납니다.
이번 실행에서 LLM 분포 설정들은 최빈 레이블을 87.5%에서 100%까지 반복한 반면, TypeSafe는 90.8%였습니다. TypeSafe는 6개 LLM 분포 조건 중 5개보다 평균 확률 변동이 낮습니다. temperature 0의 Haiku는 그보다도 덜 변합니다. 확률이 가까우면 여전히 라우팅이 바뀔 수 있습니다. TypeSafe는 8개 질문 중 2개에서 뒤집힙니다.
애플리케이션 결정에서는 최상위 확률이 최소 0.60 이상이어야 합니다. 그렇지 않으면 결과는 uncertain이 되어 사람의 검토로 넘어갑니다. 그러면 TypeSafe의 일치율은 99.2%로 올라가고, 답변의 74.2%에 자동 레이블이 붙습니다. 원시 출력을 그대로 보여주고, LLM 확률 조건에도 같은 임계값을 적용하여 기권과 변화를 눈에 보이게 유지합니다.
설정
pip install anthropic openai matplotlib ipython 'cooksafe>=0.2.0,<0.3.0'
그런 다음 TYPESAFE_API_KEY, ANTHROPIC_API_KEY, OPENAI_API_KEY를 설정합니다.
이번 실행은 프로덕션 API의 jev-latest를 사용하며, 2026-09-11에 샘플링했습니다.
import hashlib
import json
import os
import textwrap
from collections import Counter
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
from secrets import token_hex
from statistics import mean
from time import perf_counter
import anthropic
import matplotlib
import matplotlib.pyplot as plt
import numpy as np
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from matplotlib.colors import ListedColormap
from openai import OpenAI
from typesafe_sdk import Choice, TypeSafeClient
matplotlib.use("Agg") # headless render
BASE_MODELS = [
"claude-haiku-4-5",
"gpt-5.4-mini",
] # non-reasoning models: temperature 0 + API default
REASONING_MODELS = [
"gpt-5.5",
"claude-opus-4-8",
] # reasoning models: think first, no temperature
TYPESAFE_MODEL = "jev-latest" # the TypeSafe model
NUM_SAMPLES = 15 # repeated post+rubric calls per condition
MIN_CHOICE_PROBABILITY = 0.60 # illustrative automatic-action threshold
LLM_PRICES = { # $ per 1M tokens (input, output); prices + model ids as of 2026-07, see README
"claude-haiku-4-5": (1.00, 5.00),
"gpt-5.4-mini": (0.75, 4.50),
"gpt-5.5": (5.00, 30.00),
"claude-opus-4-8": (5.00, 25.00),
}
TYPESAFE_PRICE = (0.042, 0.00) # Historical TypeSafe rate, as of 2026-08
anthropic_client = anthropic.Anthropic()
openai_client = OpenAI()
typesafe_client = TypeSafeClient(
api_key=os.environ["TYPESAFE_API_KEY"],
base_url="https://api.typesafe.ai",
timeout=30.0,
)
상태: 경계선에 있는 사용자 게시물, JSON으로
아래 게시물은 경계선에 걸치도록 만들었습니다. 언어는 격하고 모욕적이며, 일부는 한 사람을, 일부는 논쟁과 커뮤니티를 겨냥합니다. 플랫폼 밖 초대(사람들을 다른 사이트로 끌어가는 링크), 계정의 이전 스트라이크 1회, 사용자 신고 4건을 담고 있으며, 위협처럼 들리는 표현은 어디에서도 명확하게 표현되지 않습니다.
여기에는 뻔한 정답 하나가 없으며, 바로 그 점이 핵심입니다. 사소한 표현 차이가 같은 게시물을 집행 경로 사이에서 무작위로 옮겨서는 안 됩니다.
LLM은 프롬프트에서 json.dumps(POST)를 받습니다. TypeSafe는 Python dict를 그대로 받습니다.
POST = {
"post_id": "P-88213",
"author": {
"user_id": "u/4471",
"account_age_days": 38,
"prior_strikes": 1,
"followers": 210,
},
"context": {
"surface": "public reply",
"in_reply_to": "another user defending a game patch",
"community": "r/gamedebates",
},
"content": {
"text": (
"Are you seriously this dense? Anyone who defends that patch is a complete clown and "
"should be embarrassed to even post here. People like you are what's ruining this "
"community and honestly you need to be dealt with. Come say it to my face, invite's "
"right here. Keep it up and I'll end your whole channel."
),
"has_link": True,
"link_domain": "discord.gg",
"language": "en",
},
"reports": {"user_reports": 4, "report_reasons": ["harassment", "spam", "threat"]},
}
루브릭: Choice 질문 8개
각 질문에는 key, 지시문 한 줄, 고정된 레이블 집합이 있습니다. 한 질문 안의 레이블들은 상호 배타적이며(정확히 하나만 적용됩니다), 각각 짧은 설명을 답니다. TypeSafe는 선택된 choice와 레이블에 대한 probabilities 분포를 돌려줍니다. LLM에게도 같은 레이블 집합을 쓰도록 요청하여, 모든 행을 비교 가능하게 유지합니다.
QUESTIONS = {
"category": (
"What is the single most applicable content-policy category for this post?",
{
"None": "No policy violation of any kind.",
"Harass": "Insults or demeans a person, with no threat of harm and no protected-class attack.",
"Hate": "Attacks a person or group over a protected characteristic (race, religion, gender, ...).",
"Violence": "Makes a credible threat of harm or incites violence against someone.",
"Spam": "Unsolicited promotion or link spam, with no personal attack.",
"Sexual": "Sexual or adult content.",
},
),
"primary_risk": (
"What is the primary moderation risk that should drive triage for this post?",
{
"Harassment": "Personal attack or targeted abuse is the main risk.",
"Violence": "A threat of harm or intimidation is the main risk.",
"LinkAbuse": "External-link or off-platform coordination risk is the main risk.",
"AccountHistory": "Prior account history or repeat behavior is the main risk.",
"LowRisk": "No meaningful moderation risk is present.",
},
),
"target": (
"Who or what is the content primarily directed at?",
{
"None": "Not directed at anyone in particular.",
"Person": "Directed at one specific individual.",
"Group": "Directed at a protected group or class.",
"Platform": "Directed at the community or platform itself, not a person.",
},
),
"action": (
"What enforcement action should be taken on this post?",
{
"Allow": "Leave the post up with no action.",
"Warn": "Leave the post up but attach a warning label.",
"Remove": "Remove the post, but do not penalize the account.",
"Strike": "Remove the post and add a strike to the account.",
"Escalate": "Take no automated action; hold for a human decision.",
},
),
"queue": (
"Which single moderation queue should own this post?",
{
"Auto": "Auto-resolve; no human queue needed.",
"General": "General moderation queue.",
"Threat": "Threat / violence response queue.",
"Spam": "Spam and platform-abuse queue.",
"TSLead": "Trust-and-safety lead / senior queue.",
},
),
"link_handling": (
"How should any external link or off-platform invite in the post be handled?",
{
"Allow": "Leave the link in place.",
"RmLink": "Strip or disable the link but keep the post.",
"Brigade": "Treat the link as coordinated brigading and action it as abuse.",
"Escalate": "Send the link to a specialist to assess before acting.",
},
),
"review_path": (
"Who should make the final call on this post?",
{
"Auto": "Automated action; no human review.",
"Human": "A frontline human moderator makes the call.",
"Senior": "A senior or specialist reviewer is required.",
"Legal": "Route to legal or law-enforcement escalation.",
},
),
"severity": (
"What is the overall severity of this post?",
{
"None": "No violation.",
"Low": "Rude or dismissive, but essentially harmless.",
"Medium": "Personal harassment with no clearly credible threat.",
"High": "Harassment together with a threat that could be read as credible.",
},
),
}
질문 방식
각 LLM 호출은 json.dumps(POST), 8개 질문 전부, 허용된 모든 레이블을 담은 프롬프트 하나입니다. 답변 형식은 두 가지입니다. 분포 모드에서 모델은 질문마다 JSON 객체 하나를 돌려주며 각 레이블에 확률을 붙입니다. 단일 선택 모드에서는 질문마다 레이블 하나만 돌려주고, 분석에서는 모든 확률 질량을 그 레이블에 둡니다.
TypeSafe 호출은 같은 게시물과 같은 8개 Choice 질문에 대한 system_one 요청 하나이며, 질문마다 분포 하나를 돌려줍니다.
또한 모든 질의에는 새 uid가 붙습니다. 게시물과 루브릭은 그대로 두고 실행마다 바뀌는 일회성 고유 값입니다. 이것은 LLM 프롬프트에 나타나고 TypeSafe 상태에 추가 필드로 들어갑니다. 이 설정은 무관한 필드에 대한 민감도와, 동일한 요청에서도 발생했을 변동을 분리할 수 없습니다.
각 헬퍼는 답변, 추정 비용, 왕복 지연 시간을 돌려줍니다.
def argmax_label(values: list, labels: list[str]) -> str | None:
"""The label with the most probability mass, or ``None`` if any value is missing or
non-numeric -- a partially parsed distribution never yields a confident-looking pick."""
numeric = [_numeric_value(value) for value in values]
if any(value is None for value in numeric):
return None
return labels[int(np.argmax(numeric))]
def choice_decision_with_uncertainty(values: list, labels: list[str]) -> str | None:
"""Abstain below the action threshold; retain invalid results as parse failures."""
label = argmax_label(values, labels)
if label is None:
return None
probabilities = [float(value) for value in values]
if any(value < 0 or value > 1 for value in probabilities):
return None
return label if max(probabilities) >= MIN_CHOICE_PROBABILITY else "uncertain"
def choice_decision_annotation(values: list, labels: list[str]) -> str:
"""Show the application decision and top probability in a heatmap cell."""
decision = choice_decision_with_uncertainty(values, labels)
if decision is None:
return ""
probability = max(float(value) for value in values)
probability_text = f"{probability:.2f}".removeprefix("0")
return f"{decision} {probability_text}"
def _numeric_value(value: object) -> float | None:
"""A finite numeric value, or ``None`` if the model emitted something unusable."""
try:
numeric = float(value)
except (TypeError, ValueError):
return None
return numeric if np.isfinite(numeric) else None
def parse_distribution(raw: object, labels: list[str]) -> list[float]:
"""Map a model's already-parsed per-question reply to per-label probabilities, in label order
(distribution-mode answers left un-normalized).
A single-pick reply is a single label string -> all the mass on that exact label; a
distribution-mode reply is a dict read label by label. Anything that doesn't match a known label
or isn't a finite number is left NaN -- we report the gap rather than massaging the reply (e.g.
stripping an echoed description) to make it fit."""
if isinstance(raw, str): # single-pick mode: a single chosen label
if raw in labels:
return [1.0 if label == raw else 0.0 for label in labels]
return [float("nan")] * len(labels)
if not isinstance(raw, dict):
return [float("nan")] * len(labels)
return [
value if (value := _numeric_value(raw.get(label))) is not None else float("nan")
for label in labels
]
def rubric_prompt(mode: str, sample_index: int, rubric_hash: str) -> str:
"""The post + all questions (with their label sets) in one prompt; ``mode`` picks the format.
``mode="dist"`` asks for a probability distribution over each question's labels; the single-pick
mode (``mode="single"``) asks for a single label per question. The uid line combines
``rubric_hash`` (which rubric version) with ``sample_index`` and a random token, so every repeat
is a distinct, independent draw and two different rubrics never share a nonce."""
lines = []
for key, (instructions, choices) in QUESTIONS.items():
labels = "\n".join(f" {label}: {desc}" for label, desc in choices.items())
lines.append(f"- {key}: {instructions}\n labels:\n{labels}")
exclusivity = (
"\n\nEach question's labels are mutually exclusive: exactly one applies. If a post could "
"arguably fit more than one, pick the single most severe / most specific label per the "
"label descriptions."
)
if mode == "single":
answer_format = (
"\n\nFor each question, pick exactly ONE label.\nRespond with ONLY a JSON object "
"mapping each question's key to one of that question's bare labels (the label only, "
"not its description), with one entry per question."
)
else:
answer_format = (
"\n\nFor each question, give a probability distribution over that question's labels "
"(values 0.00-1.00 that sum to 1).\nRespond with ONLY a JSON object mapping each "
"question's key to an object mapping that question's bare labels (the label only, "
"not its description) to probabilities, with one entry per question."
)
return (
f"uid: {rubric_hash}:{sample_index}:{token_hex(4)}\n\n"
f"Document (a reported user post):\n{json.dumps(POST, indent=2)}\n\nQuestions:\n"
+ "\n".join(lines)
+ exclusivity
+ answer_format
)
def _cost(prices: tuple[float, float], input_tokens: int, output_tokens: int) -> float:
return input_tokens / 1e6 * prices[0] + output_tokens / 1e6 * prices[1]
def _call_llm(model: str, prompt: str, temperature: float | None):
"""One LLM call -> (text, cost_usd, latency_s), routed by model name."""
reasoning = model in REASONING_MODELS
started = perf_counter()
if model.startswith("claude"):
kwargs = {
"model": model,
"max_tokens": 4096,
"messages": [{"role": "user", "content": prompt}],
}
if reasoning:
kwargs["thinking"] = {"type": "adaptive"}
elif temperature is not None:
kwargs["temperature"] = temperature
response = anthropic_client.messages.create(**kwargs)
text = next((b.text for b in response.content if b.type == "text"), "")
usage = (response.usage.input_tokens, response.usage.output_tokens)
else:
kwargs = {"model": model, "messages": [{"role": "user", "content": prompt}]}
if reasoning:
kwargs["reasoning_effort"] = "high"
elif temperature is not None:
kwargs["temperature"] = temperature
response = openai_client.chat.completions.create(**kwargs)
text = response.choices[0].message.content
usage = (response.usage.prompt_tokens, response.usage.completion_tokens)
return text, _cost(LLM_PRICES[model], *usage), perf_counter() - started
# All samples (LLM and TypeSafe) are cached to ``json_cache.json``, which ships with the cookbook, so
# re-rendering is instant and reproduces the published numbers with no API spend. ``sample_index``
# seeds the uid buster and is part of the cache key, so each of the NUM_SAMPLES repeats is its own
# entry and its own independent draw, not one draw replayed. Delete ``json_cache.json`` to re-sample
# everything live.
json_cache = JsonCache(Path("json_cache.json"))
def _rubric_fingerprint() -> str:
"""Short digest of everything that shapes the prompt/rubric: the state and every question's
text and label set. Passed into the cached calls below so that editing the post or any question
changes the cache key and forces a fresh sample, instead of silently serving a stale answer that
was generated for the old wording."""
payload = json.dumps([POST, QUESTIONS], sort_keys=True, default=str)
return hashlib.sha256(payload.encode()).hexdigest()[:12]
RUBRIC_HASH = _rubric_fingerprint()
@json_cache
def _call_typesafe(sample_index: int, rubric_hash: str, model: str):
"""Return distributions, token usage, latency, and model metadata for one call.
``rubric_hash`` and ``model`` prevent reuse across rubric or model changes.
Preserve the returned model because an alias can resolve to a different version later.
"""
questions = {
key: Choice(instructions=instructions, criteria=choices)
for key, (instructions, choices) in QUESTIONS.items()
}
started = perf_counter()
response = typesafe_client.system_one(
model=model,
state={"uid": f"{rubric_hash}:{sample_index}:{token_hex(4)}", "post": POST},
questions=questions,
)
distributions = {}
for key, (_instructions, choices) in QUESTIONS.items():
probabilities = dict(response.answers[key].probabilities)
distributions[key] = [
probabilities.get(label, float("nan")) for label in choices
]
return (
distributions,
response.usage.input_tokens,
response.usage.output_tokens,
perf_counter() - started,
{"requested_model": model, "response_model": response.model},
)
@json_cache
def ask_llm_rubric(
model: str,
mode: str,
temperature: float | None,
sample_index: int,
rubric_hash: str,
):
"""One LLM rubric query -> (per-question label distributions keyed by question key, cost_usd,
latency_s); NaNs if the reply doesn't parse.
``mode="dist"`` parses 8 label distributions; the single-pick mode (``mode="single"``) parses 8
single labels and puts all the mass on each. ``rubric_hash`` goes into the prompt's uid nonce
(and so the cache key), so an edited state/rubric busts the cache instead of serving a stale
answer."""
prompt = rubric_prompt(mode, sample_index, rubric_hash)
text, cost, latency = _call_llm(model, prompt, temperature)
# Peel a single ```json ... ``` fence (claude-haiku-4-5 sometimes adds one despite "ONLY a JSON
# object").
stripped = text.strip()
if stripped.startswith("```"):
stripped = stripped[stripped.find("\n") + 1 :] if "\n" in stripped else ""
if stripped.rstrip().endswith("```"):
stripped = stripped.rstrip()[: -len("```")]
try:
raw = json.loads(stripped)
except (ValueError, json.JSONDecodeError):
raw = {}
if not isinstance(raw, dict):
raw = {}
distributions = {
key: parse_distribution(raw.get(key), list(choices))
for key, (_instructions, choices) in QUESTIONS.items()
}
return distributions, cost, latency
실험 조건
실험 그리드
| 모델 그룹 | 모델 | 분포 (t=0) | 분포 (기본값) | 단일 선택 (t=0) |
|---|---|---|---|---|
| 비추론 모델 | claude-haiku-4-5 |
✓ | ✓ | ✓ |
| 비추론 모델 | gpt-5.4-mini |
✓ | ✓ | ✓ |
| 추론 모델 | gpt-5.5 |
— | ✓ | — |
| 추론 모델 | claude-opus-4-8 |
— | ✓ | — |
| TypeSafe | jev-latest (typesafe_choice) |
— | ✓ | — |
✓는 15회 반복으로 테스트한 조건을,—는 테스트하지 않은 조합을 나타냅니다.- 기본값 열은 temperature 인자를 보내지 않습니다. 비추론 모델은 API 기본값을 쓰고, 추론 모델과 TypeSafe는 temperature 설정 없이 실행됩니다.
- 단일 선택 조건은 질문마다 레이블 하나를 돌려줍니다.
- temperature
0은 재현성을 위해 흔히 권장되므로, API 기본값과 비교합니다.
조건마다 NUM_SAMPLES = 15회 반복을 뽑습니다. 각 반복은 자체 캐시 키를 가지며 별개의 추출로 셉니다. 캐시(json_cache.json)는 cookbook과 함께 제공되므로 다시 렌더링할 때 재사용하고 API 호출을 쓰지 않습니다. 캐시를 삭제하면 다시 실시간으로 샘플링합니다.
CONDITIONS = []
for (
model
) in BASE_MODELS: # non-reasoning models: dist at t=0 / default, then a single-pick variant
for temp_value, temp_label in ((0, "0"), (None, "default")):
CONDITIONS.append(
{
"label": f"{model} t={temp_label}",
"model": model,
"temp": temp_value,
"mode": "dist",
}
)
CONDITIONS.append(
{
"label": f"{model} single-pick t=0",
"model": model,
"temp": 0,
"mode": "single",
}
)
CONDITIONS += [ # reasoning models: one distribution condition each
{
"label": f"{model}-reasoning",
"model": model,
"temp": None,
"mode": "dist",
}
for model in REASONING_MODELS
]
LABELS = [condition["label"] for condition in CONDITIONS]
TYPESAFE_LABEL = "typesafe_choice"
ALL_LABELS = [*LABELS, TYPESAFE_LABEL]
runs: dict[
str, list
] = {} # label -> NUM_SAMPLES samples of {question key: distribution}
stats: dict[str, list] = {} # label -> NUM_SAMPLES (cost_usd, latency_s) pairs
with ThreadPoolExecutor(max_workers=16) as pool:
futures = {
condition["label"]: [
pool.submit(
ask_llm_rubric,
condition["model"],
condition["mode"],
condition["temp"],
sample_index,
RUBRIC_HASH,
)
for sample_index in range(NUM_SAMPLES)
]
for condition in CONDITIONS
}
for label, sample_futures in futures.items():
results = [future.result() for future in sample_futures]
runs[label] = [result[0] for result in results]
stats[label] = [(result[1], result[2]) for result in results]
# TypeSafe samples are drawn sequentially, after the LLM pool has closed, so each call's latency is a
# clean round trip rather than one measured under the 16-way LLM thread contention.
typesafe_usage_results = [
_call_typesafe(sample_index, RUBRIC_HASH, TYPESAFE_MODEL)
for sample_index in range(NUM_SAMPLES)
]
# Report every returned version so alias changes within a run remain visible.
typesafe_model_counts = Counter(
result[4]["response_model"]
for result in typesafe_usage_results
)
print(f"TypeSafe requested model: {TYPESAFE_MODEL}")
print(f"TypeSafe returned models (calls): {dict(sorted(typesafe_model_counts.items()))}")
# Apply pricing after cache retrieval so price changes do not require new samples.
typesafe_results = [
(distributions, _cost(TYPESAFE_PRICE, input_tokens, output_tokens), latency)
for distributions, input_tokens, output_tokens, latency, _metadata in typesafe_usage_results
]
typesafe_runs = [result[0] for result in typesafe_results]
stats[TYPESAFE_LABEL] = [(result[1], result[2]) for result in typesafe_results]
TypeSafe requested model: jev-latest
TypeSafe returned models (calls): {'jev-1.13.0': 15}
비용 + 속도 (루브릭 질의당)
아래 비용은 TypeSafe의 speed_latest 요율을 포함해, 설정에 적힌 과거 가격 가정을 사용합니다. 검증된 jev-latest 가격이나 현재 청구 금액이 아닙니다.
한 행은 8개 질문 전체를 담은 루브릭 호출 하나입니다. time/call과 cost/call은 15회 호출을 평균하고, vs ts_choice 열은 TypeSafe 수치로 나눕니다. LLM은 16방향 풀에서 실행됩니다.
typesafe_cost = mean([cost for cost, _latency in stats["typesafe_choice"]])
typesafe_latency = mean([latency for _cost, latency in stats["typesafe_choice"]])
name_w = max(len(name) for name in ALL_LABELS) + 2 # fit the longest condition label
# Stack comparison headers so the relative speed and cost columns can stay narrow.
print(
f"{'':<{name_w + 31}}{'speed vs':>11}{'cost vs':>11}\n"
f"{'condition':<{name_w}}{'calls':>7}{'time/call':>11}{'cost/call':>13}"
f"{'ts_choice':>11}{'ts_choice':>11}"
)
for name in ALL_LABELS:
costs, latencies = zip(*stats[name])
cost = mean(costs)
latency = mean(latencies)
print(
f"{name:<{name_w}}{len(costs):>7}{latency * 1000:>9.0f}ms"
f"{'$' + format(cost, '.6f'):>13}"
f"{format(latency / typesafe_latency, '.1f') + 'x':>11}"
f"{format(cost / typesafe_cost, '.1f') + 'x':>11}"
)
speed vs cost vs
condition calls time/call cost/call ts_choice ts_choice
claude-haiku-4-5 t=0 15 3853ms $0.003498 33.8x 76.1x
claude-haiku-4-5 t=default 15 3860ms $0.003494 33.8x 76.0x
claude-haiku-4-5 single-pick t=0 15 992ms $0.001527 8.7x 33.2x
gpt-5.4-mini t=0 15 2293ms $0.002299 20.1x 50.0x
gpt-5.4-mini t=default 15 1986ms $0.002164 17.4x 47.1x
gpt-5.4-mini single-pick t=0 15 826ms $0.000936 7.2x 20.3x
gpt-5.5-reasoning 15 12978ms $0.041255 113.7x 897.4x
claude-opus-4-8-reasoning 15 10376ms $0.028375 90.9x 617.2x
typesafe_choice 15 114ms $0.000046 1.0x 1.0x
이번 실행에서 typesafe_choice의 평균 왕복 지연 시간은 114ms입니다. 위 동시성 설정에서 LLM 조건은 호출당 826ms에서 13.0초까지입니다.
그림: 모든 샘플의 결정을 히트맵으로
읽는 방법은 다음과 같습니다.
- 바깥 행 그룹: 질문입니다.
- 안쪽 행: 조건입니다.
- 열: 루브릭 호출 하나 전체입니다.
- 셀 텍스트: 애플리케이션 결정과 최상위 레이블의 확률입니다.
- 셀 색: 그 질문 안에서 레이블의 위치이므로, 한 행 전체가 같은 색이면 매번 같은 결정이라는 뜻입니다.
- 회색
uncertain: 최상위 확률이0.60미만이므로 사람의 검토로 넘어갑니다. - 빗금 친
n/a: 답변이 쓸 만한 레이블로 파싱되지 않았습니다(파싱 실패). - 빈 행은 간격을 띄우는 역할만 합니다.
단일 선택 조건은 돌려준 레이블을 그대로 유지합니다. 불확실성 추정치를 제공하지 않기 때문입니다.
GAP = 1 # blank spacer row(s) between question blocks
HEAT_LABELS = ALL_LABELS
rows_per_block = len(HEAT_LABELS) # rows per question block
pooled_runs = {
**runs,
TYPESAFE_LABEL: typesafe_runs,
}
row_index_values, row_text, row_labels, blocks = [], [], [], []
for question_index, (question_key, (question_text, choices)) in enumerate(
QUESTIONS.items()
):
labels = list(choices)
if question_index: # blank spacer rows (NaN -> rendered white) separate the blocks
row_index_values.extend([np.nan] * NUM_SAMPLES for _ in range(GAP))
row_text.extend([[""] * NUM_SAMPLES for _ in range(GAP)])
row_labels.extend([""] * GAP)
blocks.append((len(row_index_values), question_key, question_text))
for label in HEAT_LABELS:
values_by_sample = [
pooled_runs[label][sample][question_key] for sample in range(NUM_SAMPLES)
]
picks = [
choice_decision_with_uncertainty(values, labels) for values in values_by_sample
]
row_index_values.append(
[
10 if pick == "uncertain" else labels.index(pick) if pick in labels else np.nan
for pick in picks
]
)
row_text.append(
[choice_decision_annotation(values, labels) for values in values_by_sample]
)
row_labels.append(label)
heatmap_matrix = np.array(row_index_values, dtype=float)
# Reserve gray for abstentions while concrete-label colors remain local to each question.
cmap = ListedColormap([*plt.get_cmap("tab10").colors, "#dddddd"])
cmap.set_bad(
"white"
) # NaN cells (spacer rows AND unparseable replies) render white here...
fig, ax = plt.subplots(figsize=(15, 0.33 * len(row_index_values) + 1))
ax.imshow(heatmap_matrix, cmap=cmap, vmin=0, vmax=10, aspect="auto")
for row in range(heatmap_matrix.shape[0]):
is_spacer_row = row_labels[row] == "" # blank separator between question blocks
for col in range(heatmap_matrix.shape[1]):
label_text = row_text[row][col]
if label_text:
ax.text(
col,
row,
label_text,
ha="center",
va="center",
fontsize=5.7,
family="monospace",
color="black",
)
elif (
not is_spacer_row
): # ...but an unparseable reply gets a hatched "n/a", not blank white
ax.add_patch(
plt.Rectangle(
(col - 0.5, row - 0.5),
1,
1,
facecolor="#e8e8e8",
edgecolor="#b0b0b0",
hatch="////",
linewidth=0,
)
)
ax.text(
col,
row,
"n/a",
ha="center",
va="center",
fontsize=5,
family="monospace",
color="#b30000",
)
ax.set_xticks(range(NUM_SAMPLES))
ax.set_xticklabels(range(1, NUM_SAMPLES + 1), fontsize=7)
ax.set_xlabel("rubric query")
ax.set_yticks(range(len(row_labels)))
ax.set_yticklabels(row_labels, fontsize=7)
ax.tick_params(length=0)
for edge in ("top", "right", "left", "bottom"):
ax.spines[edge].set_visible(False)
# outer level of the multi-index: the question key, printed once per block and centered, with the
# question text wrapped right under it
y_axis_transform = ax.get_yaxis_transform()
for start, question_key, question_text in blocks:
center = start + (rows_per_block - 1) / 2
ax.text(
-0.2,
center - 0.7,
question_key,
transform=y_axis_transform,
ha="right",
va="center",
fontsize=8,
fontweight="bold",
)
ax.text(
-0.2,
center + 0.1,
textwrap.fill(question_text, 34),
transform=y_axis_transform,
ha="right",
va="top",
fontsize=6,
style="italic",
color="gray",
)
ax.set_title(
f"Every sample's decision + top probability; gray = uncertain (< {MIN_CHOICE_PROBABILITY:.2f})\n"
f"(rows = question x condition, {NUM_SAMPLES} columns)",
pad=12,
)
fig.tight_layout()
display(fig)
더 명확한 질문은 안정적입니다. target은 모든 조건에서 Person으로, severity는 High로 읽습니다. 경계선에 있는 질문들은 조건별로 갈립니다. category, primary_risk, action, review_path, link_handling입니다. 일부 조건은 자기 15회 반복 안에서도 뒤집힙니다. 기권 전에 TypeSafe는 primary_risk(Harassment 11회, Violence 4회)와 link_handling(RmLink 8회, Brigade 7회)에서 최상위 레이블을 바꿉니다. 이제 두 행 모두 최상위 확률이 0.60 미만이라 내내 uncertain을 보입니다.
확률 표준편차
이것은 선택된 레이블만이 아니라 전체 확률 벡터를 봅니다. 조건마다 모든 질문의 15개 분포를 모아, 레이블별 확률이 반복에 걸쳐 얼마나 움직이는지(실행마다 얼마나 변하는지)의 표준편차를 구한 뒤, 그 표준편차들을 모든 레이블과 질문에 걸쳐 평균합니다. 가장 큰 단일 레이블 표준편차도 보고하고, 파싱 실패는 따로 셉니다.
표는 확률을 출력하는 모든 LLM 조건을 TypeSafe와 비교합니다. 단일 선택 행은 확률 분포가 아니라 단단한 레이블을 내므로 제외합니다.
def probability_std_stats(samples: list) -> tuple[float, float, float]:
"""Mean label std dev, max label std dev, parse-failure rate."""
label_stds = []
parse_failures = []
for question_key in QUESTIONS:
arr = np.array(
[sample[question_key] for sample in samples],
dtype=float,
)
parse_failures.extend(np.isnan(arr).any(axis=1).tolist())
label_stds.extend(np.nanstd(arr, axis=0).tolist())
return (
float(np.nanmean(label_stds)),
float(np.nanmax(label_stds)),
float(np.mean(parse_failures)),
)
PROBABILITY_OUTPUT_LABELS = [
condition["label"] for condition in CONDITIONS if condition["mode"] == "dist"
] + [TYPESAFE_LABEL]
probability_std_by_label = {
label: probability_std_stats(pooled_runs[label])
for label in PROBABILITY_OUTPUT_LABELS
}
typesafe_mean_std = probability_std_by_label[TYPESAFE_LABEL][0]
print(
f"{'condition':<{name_w}}{'mean prob std':>15}{'max prob std':>14}"
f"{'parse fail':>12}{'x TypeSafe':>12}"
)
for label in PROBABILITY_OUTPUT_LABELS:
mean_std, max_std, parse_failure_rate = probability_std_by_label[label]
relative_std = mean_std / typesafe_mean_std
print(
f"{label:<{name_w}}{mean_std:>15.4f}{max_std:>14.4f}"
f"{parse_failure_rate:>11.0%}{relative_std:>12.2f}x"
)
condition mean prob std max prob std parse fail x TypeSafe
claude-haiku-4-5 t=0 0.0012 0.0221 0% 0.12x
claude-haiku-4-5 t=default 0.0516 0.3150 1% 5.29x
gpt-5.4-mini t=0 0.0312 0.0905 0% 3.20x
gpt-5.4-mini t=default 0.0543 0.2303 0% 5.56x
gpt-5.5-reasoning 0.0305 0.1047 0% 3.12x
claude-opus-4-8-reasoning 0.0245 0.0693 0% 2.52x
typesafe_choice 0.0098 0.0515 0% 1.00x
이번 실행에서 TypeSafe의 평균 확률 표준편차는 0.0098, 최대 단일 레이블 표준편차는 0.0515입니다. temperature 0의 Haiku는 그보다 낮은 평균 표준편차 0.0012를 보입니다. 나머지 5개 LLM 확률 조건은 0.0245에서 0.0543이며, TypeSafe 평균의 약 2.5x에서 5.6x입니다. 두 레이블이 가까우면 작은 변화로도 최상위 레이블이 바뀔 수 있습니다.
그림: 불확실한 결과를 포함한 결정 일치율
최상위 확률이 0.60 미만이면 uncertain을 돌려줍니다. 확률을 출력하는 조건과 질문마다 uncertain을 포함해 가장 흔한 애플리케이션 결정을 세고, 15개 추출 전체로 나눕니다. 파싱 실패는 일치율에 불리하게 셉니다. 각 막대는 8개 질문에 걸쳐 점수를 평균하며, 일치율이 가장 높은 것이 먼저 옵니다.
단일 선택 LLM 조건은 불확실성 추정치를 제공하지 않으므로 제외합니다.
# Compute policy decisions and agreement once for both this chart and the comparison table.
decisions_by_condition = {}
policy_agreement_by_condition = {}
for label in PROBABILITY_OUTPUT_LABELS:
decisions = [
[
choice_decision_with_uncertainty(sample[key], list(choices))
for sample in pooled_runs[label]
]
for key, (_instructions, choices) in QUESTIONS.items()
]
decisions_by_condition[label] = decisions
shares = [
max(Counter(value for value in row if value is not None).values(), default=0)
/ NUM_SAMPLES
for row in decisions
]
policy_agreement_by_condition[label] = mean(shares)
# Sort by the measured agreement, keeping TypeSafe's color independent of its rank.
bar_labels = sorted(
PROBABILITY_OUTPUT_LABELS, key=policy_agreement_by_condition.__getitem__, reverse=True
)
rates = [policy_agreement_by_condition[label] for label in bar_labels]
fig_bar, bar_ax = plt.subplots(figsize=(7, 0.45 * len(bar_labels) + 1))
positions = range(len(bar_labels))
bar_ax.barh(
list(positions),
rates,
color=["#2b8cbe" if label == TYPESAFE_LABEL else "#fe9929" for label in bar_labels],
alpha=0.85,
)
for label, position, rate in zip(bar_labels, positions, rates):
marker = "*" if label == "claude-haiku-4-5 t=0" else ""
bar_ax.text(
rate + 0.01, position, f"{rate:.1%}{marker}", va="center", fontsize=8, color="gray"
)
bar_ax.set_yticks(list(positions))
bar_ax.set_yticklabels(bar_labels, fontsize=8)
bar_ax.invert_yaxis() # first condition on top
bar_ax.set_xlim(0, 1.08)
bar_ax.set_xticks(np.linspace(0, 1, 6))
bar_ax.set_xlabel("decision agreement across 15 re-runs (mean over 8 questions)")
for edge in ("top", "right", "left"):
bar_ax.spines[edge].set_visible(False)
bar_ax.tick_params(length=0)
fig_bar.suptitle("Decision agreement including uncertain outcomes", y=1.0)
# Keep the caveat inside the exported chart so it travels with the 100% annotation.
fig_bar.text(
0.01,
0.01,
"* Haiku t=0: 100% repeatability does not imply correctness.\n"
" This experiment does not measure accuracy.",
fontsize=8,
)
fig_bar.tight_layout(rect=(0, 0.11, 1, 1))
display(fig_bar)
같은 0.60 규칙에서 temperature 0의 Haiku는 100%를 기록했습니다. TypeSafe는 99.2%였고, 나머지 LLM 조건은 84.2%에서 94.2% 사이에 머물렀습니다. TypeSafe는 답변의 25.8%에서 uncertain을 돌려주고 나머지 74.2%에서는 자동으로 조치했습니다. temperature 0의 Haiku는 한 번도 기권하지 않았습니다. 이 백분율은 재현성만 측정합니다. 아래 표는 원시 일치율과 기권율을 이 그림의 정책 일치율 옆에 놓습니다.
불확실한 확률이 불확실한 결정을 낳게 하기
작은 확률 변화가 가까운 두 레이블을 맞바꿀 수 있습니다. 애플리케이션이 승자를 그대로 조치할 필요는 없습니다. 최상위 확률이 0.60 미만이면 uncertain을 돌려주고 그 사례를 사람에게 보냅니다. 정확히 0.60이면 최상위 레이블을 선택합니다. 이것은 API의 별도 confidence 필드가 아니라 돌려받은 확률을 사용하며, 모델 호출을 추가하지 않습니다.
이 임계값은 설명을 위한 애플리케이션 정책이며, 캘리브레이션된 보장도, 이번 실행의 일치율을 최대화하도록 고른 임계값도 아닙니다. 프로덕션 임계값은 레이블된 예시와 잘못된 조치 및 사람 검토의 비용을 사용해 정하십시오.
확률을 출력하는 모든 조건에 같은 규칙을 적용합니다. 단일 선택 LLM 응답에는 확률 추정치가 없습니다. 그 합성 원-핫 벡터는 불확실성을 측정할 수 없으므로, 일치율 그림과 표에서 제외합니다.
def agreement_rate(samples: list) -> float:
"""Mean over questions of the raw plurality label's share across all NUM_SAMPLES draws.
Parse failures count against agreement because a failed route is not a repeated decision.
"""
shares = []
for question_key, (_instructions, choices) in QUESTIONS.items():
labels = list(choices)
picks = [
argmax_label(samples[sample][question_key], labels)
for sample in range(NUM_SAMPLES)
]
picks = [pick for pick in picks if pick is not None]
if not picks:
shares.append(0.0)
continue
top = Counter(picks).most_common(1)[0][1]
shares.append(top / NUM_SAMPLES)
return mean(shares) if shares else float("nan")
# Keep failures separate from abstentions and count conflicting concrete actions per question.
print(
f"{'condition':<{name_w}}{'raw agree':>12}{'policy agree':>14}"
f"{'uncertain':>12}{'automatic':>12}{'conflicts':>11}"
)
for label in PROBABILITY_OUTPUT_LABELS:
decisions = decisions_by_condition[label]
flat = [value for row in decisions for value in row]
uncertain_rate = mean(value == "uncertain" for value in flat)
automatic_rate = mean(value not in (None, "uncertain") for value in flat)
conflicts = sum(
len({value for value in row if value not in (None, "uncertain")}) > 1
for row in decisions
)
print(
f"{label:<{name_w}}{agreement_rate(pooled_runs[label]):>11.1%}"
f"{policy_agreement_by_condition[label]:>13.1%}{uncertain_rate:>11.1%}"
f"{automatic_rate:>11.1%}{conflicts:>11}"
)
condition raw agree policy agree uncertain automatic conflicts
claude-haiku-4-5 t=0 100.0% 100.0% 0.0% 100.0% 0
claude-haiku-4-5 t=default 87.5% 86.7% 0.8% 98.3% 2
gpt-5.4-mini t=0 99.2% 87.5% 12.5% 87.5% 0
gpt-5.4-mini t=default 90.8% 84.2% 22.5% 77.5% 2
gpt-5.5-reasoning 90.0% 93.3% 30.8% 69.2% 1
claude-opus-4-8-reasoning 92.5% 94.2% 33.3% 66.7% 0
typesafe_choice 90.8% 99.2% 25.8% 74.2% 0
policy agree는 uncertain을 결정으로 셉니다. 파싱 실패는 일치율에 불리하게 셉니다. automatic은 레이블을 선택한 전체 답변의 비율입니다. conflicts는 반복에 걸쳐 구체적인 레이블이 둘 이상 나온 질문을 세며, 기권은 무시합니다. 이 지표들은 재현성과 애플리케이션이 얼마나 자주 조치하는지를 설명할 뿐, 그 조치가 옳은지는 말하지 않습니다.
TypeSafe의 일치율은 90.8%에서 99.2%로 올랐습니다. 답변 중 25.8%는 불확실했고 74.2%는 자동이었습니다. primary_risk와 link_handling은 매 반복마다 불확실로 돌아왔습니다. category는 Violence와 uncertain 사이를 오갔고, 어떤 반복에서는 조치 임계값을 넘고 어떤 반복에서는 넘지 않았습니다. 어떤 질문도 서로 다른 두 개의 구체적인 TypeSafe 레이블을 내지 않았습니다. 이 중 어느 것도 정확도나 우월성을 보여주지 않습니다. temperature 0의 Haiku는 여기서 기권 없이 100% 일치율을 기록했습니다.
# Show every TypeSafe decision while retaining the top probability behind it.
policy_decisions = decisions_by_condition[TYPESAFE_LABEL]
policy_values = []
for row, (_key, (_instructions, choices)) in zip(policy_decisions, QUESTIONS.items()):
labels = list(choices)
policy_values.append([
10 if value == "uncertain" else labels.index(value) if value is not None else np.nan
for value in row
])
policy_cmap = ListedColormap([*plt.get_cmap("tab10").colors, "#dddddd"])
policy_cmap.set_bad("white")
fig_policy, ax_policy = plt.subplots(figsize=(13, 4))
ax_policy.imshow(policy_values, cmap=policy_cmap, vmin=0, vmax=10, aspect="auto")
for row_index, key in enumerate(QUESTIONS):
for sample_index in range(NUM_SAMPLES):
decision = policy_decisions[row_index][sample_index]
probability = max(typesafe_runs[sample_index][key])
ax_policy.text(sample_index, row_index, f"{decision or 'n/a'}\n{probability:.2f}",
ha="center", va="center", fontsize=6)
ax_policy.set_yticks(range(len(QUESTIONS)), list(QUESTIONS))
ax_policy.set_xticks(range(NUM_SAMPLES), range(1, NUM_SAMPLES + 1))
ax_policy.set_xlabel("rubric query")
ax_policy.set_title(
"TypeSafe application decisions: gray means uncertain "
f"(top probability < {MIN_CHOICE_PROBABILITY:.2f})"
)
fig_policy.tight_layout()
display(fig_policy)
이 정책이 모델을 결정론적으로 만들지는 않습니다. 기권은 경쟁하는 레이블들을 같은 사람 검토 결과로 대체할 수 있지만, 0.60 근처의 확률은 여전히 구체적인 레이블과 uncertain 사이를 오갈 수 있습니다. 확률 통계와 표의 raw agree 열은 원래의 모델 출력을 그대로 보고합니다.
TypeSafe playground에서 열기
아래 링크는 playground에서 같은 게시물과 루브릭을 엽니다. 게시물 하나, 같은 8개 Choice, TypeSafe jev-latest입니다.
playground_link = make_playground_link(
{"post": POST},
{
key: Choice(instructions=instructions, criteria=choices)
for key, (instructions, choices) in QUESTIONS.items()
},
models=[TYPESAFE_MODEL],
)
display(
Markdown(
f"🔗 [Open this post + rubric in the TypeSafe playground]({playground_link})"
)
)
TypeSafe playground에서 이 게시물 + 루브릭 열기 →