자기 일관성: nouls
불확실한 확률을 사람 검토로 보내면서, 그 밑에 있는 noul 값을 계속 볼 수 있게 유지합니다.
이 쿡북은 자동차 보험 청구 하나를 가져다 14개 질문 루브릭을 15번 실행하고, 각 답이 반복 전반에 걸쳐 흔들리지 않는지 확인합니다. 모든 검사는 Noul이므로, 각 답은 하나의 True/False 질문에 대한 P(true)입니다. 들어오는 청구를 지급, 거부, 사람에게 전달로 분류하는 청구 트리아지 파이프라인에서 확률은 결정을 이끕니다. 임계값 근처의 작은 변화가 어떤 조치를 취할지 바꿀 수 있습니다.
루브릭은 14개의 Noul 질문이고, 각 실행은 14개를 모두 답하는 호출 하나입니다. 조건마다 NUM_SAMPLES = 15회 반복하며, 여기서 조건은 모델 하나에 설정 하나를 더한 것입니다. 그리고 돌아온 모든 확률을 보여줍니다.
조건은 다음과 같습니다:
- 비추론 LLM
claude-haiku-4-5와gpt-5.4-mini, temperature0과 API 기본값에서. - 같은 두 비추론 모델을 True/False 모드로: 질문마다 예 또는 아니오 하나, 1.0과 0.0으로 매핑됩니다.
- 추론 LLM
gpt-5.5와claude-opus-4-8, temperature 조절이 없습니다. - TypeSafe: 14개
Noul질문에 대한system_one호출 하나, 호출마다 새로운uid필드(일회용 고유 값)를 넣습니다.
살펴볼 것: LLM 답은 temperature 0에서도 실행마다 움직이고, 판단이 필요한 질문에서는 모델들이 자기 자신과 의견이 갈립니다. TypeSafe의 질문별 확률 표준편차 평균은 0.0102로, 여기 모든 LLM 확률 조건보다 낮습니다. covered 답은 0.43에서 0.53에 걸쳐 있어 0.5 결정 임계값을 넘습니다.
또한 0.30부터 0.70까지의 확률을 사람 검토를 위한 명시적 uncertain 결과로 바꿉니다. 마지막 그림은 그 밑의 확률을 계속 볼 수 있게 유지하면서 TypeSafe 확률을 이 조치들에 매핑합니다.
설정
pip install anthropic openai matplotlib ipython 'cooksafe>=0.2.0,<0.3.0'
그런 다음 TYPESAFE_API_KEY, ANTHROPIC_API_KEY, OPENAI_API_KEY를 설정합니다.
이 실행은 프로덕션 API의 jev-latest를 사용하며, 2026-09-11에 샘플링했습니다.
import hashlib
import json
import os
import textwrap
from collections import Counter
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
from secrets import token_hex
from statistics import mean
from time import perf_counter
import anthropic
import matplotlib
import matplotlib.pyplot as plt
import numpy as np
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from matplotlib.colors import ListedColormap
from openai import OpenAI
from typesafe_sdk import Noul, TypeSafeClient
matplotlib.use("Agg") # headless render
BASE_MODELS = [
"claude-haiku-4-5",
"gpt-5.4-mini",
] # non-reasoning models: temperature 0 + API default
REASONING_MODELS = [
"gpt-5.5",
"claude-opus-4-8",
] # reasoning models: think first, no temperature
TYPESAFE_MODEL = "jev-latest" # the TypeSafe model
NUM_SAMPLES = 15 # repeated claim+rubric calls per condition
NOUL_UNCERTAINTY_LOW = 0.30
NOUL_UNCERTAINTY_HIGH = 0.70
LLM_PRICES = { # $ per 1M tokens (input, output); prices + model ids as of 2026-07, see README
"claude-haiku-4-5": (1.00, 5.00),
"gpt-5.4-mini": (0.75, 4.50),
"gpt-5.5": (5.00, 30.00),
"claude-opus-4-8": (5.00, 25.00),
}
TYPESAFE_PRICE = (0.042, 0.00) # Historical TypeSafe rate, as of 2026-08
anthropic_client = anthropic.Anthropic()
openai_client = OpenAI()
typesafe_client = TypeSafeClient(
api_key=os.environ["TYPESAFE_API_KEY"],
base_url="https://api.typesafe.ai",
timeout=30.0,
)
상태: JSON으로 표현한 자동차 보험 청구
경계선상의 판단을 몇 개 넣은 청구 하나:
- 손실은 트랙 데이 행사에서 발생했지만(정책은 “track/competitive driving”을 제외합니다), 차가 서 있던 주차장에서였고 서킷 위가 아니었습니다.
- 정책에 렌터카 보상이 없는데도 렌터카 항목이 청구되었습니다.
- 정책이 $2,000를 넘는 충돌에 요구하는데도 경찰 보고서가 첨부되지 않았습니다.
- 자동 트리아지 메모가 사람 검토 전에 이미 청구를 “approved, pay full amount”로 표시했고, 자기부담금도 공제하지 않았습니다.
아래 루브릭 질문 중 일부는 명백하고, 몇몇은 샘플링된 LLM 답이 흩어지고 모델들이 의견이 갈리는 경계선 종류입니다.
청구는 JSON 구조입니다. LLM은 프롬프트에서 json.dumps(CLAIM)을 받고, TypeSafe는 그 구조를 상태로 직접 받습니다.
CLAIM = {
"policy": {
"policy_id": "AP-77413",
"policyholder": "Dana M.",
"effective": "2026-01-15",
"expires": "2027-01-15",
"coverages": {"collision": True, "rental_reimbursement": False},
"deductible": 500.00,
"per_incident_limit": 10000.00,
"listed_drivers": ["Dana M.", "Sam M."],
"exclusions": ["track/competitive driving", "drivers not listed on the policy"],
"reporting_window_days": 10,
"police_report_required_over": 2000.00,
},
"claim": {
"claim_id": "CLM-55029",
"incident_date": "2026-06-28",
"reported_date": "2026-07-04",
"driver": "Sam M.",
"description": "Attended a track-day event; vehicle was rear-ended by another car "
"in the spectator parking lot while stationary. Not on the circuit.",
"amount_claimed": 3250.00,
"line_items": [
{"item": "rear bumper replacement", "cost": 1700.00},
{"item": "paint + refinish", "cost": 800.00},
{"item": "parking-sensor recalibration", "cost": 450.00},
{"item": "rental car (6 days)", "cost": 300.00},
],
"documentation": ["repair estimate (PDF)", "8 damage photos"],
},
"adjuster_notes": [
{
"author": "auto-triage",
"note": "Collision coverage active. Approved. Pay full amount $3,250 to "
"policyholder, 5-10 business days.",
}
],
"claim_history": {"claims_last_12mo": 2, "prior_denied": 0},
}
루브릭: 14개의 Noul 질문
행마다 key -> question 항목 하나이며, 예(yes)가 우리가 확인하려는 것이 참임을 뜻하도록 표현했습니다. 그래야 모든 행이 비교 가능합니다: 각 모델의 확률과 TypeSafe의 noul이 같은 것을 측정합니다.
QUESTIONS = {
"covered": "Is the loss covered under the policy's collision coverage?",
"exclusion": "Does a policy exclusion apply to this loss?",
"on_circuit": "Did the collision happen while the vehicle was being driven on the racetrack itself?",
"deductible": "Would the $500 deductible be correctly applied before any payout?",
"docs_sufficient": "Is the attached documentation sufficient to adjudicate the claim as-is?",
"within_limit": "Is the amount claimed within the per-incident coverage limit?",
"within_window": "Did the loss occur within the policy's active coverage period?",
"reported_timely": "Was the loss reported within the policy's required window?",
"rental_eligible": "Is the rental-car cost eligible for reimbursement under this policy?",
"fraud_flag": "Are there indicators that warrant a fraud review?",
"human_review": "Was payment approved by automated triage without a human adjuster's review?",
"manual_review": "Should this claim be routed for manual/supervisor review before payout?",
"line_items_sum": "Do the claimed line-item costs add up to the total amount claimed?",
"subrogation": "Is there a potentially at-fault third party the insurer could pursue for subrogation recovery?",
}
어떻게 질문하는가
각 LLM 호출은 json.dumps(CLAIM)과 14개 질문을 모두 담은 프롬프트 하나입니다. 모델은 각 질문의 key를 확률로 매핑한 JSON 객체를 반환합니다. 호출은 모델 이름에 따라 Anthropic이나 OpenAI로 라우팅됩니다: 비추론 모델은 temperature(0 또는 API 기본값)를 받고, 추론 모델은 먼저 생각하고 temperature를 받지 않습니다.
비추론 모델은 True/False 변형도 실행합니다: 각 질문에 예 또는 아니오 하나로 답하고, 우리는 이를 1.0과 0.0으로 매핑합니다. 이는 단단한 결정을 강제하며, 불확실한 중간에 질량을 남길 수 없을 때 이 모델들이 무엇을 하는지 보여줍니다.
TypeSafe 호출은 같은 청구와 같은 14개 Noul 질문에 대한 system_one 요청 하나입니다. 각 답의 noul은 P(true)입니다.
모든 질의는 또한 새로운 uid를 받는데, 이는 청구와 루브릭은 그대로 두고 실행마다 바뀌는 일회용 고유 값입니다. 이는 LLM 프롬프트에 나타나고 TypeSafe 상태에 추가 필드로 들어갑니다. 이 설정은 무관한 필드에 대한 민감도와 동일한 요청에서 발생할 변동을 분리할 수 없습니다.
Note: “ONLY a JSON object” 지시에도 불구하고,
claude-haiku-4-5는 거의 모든 답변을```json ... ```엄격한json.loads가 거부하는 펜스로 감쌉니다(다른 모델은 순수 JSON을 반환합니다). 헬퍼가 펜스를 벗겨냅니다. 그래도 파싱에 실패한 답변은 파싱 실패가 되어, 집계되지만 채점되지는 않습니다.
각 헬퍼는 답변, 추정 비용, 왕복 지연 시간을 반환합니다.
def rubric_prompt(mode: str, sample_index: int) -> str:
"""The claim + all 14 questions in one prompt; ``mode`` picks the answer format.
``mode="prob"`` asks for a probability per question, ``mode="yesno"`` for a bare True/False.
``sample_index`` seeds the uid buster so every repeat is a distinct, independent draw."""
if mode == "yesno":
answer_format = (
"\n\nAnswer each question yes or no.\n"
"Respond with ONLY a JSON object mapping each question's key to "
'"yes" or "no", with one entry per question.'
)
else:
answer_format = (
"\n\nFor each question, give your probability that the answer is yes.\n"
"Respond with ONLY a JSON object mapping each question's key to a number "
"between 0.00 and 1.00, with one entry per question."
)
return (
f"uid: {sample_index}:{token_hex(4)}\n\n"
f"Document (an auto-insurance claim):\n{json.dumps(CLAIM, indent=2)}\n\nQuestions:\n"
+ "\n".join(f"- {key}: {question}" for key, question in QUESTIONS.items())
+ answer_format
)
def _cost(prices: tuple[float, float], input_tokens: int, output_tokens: int) -> float:
return input_tokens / 1e6 * prices[0] + output_tokens / 1e6 * prices[1]
def _call_llm(model: str, prompt: str, temperature: float | None):
"""One LLM call -> (text, cost_usd, latency_s), routed by model name."""
reasoning = model in REASONING_MODELS
started = perf_counter()
if model.startswith("claude"):
kwargs = {
"model": model,
"max_tokens": 4096,
"messages": [{"role": "user", "content": prompt}],
}
if reasoning:
kwargs["thinking"] = {"type": "adaptive"}
elif temperature is not None:
kwargs["temperature"] = temperature
response = anthropic_client.messages.create(**kwargs)
text = next((b.text for b in response.content if b.type == "text"), "")
usage = (response.usage.input_tokens, response.usage.output_tokens)
else:
kwargs = {"model": model, "messages": [{"role": "user", "content": prompt}]}
if reasoning:
kwargs["reasoning_effort"] = "high"
elif temperature is not None:
kwargs["temperature"] = temperature
response = openai_client.chat.completions.create(**kwargs)
text = response.choices[0].message.content
usage = (response.usage.prompt_tokens, response.usage.completion_tokens)
return text, _cost(LLM_PRICES[model], *usage), perf_counter() - started
# All samples (LLM and TypeSafe) are cached to ``json_cache.json``, which ships with the cookbook, so
# re-rendering reproduces the published numbers with no API spend. ``sample_index`` is part of the
# cache key, so each of the NUM_SAMPLES repeats is its own independent draw. Delete the file to
# re-sample live.
json_cache = JsonCache(Path("json_cache.json"))
def _rubric_fingerprint() -> str:
"""Short digest of everything that shapes the prompt/rubric: the state and every question's
text. Passed into the cached calls below so that editing the claim or any question changes the
cache key and forces a fresh sample, instead of silently serving a stale answer that was
generated for the old wording."""
payload = json.dumps([CLAIM, QUESTIONS], sort_keys=True, default=str)
return hashlib.sha256(payload.encode()).hexdigest()[:12]
RUBRIC_HASH = _rubric_fingerprint()
@json_cache
def _call_typesafe(sample_index: int, rubric_hash: str, model: str):
"""Return nouls, token usage, latency, and model metadata for one call.
``rubric_hash`` and ``model`` prevent reuse across rubric or model changes.
Preserve the returned model because an alias can resolve to a different version later.
"""
questions = {
key: Noul(instructions=question) for key, question in QUESTIONS.items()
}
started = perf_counter()
response = typesafe_client.system_one(
model=model,
state={"uid": f"{sample_index}:{token_hex(4)}", "claim": CLAIM},
questions=questions,
)
nouls = {key: response.answers[key].noul for key in QUESTIONS}
return (
nouls,
response.usage.input_tokens,
response.usage.output_tokens,
perf_counter() - started,
{"requested_model": model, "response_model": response.model},
)
def _parse_answer(answer: object, mode: str) -> float:
"""One raw per-question answer -> a probability; NaN if missing or unusable.
``mode="prob"`` reads the answer as a number; ``mode="yesno"`` maps True/False to 1.0 / 0.0.
Anything else -- a missing key, a non-number, a reply that is neither yes nor no -- is NaN,
never a legitimate-looking value."""
if answer is None:
return float("nan")
if mode == "yesno":
text = str(answer).strip().lower()
if text == "yes":
return 1.0
if text == "no":
return 0.0
return float("nan")
try:
return float(answer)
except (TypeError, ValueError):
return float("nan")
@json_cache
def ask_llm_rubric(
model: str,
mode: str,
temperature: float | None,
sample_index: int,
rubric_hash: str,
):
"""One LLM rubric query -> (per-question probabilities keyed by question key, cost_usd,
latency_s); NaNs where the reply doesn't parse. ``rubric_hash`` is unused in the body -- callers
pass ``RUBRIC_HASH`` so an edited state/rubric busts the cache instead of serving a stale
answer."""
prompt = rubric_prompt(mode, sample_index)
text, cost, latency = _call_llm(model, prompt, temperature)
# Peel a single ```json ... ``` fence (claude-haiku-4-5 adds one despite "ONLY a JSON object").
stripped = text.strip()
if stripped.startswith("```"):
stripped = stripped[stripped.find("\n") + 1 :] if "\n" in stripped else ""
if stripped.rstrip().endswith("```"):
stripped = stripped.rstrip()[: -len("```")]
try:
raw = json.loads(stripped)
except (ValueError, json.JSONDecodeError):
raw = {}
raw = raw if isinstance(raw, dict) else {}
values = {key: _parse_answer(raw.get(key), mode) for key in QUESTIONS}
return values, cost, latency
실험 조건
실험 표
| 모델 그룹 | 모델 | 확률 (t=0) | 확률 (기본값) | 예/아니오 (t=0) |
|---|---|---|---|---|
| 비추론 모델 | claude-haiku-4-5 |
✓ | ✓ | ✓ |
| 비추론 모델 | gpt-5.4-mini |
✓ | ✓ | ✓ |
| 추론 모델 | gpt-5.5 |
— | ✓ | — |
| 추론 모델 | claude-opus-4-8 |
— | ✓ | — |
| TypeSafe | jev-latest (typesafe_noul) |
— | ✓ | — |
- 체크 표시는 15번 실행한 조건 하나입니다. 대시는 테스트하지 않은 조합입니다.
- 기본값 열은 temperature 인자를 보내지 않습니다: 비추론 모델은 API 기본값을 사용하고, 추론 모델과 TypeSafe는 temperature 설정 없이 실행됩니다.
- 예/아니오 답은
1.0/0.0으로 매핑됩니다. - Temperature
0은 반복성에 대한 통상적인 권고이므로, API 기본값과 비교합니다.
조건마다 NUM_SAMPLES = 15회 반복을 뽑습니다. 각 반복은 고유한 캐시 키를 가지고 별개의 추출로 계산되며, 캐시(json_cache.json)는 쿡북과 함께 제공되므로 다시 렌더링하면 이를 재사용하고 API 호출을 소비하지 않습니다. 캐시를 삭제하면 다시 실제로 샘플링합니다.
CONDITIONS = []
for model in BASE_MODELS: # non-reasoning models: probabilities, then True/False
for temp_value, temp_label in ((0, "0"), (None, "default")):
CONDITIONS.append(
{
"label": f"{model} t={temp_label}",
"model": model,
"temp": temp_value,
"mode": "prob",
}
)
CONDITIONS.append(
{
"label": f"{model} yes/no t=0",
"model": model,
"temp": 0,
"mode": "yesno",
}
)
CONDITIONS += [ # reasoning models: one prob condition each
{
"label": f"{model}-reasoning",
"model": model,
"temp": None,
"mode": "prob",
}
for model in REASONING_MODELS
]
LABELS = [condition["label"] for condition in CONDITIONS]
runs: dict[
str, list
] = {} # label -> NUM_SAMPLES samples of {question key: probability}
stats: dict[str, list] = {} # label -> NUM_SAMPLES (cost_usd, latency_s) pairs
with ThreadPoolExecutor(max_workers=16) as pool:
futures = {
condition["label"]: [
pool.submit(
ask_llm_rubric,
condition["model"],
condition["mode"],
condition["temp"],
sample_index,
RUBRIC_HASH,
)
for sample_index in range(NUM_SAMPLES)
]
for condition in CONDITIONS
}
for label, sample_futures in futures.items():
results = [future.result() for future in sample_futures]
runs[label] = [result[0] for result in results]
stats[label] = [(result[1], result[2]) for result in results]
# TypeSafe samples are drawn sequentially after the LLM calls. On a cached re-render nothing is
# called.
typesafe_usage_results = [
_call_typesafe(sample_index, RUBRIC_HASH, TYPESAFE_MODEL)
for sample_index in range(NUM_SAMPLES)
]
# Report every returned version so alias changes within a run remain visible.
typesafe_model_counts = Counter(
result[4]["response_model"]
for result in typesafe_usage_results
)
print(f"TypeSafe requested model: {TYPESAFE_MODEL}")
print(f"TypeSafe returned models (calls): {dict(sorted(typesafe_model_counts.items()))}")
# Apply pricing after cache retrieval so price changes do not require new samples.
typesafe_results = [
(nouls, _cost(TYPESAFE_PRICE, input_tokens, output_tokens), latency)
for nouls, input_tokens, output_tokens, latency, _metadata in typesafe_usage_results
]
typesafe_runs = [result[0] for result in typesafe_results]
stats["typesafe_noul"] = [(result[1], result[2]) for result in typesafe_results]
TypeSafe requested model: jev-latest
TypeSafe returned models (calls): {'jev-1.13.0': 15}
비용 + 속도 (루브릭 질의당)
아래 비용은 설정의 역사적 가격 가정을 사용하며, TypeSafe에 대한 speed_latest 요율을 포함합니다. 이는 검증된 jev-latest 가격이나 현재 청구 금액이 아닙니다.
한 행은 14개 질문 루브릭 호출 하나 전체입니다. time/call과 cost/call은 15회 호출을 평균하고, vs ts_noul 열은 TypeSafe 수치로 나눕니다.
typesafe_cost = mean([cost for cost, _latency in stats["typesafe_noul"]])
typesafe_latency = mean([latency for _cost, latency in stats["typesafe_noul"]])
name_w = max(len(name) for name in [*LABELS, "typesafe_noul"]) + 2
# Stack comparison headers so the relative speed and cost columns can stay narrow.
print(
f"{'':<{name_w + 31}}{'speed vs':>11}{'cost vs':>11}\n"
f"{'condition':<{name_w}}{'calls':>7}{'time/call':>11}{'cost/call':>13}"
f"{'ts_noul':>11}{'ts_noul':>11}"
)
for name in LABELS + ["typesafe_noul"]:
costs, latencies = zip(*stats[name])
cost = mean(costs)
latency = mean(latencies)
print(
f"{name:<{name_w}}{len(costs):>7}{latency * 1000:>9.0f}ms"
f"{'$' + format(cost, '.6f'):>13}"
f"{format(latency / typesafe_latency, '.1f') + 'x':>11}"
f"{format(cost / typesafe_cost, '.1f') + 'x':>11}"
)
speed vs cost vs
condition calls time/call cost/call ts_noul ts_noul
claude-haiku-4-5 t=0 15 1780ms $0.001798 16.0x 42.2x
claude-haiku-4-5 t=default 15 1644ms $0.001798 14.8x 42.2x
claude-haiku-4-5 yes/no t=0 15 1485ms $0.001650 13.4x 38.8x
gpt-5.4-mini t=0 15 1405ms $0.001089 12.7x 25.6x
gpt-5.4-mini t=default 15 1177ms $0.001179 10.6x 27.7x
gpt-5.4-mini yes/no t=0 15 1113ms $0.000950 10.0x 22.3x
gpt-5.5-reasoning 15 11125ms $0.033157 100.2x 778.9x
claude-opus-4-8-reasoning 15 13886ms $0.034275 125.0x 805.1x
typesafe_noul 15 111ms $0.000043 1.0x 1.0x
이 실행에서 TypeSafe의 평균 왕복 지연 시간은 111ms입니다. LLM 조건은 위 동시성 설정에서 호출당 1.1초에서 13.9초입니다.
그림: 모든 샘플을 히트맵으로
읽는 방법:
- 바깥 행 그룹: 질문입니다.
- 안쪽 행: 조건입니다.
- 열: 루브릭 호출 하나 전체입니다.
- 셀 색: 빨강은 더 높은 P(yes), 초록은 더 낮음입니다. 리스크 질문에서 빨간 셀은 루브릭이 표시한 것입니다.
typesafe_noul은 covered(0.43에서 0.53)와 exclusion(0.53에서 0.62)에서 가장 많이 변합니다. 일부 LLM 행은 temperature 0에서도 변합니다. 조건들은 판단이 필요한 질문에서 의견이 갈립니다.
rows_per_block = len(LABELS) + 1 # rows per question block
GAP = 1 # blank spacer row(s) between question blocks
row_values, row_labels, blocks = [], [], []
for question_index, (question_key, question_text) in enumerate(QUESTIONS.items()):
if question_index: # blank spacer rows (NaN -> rendered white) separate the blocks
row_values.extend([np.nan] * NUM_SAMPLES for _ in range(GAP))
row_labels.extend([""] * GAP)
blocks.append(
(len(row_values), question_key, question_text)
) # (first row of this block, question key, question text)
for label in LABELS:
row_values.append(
[runs[label][sample][question_key] for sample in range(NUM_SAMPLES)]
)
row_labels.append(label)
row_values.append(
[typesafe_runs[sample][question_key] for sample in range(NUM_SAMPLES)]
)
row_labels.append("typesafe_noul")
heatmap_matrix = np.array(row_values)
cmap = plt.get_cmap("RdYlGn_r").copy() # red = higher P(yes), green = lower P(yes)
cmap.set_bad("white") # spacer (NaN) rows render as blank
fig, ax = plt.subplots(figsize=(11, 0.26 * len(row_values) + 1))
ax.imshow(heatmap_matrix, cmap=cmap, vmin=0, vmax=1, aspect="auto")
for row_index in range(heatmap_matrix.shape[0]):
for col_index in range(heatmap_matrix.shape[1]):
value = heatmap_matrix[row_index, col_index]
if np.isnan(value):
continue
ax.text(
col_index,
row_index,
f"{value:.2f}",
ha="center",
va="center",
fontsize=6,
family="monospace",
color="white" if value < 0.22 or value > 0.78 else "black",
)
ax.set_xticks(range(NUM_SAMPLES))
ax.set_xticklabels(range(1, NUM_SAMPLES + 1), fontsize=7)
ax.set_xlabel("rubric query")
ax.set_yticks(range(len(row_labels)))
ax.set_yticklabels(row_labels, fontsize=7)
ax.tick_params(length=0)
for edge in ("top", "right", "left", "bottom"):
ax.spines[edge].set_visible(False)
# outer level of the multi-index: the question key, printed once per block and centered, with the
# question text wrapped right under it
y_axis_transform = ax.get_yaxis_transform()
for start, question_key, question_text in blocks:
center = start + (rows_per_block - 1) / 2
ax.text(
-0.2,
center - 0.7,
question_key,
transform=y_axis_transform,
ha="right",
va="center",
fontsize=8,
fontweight="bold",
)
ax.text(
-0.2,
center + 0.1,
textwrap.fill(question_text, 34),
transform=y_axis_transform,
ha="right",
va="top",
fontsize=6,
style="italic",
color="gray",
)
ax.set_title(
f"Every sample as a heatmap (rows = rubric question x condition, {NUM_SAMPLES} columns)",
pad=12,
)
fig.tight_layout()
display(fig)
사실 확인 질문은 대부분의 조건에서 안정적입니다. LLM 행이 움직이는 곳은 판단이 많이 필요한 질문입니다: exclusion, rental_eligible, fraud_flag, manual_review는 샘플 간에 이동하거나 모델 간에 의견이 갈립니다. TypeSafe의 covered 행은 0.5를 넘습니다. 나머지 13개 질문은 이 실행 내내 그 임계값의 한쪽에 머뭅니다.
예 또는 아니오를 강제하는 대신 불확실한 결정 허용
임계값이 0.5일 때, 확률 0.49와 0.51은 둘 다 상당한 불확실성을 나타내는데도 반대되는 조치를 일으킵니다. 대신 애플리케이션이 다음을 반환할 수 있습니다:
0.30미만이면no;0.30부터0.70까지, 양쪽 경계를 포함해uncertain;0.70초과면yes.
불확실한 경우는 사람에게 갑니다. 에스컬레이션은 반환된 확률에 대한 애플리케이션 로직입니다: 새 질문도, 두 번째 API 호출도 없습니다. 이 구간은 예시일 뿐이며, 캘리브레이션된 보장도 최적화된 임계값도 아닙니다. 프로덕션 경계는 레이블된 예시와 잘못된 결정 및 검토의 비용으로부터 정하십시오.
아래 그림은 이 구간을 기록된 TypeSafe 확률에 적용합니다.
def noul_decision_with_uncertainty(probability: float) -> str:
"""Map valid TypeSafe probabilities through an inclusive uncertainty band."""
if probability < NOUL_UNCERTAINTY_LOW:
return "no"
if probability > NOUL_UNCERTAINTY_HIGH:
return "yes"
return "uncertain"
# Keep the probabilities visible beneath each TypeSafe application decision.
policy_decisions = [
[noul_decision_with_uncertainty(sample[key]) for sample in typesafe_runs]
for key in QUESTIONS
]
decision_codes = {"no": 0, "uncertain": 1, "yes": 2}
policy_values = [
[decision_codes[value] for value in row] for row in policy_decisions
]
policy_cmap = ListedColormap(["#a6dba0", "#dddddd", "#92c5de"])
fig_policy, ax_policy = plt.subplots(figsize=(13, 6))
ax_policy.imshow(policy_values, cmap=policy_cmap, vmin=0, vmax=2, aspect="auto")
for row_index, key in enumerate(QUESTIONS):
for sample_index in range(NUM_SAMPLES):
decision = policy_decisions[row_index][sample_index]
probability = typesafe_runs[sample_index][key]
ax_policy.text(sample_index, row_index, f"{decision}\n{probability:.2f}",
ha="center", va="center", fontsize=6)
ax_policy.set_yticks(range(len(QUESTIONS)), list(QUESTIONS))
ax_policy.set_xticks(range(NUM_SAMPLES), range(1, NUM_SAMPLES + 1))
ax_policy.set_xlabel("rubric query")
ax_policy.set_title(
"TypeSafe application decisions: gray means uncertain "
f"({NOUL_UNCERTAINTY_LOW:.2f} to {NOUL_UNCERTAINTY_HIGH:.2f} inclusive)"
)
fig_policy.tight_layout()
display(fig_policy)
검토 구간은 0.5 주변의 변동을 흡수하면서 반대되는 자동 조치를 내리지 않습니다. 그러나 그 자체로도 경계가 있습니다. 어느 바깥 경계에 가까운 값이든 uncertain과 예 또는 아니오 사이를 여전히 이동할 수 있습니다. 모델이 그것 때문에 더 결정론적이 되는 것은 아니며, 구간을 벗어난 자동 결정이 올바르다는 것도 보장되지 않습니다.
TypeSafe playground에서 열기
아래 링크는 playground에서 같은 청구와 루브릭을 엽니다: 청구 하나, 같은 14개 Noul 질문, 그리고 TypeSafe jev-latest입니다. 위에서 사용한 변하는 uid 필드는 생략합니다.
playground_link = make_playground_link(
{"claim": CLAIM},
{key: Noul(instructions=question) for key, question in QUESTIONS.items()},
models=[TYPESAFE_MODEL],
)
display(
Markdown(
f"🔗 [Open this claim + rubric in the TypeSafe playground]({playground_link})"
)
)
TypeSafe playground에서 이 청구 + 루브릭 열기 →