Dokumentation

Selbstkonsistenz: choices

Fügt Moderationsentscheidungen ein unsicheres Ergebnis hinzu und vergleicht die Übereinstimmung der Labels mit dem Anteil automatischer Aktionen.

Dieses Cookbook nimmt einen grenzwertigen Nutzerbeitrag, führt 15-mal eine Moderations-Rubrik darüber aus und prüft, ob jede Antwort über die Wiederholungen hinweg stabil bleibt. Jede Prüfung ist ein Choice, jede Antwort ist also ein Label aus einer festen Menge. In einer Moderations-Pipeline ist dieses Label die Routing-Entscheidung: entfernen oder stehen lassen, eskalieren oder automatisch lösen, an die Bedrohungs-, Spam- oder allgemeine Warteschlange senden. Wenn das Label von einem Durchlauf zum nächsten schwankt, wird derselbe Beitrag ohne guten Grund an verschiedene Stellen geroutet.

Die Rubrik besteht aus 8 Choice-Fragen, und jeder Durchlauf ist ein Aufruf, der alle 8 beantwortet. Wir machen 15 Wiederholungen pro Bedingung, wobei eine Bedingung ein Modell plus eine Einstellung ist, und zeichnen jedes Label auf, das zurückkam.

Die Bedingungen:

  • LLMs ohne Reasoning claude-haiku-4-5 und gpt-5.4-mini, bei Temperatur 0 und dem API-Standard.
  • Reasoning-LLMs gpt-5.5 und claude-opus-4-8, die keinen Temperaturregler haben.
  • TypeSafe: ein system_one-Aufruf über die 8 Choice-Fragen, mit einem frischen uid- Feld (einem wegwerfbaren eindeutigen Wert) bei jedem Aufruf, passend zum Aufbau des noul-Cookbooks.

Worauf du achten solltest: Gewählte Labels können innerhalb einer einzelnen Bedingung umspringen, TypeSafe eingeschlossen, und Bedingungen widersprechen einander.

In diesem Durchlauf wiederholen die LLM-Distributions-Einstellungen ihre häufigsten Labels zu 87.5% bis 100%, verglichen mit 90.8% bei TypeSafe. TypeSafe hat eine geringere mittlere Wahrscheinlichkeitsvariation als fünf der sechs LLM-Distributions-Bedingungen; Haiku bei Temperatur 0 variiert weniger. Dicht beieinanderliegende Wahrscheinlichkeiten erlauben dennoch Routing-Änderungen: TypeSafe springt bei 2 der 8 Fragen um.

Für Anwendungsentscheidungen verlangen wir zusätzlich eine Spitzenwahrscheinlichkeit von mindestens 0.60; andernfalls ist das Ergebnis uncertain und geht zur menschlichen Prüfung. TypeSafes Übereinstimmung steigt dann auf 99.2%, mit automatischen Labels bei 74.2% der Antworten. Wir zeigen die Rohausgaben und wenden denselben Schwellenwert auf die LLM-Wahrscheinlichkeitsbedingungen an, wobei Enthaltungen und Änderungen sichtbar bleiben.

Einrichtung

pip install anthropic openai matplotlib ipython 'cooksafe>=0.2.0,<0.3.0'

Lege dann TYPESAFE_API_KEY, ANTHROPIC_API_KEY und OPENAI_API_KEY fest. Dieser Durchlauf verwendet jev-latest auf der Produktions-API, abgetastet am 2026-09-11.

import hashlib
import json
import os
import textwrap
from collections import Counter
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
from secrets import token_hex
from statistics import mean

from time import perf_counter

import anthropic
import matplotlib
import matplotlib.pyplot as plt
import numpy as np
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from matplotlib.colors import ListedColormap
from openai import OpenAI
from typesafe_sdk import Choice, TypeSafeClient

matplotlib.use("Agg")  # headless render

BASE_MODELS = [
    "claude-haiku-4-5",
    "gpt-5.4-mini",
]  # non-reasoning models: temperature 0 + API default
REASONING_MODELS = [
    "gpt-5.5",
    "claude-opus-4-8",
]  # reasoning models: think first, no temperature
TYPESAFE_MODEL = "jev-latest"  # the TypeSafe model
NUM_SAMPLES = 15  # repeated post+rubric calls per condition
MIN_CHOICE_PROBABILITY = 0.60  # illustrative automatic-action threshold

LLM_PRICES = {  # $ per 1M tokens (input, output); prices + model ids as of 2026-07, see README
    "claude-haiku-4-5": (1.00, 5.00),
    "gpt-5.4-mini": (0.75, 4.50),
    "gpt-5.5": (5.00, 30.00),
    "claude-opus-4-8": (5.00, 25.00),
}
TYPESAFE_PRICE = (0.042, 0.00)  # Historical TypeSafe rate, as of 2026-08

anthropic_client = anthropic.Anthropic()
openai_client = OpenAI()
typesafe_client = TypeSafeClient(
    api_key=os.environ["TYPESAFE_API_KEY"],
    base_url="https://api.typesafe.ai",
    timeout=30.0,
)

Der Zustand: ein grenzwertiger Nutzerbeitrag als JSON

Der folgende Beitrag ist darauf angelegt, genau auf der Grenze zu liegen. Die Sprache ist hitzig und beleidigend, teils auf eine Person gerichtet, teils auf das Argument und die Community. Er enthält eine Einladung außerhalb der Plattform (einen Link, der Leute auf eine andere Seite zieht), einen früheren Verstoß auf dem Konto und vier Nutzermeldungen, und die bedrohlich wirkende Formulierung ist nie eindeutig.

Es gibt hier keine einzige offensichtliche Antwort, und genau das ist der Punkt: kleine Formulierungsunterschiede sollten denselben Beitrag nicht zufällig zwischen verschiedenen Durchsetzungspfaden verschieben.

Die LLMs bekommen json.dumps(POST) im Prompt. TypeSafe bekommt das Python-Dict direkt.

POST = {
    "post_id": "P-88213",
    "author": {
        "user_id": "u/4471",
        "account_age_days": 38,
        "prior_strikes": 1,
        "followers": 210,
    },
    "context": {
        "surface": "public reply",
        "in_reply_to": "another user defending a game patch",
        "community": "r/gamedebates",
    },
    "content": {
        "text": (
            "Are you seriously this dense? Anyone who defends that patch is a complete clown and "
            "should be embarrassed to even post here. People like you are what's ruining this "
            "community and honestly you need to be dealt with. Come say it to my face, invite's "
            "right here. Keep it up and I'll end your whole channel."
        ),
        "has_link": True,
        "link_domain": "discord.gg",
        "language": "en",
    },
    "reports": {"user_reports": 4, "report_reasons": ["harassment", "spam", "threat"]},
}

Die Rubrik: 8 Choice-Fragen

Jede Frage hat einen key, eine Zeile mit Anweisungen und eine feste Labelmenge. Die Labels innerhalb einer Frage schließen einander aus (genau eines trifft zu), und jedes trägt eine kurze Beschreibung. TypeSafe liefert ein gewähltes choice plus eine probabilities-Verteilung über die Labels. Die LLMs werden gebeten, dieselben Labelmengen zu verwenden, was jede Zeile vergleichbar hält.

QUESTIONS = {
    "category": (
        "What is the single most applicable content-policy category for this post?",
        {
            "None": "No policy violation of any kind.",
            "Harass": "Insults or demeans a person, with no threat of harm and no protected-class attack.",
            "Hate": "Attacks a person or group over a protected characteristic (race, religion, gender, ...).",
            "Violence": "Makes a credible threat of harm or incites violence against someone.",
            "Spam": "Unsolicited promotion or link spam, with no personal attack.",
            "Sexual": "Sexual or adult content.",
        },
    ),
    "primary_risk": (
        "What is the primary moderation risk that should drive triage for this post?",
        {
            "Harassment": "Personal attack or targeted abuse is the main risk.",
            "Violence": "A threat of harm or intimidation is the main risk.",
            "LinkAbuse": "External-link or off-platform coordination risk is the main risk.",
            "AccountHistory": "Prior account history or repeat behavior is the main risk.",
            "LowRisk": "No meaningful moderation risk is present.",
        },
    ),
    "target": (
        "Who or what is the content primarily directed at?",
        {
            "None": "Not directed at anyone in particular.",
            "Person": "Directed at one specific individual.",
            "Group": "Directed at a protected group or class.",
            "Platform": "Directed at the community or platform itself, not a person.",
        },
    ),
    "action": (
        "What enforcement action should be taken on this post?",
        {
            "Allow": "Leave the post up with no action.",
            "Warn": "Leave the post up but attach a warning label.",
            "Remove": "Remove the post, but do not penalize the account.",
            "Strike": "Remove the post and add a strike to the account.",
            "Escalate": "Take no automated action; hold for a human decision.",
        },
    ),
    "queue": (
        "Which single moderation queue should own this post?",
        {
            "Auto": "Auto-resolve; no human queue needed.",
            "General": "General moderation queue.",
            "Threat": "Threat / violence response queue.",
            "Spam": "Spam and platform-abuse queue.",
            "TSLead": "Trust-and-safety lead / senior queue.",
        },
    ),
    "link_handling": (
        "How should any external link or off-platform invite in the post be handled?",
        {
            "Allow": "Leave the link in place.",
            "RmLink": "Strip or disable the link but keep the post.",
            "Brigade": "Treat the link as coordinated brigading and action it as abuse.",
            "Escalate": "Send the link to a specialist to assess before acting.",
        },
    ),
    "review_path": (
        "Who should make the final call on this post?",
        {
            "Auto": "Automated action; no human review.",
            "Human": "A frontline human moderator makes the call.",
            "Senior": "A senior or specialist reviewer is required.",
            "Legal": "Route to legal or law-enforcement escalation.",
        },
    ),
    "severity": (
        "What is the overall severity of this post?",
        {
            "None": "No violation.",
            "Low": "Rude or dismissive, but essentially harmless.",
            "Medium": "Personal harassment with no clearly credible threat.",
            "High": "Harassment together with a threat that could be read as credible.",
        },
    ),
}

Wie wir fragen

Jeder LLM-Aufruf ist ein Prompt, der json.dumps(POST), alle 8 Fragen und jedes erlaubte Label enthält. Es gibt zwei Antwortformate. Im Distributions-Modus gibt das Modell ein JSON-Objekt pro Frage zurück, mit einer Wahrscheinlichkeit für jedes Label. Im Single-Pick-Modus gibt es ein bloßes Label pro Frage zurück, und unsere Analyse legt die gesamte Wahrscheinlichkeitsmasse auf dieses Label.

Der TypeSafe-Aufruf ist eine system_one-Anfrage über denselben Beitrag und dieselben 8 Choice-Fragen, die eine Verteilung pro Frage zurückgibt.

Jede Abfrage bekommt außerdem eine frische uid, einen wegwerfbaren eindeutigen Wert, der sich bei jedem Durchlauf ändert, während Beitrag und Rubrik unverändert bleiben. Er erscheint im LLM-Prompt und als zusätzliches Feld im TypeSafe-Zustand. Dieser Aufbau kann die Empfindlichkeit gegenüber dem irrelevanten Feld nicht von der Variation trennen, die bei identischen Anfragen auftreten würde.

Jeder Helfer gibt die Antwort, geschätzte Kosten und die Roundtrip-Latenz zurück.

def argmax_label(values: list, labels: list[str]) -> str | None:
    """The label with the most probability mass, or ``None`` if any value is missing or
    non-numeric -- a partially parsed distribution never yields a confident-looking pick."""
    numeric = [_numeric_value(value) for value in values]
    if any(value is None for value in numeric):
        return None
    return labels[int(np.argmax(numeric))]

def choice_decision_with_uncertainty(values: list, labels: list[str]) -> str | None:
    """Abstain below the action threshold; retain invalid results as parse failures."""
    label = argmax_label(values, labels)
    if label is None:
        return None
    probabilities = [float(value) for value in values]
    if any(value < 0 or value > 1 for value in probabilities):
        return None
    return label if max(probabilities) >= MIN_CHOICE_PROBABILITY else "uncertain"

def choice_decision_annotation(values: list, labels: list[str]) -> str:
    """Show the application decision and top probability in a heatmap cell."""
    decision = choice_decision_with_uncertainty(values, labels)
    if decision is None:
        return ""
    probability = max(float(value) for value in values)
    probability_text = f"{probability:.2f}".removeprefix("0")
    return f"{decision} {probability_text}"

def _numeric_value(value: object) -> float | None:
    """A finite numeric value, or ``None`` if the model emitted something unusable."""
    try:
        numeric = float(value)
    except (TypeError, ValueError):
        return None
    return numeric if np.isfinite(numeric) else None

def parse_distribution(raw: object, labels: list[str]) -> list[float]:
    """Map a model's already-parsed per-question reply to per-label probabilities, in label order
    (distribution-mode answers left un-normalized).

    A single-pick reply is a single label string -> all the mass on that exact label; a
    distribution-mode reply is a dict read label by label. Anything that doesn't match a known label
    or isn't a finite number is left NaN -- we report the gap rather than massaging the reply (e.g.
    stripping an echoed description) to make it fit."""
    if isinstance(raw, str):  # single-pick mode: a single chosen label
        if raw in labels:
            return [1.0 if label == raw else 0.0 for label in labels]
        return [float("nan")] * len(labels)
    if not isinstance(raw, dict):
        return [float("nan")] * len(labels)
    return [
        value if (value := _numeric_value(raw.get(label))) is not None else float("nan")
        for label in labels
    ]

def rubric_prompt(mode: str, sample_index: int, rubric_hash: str) -> str:
    """The post + all questions (with their label sets) in one prompt; ``mode`` picks the format.

    ``mode="dist"`` asks for a probability distribution over each question's labels; the single-pick
    mode (``mode="single"``) asks for a single label per question. The uid line combines
    ``rubric_hash`` (which rubric version) with ``sample_index`` and a random token, so every repeat
    is a distinct, independent draw and two different rubrics never share a nonce."""
    lines = []
    for key, (instructions, choices) in QUESTIONS.items():
        labels = "\n".join(f"     {label}: {desc}" for label, desc in choices.items())
        lines.append(f"- {key}: {instructions}\n   labels:\n{labels}")
    exclusivity = (
        "\n\nEach question's labels are mutually exclusive: exactly one applies. If a post could "
        "arguably fit more than one, pick the single most severe / most specific label per the "
        "label descriptions."
    )
    if mode == "single":
        answer_format = (
            "\n\nFor each question, pick exactly ONE label.\nRespond with ONLY a JSON object "
            "mapping each question's key to one of that question's bare labels (the label only, "
            "not its description), with one entry per question."
        )
    else:
        answer_format = (
            "\n\nFor each question, give a probability distribution over that question's labels "
            "(values 0.00-1.00 that sum to 1).\nRespond with ONLY a JSON object mapping each "
            "question's key to an object mapping that question's bare labels (the label only, "
            "not its description) to probabilities, with one entry per question."
        )
    return (
        f"uid: {rubric_hash}:{sample_index}:{token_hex(4)}\n\n"
        f"Document (a reported user post):\n{json.dumps(POST, indent=2)}\n\nQuestions:\n"
        + "\n".join(lines)
        + exclusivity
        + answer_format
    )

def _cost(prices: tuple[float, float], input_tokens: int, output_tokens: int) -> float:
    return input_tokens / 1e6 * prices[0] + output_tokens / 1e6 * prices[1]

def _call_llm(model: str, prompt: str, temperature: float | None):
    """One LLM call -> (text, cost_usd, latency_s), routed by model name."""
    reasoning = model in REASONING_MODELS
    started = perf_counter()
    if model.startswith("claude"):
        kwargs = {
            "model": model,
            "max_tokens": 4096,
            "messages": [{"role": "user", "content": prompt}],
        }
        if reasoning:
            kwargs["thinking"] = {"type": "adaptive"}
        elif temperature is not None:
            kwargs["temperature"] = temperature
        response = anthropic_client.messages.create(**kwargs)
        text = next((b.text for b in response.content if b.type == "text"), "")
        usage = (response.usage.input_tokens, response.usage.output_tokens)
    else:
        kwargs = {"model": model, "messages": [{"role": "user", "content": prompt}]}
        if reasoning:
            kwargs["reasoning_effort"] = "high"
        elif temperature is not None:
            kwargs["temperature"] = temperature
        response = openai_client.chat.completions.create(**kwargs)
        text = response.choices[0].message.content
        usage = (response.usage.prompt_tokens, response.usage.completion_tokens)
    return text, _cost(LLM_PRICES[model], *usage), perf_counter() - started

# All samples (LLM and TypeSafe) are cached to ``json_cache.json``, which ships with the cookbook, so
# re-rendering is instant and reproduces the published numbers with no API spend. ``sample_index``
# seeds the uid buster and is part of the cache key, so each of the NUM_SAMPLES repeats is its own
# entry and its own independent draw, not one draw replayed. Delete ``json_cache.json`` to re-sample
# everything live.
json_cache = JsonCache(Path("json_cache.json"))

def _rubric_fingerprint() -> str:
    """Short digest of everything that shapes the prompt/rubric: the state and every question's
    text and label set. Passed into the cached calls below so that editing the post or any question
    changes the cache key and forces a fresh sample, instead of silently serving a stale answer that
    was generated for the old wording."""
    payload = json.dumps([POST, QUESTIONS], sort_keys=True, default=str)
    return hashlib.sha256(payload.encode()).hexdigest()[:12]

RUBRIC_HASH = _rubric_fingerprint()

@json_cache
def _call_typesafe(sample_index: int, rubric_hash: str, model: str):
    """Return distributions, token usage, latency, and model metadata for one call.

    ``rubric_hash`` and ``model`` prevent reuse across rubric or model changes.
    Preserve the returned model because an alias can resolve to a different version later.
    """
    questions = {
        key: Choice(instructions=instructions, criteria=choices)
        for key, (instructions, choices) in QUESTIONS.items()
    }
    started = perf_counter()
    response = typesafe_client.system_one(
        model=model,
        state={"uid": f"{rubric_hash}:{sample_index}:{token_hex(4)}", "post": POST},
        questions=questions,
    )
    distributions = {}
    for key, (_instructions, choices) in QUESTIONS.items():
        probabilities = dict(response.answers[key].probabilities)
        distributions[key] = [
            probabilities.get(label, float("nan")) for label in choices
        ]
    return (
        distributions,
        response.usage.input_tokens,
        response.usage.output_tokens,
        perf_counter() - started,
        {"requested_model": model, "response_model": response.model},
    )

@json_cache
def ask_llm_rubric(
    model: str,
    mode: str,
    temperature: float | None,
    sample_index: int,
    rubric_hash: str,
):
    """One LLM rubric query -> (per-question label distributions keyed by question key, cost_usd,
    latency_s); NaNs if the reply doesn't parse.

    ``mode="dist"`` parses 8 label distributions; the single-pick mode (``mode="single"``) parses 8
    single labels and puts all the mass on each. ``rubric_hash`` goes into the prompt's uid nonce
    (and so the cache key), so an edited state/rubric busts the cache instead of serving a stale
    answer."""
    prompt = rubric_prompt(mode, sample_index, rubric_hash)
    text, cost, latency = _call_llm(model, prompt, temperature)
    # Peel a single ```json ... ``` fence (claude-haiku-4-5 sometimes adds one despite "ONLY a JSON
    # object").
    stripped = text.strip()
    if stripped.startswith("```"):
        stripped = stripped[stripped.find("\n") + 1 :] if "\n" in stripped else ""
        if stripped.rstrip().endswith("```"):
            stripped = stripped.rstrip()[: -len("```")]
    try:
        raw = json.loads(stripped)
    except (ValueError, json.JSONDecodeError):
        raw = {}
    if not isinstance(raw, dict):
        raw = {}
    distributions = {
        key: parse_distribution(raw.get(key), list(choices))
        for key, (_instructions, choices) in QUESTIONS.items()
    }
    return distributions, cost, latency

Experimentelle Bedingungen

Versuchsraster

Modellgruppe Modell Distribution (t=0) Distribution (default) Single-pick (t=0)
Modelle ohne Reasoning claude-haiku-4-5 ✓ ✓ ✓
Modelle ohne Reasoning gpt-5.4-mini ✓ ✓ ✓
Reasoning-Modelle gpt-5.5 — ✓ —
Reasoning-Modelle claude-opus-4-8 — ✓ —
TypeSafe jev-latest (typesafe_choice) — ✓ —
  • Ein ✓ markiert eine Bedingung, die mit 15 Wiederholungen getestet wurde; ein — markiert eine Kombination, die nicht getestet wird.
  • Die Spalte „default“ sendet kein Temperatur-Argument: Modelle ohne Reasoning verwenden den API-Standard, und Reasoning-Modelle sowie TypeSafe laufen ohne Temperatur-Einstellung.
  • Single-Pick-Bedingungen geben ein Label pro Frage zurück.
  • Temperatur 0 wird häufig für Wiederholbarkeit empfohlen und daher mit dem API-Standard verglichen.

Wir ziehen NUM_SAMPLES = 15 Wiederholungen pro Bedingung. Jede Wiederholung hat ihren eigenen Cache-Schlüssel und zählt als eigenständige Ziehung, und der Cache (json_cache.json) wird mit dem Cookbook ausgeliefert, sodass ein erneutes Rendern ihn wiederverwendet und keine API-Aufrufe verbraucht. Lösche den Cache, um erneut live zu sampeln.

CONDITIONS = []
for (
    model
) in BASE_MODELS:  # non-reasoning models: dist at t=0 / default, then a single-pick variant
    for temp_value, temp_label in ((0, "0"), (None, "default")):
        CONDITIONS.append(
            {
                "label": f"{model} t={temp_label}",
                "model": model,
                "temp": temp_value,
                "mode": "dist",
            }
        )
    CONDITIONS.append(
        {
            "label": f"{model} single-pick t=0",
            "model": model,
            "temp": 0,
            "mode": "single",
        }
    )
CONDITIONS += [  # reasoning models: one distribution condition each
    {
        "label": f"{model}-reasoning",
        "model": model,
        "temp": None,
        "mode": "dist",
    }
    for model in REASONING_MODELS
]
LABELS = [condition["label"] for condition in CONDITIONS]
TYPESAFE_LABEL = "typesafe_choice"
ALL_LABELS = [*LABELS, TYPESAFE_LABEL]

runs: dict[
    str, list
] = {}  # label -> NUM_SAMPLES samples of {question key: distribution}
stats: dict[str, list] = {}  # label -> NUM_SAMPLES (cost_usd, latency_s) pairs
with ThreadPoolExecutor(max_workers=16) as pool:
    futures = {
        condition["label"]: [
            pool.submit(
                ask_llm_rubric,
                condition["model"],
                condition["mode"],
                condition["temp"],
                sample_index,
                RUBRIC_HASH,
            )
            for sample_index in range(NUM_SAMPLES)
        ]
        for condition in CONDITIONS
    }
    for label, sample_futures in futures.items():
        results = [future.result() for future in sample_futures]
        runs[label] = [result[0] for result in results]
        stats[label] = [(result[1], result[2]) for result in results]

# TypeSafe samples are drawn sequentially, after the LLM pool has closed, so each call's latency is a
# clean round trip rather than one measured under the 16-way LLM thread contention.
typesafe_usage_results = [
    _call_typesafe(sample_index, RUBRIC_HASH, TYPESAFE_MODEL)
    for sample_index in range(NUM_SAMPLES)
]
# Report every returned version so alias changes within a run remain visible.
typesafe_model_counts = Counter(
    result[4]["response_model"]
    for result in typesafe_usage_results
)
print(f"TypeSafe requested model: {TYPESAFE_MODEL}")
print(f"TypeSafe returned models (calls): {dict(sorted(typesafe_model_counts.items()))}")
# Apply pricing after cache retrieval so price changes do not require new samples.
typesafe_results = [
    (distributions, _cost(TYPESAFE_PRICE, input_tokens, output_tokens), latency)
    for distributions, input_tokens, output_tokens, latency, _metadata in typesafe_usage_results
]
typesafe_runs = [result[0] for result in typesafe_results]
stats[TYPESAFE_LABEL] = [(result[1], result[2]) for result in typesafe_results]
TypeSafe requested model: jev-latest
TypeSafe returned models (calls): {'jev-1.13.0': 15}

Kosten + Geschwindigkeit (pro Rubrik-Abfrage)

Die Kosten unten verwenden die historischen Preisannahmen aus der Einrichtung, einschließlich des speed_latest-Satzes für TypeSafe. Es sind keine verifizierten jev-latest-Preise oder aktuellen Abrechnungsbeträge.

Eine Zeile ist ein vollständiger Rubrik-Aufruf mit 8 Fragen. time/call und cost/call mitteln über die 15 Aufrufe, und die Spalten vs ts_choice teilen durch die TypeSafe-Werte. Die LLMs laufen in einem 16-fachen Pool.

typesafe_cost = mean([cost for cost, _latency in stats["typesafe_choice"]])
typesafe_latency = mean([latency for _cost, latency in stats["typesafe_choice"]])
name_w = max(len(name) for name in ALL_LABELS) + 2  # fit the longest condition label
# Stack comparison headers so the relative speed and cost columns can stay narrow.
print(
    f"{'':<{name_w + 31}}{'speed vs':>11}{'cost vs':>11}\n"
    f"{'condition':<{name_w}}{'calls':>7}{'time/call':>11}{'cost/call':>13}"
    f"{'ts_choice':>11}{'ts_choice':>11}"
)
for name in ALL_LABELS:
    costs, latencies = zip(*stats[name])
    cost = mean(costs)
    latency = mean(latencies)
    print(
        f"{name:<{name_w}}{len(costs):>7}{latency * 1000:>9.0f}ms"
        f"{'$' + format(cost, '.6f'):>13}"
        f"{format(latency / typesafe_latency, '.1f') + 'x':>11}"
        f"{format(cost / typesafe_cost, '.1f') + 'x':>11}"
    )
                                                                    speed vs    cost vs
condition                           calls  time/call    cost/call  ts_choice  ts_choice
claude-haiku-4-5 t=0                   15     3853ms    $0.003498      33.8x      76.1x
claude-haiku-4-5 t=default             15     3860ms    $0.003494      33.8x      76.0x
claude-haiku-4-5 single-pick t=0       15      992ms    $0.001527       8.7x      33.2x
gpt-5.4-mini t=0                       15     2293ms    $0.002299      20.1x      50.0x
gpt-5.4-mini t=default                 15     1986ms    $0.002164      17.4x      47.1x
gpt-5.4-mini single-pick t=0           15      826ms    $0.000936       7.2x      20.3x
gpt-5.5-reasoning                      15    12978ms    $0.041255     113.7x     897.4x
claude-opus-4-8-reasoning              15    10376ms    $0.028375      90.9x     617.2x
typesafe_choice                        15      114ms    $0.000046       1.0x       1.0x

In diesem Durchlauf hat typesafe_choice eine mittlere Roundtrip-Latenz von 114ms. Die LLM-Bedingungen reichen von 826ms bis 13.0 Sekunden pro Aufruf unter den obigen Nebenläufigkeits-Einstellungen.

Diagramm: die Entscheidung jeder Stichprobe als Heatmap

So liest du es:

  • Äußere Zeilengruppe: die Frage.
  • Innere Zeile: die Bedingung.
  • Spalte: ein vollständiger Rubrik-Aufruf.
  • Zelltext: die Anwendungsentscheidung plus die Wahrscheinlichkeit des höchsten Labels.
  • Zellfarbe: die Position des Labels innerhalb dieser Frage; dieselbe Farbe über eine ganze Zeile hinweg bedeutet also jedes Mal dieselbe Entscheidung.
  • Grau uncertain: die Spitzenwahrscheinlichkeit liegt unter 0.60, der Fall geht also zur menschlichen Prüfung.
  • Schraffiert n/a: die Antwort ließ sich nicht in brauchbare Labels parsen (ein Parse-Fehler).
  • Leere Zeilen sind nur Abstandshalter.

Single-Pick-Bedingungen behalten ihre zurückgegebenen Labels: Sie liefern keine Unsicherheitsschätzung.

GAP = 1  # blank spacer row(s) between question blocks
HEAT_LABELS = ALL_LABELS
rows_per_block = len(HEAT_LABELS)  # rows per question block
pooled_runs = {
    **runs,
    TYPESAFE_LABEL: typesafe_runs,
}

row_index_values, row_text, row_labels, blocks = [], [], [], []
for question_index, (question_key, (question_text, choices)) in enumerate(
    QUESTIONS.items()
):
    labels = list(choices)
    if question_index:  # blank spacer rows (NaN -> rendered white) separate the blocks
        row_index_values.extend([np.nan] * NUM_SAMPLES for _ in range(GAP))
        row_text.extend([[""] * NUM_SAMPLES for _ in range(GAP)])
        row_labels.extend([""] * GAP)
    blocks.append((len(row_index_values), question_key, question_text))
    for label in HEAT_LABELS:
        values_by_sample = [
            pooled_runs[label][sample][question_key] for sample in range(NUM_SAMPLES)
        ]
        picks = [
            choice_decision_with_uncertainty(values, labels) for values in values_by_sample
        ]
        row_index_values.append(
            [
                10 if pick == "uncertain" else labels.index(pick) if pick in labels else np.nan
                for pick in picks
            ]
        )
        row_text.append(
            [choice_decision_annotation(values, labels) for values in values_by_sample]
        )
        row_labels.append(label)

heatmap_matrix = np.array(row_index_values, dtype=float)
# Reserve gray for abstentions while concrete-label colors remain local to each question.
cmap = ListedColormap([*plt.get_cmap("tab10").colors, "#dddddd"])
cmap.set_bad(
    "white"
)  # NaN cells (spacer rows AND unparseable replies) render white here...

fig, ax = plt.subplots(figsize=(15, 0.33 * len(row_index_values) + 1))
ax.imshow(heatmap_matrix, cmap=cmap, vmin=0, vmax=10, aspect="auto")
for row in range(heatmap_matrix.shape[0]):
    is_spacer_row = row_labels[row] == ""  # blank separator between question blocks
    for col in range(heatmap_matrix.shape[1]):
        label_text = row_text[row][col]
        if label_text:
            ax.text(
                col,
                row,
                label_text,
                ha="center",
                va="center",
                fontsize=5.7,
                family="monospace",
                color="black",
            )
        elif (
            not is_spacer_row
        ):  # ...but an unparseable reply gets a hatched "n/a", not blank white
            ax.add_patch(
                plt.Rectangle(
                    (col - 0.5, row - 0.5),
                    1,
                    1,
                    facecolor="#e8e8e8",
                    edgecolor="#b0b0b0",
                    hatch="////",
                    linewidth=0,
                )
            )
            ax.text(
                col,
                row,
                "n/a",
                ha="center",
                va="center",
                fontsize=5,
                family="monospace",
                color="#b30000",
            )

ax.set_xticks(range(NUM_SAMPLES))
ax.set_xticklabels(range(1, NUM_SAMPLES + 1), fontsize=7)
ax.set_xlabel("rubric query")
ax.set_yticks(range(len(row_labels)))
ax.set_yticklabels(row_labels, fontsize=7)
ax.tick_params(length=0)
for edge in ("top", "right", "left", "bottom"):
    ax.spines[edge].set_visible(False)

# outer level of the multi-index: the question key, printed once per block and centered, with the
# question text wrapped right under it
y_axis_transform = ax.get_yaxis_transform()
for start, question_key, question_text in blocks:
    center = start + (rows_per_block - 1) / 2
    ax.text(
        -0.2,
        center - 0.7,
        question_key,
        transform=y_axis_transform,
        ha="right",
        va="center",
        fontsize=8,
        fontweight="bold",
    )
    ax.text(
        -0.2,
        center + 0.1,
        textwrap.fill(question_text, 34),
        transform=y_axis_transform,
        ha="right",
        va="top",
        fontsize=6,
        style="italic",
        color="gray",
    )

ax.set_title(
    f"Every sample's decision + top probability; gray = uncertain (< {MIN_CHOICE_PROBABILITY:.2f})\n"
    f"(rows = question x condition, {NUM_SAMPLES} columns)",
    pad=12,
)
fig.tight_layout()
display(fig)
Ausgabe

Die eindeutigeren Fragen bleiben stabil: target liest durchweg Person und severity durchweg High. Die grenzwertigen verteilen sich über die Bedingungen: category, primary_risk, action, review_path und link_handling. Manche Bedingungen springen auch innerhalb ihrer eigenen 15 Wiederholungen um. Vor der Enthaltung wechselt TypeSafe sein höchstes Label bei primary_risk (Harassment 11-mal, Violence 4-mal) und link_handling (RmLink 8-mal, Brigade 7-mal). Beide Zeilen zeigen jetzt durchgängig uncertain, weil ihre Spitzenwahrscheinlichkeiten unter 0.60 liegen.

Standardabweichung der Wahrscheinlichkeit

Dies betrachtet die vollständigen Wahrscheinlichkeitsvektoren, nicht nur das gewählte Label. Für jede Bedingung sammeln wir alle 15 Verteilungen für jede Frage, nehmen die Standardabweichung der Wahrscheinlichkeit jedes Labels über die Wiederholungen (wie stark sie sich von Durchlauf zu Durchlauf bewegt) und mitteln diese Standardabweichungen dann über alle Labels und Fragen. Wir berichten außerdem die einzelne größte Label-Standardabweichung und zählen Parse-Fehler separat.

Die Tabelle vergleicht jede LLM-Bedingung mit Wahrscheinlichkeitsausgabe gegen TypeSafe. Die Single-Pick-Zeilen bleiben außen vor, da sie harte Labels statt Wahrscheinlichkeitsverteilungen ausgeben.

def probability_std_stats(samples: list) -> tuple[float, float, float]:
    """Mean label std dev, max label std dev, parse-failure rate."""
    label_stds = []
    parse_failures = []
    for question_key in QUESTIONS:
        arr = np.array(
            [sample[question_key] for sample in samples],
            dtype=float,
        )
        parse_failures.extend(np.isnan(arr).any(axis=1).tolist())
        label_stds.extend(np.nanstd(arr, axis=0).tolist())
    return (
        float(np.nanmean(label_stds)),
        float(np.nanmax(label_stds)),
        float(np.mean(parse_failures)),
    )

PROBABILITY_OUTPUT_LABELS = [
    condition["label"] for condition in CONDITIONS if condition["mode"] == "dist"
] + [TYPESAFE_LABEL]
probability_std_by_label = {
    label: probability_std_stats(pooled_runs[label])
    for label in PROBABILITY_OUTPUT_LABELS
}
typesafe_mean_std = probability_std_by_label[TYPESAFE_LABEL][0]

print(
    f"{'condition':<{name_w}}{'mean prob std':>15}{'max prob std':>14}"
    f"{'parse fail':>12}{'x TypeSafe':>12}"
)
for label in PROBABILITY_OUTPUT_LABELS:
    mean_std, max_std, parse_failure_rate = probability_std_by_label[label]
    relative_std = mean_std / typesafe_mean_std
    print(
        f"{label:<{name_w}}{mean_std:>15.4f}{max_std:>14.4f}"
        f"{parse_failure_rate:>11.0%}{relative_std:>12.2f}x"
    )
condition                           mean prob std  max prob std  parse fail  x TypeSafe
claude-haiku-4-5 t=0                       0.0012        0.0221         0%        0.12x
claude-haiku-4-5 t=default                 0.0516        0.3150         1%        5.29x
gpt-5.4-mini t=0                           0.0312        0.0905         0%        3.20x
gpt-5.4-mini t=default                     0.0543        0.2303         0%        5.56x
gpt-5.5-reasoning                          0.0305        0.1047         0%        3.12x
claude-opus-4-8-reasoning                  0.0245        0.0693         0%        2.52x
typesafe_choice                            0.0098        0.0515         0%        1.00x

In diesem Durchlauf hat TypeSafe eine mittlere Wahrscheinlichkeits-Standardabweichung von 0.0098 und eine maximale Einzel-Label-Standardabweichung von 0.0515. Haiku bei Temperatur 0 hat eine niedrigere mittlere Standardabweichung von 0.0012. Die übrigen fünf LLM-Wahrscheinlichkeitsbedingungen reichen von 0.0245 bis 0.0543, etwa 2.5x bis 5.6x des TypeSafe-Mittels. Kleine Änderungen können das höchste Label trotzdem umschalten, wenn zwei Labels dicht beieinanderliegen.

Diagramm: Entscheidungsübereinstimmung mit einem unsicheren Ergebnis

Gib uncertain zurück, wenn die Spitzenwahrscheinlichkeit unter 0.60 liegt. Zähle für jede Bedingung mit Wahrscheinlichkeitsausgabe und jede Frage die häufigste Anwendungsentscheidung, uncertain eingeschlossen, und teile durch alle 15 Ziehungen. Parse-Fehler zählen gegen die Übereinstimmung. Jeder Balken mittelt die Punktzahl über alle 8 Fragen, mit der höchsten Übereinstimmung zuerst.

Single-Pick-LLM-Bedingungen sind ausgeschlossen, weil sie keine Unsicherheitsschätzung liefern.

# Compute policy decisions and agreement once for both this chart and the comparison table.
decisions_by_condition = {}
policy_agreement_by_condition = {}
for label in PROBABILITY_OUTPUT_LABELS:
    decisions = [
        [
            choice_decision_with_uncertainty(sample[key], list(choices))
            for sample in pooled_runs[label]
        ]
        for key, (_instructions, choices) in QUESTIONS.items()
    ]
    decisions_by_condition[label] = decisions
    shares = [
        max(Counter(value for value in row if value is not None).values(), default=0)
        / NUM_SAMPLES
        for row in decisions
    ]
    policy_agreement_by_condition[label] = mean(shares)

# Sort by the measured agreement, keeping TypeSafe's color independent of its rank.
bar_labels = sorted(
    PROBABILITY_OUTPUT_LABELS, key=policy_agreement_by_condition.__getitem__, reverse=True
)
rates = [policy_agreement_by_condition[label] for label in bar_labels]

fig_bar, bar_ax = plt.subplots(figsize=(7, 0.45 * len(bar_labels) + 1))
positions = range(len(bar_labels))
bar_ax.barh(
    list(positions),
    rates,
    color=["#2b8cbe" if label == TYPESAFE_LABEL else "#fe9929" for label in bar_labels],
    alpha=0.85,
)
for label, position, rate in zip(bar_labels, positions, rates):
    marker = "*" if label == "claude-haiku-4-5 t=0" else ""
    bar_ax.text(
        rate + 0.01, position, f"{rate:.1%}{marker}", va="center", fontsize=8, color="gray"
    )
bar_ax.set_yticks(list(positions))
bar_ax.set_yticklabels(bar_labels, fontsize=8)
bar_ax.invert_yaxis()  # first condition on top
bar_ax.set_xlim(0, 1.08)
bar_ax.set_xticks(np.linspace(0, 1, 6))
bar_ax.set_xlabel("decision agreement across 15 re-runs (mean over 8 questions)")
for edge in ("top", "right", "left"):
    bar_ax.spines[edge].set_visible(False)
bar_ax.tick_params(length=0)
fig_bar.suptitle("Decision agreement including uncertain outcomes", y=1.0)
# Keep the caveat inside the exported chart so it travels with the 100% annotation.
fig_bar.text(
    0.01,
    0.01,
    "* Haiku t=0: 100% repeatability does not imply correctness.\n"
    "  This experiment does not measure accuracy.",
    fontsize=8,
)
fig_bar.tight_layout(rect=(0, 0.11, 1, 1))
display(fig_bar)
Ausgabe

Unter derselben 0.60-Regel erreichte Haiku bei Temperatur 0 100%. TypeSafe erreichte 99.2%, und die übrigen LLM-Bedingungen landeten zwischen 84.2% und 94.2%. TypeSafe gab bei 25.8% der Antworten uncertain zurück und handelte bei den anderen 74.2% automatisch; Haiku bei Temperatur 0 enthielt sich nie. Diese Prozentsätze messen nur die Wiederholbarkeit. Die Tabelle unten stellt die rohe Übereinstimmung und die Enthaltungsraten neben die Policy-Übereinstimmung in diesem Diagramm.

Lass unsichere Wahrscheinlichkeiten eine unsichere Entscheidung erzeugen

Eine kleine Wahrscheinlichkeitsänderung kann zwei dicht beieinanderliegende Labels vertauschen. Die Anwendung muss nicht auf den Gewinner reagieren: Gib uncertain zurück, wenn die Spitzenwahrscheinlichkeit unter 0.60 liegt, und schicke diesen Fall an einen Menschen. Bei genau 0.60 wähle das höchste Label. Dies nutzt die zurückgegebenen Wahrscheinlichkeiten, nicht das separate confidence-Feld der API, und fügt keine Modellaufrufe hinzu.

Der Schwellenwert ist eine beispielhafte Anwendungsrichtlinie, keine kalibrierte Garantie und kein Schwellenwert, der gewählt wurde, um die Übereinstimmung dieses Durchlaufs zu maximieren. Wähle Produktions-Schwellenwerte anhand gekennzeichneter Beispiele sowie der Kosten fehlerhafter Aktionen und menschlicher Prüfung.

Wir wenden dieselbe Regel auf jede Bedingung mit Wahrscheinlichkeitsausgabe an. Single-Pick-LLM-Antworten haben keine Wahrscheinlichkeitsschätzung; ihre synthetischen One-Hot-Vektoren können keine Unsicherheit messen, daher sind sie aus dem Übereinstimmungs-Diagramm und der Tabelle ausgeschlossen.

def agreement_rate(samples: list) -> float:
    """Mean over questions of the raw plurality label's share across all NUM_SAMPLES draws.

    Parse failures count against agreement because a failed route is not a repeated decision.
    """
    shares = []
    for question_key, (_instructions, choices) in QUESTIONS.items():
        labels = list(choices)
        picks = [
            argmax_label(samples[sample][question_key], labels)
            for sample in range(NUM_SAMPLES)
        ]
        picks = [pick for pick in picks if pick is not None]
        if not picks:
            shares.append(0.0)
            continue
        top = Counter(picks).most_common(1)[0][1]
        shares.append(top / NUM_SAMPLES)
    return mean(shares) if shares else float("nan")

# Keep failures separate from abstentions and count conflicting concrete actions per question.
print(
    f"{'condition':<{name_w}}{'raw agree':>12}{'policy agree':>14}"
    f"{'uncertain':>12}{'automatic':>12}{'conflicts':>11}"
)
for label in PROBABILITY_OUTPUT_LABELS:
    decisions = decisions_by_condition[label]
    flat = [value for row in decisions for value in row]
    uncertain_rate = mean(value == "uncertain" for value in flat)
    automatic_rate = mean(value not in (None, "uncertain") for value in flat)
    conflicts = sum(
        len({value for value in row if value not in (None, "uncertain")}) > 1
        for row in decisions
    )
    print(
        f"{label:<{name_w}}{agreement_rate(pooled_runs[label]):>11.1%}"
        f"{policy_agreement_by_condition[label]:>13.1%}{uncertain_rate:>11.1%}"
        f"{automatic_rate:>11.1%}{conflicts:>11}"
    )
condition                            raw agree  policy agree   uncertain   automatic  conflicts
claude-haiku-4-5 t=0                   100.0%       100.0%       0.0%     100.0%          0
claude-haiku-4-5 t=default              87.5%        86.7%       0.8%      98.3%          2
gpt-5.4-mini t=0                        99.2%        87.5%      12.5%      87.5%          0
gpt-5.4-mini t=default                  90.8%        84.2%      22.5%      77.5%          2
gpt-5.5-reasoning                       90.0%        93.3%      30.8%      69.2%          1
claude-opus-4-8-reasoning               92.5%        94.2%      33.3%      66.7%          0
typesafe_choice                         90.8%        99.2%      25.8%      74.2%          0

policy agree zählt uncertain als Entscheidung; Parse-Fehler zählen gegen die Übereinstimmung. automatic ist der Anteil aller Antworten, die ein Label auswählen. conflicts zählt Fragen mit mehr als einem konkreten Label über die Wiederholungen hinweg, Enthaltungen ausgenommen. Diese Maße beschreiben die Wiederholbarkeit und wie oft die Anwendung handelt, nicht ob ihre Aktionen richtig sind.

TypeSafes Übereinstimmung stieg von 90.8% auf 99.2%. Von den Antworten waren 25.8% unsicher und 74.2% automatisch. primary_risk und link_handling kamen bei jeder Wiederholung unsicher zurück; category wechselte zwischen Violence und uncertain und überschritt den Aktions-Schwellenwert mal, mal nicht. Keine Frage erzeugte zwei verschiedene konkrete TypeSafe-Labels. Nichts davon zeigt Genauigkeit oder Überlegenheit: Haiku bei Temperatur 0 hatte hier 100% Übereinstimmung, ohne Enthaltungen.

# Show every TypeSafe decision while retaining the top probability behind it.
policy_decisions = decisions_by_condition[TYPESAFE_LABEL]
policy_values = []
for row, (_key, (_instructions, choices)) in zip(policy_decisions, QUESTIONS.items()):
    labels = list(choices)
    policy_values.append([
        10 if value == "uncertain" else labels.index(value) if value is not None else np.nan
        for value in row
    ])
policy_cmap = ListedColormap([*plt.get_cmap("tab10").colors, "#dddddd"])
policy_cmap.set_bad("white")
fig_policy, ax_policy = plt.subplots(figsize=(13, 4))
ax_policy.imshow(policy_values, cmap=policy_cmap, vmin=0, vmax=10, aspect="auto")
for row_index, key in enumerate(QUESTIONS):
    for sample_index in range(NUM_SAMPLES):
        decision = policy_decisions[row_index][sample_index]
        probability = max(typesafe_runs[sample_index][key])
        ax_policy.text(sample_index, row_index, f"{decision or 'n/a'}\n{probability:.2f}",
                       ha="center", va="center", fontsize=6)
ax_policy.set_yticks(range(len(QUESTIONS)), list(QUESTIONS))
ax_policy.set_xticks(range(NUM_SAMPLES), range(1, NUM_SAMPLES + 1))
ax_policy.set_xlabel("rubric query")
ax_policy.set_title(
    "TypeSafe application decisions: gray means uncertain "
    f"(top probability < {MIN_CHOICE_PROBABILITY:.2f})"
)
fig_policy.tight_layout()
display(fig_policy)
Ausgabe

Diese Richtlinie macht das Modell nicht deterministisch. Enthalten kann konkurrierende Labels durch dasselbe Ergebnis „menschliche Prüfung“ ersetzen, aber eine Wahrscheinlichkeit nahe 0.60 kann noch zwischen einem konkreten Label und uncertain wechseln. Die Wahrscheinlichkeitsstatistik und die Spalte raw agree der Tabelle geben weiterhin die ursprünglichen Modellausgaben wieder.

Im TypeSafe-Playground öffnen

Der Link unten öffnet denselben Beitrag und dieselbe Rubrik im Playground: ein Beitrag, dieselben 8 Choices und TypeSafe jev-latest.

playground_link = make_playground_link(
    {"post": POST},
    {
        key: Choice(instructions=instructions, criteria=choices)
        for key, (instructions, choices) in QUESTIONS.items()
    },
    models=[TYPESAFE_MODEL],
)
display(
    Markdown(
        f"🔗 [Open this post + rubric in the TypeSafe playground]({playground_link})"
    )
)
Öffne diesen Beitrag + diese Rubrik im TypeSafe-Playground →