Dokumentation

Zitate gegenprüfen

Erkennt falsche oder halluzinierte Zitate durch Abgleich mit dem Quelldokument. Eine Choice-Frage entscheidet, ob der Kontext des Zitats die Aussage stützt.

Ein LLM beantwortet eine Frage und hängt Zitate an: für jede Aussage einen Abschnitt eines Quelldokuments und das Zitat, auf dem sie beruht. Einige dieser Zitate sind falsch oder halluziniert: Das Zitat kann im Dokument völlig fehlen, oder wörtlich darin stehen, während sein Kontext das Gegenteil der Aussage sagt.

Eines von Hand zu prüfen ist langsam: das Dokument finden, das Zitat darin finden und dann genug von seinem Kontext lesen, um zu erkennen, ob es die Aussage stützt.

Um diese Prüfung zu automatisieren, suchen wir zuerst mit einem gewöhnlichen String-Vergleich nach fehlenden Zitaten und verwenden dann eine Choice-Frage, um den Kontext jedes verbleibenden Zitats zu lesen und zu entscheiden, ob er die Aussage stützt.

  %%{init: {"flowchart": {"wrappingWidth": 330}}}%%
flowchart LR
    cite["source document + citation"]

    match{"is the quote<br/>in the source?"}
    fab["mark <b>fabricated</b>"]

    subgraph request[" "]
        q["Choice &mdash; how does the<br/>section relate to the claim?<br/>supports &rarr; mark <b>verified</b><br/>contradicts &rarr; mark <b>contradicted</b><br/>says nothing &rarr; mark <b>unsupported</b>"]
    end

    gate{"confidence<br/>&ge; 0.8?"}
    stand["let the verdict stand"]
    review["a human confirms it"]

    cite --> match
    %% the two edges that reach the call come first, so they stay adjacent; the
    %% string match's own verdict is declared last and lands below them
    match -- "found" --> request
    match -- "no quote" --> request
    match -- "not found" --> fab
    request --> gate
    gate --> stand
    gate --> review

Unten durchlaufen acht Zitate aus der Antwort eines LLM über RFC 7519 (JSON Web Token) die Prüfung. Die vier korrekten kamen als verified mit einer Konfidenz von 0.93 oder höher zurück. Alle vier eingepflanzten Fehler wurden erkannt: ein erfundenes Zitat, eine widersprochene Aussage und zwei nicht gestützte Zitate, die an einen Menschen gingen.

check_citation(), die Funktion, die du hier baust, nimmt ein Quelldokument und ein Zitat und gibt eines von vier Urteilen zurück: verified, unsupported, contradicted oder fabricated. Außerdem gibt sie eine Konfidenz zurück, die diejenigen markiert, die ein Mensch ansehen sollte.

Einrichtung

pip install ipython 'cooksafe>=0.2.0,<0.3.0'

Setze dann TYPESAFE_API_KEY. Jeder API-Aufruf wird in json_cache.json zwischengespeichert, das mit dem Cookbook ausgeliefert wird, sodass ein erneutes Ausführen die veröffentlichten Zahlen wiedergibt, statt die API aufzurufen. Lösche diese Datei, um alles live auszuführen.

Die Zahlen unten stammen aus jev-1.12 vom 2026-08-16.

import json
import os
import re
from pathlib import Path
from time import perf_counter

from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, TypeSafeClient

TYPESAFE_MODEL = "jev-1.12"
AUTO_ACCEPT = 0.8  # start high for more human review as you build trust in the model

client = TypeSafeClient(
    api_key=os.environ.get("TYPESAFE_API_KEY", "cache-only"),
    base_url=os.environ.get("TYPESAFE_ENDPOINT"),
    timeout=120.0,
)
json_cache = JsonCache(Path("json_cache.json"))

Lade die Quelle und die Zitate

Die Quelle ist RFC 7519 (JSON Web Token), von rfc-editor.org geholt und neben diesem Cookbook als rfc7519.txt abgelegt. Der Code unten entfernt die Seitenkopf- und -fußzeilen und teilt den Text dann in nummerierte Abschnitte.

Die acht Zitate in citations.json wurden von einem LLM gegen den RFC geschrieben. Vier sind korrekt; die anderen vier haben wir so bearbeitet, dass sie die Prüfung nicht bestehen.

def load_source() -> str:
    """RFC 7519 verbatim, minus the page headers and footers that interrupt its paragraphs."""
    lines = []
    for line in Path("rfc7519.txt").read_text().splitlines():
        bare = line.lstrip("\f")
        if re.match(r"Jones, et al\.\s.*\[Page \d+\]$", bare):
            continue
        if re.match(r"RFC 7519\s+JSON Web Token \(JWT\)\s+May 2015$", bare):
            continue
        lines.append(bare)
    return re.sub(r"\n{3,}", "\n\n", "\n".join(lines))

def split_sections(source: str) -> dict[str, str]:
    """Map each numbered section ("4.1.3") to its text, split on the RFC's header lines."""
    boundary = re.compile(r"(?m)^(?:(\d+(?:\.\d+)*)\.  .+|Appendix [A-Z]\..*)$")
    marks = list(boundary.finditer(source))
    sections = {}
    for mark, nxt in zip(marks, marks[1:] + [None]):
        if mark.group(1) is None:  # an appendix header only terminates the section before it
            continue
        sections[mark.group(1)] = source[mark.start() : nxt.start() if nxt else len(source)].strip()
    return sections

SOURCE = load_source()
SECTIONS = split_sections(SOURCE)
CITATIONS = json.loads(Path("citations.json").read_text())

print(f"{len(SOURCE):,} characters, {len(SECTIONS)} numbered sections, {len(CITATIONS)} citations")
print("\nA citation with a quote:")
print(json.dumps(CITATIONS[1], indent=2))
print("\nA claim-only citation:")
print(json.dumps(next(c for c in CITATIONS if c["quote"] is None), indent=2))
58,365 characters, 45 numbered sections, 8 citations

A citation with a quote:
{
  "id": "aud_reject",
  "claim": "If a validator does not find itself in a token's audience list, it has to reject the token.",
  "quote": "If the principal processing the claim does not identify itself with a value in the \"aud\" claim when this claim is present, then the JWT MUST be rejected.",
  "section": "4.1.3"
}

A claim-only citation:
{
  "id": "iat_future",
  "claim": "The \"iat\" claim requires validators to reject tokens whose issue time is in the future.",
  "quote": null,
  "section": "4.1.6"
}

Finde jedes Zitat in der Quelle

Ein Zitat, das nicht in der Quelle steht, ist erfunden, und dafür ist kein Modell nötig. Normalisiere Leerzeichen und typografische Anführungszeichen, damit ein Zitat über die Zeilenumbrüche des RFC hinweg weiterhin übereinstimmt, und suche es dann als Teilstring. Eine Übereinstimmung sagt außerdem, aus welchem Abschnitt das Zitat stammt, und dieser Abschnitt ist der Text, den das Modell im nächsten Schritt liest.

Ein Zitat kann einen Abschnitt nennen, ohne etwas daraus zu zitieren. In diesem Fall gibt es nichts abzugleichen, also nimm den Abschnitt, den das Zitat nennt, und gehe direkt zum Modell.

def normalize(text: str) -> str:
    """Collapse whitespace and fold curly quotes, so a quote matches across line wraps."""
    table = str.maketrans({"“": '"', "”": '"', "‘": "'", "’": "'"})
    return re.sub(r"\s+", " ", text.translate(table)).strip()

def find_quote(sections: dict[str, str], quote: str) -> str | None:
    """The number of the section that contains the quote verbatim, or None."""
    needle = normalize(quote)
    for number in sorted(sections, key=lambda n: [int(p) for p in n.split(".")]):
        if needle in normalize(sections[number]):
            return number
    return None

def locate(sections: dict[str, str], citation: dict) -> tuple[str, str | None]:
    """Step 1 for one citation: a status, plus the section step 2 will read."""
    if citation["quote"] is None:
        return "section-only", sections[citation["section"]]
    number = find_quote(sections, citation["quote"])
    if number is None:
        return "missing", None
    return "found", sections[number]

for citation in CITATIONS:
    status, section = locate(SECTIONS, citation)
    where = f"section of {len(section):,} chars" if section else "not in the source"
    print(f"{citation['id']:<18}{status:<14}{where}")
epoch_seconds     found         section of 3,122 chars
aud_reject        found         section of 761 chars
sig_reporting     missing       not in the source
clock_skew        found         section of 529 chars
exp_required      found         section of 529 chars
pii_encryption    found         section of 1,653 chars
iat_future        section-only  section of 270 chars
duplicate_names   found         section of 918 chars

Prüfe, ob die Quelle die Aussage stützt

Ein Zitat, das an dieser Stelle noch einen Belegtext hat, stimmt wörtlich mit der Quelle überein. Das genügt nicht: Das Zitat kann korrekt sein und die darauf aufgebaute Aussage trotzdem falsch. Um das zu entscheiden, braucht es den Kontext des Zitats, den Abschnitt, den Schritt 1 gefunden hat.

Eine Choice-Frage pro verbleibendem Zitat deckt die drei Arten ab, wie ein Abschnitt zu einer Aussage stehen kann. Die Option mit der höchsten Wahrscheinlichkeit ist das Urteil, und AUTO_ACCEPT (0.8 im Code oben) entscheidet, was damit geschieht:

  • Konfidenz von 0.8 oder höher: das Urteil steht für sich;
  • unter 0.8: ein Mensch bestätigt das Urteil, bevor etwas darauf reagiert.

Starte hoch und senke den Schwellenwert, sobald du siehst, wie das Modell mit deinen eigenen Dokumenten umgeht.

QUESTIONS = {
    "relation": Choice(
        instructions="How does the section relate to the claim?",
        criteria={
            "supports": "The section states the claim or directly implies that it is true",
            "contradicts": "The section states the opposite of the claim or implies it is false",
            "says_nothing": "The section does not address what the claim asserts, either way",
        },
    ),
}

RELATION_TO_VERDICT = {
    "supports": "verified",
    "contradicts": "contradicted",
    "says_nothing": "unsupported",
}

@json_cache
def ask(claim: str, section: str) -> dict:
    started = perf_counter()
    response = client.system_one(
        state={"claim": claim, "section": section},
        questions=QUESTIONS,
        model=TYPESAFE_MODEL,
    )
    answer = response.answers["relation"]
    return {
        "choice": answer.choice,
        "probabilities": answer.probabilities,
        "confidence": answer.confidence,
        "seconds": round(perf_counter() - started, 2),
        "input_tokens": response.usage.input_tokens or 0,
        "output_tokens": response.usage.output_tokens or 0,
    }

def verdict(status: str, answer: dict | None) -> dict:
    """Fold step 1 and step 2 into one of the four labels, plus an auto-or-review flag."""
    if status == "missing":
        # confidence None: no model was called, so there is no model confidence to report
        return {"verdict": "fabricated", "confidence": None, "auto": True}
    return {
        "verdict": RELATION_TO_VERDICT[answer["choice"]],
        "confidence": answer["confidence"],
        "auto": answer["confidence"] >= AUTO_ACCEPT,
    }

def check_citation(sections: dict[str, str], citation: dict) -> dict:
    status, section = locate(sections, citation)
    answer = ask(citation["claim"], section) if section is not None else None
    return {"id": citation["id"], "status": status, "answer": answer, **verdict(status, answer)}

Prüfe jedes Zitat

Alle acht Zitate durch dieselbe Prüfung:

print(f"{'citation':<18}{'quote':<14}{'relation':<14}{'conf':>6}  {'verdict':<13}{'action':>7}")
for citation in CITATIONS:
    result = check_citation(SECTIONS, citation)
    answer = result["answer"]
    relation = answer["choice"] if answer else "-"
    conf = f"{answer['confidence']:.2f}" if answer else "-"
    action = "auto" if result["auto"] else "review"
    print(
        f"{result['id']:<18}{result['status']:<14}{relation:<14}{conf:>6}"
        f"  {result['verdict']:<13}{action:>7}"
    )
citation          quote         relation        conf  verdict       action
epoch_seconds     found         supports        0.93  verified        auto
aud_reject        found         supports        0.95  verified        auto
sig_reporting     missing       -                  -  fabricated      auto
clock_skew        found         supports        0.99  verified        auto
exp_required      found         contradicts     0.99  contradicted    auto
pii_encryption    found         says_nothing    0.27  unsupported   review
iat_future        section-only  says_nothing    0.56  unsupported   review
duplicate_names   found         supports        0.99  verified        auto

Vier Zitate kamen als verified zurück, eines als fabricated, eines als contradicted und zwei als unsupported.

  • epoch_seconds, aud_reject, clock_skew und duplicate_names sind die vier korrekten. Alle kamen als verified mit einer Konfidenz von 0.93 oder höher zurück, weit über AUTO_ACCEPT.
  • sig_reporting erreichte das Modell nie. Sein Zitat steht nicht im RFC, also markiert der String-Vergleich allein es als fabricated.
  • exp_required zitiert Abschnitt 4.1.4 wörtlich, und derselbe Abschnitt sagt „Use of this claim is OPTIONAL“, also ist es contradicted, mit einer Konfidenz von 0.99.
  • pii_encryption und iat_future kamen als unsupported mit 0.27 und 0.56 zurück, beide unter dem Schwellenwert, also gingen beide an einen Menschen. pii_encryption zeigt, warum der String-Vergleich für sich allein nicht ausreicht: Sein Zitat steht wörtlich in der Quelle, und der Abschnitt, aus dem es stammt, sagt nichts über die Aussage.

Um das auf deine eigenen Daten zu richten, ersetze rfc7519.txt und citations.json. load_source() und split_sections() sind für das Layout eines RFC geschrieben, ein Dokument anderer Form braucht also seine eigene Zerlegung.

Der String-Vergleich ist nach der Normalisierung exakt: Ein Zitat, das abgeschnitten oder leicht umformuliert ist, kommt als fabricated zurück. Ein Produktionssystem, das unsauberes Zitieren toleriert, bräuchte stattdessen Fuzzy-Matching.

Im Playground öffnen

Der Link enthält die Aussage und den Abschnitt eines Zitats sowie die Frage. Öffne ihn, um denselben Aufruf live im Browser auszuführen.

example = next(c for c in CITATIONS if c["id"] == "exp_required")
_, example_section = locate(SECTIONS, example)
playground_link = make_playground_link(
    {"claim": example["claim"], "section": example_section}, QUESTIONS, models=[TYPESAFE_MODEL]
)
display(Markdown(f"🔗 [Open one citation's claim + section in the TypeSafe playground]({playground_link})"))
Öffne die Aussage + den Abschnitt eines Zitats im TypeSafe-Playground →