Zitate gegenprüfen
Erkennt falsche oder halluzinierte Zitate durch Abgleich mit dem Quelldokument. Eine Choice-Frage entscheidet, ob der Kontext des Zitats die Aussage stützt.
Ein LLM beantwortet eine Frage und hängt Zitate an: für jede Aussage einen Abschnitt eines Quelldokuments und das Zitat, auf dem sie beruht. Einige dieser Zitate sind falsch oder halluziniert: Das Zitat kann im Dokument völlig fehlen, oder wörtlich darin stehen, während sein Kontext das Gegenteil der Aussage sagt.
Eines von Hand zu prüfen ist langsam: das Dokument finden, das Zitat darin finden und dann genug von seinem Kontext lesen, um zu erkennen, ob es die Aussage stützt.
Um diese Prüfung zu automatisieren, suchen wir zuerst mit einem gewöhnlichen
String-Vergleich nach fehlenden Zitaten und verwenden dann eine Choice-Frage, um den
Kontext jedes verbleibenden Zitats zu lesen und zu entscheiden, ob er die Aussage stützt.
%%{init: {"flowchart": {"wrappingWidth": 330}}}%%
flowchart LR
cite["source document + citation"]
match{"is the quote<br/>in the source?"}
fab["mark <b>fabricated</b>"]
subgraph request[" "]
q["Choice — how does the<br/>section relate to the claim?<br/>supports → mark <b>verified</b><br/>contradicts → mark <b>contradicted</b><br/>says nothing → mark <b>unsupported</b>"]
end
gate{"confidence<br/>≥ 0.8?"}
stand["let the verdict stand"]
review["a human confirms it"]
cite --> match
%% the two edges that reach the call come first, so they stay adjacent; the
%% string match's own verdict is declared last and lands below them
match -- "found" --> request
match -- "no quote" --> request
match -- "not found" --> fab
request --> gate
gate --> stand
gate --> review
Unten durchlaufen acht Zitate aus der Antwort eines LLM über RFC 7519 (JSON Web Token) die
Prüfung. Die vier korrekten kamen als verified mit einer Konfidenz von 0.93 oder höher
zurück. Alle vier eingepflanzten Fehler wurden erkannt: ein erfundenes Zitat, eine
widersprochene Aussage und zwei nicht gestützte Zitate, die an einen Menschen gingen.
check_citation(), die Funktion, die du hier baust, nimmt ein Quelldokument und ein Zitat
und gibt eines von vier Urteilen zurück: verified, unsupported, contradicted oder
fabricated. Außerdem gibt sie eine Konfidenz zurück, die diejenigen markiert, die ein
Mensch ansehen sollte.
Einrichtung
pip install ipython 'cooksafe>=0.2.0,<0.3.0'
Setze dann TYPESAFE_API_KEY. Jeder API-Aufruf wird in json_cache.json
zwischengespeichert, das mit dem Cookbook ausgeliefert wird, sodass ein erneutes Ausführen
die veröffentlichten Zahlen wiedergibt, statt die API aufzurufen. Lösche diese Datei, um
alles live auszuführen.
Die Zahlen unten stammen aus jev-1.12 vom 2026-08-16.
import json
import os
import re
from pathlib import Path
from time import perf_counter
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, TypeSafeClient
TYPESAFE_MODEL = "jev-1.12"
AUTO_ACCEPT = 0.8 # start high for more human review as you build trust in the model
client = TypeSafeClient(
api_key=os.environ.get("TYPESAFE_API_KEY", "cache-only"),
base_url=os.environ.get("TYPESAFE_ENDPOINT"),
timeout=120.0,
)
json_cache = JsonCache(Path("json_cache.json"))
Lade die Quelle und die Zitate
Die Quelle ist RFC 7519 (JSON Web Token), von
rfc-editor.org geholt und neben diesem Cookbook als rfc7519.txt abgelegt. Der Code unten
entfernt die Seitenkopf- und -fußzeilen und teilt den Text dann in nummerierte Abschnitte.
Die acht Zitate in citations.json wurden von einem LLM gegen den RFC geschrieben. Vier
sind korrekt; die anderen vier haben wir so bearbeitet, dass sie die Prüfung nicht bestehen.
def load_source() -> str:
"""RFC 7519 verbatim, minus the page headers and footers that interrupt its paragraphs."""
lines = []
for line in Path("rfc7519.txt").read_text().splitlines():
bare = line.lstrip("\f")
if re.match(r"Jones, et al\.\s.*\[Page \d+\]$", bare):
continue
if re.match(r"RFC 7519\s+JSON Web Token \(JWT\)\s+May 2015$", bare):
continue
lines.append(bare)
return re.sub(r"\n{3,}", "\n\n", "\n".join(lines))
def split_sections(source: str) -> dict[str, str]:
"""Map each numbered section ("4.1.3") to its text, split on the RFC's header lines."""
boundary = re.compile(r"(?m)^(?:(\d+(?:\.\d+)*)\. .+|Appendix [A-Z]\..*)$")
marks = list(boundary.finditer(source))
sections = {}
for mark, nxt in zip(marks, marks[1:] + [None]):
if mark.group(1) is None: # an appendix header only terminates the section before it
continue
sections[mark.group(1)] = source[mark.start() : nxt.start() if nxt else len(source)].strip()
return sections
SOURCE = load_source()
SECTIONS = split_sections(SOURCE)
CITATIONS = json.loads(Path("citations.json").read_text())
print(f"{len(SOURCE):,} characters, {len(SECTIONS)} numbered sections, {len(CITATIONS)} citations")
print("\nA citation with a quote:")
print(json.dumps(CITATIONS[1], indent=2))
print("\nA claim-only citation:")
print(json.dumps(next(c for c in CITATIONS if c["quote"] is None), indent=2))
58,365 characters, 45 numbered sections, 8 citations
A citation with a quote:
{
"id": "aud_reject",
"claim": "If a validator does not find itself in a token's audience list, it has to reject the token.",
"quote": "If the principal processing the claim does not identify itself with a value in the \"aud\" claim when this claim is present, then the JWT MUST be rejected.",
"section": "4.1.3"
}
A claim-only citation:
{
"id": "iat_future",
"claim": "The \"iat\" claim requires validators to reject tokens whose issue time is in the future.",
"quote": null,
"section": "4.1.6"
}
Finde jedes Zitat in der Quelle
Ein Zitat, das nicht in der Quelle steht, ist erfunden, und dafür ist kein Modell nötig. Normalisiere Leerzeichen und typografische Anführungszeichen, damit ein Zitat über die Zeilenumbrüche des RFC hinweg weiterhin übereinstimmt, und suche es dann als Teilstring. Eine Übereinstimmung sagt außerdem, aus welchem Abschnitt das Zitat stammt, und dieser Abschnitt ist der Text, den das Modell im nächsten Schritt liest.
Ein Zitat kann einen Abschnitt nennen, ohne etwas daraus zu zitieren. In diesem Fall gibt es nichts abzugleichen, also nimm den Abschnitt, den das Zitat nennt, und gehe direkt zum Modell.
def normalize(text: str) -> str:
"""Collapse whitespace and fold curly quotes, so a quote matches across line wraps."""
table = str.maketrans({"“": '"', "”": '"', "‘": "'", "’": "'"})
return re.sub(r"\s+", " ", text.translate(table)).strip()
def find_quote(sections: dict[str, str], quote: str) -> str | None:
"""The number of the section that contains the quote verbatim, or None."""
needle = normalize(quote)
for number in sorted(sections, key=lambda n: [int(p) for p in n.split(".")]):
if needle in normalize(sections[number]):
return number
return None
def locate(sections: dict[str, str], citation: dict) -> tuple[str, str | None]:
"""Step 1 for one citation: a status, plus the section step 2 will read."""
if citation["quote"] is None:
return "section-only", sections[citation["section"]]
number = find_quote(sections, citation["quote"])
if number is None:
return "missing", None
return "found", sections[number]
for citation in CITATIONS:
status, section = locate(SECTIONS, citation)
where = f"section of {len(section):,} chars" if section else "not in the source"
print(f"{citation['id']:<18}{status:<14}{where}")
epoch_seconds found section of 3,122 chars
aud_reject found section of 761 chars
sig_reporting missing not in the source
clock_skew found section of 529 chars
exp_required found section of 529 chars
pii_encryption found section of 1,653 chars
iat_future section-only section of 270 chars
duplicate_names found section of 918 chars
Prüfe, ob die Quelle die Aussage stützt
Ein Zitat, das an dieser Stelle noch einen Belegtext hat, stimmt wörtlich mit der Quelle überein. Das genügt nicht: Das Zitat kann korrekt sein und die darauf aufgebaute Aussage trotzdem falsch. Um das zu entscheiden, braucht es den Kontext des Zitats, den Abschnitt, den Schritt 1 gefunden hat.
Eine Choice-Frage pro verbleibendem Zitat deckt die drei Arten ab, wie ein Abschnitt zu
einer Aussage stehen kann.
Die Option mit der höchsten Wahrscheinlichkeit ist das Urteil, und AUTO_ACCEPT (0.8 im
Code oben) entscheidet, was damit geschieht:
- Konfidenz von 0.8 oder höher: das Urteil steht für sich;
- unter 0.8: ein Mensch bestätigt das Urteil, bevor etwas darauf reagiert.
Starte hoch und senke den Schwellenwert, sobald du siehst, wie das Modell mit deinen eigenen Dokumenten umgeht.
QUESTIONS = {
"relation": Choice(
instructions="How does the section relate to the claim?",
criteria={
"supports": "The section states the claim or directly implies that it is true",
"contradicts": "The section states the opposite of the claim or implies it is false",
"says_nothing": "The section does not address what the claim asserts, either way",
},
),
}
RELATION_TO_VERDICT = {
"supports": "verified",
"contradicts": "contradicted",
"says_nothing": "unsupported",
}
@json_cache
def ask(claim: str, section: str) -> dict:
started = perf_counter()
response = client.system_one(
state={"claim": claim, "section": section},
questions=QUESTIONS,
model=TYPESAFE_MODEL,
)
answer = response.answers["relation"]
return {
"choice": answer.choice,
"probabilities": answer.probabilities,
"confidence": answer.confidence,
"seconds": round(perf_counter() - started, 2),
"input_tokens": response.usage.input_tokens or 0,
"output_tokens": response.usage.output_tokens or 0,
}
def verdict(status: str, answer: dict | None) -> dict:
"""Fold step 1 and step 2 into one of the four labels, plus an auto-or-review flag."""
if status == "missing":
# confidence None: no model was called, so there is no model confidence to report
return {"verdict": "fabricated", "confidence": None, "auto": True}
return {
"verdict": RELATION_TO_VERDICT[answer["choice"]],
"confidence": answer["confidence"],
"auto": answer["confidence"] >= AUTO_ACCEPT,
}
def check_citation(sections: dict[str, str], citation: dict) -> dict:
status, section = locate(sections, citation)
answer = ask(citation["claim"], section) if section is not None else None
return {"id": citation["id"], "status": status, "answer": answer, **verdict(status, answer)}
Prüfe jedes Zitat
Alle acht Zitate durch dieselbe Prüfung:
print(f"{'citation':<18}{'quote':<14}{'relation':<14}{'conf':>6} {'verdict':<13}{'action':>7}")
for citation in CITATIONS:
result = check_citation(SECTIONS, citation)
answer = result["answer"]
relation = answer["choice"] if answer else "-"
conf = f"{answer['confidence']:.2f}" if answer else "-"
action = "auto" if result["auto"] else "review"
print(
f"{result['id']:<18}{result['status']:<14}{relation:<14}{conf:>6}"
f" {result['verdict']:<13}{action:>7}"
)
citation quote relation conf verdict action
epoch_seconds found supports 0.93 verified auto
aud_reject found supports 0.95 verified auto
sig_reporting missing - - fabricated auto
clock_skew found supports 0.99 verified auto
exp_required found contradicts 0.99 contradicted auto
pii_encryption found says_nothing 0.27 unsupported review
iat_future section-only says_nothing 0.56 unsupported review
duplicate_names found supports 0.99 verified auto
Vier Zitate kamen als verified zurück, eines als fabricated, eines als contradicted
und zwei als unsupported.
epoch_seconds,aud_reject,clock_skewundduplicate_namessind die vier korrekten. Alle kamen alsverifiedmit einer Konfidenz von 0.93 oder höher zurück, weit überAUTO_ACCEPT.sig_reportingerreichte das Modell nie. Sein Zitat steht nicht im RFC, also markiert der String-Vergleich allein es alsfabricated.exp_requiredzitiert Abschnitt 4.1.4 wörtlich, und derselbe Abschnitt sagt „Use of this claim is OPTIONAL“, also ist escontradicted, mit einer Konfidenz von 0.99.pii_encryptionundiat_futurekamen alsunsupportedmit 0.27 und 0.56 zurück, beide unter dem Schwellenwert, also gingen beide an einen Menschen.pii_encryptionzeigt, warum der String-Vergleich für sich allein nicht ausreicht: Sein Zitat steht wörtlich in der Quelle, und der Abschnitt, aus dem es stammt, sagt nichts über die Aussage.
Um das auf deine eigenen Daten zu richten, ersetze rfc7519.txt und citations.json.
load_source() und split_sections() sind für das Layout eines RFC geschrieben, ein
Dokument anderer Form braucht also seine eigene Zerlegung.
Der String-Vergleich ist nach der Normalisierung exakt: Ein Zitat, das abgeschnitten oder
leicht umformuliert ist, kommt als fabricated zurück. Ein Produktionssystem, das
unsauberes Zitieren toleriert, bräuchte stattdessen Fuzzy-Matching.
Im Playground öffnen
Der Link enthält die Aussage und den Abschnitt eines Zitats sowie die Frage. Öffne ihn, um denselben Aufruf live im Browser auszuführen.
example = next(c for c in CITATIONS if c["id"] == "exp_required")
_, example_section = locate(SECTIONS, example)
playground_link = make_playground_link(
{"claim": example["claim"], "section": example_section}, QUESTIONS, models=[TYPESAFE_MODEL]
)
display(Markdown(f"🔗 [Open one citation's claim + section in the TypeSafe playground]({playground_link})"))
Öffne die Aussage + den Abschnitt eines Zitats im TypeSafe-Playground →