Перепроверка цитат
Ловит неверные или выдуманные цитаты, сверяясь с исходным документом. Один вопрос Choice решает, подтверждает ли контекст цитаты утверждение.
LLM отвечает на вопрос и прикладывает цитаты: к каждому утверждению — раздел исходного документа и цитату, на которой оно держится. Некоторые из этих цитат неверны или выдуманы: цитаты может вообще не быть в документе, либо она стоит в нём слово в слово, но её контекст говорит противоположное утверждению.
Проверять вручную долго: найти документ, найти в нём цитату, затем прочитать достаточно контекста, чтобы понять, подкрепляет ли она утверждение.
Чтобы автоматизировать эту проверку, мы сначала ищем отсутствующие цитаты обычным
сопоставлением строк, а затем используем вопрос Choice, чтобы прочитать контекст
каждой уцелевшей цитаты и решить, подтверждает ли она утверждение.
%%{init: {"flowchart": {"wrappingWidth": 330}}}%%
flowchart LR
cite["source document + citation"]
match{"is the quote<br/>in the source?"}
fab["mark <b>fabricated</b>"]
subgraph request[" "]
q["Choice — how does the<br/>section relate to the claim?<br/>supports → mark <b>verified</b><br/>contradicts → mark <b>contradicted</b><br/>says nothing → mark <b>unsupported</b>"]
end
gate{"confidence<br/>≥ 0.8?"}
stand["let the verdict stand"]
review["a human confirms it"]
cite --> match
%% the two edges that reach the call come first, so they stay adjacent; the
%% string match's own verdict is declared last and lands below them
match -- "found" --> request
match -- "no quote" --> request
match -- "not found" --> fab
request --> gate
gate --> stand
gate --> review
Ниже через проверку проходят восемь цитат из ответа LLM про RFC 7519 (JSON Web Token).
Четыре верные вернулись как verified с уверенностью 0.93 или выше. Все четыре
подложенных ошибки были пойманы: выдуманная цитата, опровергнутое утверждение и две
неподтверждённые цитаты, отправленные человеку.
check_citation(), функция, которую вы здесь строите, принимает исходный документ и
одну цитату и возвращает один из четырёх вердиктов: verified, unsupported,
contradicted или fabricated. Она также возвращает уверенность, помечающую те
случаи, на которые стоит взглянуть человеку.
Установка
pip install ipython 'cooksafe>=0.2.0,<0.3.0'
затем задайте TYPESAFE_API_KEY. Каждый вызов API кэшируется в json_cache.json, а он
поставляется вместе с cookbook, поэтому повторный запуск воспроизводит опубликованные
числа вместо обращения к API. Удалите этот файл, чтобы запустить всё вживую.
Числа ниже получены на jev-1.12 2026-08-16.
import json
import os
import re
from pathlib import Path
from time import perf_counter
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, TypeSafeClient
TYPESAFE_MODEL = "jev-1.12"
AUTO_ACCEPT = 0.8 # start high for more human review as you build trust in the model
client = TypeSafeClient(
api_key=os.environ.get("TYPESAFE_API_KEY", "cache-only"),
base_url=os.environ.get("TYPESAFE_ENDPOINT"),
timeout=120.0,
)
json_cache = JsonCache(Path("json_cache.json"))
Загрузка источника и цитат
Источник — RFC 7519 (JSON Web Token),
скачанный с rfc-editor.org и закоммиченный рядом с этим cookbook как rfc7519.txt. Код
ниже убирает колонтитулы страниц и разбивает текст на пронумерованные разделы.
Восемь цитат в citations.json написаны LLM по этому RFC. Четыре верны; остальные
четыре мы отредактировали так, чтобы они провалили проверку.
def load_source() -> str:
"""RFC 7519 verbatim, minus the page headers and footers that interrupt its paragraphs."""
lines = []
for line in Path("rfc7519.txt").read_text().splitlines():
bare = line.lstrip("\f")
if re.match(r"Jones, et al\.\s.*\[Page \d+\]$", bare):
continue
if re.match(r"RFC 7519\s+JSON Web Token \(JWT\)\s+May 2015$", bare):
continue
lines.append(bare)
return re.sub(r"\n{3,}", "\n\n", "\n".join(lines))
def split_sections(source: str) -> dict[str, str]:
"""Map each numbered section ("4.1.3") to its text, split on the RFC's header lines."""
boundary = re.compile(r"(?m)^(?:(\d+(?:\.\d+)*)\. .+|Appendix [A-Z]\..*)$")
marks = list(boundary.finditer(source))
sections = {}
for mark, nxt in zip(marks, marks[1:] + [None]):
if mark.group(1) is None: # an appendix header only terminates the section before it
continue
sections[mark.group(1)] = source[mark.start() : nxt.start() if nxt else len(source)].strip()
return sections
SOURCE = load_source()
SECTIONS = split_sections(SOURCE)
CITATIONS = json.loads(Path("citations.json").read_text())
print(f"{len(SOURCE):,} characters, {len(SECTIONS)} numbered sections, {len(CITATIONS)} citations")
print("\nA citation with a quote:")
print(json.dumps(CITATIONS[1], indent=2))
print("\nA claim-only citation:")
print(json.dumps(next(c for c in CITATIONS if c["quote"] is None), indent=2))
58,365 characters, 45 numbered sections, 8 citations
A citation with a quote:
{
"id": "aud_reject",
"claim": "If a validator does not find itself in a token's audience list, it has to reject the token.",
"quote": "If the principal processing the claim does not identify itself with a value in the \"aud\" claim when this claim is present, then the JWT MUST be rejected.",
"section": "4.1.3"
}
A claim-only citation:
{
"id": "iat_future",
"claim": "The \"iat\" claim requires validators to reject tokens whose issue time is in the future.",
"quote": null,
"section": "4.1.6"
}
Поиск каждой цитаты в источнике
Цитата, которой нет в источнике, выдумана, и чтобы это выяснить, модель не нужна. Нормализуйте пробелы и фигурные кавычки, чтобы цитата по-прежнему совпадала вопреки переносам строк в RFC, затем ищите её как подстроку. Совпадение также говорит, из какого раздела пришла цитата, и этот раздел — тот текст, который модель читает на следующем шаге.
Цитата может называть раздел, ничего из него не цитируя. В этом случае сопоставлять нечего, поэтому берётся раздел, названный в цитате, и мы идём прямо к модели.
def normalize(text: str) -> str:
"""Collapse whitespace and fold curly quotes, so a quote matches across line wraps."""
table = str.maketrans({"“": '"', "”": '"', "‘": "'", "’": "'"})
return re.sub(r"\s+", " ", text.translate(table)).strip()
def find_quote(sections: dict[str, str], quote: str) -> str | None:
"""The number of the section that contains the quote verbatim, or None."""
needle = normalize(quote)
for number in sorted(sections, key=lambda n: [int(p) for p in n.split(".")]):
if needle in normalize(sections[number]):
return number
return None
def locate(sections: dict[str, str], citation: dict) -> tuple[str, str | None]:
"""Step 1 for one citation: a status, plus the section step 2 will read."""
if citation["quote"] is None:
return "section-only", sections[citation["section"]]
number = find_quote(sections, citation["quote"])
if number is None:
return "missing", None
return "found", sections[number]
for citation in CITATIONS:
status, section = locate(SECTIONS, citation)
where = f"section of {len(section):,} chars" if section else "not in the source"
print(f"{citation['id']:<18}{status:<14}{where}")
epoch_seconds found section of 3,122 chars
aud_reject found section of 761 chars
sig_reporting missing not in the source
clock_skew found section of 529 chars
exp_required found section of 529 chars
pii_encryption found section of 1,653 chars
iat_future section-only section of 270 chars
duplicate_names found section of 918 chars
Проверка, подтверждает ли источник утверждение
Если у цитаты к этому моменту всё ещё есть процитированный фрагмент, он совпадает с источником слово в слово. Этого недостаточно: сам фрагмент может быть точным, а утверждение, построенное поверх него, — всё ещё неверным. Чтобы это решить, нужен контекст фрагмента — раздел, найденный на шаге 1.
По одному вопросу Choice на каждую уцелевшую цитату покрывает три способа, которыми
раздел может относиться к утверждению.
Вариант с наибольшей вероятностью и есть вердикт, а AUTO_ACCEPT (0.8 в коде выше)
решает, что с ним будет дальше:
- уверенность 0.8 или выше: вердикт стоит сам по себе;
- ниже 0.8: человек подтверждает вердикт, прежде чем на него что-то отреагирует.
Начинайте с высокого порога и снижайте его по мере того, как увидите, как модель показывает себя на ваших собственных документах.
QUESTIONS = {
"relation": Choice(
instructions="How does the section relate to the claim?",
criteria={
"supports": "The section states the claim or directly implies that it is true",
"contradicts": "The section states the opposite of the claim or implies it is false",
"says_nothing": "The section does not address what the claim asserts, either way",
},
),
}
RELATION_TO_VERDICT = {
"supports": "verified",
"contradicts": "contradicted",
"says_nothing": "unsupported",
}
@json_cache
def ask(claim: str, section: str) -> dict:
started = perf_counter()
response = client.system_one(
state={"claim": claim, "section": section},
questions=QUESTIONS,
model=TYPESAFE_MODEL,
)
answer = response.answers["relation"]
return {
"choice": answer.choice,
"probabilities": answer.probabilities,
"confidence": answer.confidence,
"seconds": round(perf_counter() - started, 2),
"input_tokens": response.usage.input_tokens or 0,
"output_tokens": response.usage.output_tokens or 0,
}
def verdict(status: str, answer: dict | None) -> dict:
"""Fold step 1 and step 2 into one of the four labels, plus an auto-or-review flag."""
if status == "missing":
# confidence None: no model was called, so there is no model confidence to report
return {"verdict": "fabricated", "confidence": None, "auto": True}
return {
"verdict": RELATION_TO_VERDICT[answer["choice"]],
"confidence": answer["confidence"],
"auto": answer["confidence"] >= AUTO_ACCEPT,
}
def check_citation(sections: dict[str, str], citation: dict) -> dict:
status, section = locate(sections, citation)
answer = ask(citation["claim"], section) if section is not None else None
return {"id": citation["id"], "status": status, "answer": answer, **verdict(status, answer)}
Проверка всех цитат
Все восемь цитат через одну и ту же проверку:
print(f"{'citation':<18}{'quote':<14}{'relation':<14}{'conf':>6} {'verdict':<13}{'action':>7}")
for citation in CITATIONS:
result = check_citation(SECTIONS, citation)
answer = result["answer"]
relation = answer["choice"] if answer else "-"
conf = f"{answer['confidence']:.2f}" if answer else "-"
action = "auto" if result["auto"] else "review"
print(
f"{result['id']:<18}{result['status']:<14}{relation:<14}{conf:>6}"
f" {result['verdict']:<13}{action:>7}"
)
citation quote relation conf verdict action
epoch_seconds found supports 0.93 verified auto
aud_reject found supports 0.95 verified auto
sig_reporting missing - - fabricated auto
clock_skew found supports 0.99 verified auto
exp_required found contradicts 0.99 contradicted auto
pii_encryption found says_nothing 0.27 unsupported review
iat_future section-only says_nothing 0.56 unsupported review
duplicate_names found supports 0.99 verified auto
Четыре цитаты вернулись как verified, одна — как fabricated, одна — как
contradicted и две — как unsupported.
epoch_seconds,aud_reject,clock_skewиduplicate_names— те самые четыре верные. Все они вернулись какverifiedс уверенностью 0.93 или выше, намного вышеAUTO_ACCEPT.sig_reportingвообще не дошла до модели. Её цитаты нет в RFC, поэтому одного сопоставления строк хватает, чтобы пометить её какfabricated.exp_requiredцитирует раздел 4.1.4 слово в слово, а тот же раздел говорит: «Use of this claim is OPTIONAL», поэтому онаcontradicted, с уверенностью 0.99.pii_encryptionиiat_futureвернулись какunsupportedс 0.27 и 0.56, обе ниже порога, поэтому обе ушли к человеку.pii_encryptionпоказывает, почему одного сопоставления строк недостаточно: её цитата есть в источнике слово в слово, а раздел, из которого она пришла, ничего не говорит об утверждении.
Чтобы нацелить это на свои данные, замените rfc7519.txt и citations.json.
load_source() и split_sections() написаны под вёрстку RFC, поэтому документ другой
формы нуждается в собственном разборе.
Сопоставление строк после нормализации точное: цитата, которая обрезана или слегка
переформулирована, возвращается как fabricated. Продакшн-системе, которая терпит
неаккуратное цитирование, вместо этого понадобилось бы нечёткое сопоставление.
Открыть в playground
Ссылка хранит утверждение и раздел одной цитаты, а также вопрос. Откройте её, чтобы выполнить тот же вызов вживую в браузере.
example = next(c for c in CITATIONS if c["id"] == "exp_required")
_, example_section = locate(SECTIONS, example)
playground_link = make_playground_link(
{"claim": example["claim"], "section": example_section}, QUESTIONS, models=[TYPESAFE_MODEL]
)
display(Markdown(f"🔗 [Open one citation's claim + section in the TypeSafe playground]({playground_link})"))
Открыть утверждение и раздел одной цитаты в playground TypeSafe →