인용 재검증
원본 문서와 대조하여 틀렸거나 환각된 인용을 잡아냅니다. Choice 질문 하나로 인용문이 놓인 컨텍스트가 해당 주장을 뒷받침하는지 판단합니다.
LLM은 질문에 답하면서 인용을 붙입니다. 즉 각 주장마다 원본 문서의 한 절과 그 주장이 근거로 삼는 인용문을 답니다. 그중 일부 인용은 틀렸거나 환각입니다. 인용문이 문서에 아예 없을 수도 있고, 문서에 한 글자도 틀리지 않고 들어 있지만 그것이 놓인 컨텍스트가 주장과 반대되는 내용을 말할 수도 있습니다.
이런 인용을 사람이 하나씩 확인하는 것은 느립니다. 문서를 찾고, 그 안에서 인용문을 찾고, 그다음 그 주장을 뒷받침하는지 판단할 만큼 컨텍스트를 충분히 읽어야 합니다.
이 검사를 자동화하기 위해, 먼저 평범한 문자열 매칭으로 문서에 없는 인용문을 찾아냅니다. 그런 다음 Choice 질문으로 남은 각 인용문의 컨텍스트를 읽어 그것이 주장을 뒷받침하는지 판단합니다.
%%{init: {"flowchart": {"wrappingWidth": 330}}}%%
flowchart LR
cite["source document + citation"]
match{"is the quote<br/>in the source?"}
fab["mark <b>fabricated</b>"]
subgraph request[" "]
q["Choice — how does the<br/>section relate to the claim?<br/>supports → mark <b>verified</b><br/>contradicts → mark <b>contradicted</b><br/>says nothing → mark <b>unsupported</b>"]
end
gate{"confidence<br/>≥ 0.8?"}
stand["let the verdict stand"]
review["a human confirms it"]
cite --> match
%% the two edges that reach the call come first, so they stay adjacent; the
%% string match's own verdict is declared last and lands below them
match -- "found" --> request
match -- "no quote" --> request
match -- "not found" --> fab
request --> gate
gate --> stand
gate --> review
아래에서는 RFC 7519(JSON Web Token)에 관한 어떤 LLM의 답변에서 나온 인용 8개가 이 검사를 거칩니다. 정확한 4개는 신뢰도 0.93 이상으로 verified를 반환했습니다. 심어 둔 4개의 오류도 모두 잡혔습니다. 날조된 인용 하나, 반박된 주장 하나, 그리고 사람에게 넘겨진 뒷받침되지 않는 인용 두 개입니다.
여기서 여러분이 만들 함수 check_citation()은 원본 문서 하나와 인용 하나를 받아 네 가지 판정 중 하나를 반환합니다. verified, unsupported, contradicted, fabricated 중 하나입니다. 또한 사람이 살펴봐야 할 항목을 표시하는 신뢰도도 반환합니다.
준비
pip install ipython 'cooksafe>=0.2.0,<0.3.0'
그다음 TYPESAFE_API_KEY를 설정하십시오. 모든 API 호출은 json_cache.json에 캐시되며, 이 파일은 cookbook과 함께 제공되므로 다시 실행하면 API를 실제로 호출하는 대신 게시된 숫자를 재생합니다. 이 파일을 삭제하면 전부 실제로 호출합니다.
아래 숫자는 2026-08-16의 jev-1.12에서 나왔습니다.
import json
import os
import re
from pathlib import Path
from time import perf_counter
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, TypeSafeClient
TYPESAFE_MODEL = "jev-1.12"
AUTO_ACCEPT = 0.8 # start high for more human review as you build trust in the model
client = TypeSafeClient(
api_key=os.environ.get("TYPESAFE_API_KEY", "cache-only"),
base_url=os.environ.get("TYPESAFE_ENDPOINT"),
timeout=120.0,
)
json_cache = JsonCache(Path("json_cache.json"))
소스와 인용 불러오기
소스는 RFC 7519(JSON Web Token)이며, rfc-editor.org에서 받아 이 cookbook과 같은 위치에 rfc7519.txt로 커밋되어 있습니다. 아래 코드는 페이지의 머리말과 꼬리말을 제거한 뒤 텍스트를 번호가 붙은 절로 나눕니다.
citations.json의 인용 8개는 어떤 LLM이 이 RFC를 보고 작성한 것입니다. 4개는 정확하고, 나머지 4개는 검사에서 실패하도록 우리가 고쳐 두었습니다.
def load_source() -> str:
"""RFC 7519 verbatim, minus the page headers and footers that interrupt its paragraphs."""
lines = []
for line in Path("rfc7519.txt").read_text().splitlines():
bare = line.lstrip("\f")
if re.match(r"Jones, et al\.\s.*\[Page \d+\]$", bare):
continue
if re.match(r"RFC 7519\s+JSON Web Token \(JWT\)\s+May 2015$", bare):
continue
lines.append(bare)
return re.sub(r"\n{3,}", "\n\n", "\n".join(lines))
def split_sections(source: str) -> dict[str, str]:
"""Map each numbered section ("4.1.3") to its text, split on the RFC's header lines."""
boundary = re.compile(r"(?m)^(?:(\d+(?:\.\d+)*)\. .+|Appendix [A-Z]\..*)$")
marks = list(boundary.finditer(source))
sections = {}
for mark, nxt in zip(marks, marks[1:] + [None]):
if mark.group(1) is None: # an appendix header only terminates the section before it
continue
sections[mark.group(1)] = source[mark.start() : nxt.start() if nxt else len(source)].strip()
return sections
SOURCE = load_source()
SECTIONS = split_sections(SOURCE)
CITATIONS = json.loads(Path("citations.json").read_text())
print(f"{len(SOURCE):,} characters, {len(SECTIONS)} numbered sections, {len(CITATIONS)} citations")
print("\nA citation with a quote:")
print(json.dumps(CITATIONS[1], indent=2))
print("\nA claim-only citation:")
print(json.dumps(next(c for c in CITATIONS if c["quote"] is None), indent=2))
58,365 characters, 45 numbered sections, 8 citations
A citation with a quote:
{
"id": "aud_reject",
"claim": "If a validator does not find itself in a token's audience list, it has to reject the token.",
"quote": "If the principal processing the claim does not identify itself with a value in the \"aud\" claim when this claim is present, then the JWT MUST be rejected.",
"section": "4.1.3"
}
A claim-only citation:
{
"id": "iat_future",
"claim": "The \"iat\" claim requires validators to reject tokens whose issue time is in the future.",
"quote": null,
"section": "4.1.6"
}
소스에서 각 인용문 찾기
소스에 없는 인용문은 날조된 것이며, 이를 알아내는 데 모델은 필요하지 않습니다. 공백과 곱은따옴표(스마트 따옴표)를 정규화하여 인용문이 RFC의 줄바꿈을 넘어서도 매칭되게 만든 뒤, 부분 문자열로 찾습니다. 매칭에 성공하면 그 인용문이 어느 절에서 왔는지도 알 수 있고, 그 절이 다음 단계에서 모델이 읽는 텍스트입니다.
인용이 절 하나를 지목하면서 그 내용은 전혀 인용하지 않을 수도 있습니다. 이 경우에는 매칭할 것이 없으므로, 인용이 지목한 절을 그대로 가져다 모델로 넘어갑니다.
def normalize(text: str) -> str:
"""Collapse whitespace and fold curly quotes, so a quote matches across line wraps."""
table = str.maketrans({"“": '"', "”": '"', "‘": "'", "’": "'"})
return re.sub(r"\s+", " ", text.translate(table)).strip()
def find_quote(sections: dict[str, str], quote: str) -> str | None:
"""The number of the section that contains the quote verbatim, or None."""
needle = normalize(quote)
for number in sorted(sections, key=lambda n: [int(p) for p in n.split(".")]):
if needle in normalize(sections[number]):
return number
return None
def locate(sections: dict[str, str], citation: dict) -> tuple[str, str | None]:
"""Step 1 for one citation: a status, plus the section step 2 will read."""
if citation["quote"] is None:
return "section-only", sections[citation["section"]]
number = find_quote(sections, citation["quote"])
if number is None:
return "missing", None
return "found", sections[number]
for citation in CITATIONS:
status, section = locate(SECTIONS, citation)
where = f"section of {len(section):,} chars" if section else "not in the source"
print(f"{citation['id']:<18}{status:<14}{where}")
epoch_seconds found section of 3,122 chars
aud_reject found section of 761 chars
sig_reporting missing not in the source
clock_skew found section of 529 chars
exp_required found section of 529 chars
pii_encryption found section of 1,653 chars
iat_future section-only section of 270 chars
duplicate_names found section of 918 chars
소스가 주장을 뒷받침하는지 검증하기
여기까지 와서 아직 인용문을 가진 인용은 소스와 한 글자도 틀리지 않고 일치합니다. 하지만 그것으로 충분하지 않습니다. 인용문은 정확한데 그 위에 세운 주장은 여전히 틀릴 수 있습니다. 이를 판단하려면 인용문이 놓인 컨텍스트, 즉 1단계에서 찾은 그 절을 봐야 합니다.
남은 인용마다 Choice 질문 하나로 절과 주장 사이에 가능한 세 가지 관계를 다룹니다. 확률이 가장 높은 선택지가 판정이며, AUTO_ACCEPT(위 코드에서는 0.8)가 그 판정을 어떻게 처리할지 결정합니다.
- 신뢰도 0.8 이상: 판정이 그대로 성립합니다.
- 0.8 미만: 사람이 그 판정을 확인한 뒤에야 누군가 그것에 따라 행동합니다.
처음에는 높게 잡고, 모델이 여러분의 문서에서 어떻게 하는지 지켜보면서 임계값을 낮추십시오.
QUESTIONS = {
"relation": Choice(
instructions="How does the section relate to the claim?",
criteria={
"supports": "The section states the claim or directly implies that it is true",
"contradicts": "The section states the opposite of the claim or implies it is false",
"says_nothing": "The section does not address what the claim asserts, either way",
},
),
}
RELATION_TO_VERDICT = {
"supports": "verified",
"contradicts": "contradicted",
"says_nothing": "unsupported",
}
@json_cache
def ask(claim: str, section: str) -> dict:
started = perf_counter()
response = client.system_one(
state={"claim": claim, "section": section},
questions=QUESTIONS,
model=TYPESAFE_MODEL,
)
answer = response.answers["relation"]
return {
"choice": answer.choice,
"probabilities": answer.probabilities,
"confidence": answer.confidence,
"seconds": round(perf_counter() - started, 2),
"input_tokens": response.usage.input_tokens or 0,
"output_tokens": response.usage.output_tokens or 0,
}
def verdict(status: str, answer: dict | None) -> dict:
"""Fold step 1 and step 2 into one of the four labels, plus an auto-or-review flag."""
if status == "missing":
# confidence None: no model was called, so there is no model confidence to report
return {"verdict": "fabricated", "confidence": None, "auto": True}
return {
"verdict": RELATION_TO_VERDICT[answer["choice"]],
"confidence": answer["confidence"],
"auto": answer["confidence"] >= AUTO_ACCEPT,
}
def check_citation(sections: dict[str, str], citation: dict) -> dict:
status, section = locate(sections, citation)
answer = ask(citation["claim"], section) if section is not None else None
return {"id": citation["id"], "status": status, "answer": answer, **verdict(status, answer)}
모든 인용 검사하기
인용 8개 전부 같은 검사를 거칩니다.
print(f"{'citation':<18}{'quote':<14}{'relation':<14}{'conf':>6} {'verdict':<13}{'action':>7}")
for citation in CITATIONS:
result = check_citation(SECTIONS, citation)
answer = result["answer"]
relation = answer["choice"] if answer else "-"
conf = f"{answer['confidence']:.2f}" if answer else "-"
action = "auto" if result["auto"] else "review"
print(
f"{result['id']:<18}{result['status']:<14}{relation:<14}{conf:>6}"
f" {result['verdict']:<13}{action:>7}"
)
citation quote relation conf verdict action
epoch_seconds found supports 0.93 verified auto
aud_reject found supports 0.95 verified auto
sig_reporting missing - - fabricated auto
clock_skew found supports 0.99 verified auto
exp_required found contradicts 0.99 contradicted auto
pii_encryption found says_nothing 0.27 unsupported review
iat_future section-only says_nothing 0.56 unsupported review
duplicate_names found supports 0.99 verified auto
인용 4개가 verified를 반환했고, 1개가 fabricated, 1개가 contradicted, 2개가 unsupported를 반환했습니다.
epoch_seconds,aud_reject,clock_skew,duplicate_names가 바로 그 정확한 4개입니다. 모두 신뢰도 0.93 이상으로verified를 반환했으며, 이는AUTO_ACCEPT를 훨씬 웃도는 값입니다.sig_reporting은 모델에 전달되지도 않았습니다. 그 인용문이 RFC에 없으므로 문자열 매칭만으로fabricated로 판정됩니다.exp_required는 4.1.4절을 한 글자도 틀리지 않고 인용했는데, 같은 절에 “Use of this claim is OPTIONAL”이라고 적혀 있으므로 신뢰도 0.99로contradicted입니다.pii_encryption과iat_future는 각각 0.27과 0.56으로unsupported를 반환했고, 둘 다 임계값 아래여서 사람에게 넘겨졌습니다.pii_encryption은 문자열 매칭만으로는 충분하지 않은 이유를 보여줍니다. 그 인용문은 소스에 한 글자도 틀리지 않고 들어 있지만, 그것이 나온 절은 이 주장에 대해 아무 말도 하지 않습니다.
이것을 여러분의 데이터에 적용하려면 rfc7519.txt와 citations.json을 교체하십시오. load_source()와 split_sections()는 RFC의 판형에 맞춰 작성되었으므로, 형태가 다른 문서라면 파서를 직접 써야 합니다.
정규화 후의 문자열 매칭은 정확 매칭입니다. 잘리거나 표현이 조금 바뀐 인용문은 fabricated로 반환됩니다. 불완전한 인용 방식을 용인하는 프로덕션 시스템이라면 퍼지 매칭을 대신 써야 합니다.
Playground에서 열기
이 링크에는 어떤 인용의 주장과 절, 그리고 그 질문이 담겨 있습니다. 열면 브라우저에서 같은 호출을 실시간으로 실행할 수 있습니다.
example = next(c for c in CITATIONS if c["id"] == "exp_required")
_, example_section = locate(SECTIONS, example)
playground_link = make_playground_link(
{"claim": example["claim"], "section": example_section}, QUESTIONS, models=[TYPESAFE_MODEL]
)
display(Markdown(f"🔗 [Open one citation's claim + section in the TypeSafe playground]({playground_link})"))
TypeSafe Playground에서 어떤 인용의 주장 + 절 열기 →