Verificación doble de citas
Detecta citas erróneas o inventadas comprobándolas contra el documento fuente. Una pregunta de Choice decide si el contexto de la cita respalda la afirmación.
Un LLM responde a una pregunta y adjunta citas: por cada afirmación, una sección de un documento fuente y la cita en la que se apoya. Algunas de esas citas están mal o son inventadas: la cita puede faltar del documento por completo, o estar en él palabra por palabra mientras su contexto dice lo contrario de la afirmación.
Comprobar una a mano es lento: encontrar el documento, encontrar la cita dentro de él y luego leer suficiente contexto para saber si respalda la afirmación.
Para automatizar esa comprobación, primero buscamos las citas que faltan con una comparación de cadenas normal, y luego usamos una pregunta Choice para leer el contexto de cada cita que queda y decidir si respalda la afirmación.
%%{init: {"flowchart": {"wrappingWidth": 330}}}%%
flowchart LR
cite["source document + citation"]
match{"is the quote<br/>in the source?"}
fab["mark <b>fabricated</b>"]
subgraph request[" "]
q["Choice — how does the<br/>section relate to the claim?<br/>supports → mark <b>verified</b><br/>contradicts → mark <b>contradicted</b><br/>says nothing → mark <b>unsupported</b>"]
end
gate{"confidence<br/>≥ 0.8?"}
stand["let the verdict stand"]
review["a human confirms it"]
cite --> match
%% the two edges that reach the call come first, so they stay adjacent; the
%% string match's own verdict is declared last and lands below them
match -- "found" --> request
match -- "no quote" --> request
match -- "not found" --> fab
request --> gate
gate --> stand
gate --> review
Más abajo, ocho citas de la respuesta de un LLM sobre el RFC 7519 (JSON Web Token) pasan por la comprobación. Las cuatro correctas volvieron como verified con una confianza de 0.93 o más. Se detectaron los cuatro fallos plantados: una cita fabricada, una afirmación contradicha y dos citas no respaldadas enviadas a una persona.
check_citation(), la función que construyes aquí, toma un documento fuente y una cita y devuelve uno de cuatro veredictos: verified, unsupported, contradicted o fabricated. También devuelve una confianza que marca las que debería revisar una persona.
Configuración
pip install ipython 'cooksafe>=0.2.0,<0.3.0'
luego define TYPESAFE_API_KEY. Cada llamada a la API se guarda en caché en json_cache.json, que se distribuye con el cookbook, así que volver a ejecutar reproduce los números publicados en lugar de llamar a la API. Elimina ese archivo para ejecutarlo todo en vivo.
Los números de abajo provienen de jev-1.12 el 2026-08-16.
import json
import os
import re
from pathlib import Path
from time import perf_counter
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, TypeSafeClient
TYPESAFE_MODEL = "jev-1.12"
AUTO_ACCEPT = 0.8 # start high for more human review as you build trust in the model
client = TypeSafeClient(
api_key=os.environ.get("TYPESAFE_API_KEY", "cache-only"),
base_url=os.environ.get("TYPESAFE_ENDPOINT"),
timeout=120.0,
)
json_cache = JsonCache(Path("json_cache.json"))
Carga la fuente y las citas
La fuente es el RFC 7519 (JSON Web Token), obtenido de rfc-editor.org e incluido junto a este cookbook como rfc7519.txt. El código de abajo elimina las cabeceras y los pies de página y luego divide el texto en secciones numeradas.
Las ocho citas de citations.json las escribió un LLM contra el RFC. Cuatro son correctas; editamos las otras cuatro para que fallaran la comprobación.
def load_source() -> str:
"""RFC 7519 verbatim, minus the page headers and footers that interrupt its paragraphs."""
lines = []
for line in Path("rfc7519.txt").read_text().splitlines():
bare = line.lstrip("\f")
if re.match(r"Jones, et al\.\s.*\[Page \d+\]$", bare):
continue
if re.match(r"RFC 7519\s+JSON Web Token \(JWT\)\s+May 2015$", bare):
continue
lines.append(bare)
return re.sub(r"\n{3,}", "\n\n", "\n".join(lines))
def split_sections(source: str) -> dict[str, str]:
"""Map each numbered section ("4.1.3") to its text, split on the RFC's header lines."""
boundary = re.compile(r"(?m)^(?:(\d+(?:\.\d+)*)\. .+|Appendix [A-Z]\..*)$")
marks = list(boundary.finditer(source))
sections = {}
for mark, nxt in zip(marks, marks[1:] + [None]):
if mark.group(1) is None: # an appendix header only terminates the section before it
continue
sections[mark.group(1)] = source[mark.start() : nxt.start() if nxt else len(source)].strip()
return sections
SOURCE = load_source()
SECTIONS = split_sections(SOURCE)
CITATIONS = json.loads(Path("citations.json").read_text())
print(f"{len(SOURCE):,} characters, {len(SECTIONS)} numbered sections, {len(CITATIONS)} citations")
print("\nA citation with a quote:")
print(json.dumps(CITATIONS[1], indent=2))
print("\nA claim-only citation:")
print(json.dumps(next(c for c in CITATIONS if c["quote"] is None), indent=2))
58,365 characters, 45 numbered sections, 8 citations
A citation with a quote:
{
"id": "aud_reject",
"claim": "If a validator does not find itself in a token's audience list, it has to reject the token.",
"quote": "If the principal processing the claim does not identify itself with a value in the \"aud\" claim when this claim is present, then the JWT MUST be rejected.",
"section": "4.1.3"
}
A claim-only citation:
{
"id": "iat_future",
"claim": "The \"iat\" claim requires validators to reject tokens whose issue time is in the future.",
"quote": null,
"section": "4.1.6"
}
Encuentra cada cita en la fuente
Una cita que no está en la fuente está fabricada, y no hace falta ningún modelo para descubrirlo. Normaliza los espacios y las comillas tipográficas para que una cita siga coincidiendo a través de los saltos de línea del RFC, y luego búscala como subcadena. Una coincidencia también dice de qué sección vino la cita, y esa sección es el texto que lee el modelo en el siguiente paso.
Una cita puede nombrar una sección sin citar nada de ella. En ese caso no hay nada que comparar, así que toma la sección que nombra la cita y ve directamente al modelo.
def normalize(text: str) -> str:
"""Collapse whitespace and fold curly quotes, so a quote matches across line wraps."""
table = str.maketrans({"“": '"', "”": '"', "‘": "'", "’": "'"})
return re.sub(r"\s+", " ", text.translate(table)).strip()
def find_quote(sections: dict[str, str], quote: str) -> str | None:
"""The number of the section that contains the quote verbatim, or None."""
needle = normalize(quote)
for number in sorted(sections, key=lambda n: [int(p) for p in n.split(".")]):
if needle in normalize(sections[number]):
return number
return None
def locate(sections: dict[str, str], citation: dict) -> tuple[str, str | None]:
"""Step 1 for one citation: a status, plus the section step 2 will read."""
if citation["quote"] is None:
return "section-only", sections[citation["section"]]
number = find_quote(sections, citation["quote"])
if number is None:
return "missing", None
return "found", sections[number]
for citation in CITATIONS:
status, section = locate(SECTIONS, citation)
where = f"section of {len(section):,} chars" if section else "not in the source"
print(f"{citation['id']:<18}{status:<14}{where}")
epoch_seconds found section of 3,122 chars
aud_reject found section of 761 chars
sig_reporting missing not in the source
clock_skew found section of 529 chars
exp_required found section of 529 chars
pii_encryption found section of 1,653 chars
iat_future section-only section of 270 chars
duplicate_names found section of 918 chars
Verifica si la fuente respalda la afirmación
Una cita que todavía tiene una cita textual en este punto coincide con la fuente palabra por palabra. Eso no basta: la cita puede ser exacta y la afirmación construida sobre ella seguir siendo falsa. Decidir eso requiere el contexto de la cita, la sección que encontró el paso 1.
Una pregunta Choice por cada cita que queda cubre las tres formas en que una sección puede relacionarse con una afirmación.
La opción con la mayor probabilidad es el veredicto, y AUTO_ACCEPT (0.8 en el código de arriba) decide qué pasa con él:
- confianza de 0.8 o más: el veredicto se mantiene por sí solo;
- por debajo de 0.8: una persona confirma el veredicto antes de que nada actúe sobre él.
Empieza alto y baja el umbral a medida que veas cómo se comporta el modelo con tus propios documentos.
QUESTIONS = {
"relation": Choice(
instructions="How does the section relate to the claim?",
criteria={
"supports": "The section states the claim or directly implies that it is true",
"contradicts": "The section states the opposite of the claim or implies it is false",
"says_nothing": "The section does not address what the claim asserts, either way",
},
),
}
RELATION_TO_VERDICT = {
"supports": "verified",
"contradicts": "contradicted",
"says_nothing": "unsupported",
}
@json_cache
def ask(claim: str, section: str) -> dict:
started = perf_counter()
response = client.system_one(
state={"claim": claim, "section": section},
questions=QUESTIONS,
model=TYPESAFE_MODEL,
)
answer = response.answers["relation"]
return {
"choice": answer.choice,
"probabilities": answer.probabilities,
"confidence": answer.confidence,
"seconds": round(perf_counter() - started, 2),
"input_tokens": response.usage.input_tokens or 0,
"output_tokens": response.usage.output_tokens or 0,
}
def verdict(status: str, answer: dict | None) -> dict:
"""Fold step 1 and step 2 into one of the four labels, plus an auto-or-review flag."""
if status == "missing":
# confidence None: no model was called, so there is no model confidence to report
return {"verdict": "fabricated", "confidence": None, "auto": True}
return {
"verdict": RELATION_TO_VERDICT[answer["choice"]],
"confidence": answer["confidence"],
"auto": answer["confidence"] >= AUTO_ACCEPT,
}
def check_citation(sections: dict[str, str], citation: dict) -> dict:
status, section = locate(sections, citation)
answer = ask(citation["claim"], section) if section is not None else None
return {"id": citation["id"], "status": status, "answer": answer, **verdict(status, answer)}
Comprueba cada cita
Las ocho citas por la misma comprobación:
print(f"{'citation':<18}{'quote':<14}{'relation':<14}{'conf':>6} {'verdict':<13}{'action':>7}")
for citation in CITATIONS:
result = check_citation(SECTIONS, citation)
answer = result["answer"]
relation = answer["choice"] if answer else "-"
conf = f"{answer['confidence']:.2f}" if answer else "-"
action = "auto" if result["auto"] else "review"
print(
f"{result['id']:<18}{result['status']:<14}{relation:<14}{conf:>6}"
f" {result['verdict']:<13}{action:>7}"
)
citation quote relation conf verdict action
epoch_seconds found supports 0.93 verified auto
aud_reject found supports 0.95 verified auto
sig_reporting missing - - fabricated auto
clock_skew found supports 0.99 verified auto
exp_required found contradicts 0.99 contradicted auto
pii_encryption found says_nothing 0.27 unsupported review
iat_future section-only says_nothing 0.56 unsupported review
duplicate_names found supports 0.99 verified auto
Cuatro citas volvieron como verified, una como fabricated, una como contradicted y dos como unsupported.
epoch_seconds,aud_reject,clock_skewyduplicate_namesson las cuatro correctas. Todas volvieron comoverifiedcon una confianza de 0.93 o más, muy por encima deAUTO_ACCEPT.sig_reportingnunca llegó al modelo. Su cita no está en el RFC, así que la sola comparación de cadenas la marca comofabricated.exp_requiredcita la sección 4.1.4 palabra por palabra, y esa misma sección dice «Use of this claim is OPTIONAL», así que escontradicted, con una confianza de 0.99.pii_encryptioneiat_futurevolvieron comounsupportedcon 0.27 y 0.56, ambas por debajo del umbral, así que las dos fueron a una persona.pii_encryptionmuestra por qué la comparación de cadenas no basta por sí sola: su cita está en la fuente palabra por palabra, y la sección de la que viene no dice nada sobre la afirmación.
Para apuntar esto a tus propios datos, sustituye rfc7519.txt y citations.json. load_source() y split_sections() están escritas para la maquetación de un RFC, así que un documento con otra forma necesita su propio análisis.
La comparación de cadenas es exacta tras la normalización: una cita truncada o ligeramente reformulada vuelve como fabricated. Un sistema en producción que tolere citas descuidadas necesitaría en cambio una coincidencia difusa.
Ábrelo en el playground
El enlace contiene la afirmación y la sección de una cita, más la pregunta. Ábrelo para ejecutar la misma llamada en vivo en el navegador.
example = next(c for c in CITATIONS if c["id"] == "exp_required")
_, example_section = locate(SECTIONS, example)
playground_link = make_playground_link(
{"claim": example["claim"], "section": example_section}, QUESTIONS, models=[TYPESAFE_MODEL]
)
display(Markdown(f"🔗 [Open one citation's claim + section in the TypeSafe playground]({playground_link})"))
Abre la afirmación + la sección de una cita en el playground de TypeSafe →