ドキュメント

引用のダブルチェック

引用のダブルチェック

原文ドキュメントと照合して、誤った引用や幻覚による引用を見つけ出します。Choice の質問 1 つで、引用の文脈が主張を支えているかを判定できます。

LLM が質問に答えるとき、引用を添えます。主張ごとに、原文ドキュメントの一節と、その根拠となる引用文です。その引用の一部は誤っていたり、幻覚だったりします。引用文がドキュメントにまったく存在しないこともあれば、一字一句そのまま載っていながら、その文脈が主張と反対のことを言っていることもあります。

1 件を手作業で確認するのは手間がかかります。ドキュメントを探し、その中の引用文を探し、さらに主張を支えているかどうかを判断できるだけの文脈を読みます。

この確認を自動化するため、まず通常の文字列マッチで見つからない引用を探し、次に Choice の質問で残った引用の文脈を読み、主張を支えているかを判定します。

  %%{init: {"flowchart": {"wrappingWidth": 330}}}%%
flowchart LR
    cite["source document + citation"]

    match{"is the quote<br/>in the source?"}
    fab["mark <b>fabricated</b>"]

    subgraph request[" "]
        q["Choice &mdash; how does the<br/>section relate to the claim?<br/>supports &rarr; mark <b>verified</b><br/>contradicts &rarr; mark <b>contradicted</b><br/>says nothing &rarr; mark <b>unsupported</b>"]
    end

    gate{"confidence<br/>&ge; 0.8?"}
    stand["let the verdict stand"]
    review["a human confirms it"]

    cite --> match
    %% the two edges that reach the call come first, so they stay adjacent; the
    %% string match's own verdict is declared last and lands below them
    match -- "found" --> request
    match -- "no quote" --> request
    match -- "not found" --> fab
    request --> gate
    gate --> stand
    gate --> review

以下では、RFC 7519(JSON Web Token)に関する LLM の回答から 8 件の引用を取り出し、この確認にかけます。正確な 4 件は信頼度 0.93 以上で verified を返しました。仕込んだ 4 件の誤りもすべて検出できました。捏造された引用文、否定された主張、そして人手に回された 2 件の未対応の引用です。

ここで作る check_citation() は、原文ドキュメントと引用 1 件を受け取り、4 つの判定のいずれかを返します。verified、unsupported、contradicted、fabricated です。あわせて、人が見るべきものを示す信頼度も返します。

準備

pip install ipython 'cooksafe>=0.2.0,<0.3.0'

次に TYPESAFE_API_KEY を設定します。API 呼び出しはすべて json_cache.json にキャッシュされます。このファイルは cookbook に同梱されているので、再実行時は API を呼ばずに公開済みの数値を再生します。ファイルを削除すれば、すべて実際に呼び出せます。

以下の数値は 2026-08-16 の jev-1.12 によるものです。

import json
import os
import re
from pathlib import Path
from time import perf_counter

from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, TypeSafeClient

TYPESAFE_MODEL = "jev-1.12"
AUTO_ACCEPT = 0.8  # start high for more human review as you build trust in the model

client = TypeSafeClient(
    api_key=os.environ.get("TYPESAFE_API_KEY", "cache-only"),
    base_url=os.environ.get("TYPESAFE_ENDPOINT"),
    timeout=120.0,
)
json_cache = JsonCache(Path("json_cache.json"))

原文と引用を読み込む

原文は RFC 7519(JSON Web Token)です。rfc-editor.org から取得し、この cookbook と同じ場所に rfc7519.txt として配置します。以下のコードはページのヘッダーとフッターを除去し、テキストを番号付きの節に分割します。

citations.json の 8 件の引用は、LLM が RFC に基づいて書いたものです。4 件は正確で、残りの 4 件は確認に通らないよう手を加えています。

def load_source() -> str:
    """RFC 7519 verbatim, minus the page headers and footers that interrupt its paragraphs."""
    lines = []
    for line in Path("rfc7519.txt").read_text().splitlines():
        bare = line.lstrip("\f")
        if re.match(r"Jones, et al\.\s.*\[Page \d+\]$", bare):
            continue
        if re.match(r"RFC 7519\s+JSON Web Token \(JWT\)\s+May 2015$", bare):
            continue
        lines.append(bare)
    return re.sub(r"\n{3,}", "\n\n", "\n".join(lines))

def split_sections(source: str) -> dict[str, str]:
    """Map each numbered section ("4.1.3") to its text, split on the RFC's header lines."""
    boundary = re.compile(r"(?m)^(?:(\d+(?:\.\d+)*)\.  .+|Appendix [A-Z]\..*)$")
    marks = list(boundary.finditer(source))
    sections = {}
    for mark, nxt in zip(marks, marks[1:] + [None]):
        if mark.group(1) is None:  # an appendix header only terminates the section before it
            continue
        sections[mark.group(1)] = source[mark.start() : nxt.start() if nxt else len(source)].strip()
    return sections

SOURCE = load_source()
SECTIONS = split_sections(SOURCE)
CITATIONS = json.loads(Path("citations.json").read_text())

print(f"{len(SOURCE):,} characters, {len(SECTIONS)} numbered sections, {len(CITATIONS)} citations")
print("\nA citation with a quote:")
print(json.dumps(CITATIONS[1], indent=2))
print("\nA claim-only citation:")
print(json.dumps(next(c for c in CITATIONS if c["quote"] is None), indent=2))
58,365 characters, 45 numbered sections, 8 citations

A citation with a quote:
{
  "id": "aud_reject",
  "claim": "If a validator does not find itself in a token's audience list, it has to reject the token.",
  "quote": "If the principal processing the claim does not identify itself with a value in the \"aud\" claim when this claim is present, then the JWT MUST be rejected.",
  "section": "4.1.3"
}

A claim-only citation:
{
  "id": "iat_future",
  "claim": "The \"iat\" claim requires validators to reject tokens whose issue time is in the future.",
  "quote": null,
  "section": "4.1.6"
}

原文の中で各引用を探す

原文に存在しない引用文は捏造であり、それを見つけるのにモデルは要りません。空白とカーリークォートを正規化して、RFC の改行をまたいでも引用が一致するようにし、部分文字列として探します。一致すれば、その引用がどの節に由来するかもわかります。その節が、次のステップでモデルが読むテキストです。

引用は、何も引用せずに節だけを指定することもあります。その場合は一致させるものがありません。指定された節をそのまま取り、モデルに回します。

def normalize(text: str) -> str:
    """Collapse whitespace and fold curly quotes, so a quote matches across line wraps."""
    table = str.maketrans({"“": '"', "”": '"', "‘": "'", "’": "'"})
    return re.sub(r"\s+", " ", text.translate(table)).strip()

def find_quote(sections: dict[str, str], quote: str) -> str | None:
    """The number of the section that contains the quote verbatim, or None."""
    needle = normalize(quote)
    for number in sorted(sections, key=lambda n: [int(p) for p in n.split(".")]):
        if needle in normalize(sections[number]):
            return number
    return None

def locate(sections: dict[str, str], citation: dict) -> tuple[str, str | None]:
    """Step 1 for one citation: a status, plus the section step 2 will read."""
    if citation["quote"] is None:
        return "section-only", sections[citation["section"]]
    number = find_quote(sections, citation["quote"])
    if number is None:
        return "missing", None
    return "found", sections[number]

for citation in CITATIONS:
    status, section = locate(SECTIONS, citation)
    where = f"section of {len(section):,} chars" if section else "not in the source"
    print(f"{citation['id']:<18}{status:<14}{where}")
epoch_seconds     found         section of 3,122 chars
aud_reject        found         section of 761 chars
sig_reporting     missing       not in the source
clock_skew        found         section of 529 chars
exp_required      found         section of 529 chars
pii_encryption    found         section of 1,653 chars
iat_future        section-only  section of 270 chars
duplicate_names   found         section of 918 chars

原文が主張を支えているかを検証する

ここまででまだ引用文が残っている引用は、原文と一字一句一致しています。しかし十分ではありません。引用文は正確でも、その上に組み立てた主張が誤っていることがあります。それを判断するには、引用の文脈、つまりステップ 1 で見つけた節が必要です。

残った引用それぞれに Choice の質問を 1 つ投げ、節と主張が取りうる 3 つの関係を網羅します。確率が最も高い選択肢が判定結果で、AUTO_ACCEPT(上のコードでは 0.8)がその後の扱いを決めます。

  • 信頼度 0.8 以上:判定はそのまま有効です。
  • 0.8 未満:何かがそれに基づいて動く前に、人が判定を確認します。

最初は高めに設定し、モデルが自分のドキュメントでどう振る舞うかを見ながら、しきい値を下げていきます。

QUESTIONS = {
    "relation": Choice(
        instructions="How does the section relate to the claim?",
        criteria={
            "supports": "The section states the claim or directly implies that it is true",
            "contradicts": "The section states the opposite of the claim or implies it is false",
            "says_nothing": "The section does not address what the claim asserts, either way",
        },
    ),
}

RELATION_TO_VERDICT = {
    "supports": "verified",
    "contradicts": "contradicted",
    "says_nothing": "unsupported",
}

@json_cache
def ask(claim: str, section: str) -> dict:
    started = perf_counter()
    response = client.system_one(
        state={"claim": claim, "section": section},
        questions=QUESTIONS,
        model=TYPESAFE_MODEL,
    )
    answer = response.answers["relation"]
    return {
        "choice": answer.choice,
        "probabilities": answer.probabilities,
        "confidence": answer.confidence,
        "seconds": round(perf_counter() - started, 2),
        "input_tokens": response.usage.input_tokens or 0,
        "output_tokens": response.usage.output_tokens or 0,
    }

def verdict(status: str, answer: dict | None) -> dict:
    """Fold step 1 and step 2 into one of the four labels, plus an auto-or-review flag."""
    if status == "missing":
        # confidence None: no model was called, so there is no model confidence to report
        return {"verdict": "fabricated", "confidence": None, "auto": True}
    return {
        "verdict": RELATION_TO_VERDICT[answer["choice"]],
        "confidence": answer["confidence"],
        "auto": answer["confidence"] >= AUTO_ACCEPT,
    }

def check_citation(sections: dict[str, str], citation: dict) -> dict:
    status, section = locate(sections, citation)
    answer = ask(citation["claim"], section) if section is not None else None
    return {"id": citation["id"], "status": status, "answer": answer, **verdict(status, answer)}

すべての引用を確認する

8 件の引用をすべて同じ確認にかけます。

print(f"{'citation':<18}{'quote':<14}{'relation':<14}{'conf':>6}  {'verdict':<13}{'action':>7}")
for citation in CITATIONS:
    result = check_citation(SECTIONS, citation)
    answer = result["answer"]
    relation = answer["choice"] if answer else "-"
    conf = f"{answer['confidence']:.2f}" if answer else "-"
    action = "auto" if result["auto"] else "review"
    print(
        f"{result['id']:<18}{result['status']:<14}{relation:<14}{conf:>6}"
        f"  {result['verdict']:<13}{action:>7}"
    )
citation          quote         relation        conf  verdict       action
epoch_seconds     found         supports        0.93  verified        auto
aud_reject        found         supports        0.95  verified        auto
sig_reporting     missing       -                  -  fabricated      auto
clock_skew        found         supports        0.99  verified        auto
exp_required      found         contradicts     0.99  contradicted    auto
pii_encryption    found         says_nothing    0.27  unsupported   review
iat_future        section-only  says_nothing    0.56  unsupported   review
duplicate_names   found         supports        0.99  verified        auto

4 件の引用が verified を返し、1 件が fabricated、1 件が contradicted、2 件が unsupported でした。

  • epoch_seconds、aud_reject、clock_skew、duplicate_names が正確な 4 件です。いずれも信頼度 0.93 以上で verified を返し、AUTO_ACCEPT を大きく上回っています。
  • sig_reporting はモデルに届きませんでした。引用文が RFC にないため、文字列マッチだけで fabricated と判定されます。
  • exp_required は 4.1.4 節を一字一句引用していますが、同じ節に “Use of this claim is OPTIONAL” とあるため、contradicted、信頼度 0.99 です。
  • pii_encryption と iat_future は 0.27 と 0.56 で unsupported を返し、どちらもしきい値を下回るため、両方とも人に回されました。pii_encryption は、文字列マッチだけでは足りない理由を示しています。引用文は原文に一字一句ありますが、それが由来する節は主張について何も述べていません。

自分のデータに使うには、rfc7519.txt と citations.json を差し替えます。load_source() と split_sections() は RFC のレイアウト向けに書かれているので、形の異なるドキュメントには独自のパースが必要です。

正規化後の文字列マッチは完全一致です。切り詰められた引用や少し言い換えられた引用は fabricated になります。引用の不正確さを許容する本番システムなら、代わりにあいまい一致が必要です。

Playground で開く

このリンクには、ある引用の主張とその節、そして質問が入っています。開くと、同じ呼び出しをブラウザで実際に実行できます。

example = next(c for c in CITATIONS if c["id"] == "exp_required")
_, example_section = locate(SECTIONS, example)
playground_link = make_playground_link(
    {"claim": example["claim"], "section": example_section}, QUESTIONS, models=[TYPESAFE_MODEL]
)
display(Markdown(f"🔗 [Open one citation's claim + section in the TypeSafe playground]({playground_link})"))
TypeSafe Playground で、ある引用の主張 + 節を開く →