ドキュメント

SDE カスケード

SDE カスケード

2 段階の構造化データ抽出カスケード(mini → verify → reasoning)で、大規模な推論モデルの品質の大半を、ごく一部のコストで得ます。

  • 概要
    • 大規模な推論モデルは構造化データをうまく抽出しますが、遅くて高コストです
    • 小さいモデルは安価ですが、誤りを犯します
    • カスケードなら、品質の大半をごく一部のコストで得られます
    • 使用するモデルとその価格($ / 100 万 token、入力 / 出力。標準料金、2026 年 9 月 15 日時点):
  • アルゴリズム
    1. 安価で小さなモデルで抽出します。
    2. TypeSafe のプリミティブで検証します。フィールドごとのイエス・ノー(「Noul の質問」)です
      • (例:「この値はソースに存在しないか?」「無関係なテキストから持ち込まれたか?」)。それぞれ P(何かが誤っている) を返します。
    3. 検証器のシグナルが発火したら、高コストな推論モデルへエスカレーションします。そうでなければ安価な答えを保持します。
  • この Cookbook
    • 実例を 1 つ端から端までたどり、その後 100 件のプロンプトでのトレードオフを示します
    • 注意:2 つの抽出段はテキストモードの OpenAI を使います
    • 構造化出力、tool 呼び出し、json モードは使いません。理由は:
      • schema 追従の誤りは、LLM が犯すと想定する誤りではありません(この合成データを作るのは簡単です)
      • LLM が実際に schema に従えない場合、ほぼ必ず非常に混乱しているので、制約付きデコーディングでは根本的な問題は解決しません
      • ただし、試してみることをおすすめします!

準備

  • 依存関係をインストールします(TypeSafe の検証器クライアントは TypeSafe のパッケージインデックスから配信されます):
pip install openai datasets jsonschema ipython 'cooksafe>=0.2.0,<0.3.0'
  • 次に、環境に OPENAI_API_KEY と TYPESAFE_API_KEY を設定します
import json
import os
from pathlib import Path

import jsonschema
from cooksafe import JsonCache, make_playground_link
from datasets import load_dataset
from IPython.display import Markdown, display
from openai import OpenAI
from typesafe_sdk import Noul, NoulCriteria, TypeSafeClient

MINI = "gpt-5.4-mini"  # rung 0: cheap + fast
REASONING = "gpt-5.5"  # rung 1: strong, run with reasoning_effort="high"
TS_MODEL = "jev-1.12"  # the TypeSafe verifier model
FIRE_T = 0.7  # escalate if any per-field P(wrong) exceeds this; also the "<== FIRES" display marker

oai = OpenAI()

ts = TypeSafeClient(api_key=os.environ["TYPESAFE_API_KEY"], timeout=30.0)

ステップ 1:データ

scrapegraphai という HuggingFace データセットを選びます

SCRAPEGRAPHAI_REVISION = "4bb9fba1dff9181c5acdb60a5a26fea62fa54fe9"
row = load_dataset(
    "scrapegraphai/scrapegraphai-100k",
    revision=SCRAPEGRAPHAI_REVISION,
    split="train",
)[516]
schema = json.loads(row["schema"])
prompt = row["prompt"]
content = row["content"]

print(
    f"""
PROMPT
===========
{prompt}

SCHEMA
===========
{json.dumps(schema, indent=2)}

CONTENT
===========
{content}
""".strip()
)
PROMPT
===========
Find registration open date fall semester for New York University in New York, NY for the 2024-2025 school year.

SCHEMA
===========
{
  "properties": {
    "registration_open_date": {
      "description": "The date that registration opens for the fall semester. MUST be in the format mm/dd/yyyy. For example, for a college in the 2024-2025 school year, it might be something like 09/05/2024. Return a blank string if you are unsure.",
      "title": "Registration Open Date",
      "type": "string"
    },
    "description": {
      "description": "A brief description of the registration open date. For example, 'Registration opens for the fall semester'.",
      "title": "Description",
      "type": "string"
    }
  },
  "required": [
    "registration_open_date",
    "description"
  ],
  "title": "RegistrationOpen",
  "type": "object"
}

CONTENT
===========
Skip to content Skip to current page navigation

[ ](https://www.nyu.edu/)

Search Site

[ ](https://www.nyu.edu/)

  * [ Academics](https://www.nyu.edu/academics.html)
  * [ Admissions](https://www.nyu.edu/admissions.html)
  * [ Research](https://www.nyu.edu/research.html)
  * [ University Life](https://www.nyu.edu/life.html)
  * [ About](https://www.nyu.edu/about.html)

All NYU

#  Mobile Navigation

[ ](https://www.nyu.edu/)

Search Site

  * [Academics](https://www.nyu.edu/academics.html)
  * [Admissions](https://www.nyu.edu/admissions.html)
  * [Research](https://www.nyu.edu/research.html)
  * [University Life](https://www.nyu.edu/life.html)
  * [About](https://www.nyu.edu/about.html)

All NYU

Info for

  * Back to main menu
  * Info for

    * [Students](https://www.nyu.edu/students.html)
    * [Faculty](https://www.nyu.edu/faculty.html)
    * [Alumni](https://www.nyu.edu/alumni.html)
    * [Employees](https://www.nyu.edu/employees.html)
    * [Community](https://www.nyu.edu/community.html)

[Log In](http://home.nyu.edu/)

Info for

  * [Students](https://www.nyu.edu/students.html)
  * [Faculty](https://www.nyu.edu/faculty.html)
  * [Alumni](https://www.nyu.edu/alumni.html)
  * [Employees](https://www.nyu.edu/employees.html)
  * [Community](https://www.nyu.edu/community.html)

[Log In](https://home.nyu.edu/)

Search Site Search

#  Events Calendar

Search Events

Apply Reset

  * [About the Events Calendar ](https://www.nyu.edu/employees/resources-and-services/media-and-communications/digital-communications/university-events-calendar.html)
  * [Events Calendar Tutorial ](https://www.nyu.edu/employees/resources-and-services/media-and-communications/digital-communications/university-events-calendar/tutorials.html)
  * [Report issue or provide feedback ](https://nyu.service-now.com/sp?id=sc_cat_item&sys_id=7698dd2a98bcf4004c8c03063d84e274)

Search Filters Calendar

New York University

Equal Opportunity and Non-Discrimination at NYU - New York University is committed to maintaining an environment that encourages and fosters respect for individual values and appropriate conduct among all persons. In all University spaces--physical and digital--programming, activities, and events are carried out in accordance with applicable law as well as University policy, which includes but is not limited to its Non-Discrimination and Anti-Harassment Policy.

Unless otherwise noted, all content copyright New York University. All rights reserved.

  * [Search](https://search.nyu.edu/)
  * [Campus Map](https://www.nyu.edu/map.html)
  * [Events](https://events.nyu.edu/)
  * [Contact Us](https://www.nyu.edu/contact-us.html)
  * [Give](https://www.nyu.edu/about/giving.html)
  * [Copyright & Fair Use](https://www.nyu.edu/copyright-and-fair-use.html)
  * [Privacy](https://www.nyu.edu/privacy.html)
  * [Accessibility](https://www.nyu.edu/accessibility.html)
  * [Feedback](https://www.nyu.edu/#feedback.html)

  * [New York Campus](https://www.nyu.edu/)
  * [Abu Dhabi Campus](https://nyuad.nyu.edu/)
  * [Shanghai Campus](https://shanghai.nyu.edu/)

  * [![](https://events.nyu.edu/live/resource/image/_i/themes/global/images/icons/facebook.rev.1773448757.svg)](https://facebook.com/)
  * [![](https://events.nyu.edu/live/resource/image/_i/themes/global/images/icons/linkedin.rev.1773448758.svg)](https://linkedin.com/)
  * [![](https://events.nyu.edu/live/resource/image/_i/themes/global/images/icons/x.rev.1773448757.svg)](https://x.com/)
  * [![](https://events.nyu.edu/live/resource/image/_i/themes/global/images/icons/instagram.rev.1773448757.svg)](https://instagram.com/)
  * [![](https://events.nyu.edu/live/resource/image/_i/themes/global/images/icons/youtube.rev.1773448758.svg)](https://youtube.com/)
  • この行は NYU のイベントカレンダーページ(「Fall 2024 Census Date」)です:
    • schema は 2 つのフィールドだけを要求します:registration_open_date と description
    • プロンプトのスクレイプはカレンダーのナビと定型文しか捉えていません:登録日も説明もありません
    • schema の description フィールドは、自身のフィールド説明に例の値(「Registration opens for the fall semester」)まで載せている点に注意してください
  • したがって、行儀のよい抽出器はページに含まれないフィールドをでっち上げるのを拒むはずです
  • 小さいモデルが正しいことをするか見てみましょう!

ステップ 2:mini モデルで抽出する(テキストモード)

  • 注意:gpt-5.4-mini はこの入力に対して非常に確率的です —— temperature=0 でも、ほぼ毎回異なる description をでっち上げます。再現可能なウォークスルーのために、このノートブックの残りで説明する 1 つの典型的な捏造(検証器が P(wrong) > 0.8 としてフラグを立てるもの)をハードコードします。実際のパイプラインなら、extract(MINI, prompt, schema, content, temperature=0) を直接呼ぶだけです。
EXTRACT_SYSTEM = (
    "You extract structured data from documents. Return only values supported by the text. "
    "Follow any value format specified by the schema or its field descriptions."
)

# LLM and TypeSafe calls are cached to ``json_cache.json``, which ships with the cookbook, so
# re-rendering reproduces the published results with no API spend; delete the file to re-run live.
json_cache = JsonCache(Path("json_cache.json"))

@json_cache
def extract(
    model: str,
    prompt: str,
    schema: dict,
    content: str,
    *,
    reasoning_effort: str | None = None,
    temperature: float | None = None,
) -> dict:
    user = (
        f"{prompt}\n\nReturn ONLY a JSON object matching this JSON Schema:\n"
        f"{json.dumps(schema, indent=2)}\n\nDocument:\n{content}"
    )
    kwargs = {
        "model": model,
        "messages": [
            {"role": "system", "content": EXTRACT_SYSTEM},
            {"role": "user", "content": user},
        ],
    }
    if reasoning_effort:
        kwargs["reasoning_effort"] = reasoning_effort
    if temperature is not None:
        kwargs["temperature"] = temperature
    text = oai.chat.completions.create(**kwargs).choices[0].message.content
    # The prompt asks for ONLY a JSON object, so parse the reply as-is -- no regex fishing a
    # substring out of a malformed reply. If ``json.loads`` fails, treat it as an empty extraction
    # (the record-level analog of NaN): every field reads as absent, which the verifier flags and the
    # gate escalates -- the safe direction. Schema-following errors are rare here (see the overview).
    try:
        return json.loads(text)
    except (ValueError, json.JSONDecodeError):
        return {}

# Hard-coded canonical fabrication (see note above); a real pipeline would use extract(MINI, prompt, schema, content, temperature=0).
mini_record = {
    "registration_open_date": "",
    "description": "Registration opens for the fall semester",
}
print("mini extraction:\n", json.dumps(mini_record, indent=2))

# The record is a perfect fit for the JSON Schema -- and still wrong. Schema validation is necessary
# but not sufficient: it catches structural errors, never semantic ones. That gap is the whole point.
print("\nschema-valid:", jsonschema.Draft202012Validator(schema).is_valid(mini_record))
mini extraction:
 {
  "registration_open_date": "",
  "description": "Registration opens for the fall semester"
}

schema-valid: True
  • このレコードは schema に適合しています(上の行は True を出力します)が、それでも誤っています:
    • registration_open_date は空のままです。これはページと一致します —— ページは日付を示していません
    • しかし description は捏造です:ページは登録日を一切説明していないので、mini はもっともらしいものをでっち上げます。schema 自身の例「Registration opens for the fall semester」をそのまま繰り返したり、「…was not found in the document」と述べたりします
    • JSON-Schema のチェックではこれは見抜けません。安価なモデルはこの種の自信に満ちた schema 適合の捏造を生み、それを見つけるのが意味検証器の役割です

ステップ 3:TypeSafe で検証する

  • 検証器は TypeSafe です。フィールドごとに Noul の質問を組み立てます:
    • 狭いイエス・ノーで、true = 何かが誤っている(エスカレーション)となるように組み立てます
  • TypeSafe は 1 回の system_one 呼び出しで、質問ごとに較正された noul = P(true) を返します
  • 質問セット:
    • 1 つの全体的な __overall__::judge ヘッド(「このレコードはエスカレーションすべきか?」)。これは、レコード全体の判断とフィールドごとのヘッドを対比するために計算・表示しますが、ステップ 4 のゲートはこれを使いません —— エスカレーションはフィールドごとのバッテリーが駆動します。
    • フィールドごとのバッテリー
      • 空でないフィールドにはすべてのヘッドが付きます
      • 空のフィールド(null / “” / [])には absence_wrong ヘッドだけが付きます
    • (完全なパイプラインには、コンテナ全体に対する spurious ヘッドと、全体の difficulty score もありますが、このウォークスルーを 2 つのゲーティングヘッドに絞るため、ここでは示しません)
  • TypeSafe の流儀:分解
    • すべてがプログラム的に分解されている点に注目してください。これが TypeSafe の流儀です。
    • 分解はすべてのプロンプトの知能を最大化し、アルゴリズムを調整可能で解釈可能にします。
    • これが流儀だ
# metric -> (question, NoulCriteria)
MAIN_QUESTIONS = {
    "name_desc_mismatch": (
        "Does the `extracted_field` fail to match the field at `path` or the `description` in the "
        "`field_spec`? If the `description` is empty, judge against the `path` alone.",
        NoulCriteria(
            true="the `extracted_field` does not match the field name or its `description`",
            false="the `extracted_field` matches the field name and `description`",
        ),
    ),
    "type_mismatch": (
        "Does the `extracted_field` violate the `type` declared in the `field_spec`?",
        NoulCriteria(
            true="the `extracted_field` violates the declared `type`",
            false="the `extracted_field` conforms to the declared `type`",
        ),
    ),
    "unreasonable": (
        "Is the `extracted_field` one that a reasonable person would not have extracted for this "
        "`field_spec`?",
        NoulCriteria(
            true="a reasonable person would not have extracted this value",
            false="the extraction is reasonable",
        ),
    ),
    "hallucinated": (
        "Is the `extracted_field` unsupported by, or absent from, the source text?",
        NoulCriteria(
            true="the `extracted_field` is a hallucination -- not supported by, or absent "
            "from, the source text",
            false="the `extracted_field` is supported by the source text",
        ),
    ),
    "off_target": (
        "Does the source text fail to genuinely report the thing the `field_spec` describes, so the "
        "value was pulled from incidental text?",
        NoulCriteria(
            true="the source does not genuinely provide this field -- the value was pulled "
            "from incidental text",
            false="the source genuinely reports this field",
        ),
    ),
    "incomplete": (
        "Does the `extracted_field` fail to capture a value the source supports (note whether the "
        "`field_spec` is `required`)?",
        NoulCriteria(
            true="the field is wrongly empty, null, or missing a value the source supports",
            false="the field captures the value the source supports",
        ),
    ),
    "format_violation": (
        "Does the `extracted_field` violate the format or constraints implied by the `description`, "
        "the schema `type`, and the extraction instructions (e.g. date format, units, enum membership)?",
        NoulCriteria(
            true="the `extracted_field` violates the implied format or constraints",
            false="the `extracted_field` satisfies the format and constraints",
        ),
    ),
}
ABSENCE_QUESTION = (
    "The `extracted_field` is empty, null, or an empty collection. Does the source text contain the "
    "information the `field_spec` describes, making the empty result wrong?"
)
ABSENCE_CRITERIA = NoulCriteria(
    true="a value was wrongly omitted", false="returning nothing is correct"
)

# The pipeline also asks one holistic, whole-record head: "should this be escalated?"
OVERALL_JUDGE = (
    "Is this extracted record an incorrect extraction -- some value unsupported by the source or "
    "not conforming to the schema, required information missing or wrong, or some field hallucinated -- "
    "so it should be escalated to a smarter model?"
)
OVERALL_JUDGE_CRITERIA = NoulCriteria(
    true="the record is an incorrect extraction",
    false="the record is a correct extraction",
)

def is_empty(v) -> bool:
    return v is None or (isinstance(v, (str, list, dict)) and len(v) == 0)

def field_spec(name: str) -> dict:
    """Minimal spec pulled from the schema (unwrapping anyOf/null for optional fields)."""
    p = schema["properties"][name]
    branches = p.get("anyOf") or []
    typ = p.get("type") or next(
        (b["type"] for b in branches if b.get("type") != "null"), "unknown"
    )
    return {
        "path": name,
        "type": typ,
        "description": p.get("description", ""),
        "required": name in schema.get("required", []),
    }

def build_questions(record: dict) -> dict[str, Noul]:
    """The verify question set: one holistic ``__overall__::judge`` head plus a per-field battery,
    keyed ``field::metric`` (mirrors build_verify_prompts)."""
    questions: dict[str, Noul] = {
        "__overall__::judge": Noul(
            instructions=OVERALL_JUDGE, criteria=OVERALL_JUDGE_CRITERIA
        ),
    }
    for name, value in record.items():
        spec = field_spec(name)
        if is_empty(value):
            questions[f"{name}::absence_wrong"] = Noul(
                instructions={
                    "field_spec": spec,
                    "extracted_field": value,
                    "main_question": ABSENCE_QUESTION,
                },
                criteria=ABSENCE_CRITERIA,
            )
            continue
        for metric, (question, criteria) in MAIN_QUESTIONS.items():
            if metric == "type_mismatch" and spec["type"] == "unknown":
                continue
            questions[f"{name}::{metric}"] = Noul(
                instructions={
                    "field_spec": spec,
                    "extracted_field": value,
                    "main_question": question,
                },
                criteria=criteria,
            )
    return questions

@json_cache
def verify(record: dict) -> dict[str, float | str]:
    """Run the whole Noul battery over a record in one TypeSafe call; return ``{field::metric: P(true)}``."""
    state = {
        "system_message": EXTRACT_SYSTEM,
        "instruction": "Extract the structured record from this document",
        "source_text": row["content"],
        "schema": schema,
        "extraction": record,
    }
    questions = build_questions(record)
    answers = ts.system_one(state=state, questions=questions, model=TS_MODEL).answers
    return {qid: ans.noul for qid, ans in answers.items()} | {
        "playground_link": make_playground_link(state, questions)
    }

mini の抽出結果に対してバッテリー全体を実行する

checks = verify(mini_record)
playground_link = checks.pop("playground_link")
display(
    Markdown(
        f"🔗 [Open this verification in the TypeSafe playground]({playground_link})"
    )
)

print(f"{'qid':<40}{'P(wrong)':>9}")
print("-" * 50)
for fld, p in sorted(checks.items(), key=lambda c: -c[-1]):
    flag = "  <== FIRES" if p > FIRE_T else ""
    print(f"{fld:<40}{p:>9.2f}{flag}")
qid                                      P(wrong)
--------------------------------------------------
description::hallucinated                    0.95  <== FIRES
description::off_target                      0.85  <== FIRES
description::unreasonable                    0.58
__overall__::judge                           0.56
description::incomplete                      0.16
registration_open_date::absence_wrong        0.14
description::format_violation                0.10
description::name_desc_mismatch              0.08
description::type_mismatch                   0.02
TypeSafe playground でこの検証を開く →
  • TypeSafe は、実際に誤っているフィールドにシグナルを集中させます。
  • 結果は較正されています:誤っているフィールドでは高く、正しいフィールドでは低く、明らかに誤りとは言えないが怪しいフィールドでは中程度です
  • これが、大雑把な「全体として良いか?」という判定器に対する、typesafe 検証器の利点です

ステップ 4:エスカレーションゲート

  • ここでは any_flag でゲートします:いずれかのフィールドのフラグが FIRE_T(0.7、上で設定し、ステップ 3 の <== FIRES マーカーと共有)を超えたらエスカレーションします
  • これは max 型のゲート(いずれかのフィールドが発火したらエスカレーション)であり、平均ではないので、1 つの自信のあるレッドフラグで十分で、平均化されて無音になることはありません
# any_flag is a per-field gate: the holistic __overall__ head is shown above but not part of it
fired = {
    qid: p
    for qid, p in checks.items()
    if not qid.startswith("__overall__") and p > FIRE_T
}
escalate = bool(fired)

print(
    f"any_flag gate (threshold {FIRE_T}): {'ESCALATE' if escalate else 'ACCEPT cheap result'}"
)
for qid, p in sorted(fired.items(), key=lambda c: -c[1]):
    print(f"  fired: {qid}  (P={p:.2f})")
any_flag gate (threshold 0.7): ESCALATE
  fired: description::hallucinated  (P=0.95)
  fired: description::off_target  (P=0.85)

ステップ 5:推論モデルへエスカレーションする

シグナルが発火したので、強力なモデル(gpt-5.5、reasoning_effort="high")にコストを払います

final_record = (
    extract(REASONING, prompt, schema, content, reasoning_effort="high")
    if escalate
    else mini_record
)

print("mini      :", json.dumps(mini_record))
print("reasoning :", json.dumps(final_record))
print("\nfield-level diff (mini -> final):")
for name in mini_record:
    if mini_record[name] != final_record.get(name):
        print(f"  {name}: {mini_record[name]!r}  ->  {final_record.get(name)!r}")
mini      : {"registration_open_date": "", "description": "Registration opens for the fall semester"}
reasoning : {"description": "", "registration_open_date": ""}

field-level diff (mini -> final):
  description: 'Registration opens for the fall semester'  ->  ''
  • 改善点
    • 推論モデルは捏造された description を捨て、"" を返します
    • ページが登録日を一切説明していないことを認識し、でっち上げるのを拒みました
    • カスケードは、自信に満ちた schema 適合の捏造を、正直な空のフィールドに変えました
    • そして、この 1 項目にだけ推論モデルのコストをかけたのは、検証器がそう指示したからです

ステップ 6:100 件のプロンプトでの挙動

  • これらは TypeSafe の内部結果で、上の一般的な方法で生成したものです:
    • 同じ extract → verify → escalate ループ、gpt-5.4-mini → gpt-5.5-reasoning、フィールドごとのヘッドに対する any_flag ゲートを、100 件の scrapegraphai プロンプトで実行
    • 各項目の安価な段の抽出は TypeSafe が採点します。ゲートのしきい値(「cut」)を 0→1 に掃引し、得られた各設定を(コスト、品質)空間にプロットします
    • このグラフは過去のスナップショットです。そのコストは、上に示した現在の Jev 料金では再計算されていません
内部結果:100 件のプロンプトにおけるコスト/品質フロンティア
  • 読み方:
    • 黒い菱形 = 4 つのモデルをそれぞれ単独で実行したもの(コストは能力とともに上昇します。最強の gpt-5.5-reasoning は右上に位置し、品質 ≈0.81、コスト ≈$0.10 / 抽出)
    • 青い点 = 様々なゲートしきい値でのカスケード。破線は pareto フロンティアです
    • カスケードのフロンティアはすべての単独モデルの左上に位置します:ゲートを掃引すれば、最上位モデルの品質の大半をごく一部のコストで得られます
    • 安価な段が簡単な項目をほぼ無料で処理し、フラグが立った項目だけが推論モデルのコストを払います

付録 A:良い検証器シグナルとは

  • カスケードの良し悪しは検証器次第です。有用なシグナルと無用なシグナルを分けるものは:
    • 狭く、根拠がある。
      • 1 つのフィールドについてソースと照らす 1 つの検証可能なイエス・ノー(例:「この値はソースに存在しないか?」)であり、曖昧な「この抽出は良いか?」ではありません
      • 曖昧な質問は、ぼんやりした未較正のスコアを返します
    • Bad = TRUE、criteria を明示する。
      • 各質問を、エスカレーションの場合が true の場合となるように組み立て、true/false の意味を明示します
    • フィールドごとに、max で集約する。
      • フィールドごとのフラグは誤りを局所化し、疎で強いままになります
      • max(「いずれかのフラグが発火」)は、1 つの自信のあるレッドフラグでエスカレーションすることを保証し、平均化されて無音になることはありません
    • 独立していて安価。
      • 専用の検証器(ここでは TypeSafe)が出力を判定することで、抽出器自身の盲点を捉えます
      • 安価でなければ、削減できるコストが残りません
    • 分離性がある / 較正されている。
      • 良いシグナルは実際の誤りで高く、正しいものでは低いので、単一のしきい値が「受け入れ」と「エスカレーション」をきれいに分けます
      • その分離が、pareto 曲線を左上へ押し上げます