SDE 캐스케이드
2단계 구조화 데이터 추출 캐스케이드(mini → verify → reasoning)로 큰 추론 모델의 품질 대부분을 적은 비용으로 얻습니다.
- 개요
- 큰 추론 모델은 구조화 데이터를 잘 추출하지만 느리고 비쌉니다
- 작은 모델은 저렴하지만 실수를 합니다
- 캐스케이드는 품질 대부분을 적은 비용으로 얻습니다
- 사용하는 모델과 가격($ / 100만 토큰, 입력 / 출력, 2026년 9월 15일 확인한 표준 요율):
- 0단계(mini):
gpt-5.4-mini, $0.75 / $4.50 - 1단계(reasoning):
gpt-5.5, $5.00 / $30.00 (mini의 약 7배) - 검증기: TypeSafe
jev-1.12, $0.042 / $0.00 (출력 토큰은 무료, 공개된 Jev 가격 참고)
- 0단계(mini):
- 알고리즘
- 저렴하고 작은 모델로 추출합니다.
- TypeSafe 프리미티브로 검증합니다: 필드별 예/아니오(“Noul 질문”) 질문입니다
- (예: “이 값이 소스에 없는가?”, “무관한 텍스트에서 가져온 것인가?”), 각각 P(오류가 있음)을 반환합니다.
- 검증 신호가 발생하면 비싼 추론 모델로 에스컬레이션합니다. 그렇지 않으면 저렴한 답을 유지합니다.
- 이 Cookbook
- 실제 예시 하나를 처음부터 끝까지 살펴본 다음, 100개 프롬프트에 걸친 트레이드오프를 보여줍니다
- 참고: 두 추출 단계 모두 텍스트 모드 OpenAI를 사용합니다
- 우리는 구조화 출력, 도구 호출, json 모드를 사용하지 않습니다. 이유는 다음과 같습니다:
- 스키마 준수 실수는 우리가 LLM이 범할 것이라 예상하는 실수가 아닙니다(이를 위한 합성 데이터를 만드는 것은 쉽습니다)
- LLM이 실제로 스키마를 따르지 못한다면 거의 항상 매우 혼란스러운 상태이므로, 제약 디코딩이 근본 문제를 해결하지 못합니다
- 그래도 시도해 보시길 권합니다!
설정
- 의존성을 설치합니다(TypeSafe 검증기 클라이언트는 TypeSafe의 패키지 인덱스에서 제공됩니다):
pip install openai datasets jsonschema ipython 'cooksafe>=0.2.0,<0.3.0'
- 그런 다음 환경에
OPENAI_API_KEY와TYPESAFE_API_KEY를 설정합니다
import json
import os
from pathlib import Path
import jsonschema
from cooksafe import JsonCache, make_playground_link
from datasets import load_dataset
from IPython.display import Markdown, display
from openai import OpenAI
from typesafe_sdk import Noul, NoulCriteria, TypeSafeClient
MINI = "gpt-5.4-mini" # rung 0: cheap + fast
REASONING = "gpt-5.5" # rung 1: strong, run with reasoning_effort="high"
TS_MODEL = "jev-1.12" # the TypeSafe verifier model
FIRE_T = 0.7 # escalate if any per-field P(wrong) exceeds this; also the "<== FIRES" display marker
oai = OpenAI()
ts = TypeSafeClient(api_key=os.environ["TYPESAFE_API_KEY"], timeout=30.0)
1단계: 데이터
scrapegraphai라는 HuggingFace 데이터셋을 선택합니다
SCRAPEGRAPHAI_REVISION = "4bb9fba1dff9181c5acdb60a5a26fea62fa54fe9"
row = load_dataset(
"scrapegraphai/scrapegraphai-100k",
revision=SCRAPEGRAPHAI_REVISION,
split="train",
)[516]
schema = json.loads(row["schema"])
prompt = row["prompt"]
content = row["content"]
print(
f"""
PROMPT
===========
{prompt}
SCHEMA
===========
{json.dumps(schema, indent=2)}
CONTENT
===========
{content}
""".strip()
)
PROMPT
===========
Find registration open date fall semester for New York University in New York, NY for the 2024-2025 school year.
SCHEMA
===========
{
"properties": {
"registration_open_date": {
"description": "The date that registration opens for the fall semester. MUST be in the format mm/dd/yyyy. For example, for a college in the 2024-2025 school year, it might be something like 09/05/2024. Return a blank string if you are unsure.",
"title": "Registration Open Date",
"type": "string"
},
"description": {
"description": "A brief description of the registration open date. For example, 'Registration opens for the fall semester'.",
"title": "Description",
"type": "string"
}
},
"required": [
"registration_open_date",
"description"
],
"title": "RegistrationOpen",
"type": "object"
}
CONTENT
===========
Skip to content Skip to current page navigation
[ ](https://www.nyu.edu/)
Search Site
[ ](https://www.nyu.edu/)
* [ Academics](https://www.nyu.edu/academics.html)
* [ Admissions](https://www.nyu.edu/admissions.html)
* [ Research](https://www.nyu.edu/research.html)
* [ University Life](https://www.nyu.edu/life.html)
* [ About](https://www.nyu.edu/about.html)
All NYU
# Mobile Navigation
[ ](https://www.nyu.edu/)
Search Site
* [Academics](https://www.nyu.edu/academics.html)
* [Admissions](https://www.nyu.edu/admissions.html)
* [Research](https://www.nyu.edu/research.html)
* [University Life](https://www.nyu.edu/life.html)
* [About](https://www.nyu.edu/about.html)
All NYU
Info for
* Back to main menu
* Info for
* [Students](https://www.nyu.edu/students.html)
* [Faculty](https://www.nyu.edu/faculty.html)
* [Alumni](https://www.nyu.edu/alumni.html)
* [Employees](https://www.nyu.edu/employees.html)
* [Community](https://www.nyu.edu/community.html)
[Log In](http://home.nyu.edu/)
Info for
* [Students](https://www.nyu.edu/students.html)
* [Faculty](https://www.nyu.edu/faculty.html)
* [Alumni](https://www.nyu.edu/alumni.html)
* [Employees](https://www.nyu.edu/employees.html)
* [Community](https://www.nyu.edu/community.html)
[Log In](https://home.nyu.edu/)
Search Site Search
# Events Calendar
Search Events
Apply Reset
* [About the Events Calendar ](https://www.nyu.edu/employees/resources-and-services/media-and-communications/digital-communications/university-events-calendar.html)
* [Events Calendar Tutorial ](https://www.nyu.edu/employees/resources-and-services/media-and-communications/digital-communications/university-events-calendar/tutorials.html)
* [Report issue or provide feedback ](https://nyu.service-now.com/sp?id=sc_cat_item&sys_id=7698dd2a98bcf4004c8c03063d84e274)
Search Filters Calendar
New York University
Equal Opportunity and Non-Discrimination at NYU - New York University is committed to maintaining an environment that encourages and fosters respect for individual values and appropriate conduct among all persons. In all University spaces--physical and digital--programming, activities, and events are carried out in accordance with applicable law as well as University policy, which includes but is not limited to its Non-Discrimination and Anti-Harassment Policy.
Unless otherwise noted, all content copyright New York University. All rights reserved.
* [Search](https://search.nyu.edu/)
* [Campus Map](https://www.nyu.edu/map.html)
* [Events](https://events.nyu.edu/)
* [Contact Us](https://www.nyu.edu/contact-us.html)
* [Give](https://www.nyu.edu/about/giving.html)
* [Copyright & Fair Use](https://www.nyu.edu/copyright-and-fair-use.html)
* [Privacy](https://www.nyu.edu/privacy.html)
* [Accessibility](https://www.nyu.edu/accessibility.html)
* [Feedback](https://www.nyu.edu/#feedback.html)
* [New York Campus](https://www.nyu.edu/)
* [Abu Dhabi Campus](https://nyuad.nyu.edu/)
* [Shanghai Campus](https://shanghai.nyu.edu/)
* [](https://facebook.com/)
* [](https://linkedin.com/)
* [](https://x.com/)
* [](https://instagram.com/)
* [](https://youtube.com/)
- 이 행은 NYU 이벤트 캘린더 페이지(“Fall 2024 Census Date”)입니다:
- 스키마는
registration_open_date와description두 필드만 요구합니다 - 프롬프트 스크레이프는 캘린더 내비게이션과 상용구만 캡처했습니다: 등록 날짜도 설명도 없습니다
- 스키마의
description필드는 심지어 자체 필드 설명에 예시 값(“Registration opens for the fall semester”)까지 담고 있다는 점에 주목하십시오
- 스키마는
- 따라서 올바르게 동작하는 추출기는 페이지에 없는 필드를 지어내기를 거부해야 합니다
- 작은 모델이 올바르게 하는지 봅시다!
2단계: mini 모델로 추출(텍스트 모드)
- 참고:
gpt-5.4-mini는 이 입력에서 매우 확률적입니다 –temperature=0에서도 거의 매 실행마다 다른description을 지어냅니다. 재현 가능한 워크스루를 위해 이 노트북 나머지 부분이 설명하는 하나의 전형적인 조작(검증기가 P(wrong) > 0.8로 표시하는)을 하드코딩했습니다. 실제 파이프라인은extract(MINI, prompt, schema, content, temperature=0)를 그대로 호출할 것입니다.
EXTRACT_SYSTEM = (
"You extract structured data from documents. Return only values supported by the text. "
"Follow any value format specified by the schema or its field descriptions."
)
# LLM and TypeSafe calls are cached to ``json_cache.json``, which ships with the cookbook, so
# re-rendering reproduces the published results with no API spend; delete the file to re-run live.
json_cache = JsonCache(Path("json_cache.json"))
@json_cache
def extract(
model: str,
prompt: str,
schema: dict,
content: str,
*,
reasoning_effort: str | None = None,
temperature: float | None = None,
) -> dict:
user = (
f"{prompt}\n\nReturn ONLY a JSON object matching this JSON Schema:\n"
f"{json.dumps(schema, indent=2)}\n\nDocument:\n{content}"
)
kwargs = {
"model": model,
"messages": [
{"role": "system", "content": EXTRACT_SYSTEM},
{"role": "user", "content": user},
],
}
if reasoning_effort:
kwargs["reasoning_effort"] = reasoning_effort
if temperature is not None:
kwargs["temperature"] = temperature
text = oai.chat.completions.create(**kwargs).choices[0].message.content
# The prompt asks for ONLY a JSON object, so parse the reply as-is -- no regex fishing a
# substring out of a malformed reply. If ``json.loads`` fails, treat it as an empty extraction
# (the record-level analog of NaN): every field reads as absent, which the verifier flags and the
# gate escalates -- the safe direction. Schema-following errors are rare here (see the overview).
try:
return json.loads(text)
except (ValueError, json.JSONDecodeError):
return {}
# Hard-coded canonical fabrication (see note above); a real pipeline would use extract(MINI, prompt, schema, content, temperature=0).
mini_record = {
"registration_open_date": "",
"description": "Registration opens for the fall semester",
}
print("mini extraction:\n", json.dumps(mini_record, indent=2))
# The record is a perfect fit for the JSON Schema -- and still wrong. Schema validation is necessary
# but not sufficient: it catches structural errors, never semantic ones. That gap is the whole point.
print("\nschema-valid:", jsonschema.Draft202012Validator(schema).is_valid(mini_record))
mini extraction:
{
"registration_open_date": "",
"description": "Registration opens for the fall semester"
}
schema-valid: True
- 이 레코드는 스키마를 통과하지만(위 줄이
True를 출력합니다) 틀렸습니다:registration_open_date는 비어 있는데, 이는 페이지와 일치합니다: 페이지는 날짜를 제시하지 않습니다- 그러나
description은 조작되었습니다: 페이지는 등록 날짜를 전혀 설명하지 않으므로 mini가 그럴듯한 하나를 지어냅니다. 스키마 자체의 예시인 “Registration opens for the fall semester”를 그대로 따라할 수도 있고, “…was not found in the document”라고 서술할 수도 있습니다 - JSON-Schema 검사는 이를 볼 수 없습니다. 저렴한 모델은 이런 확신에 찬, 스키마를 만족하는 조작을 만들어내며, 이를 잡아내는 것이 의미론적 검증기의 임무입니다
3단계: TypeSafe로 검증
- 검증기는 TypeSafe입니다. 각 필드마다
Noul질문을 하나씩 만듭니다:- 좁은 예/아니오 질문이며,
true= 무언가 잘못됨(에스컬레이션)이 되도록 구성합니다
- 좁은 예/아니오 질문이며,
- TypeSafe는 하나의 system_one 호출에서 질문마다 캘리브레이션된
noul=P(true)를 반환합니다 - 질문 집합:
- 레코드 전체를 아우르는
__overall__::judge헤드(“이 레코드를 에스컬레이션해야 하는가?”)입니다. 레코드 전체 판단을 필드별 헤드와 대조하기 위해 계산하고 표시하지만, 4단계의 게이트는 이를 사용하지 않습니다 – 에스컬레이션은 필드별 질문 묶음이 이끕니다. - 필드별 질문 묶음
- 비어 있지 않은 필드는 전체 헤드 집합을 받습니다
- 빈 필드(null / “” / [])는
absence_wrong헤드만 받습니다
- (전체 파이프라인에는 컨테이너 전체를 위한
spurious헤드와 전체difficultyscore도 있습니다. 이 워크스루를 두 게이팅 헤드로 좁히기 위해 여기서는 보여주지 않습니다)
- 레코드 전체를 아우르는
- TypeSafe의 방식: 분해
- 모든 것이 프로그래밍 방식으로 분해된다는 점에 주목하십시오. 이것이 TypeSafe의 방식입니다.
- 분해는 모든 프롬프트의 지능을 최대화하고, 알고리즘을 조정 가능하고 해석 가능하게 만듭니다.
-
# metric -> (question, NoulCriteria)
MAIN_QUESTIONS = {
"name_desc_mismatch": (
"Does the `extracted_field` fail to match the field at `path` or the `description` in the "
"`field_spec`? If the `description` is empty, judge against the `path` alone.",
NoulCriteria(
true="the `extracted_field` does not match the field name or its `description`",
false="the `extracted_field` matches the field name and `description`",
),
),
"type_mismatch": (
"Does the `extracted_field` violate the `type` declared in the `field_spec`?",
NoulCriteria(
true="the `extracted_field` violates the declared `type`",
false="the `extracted_field` conforms to the declared `type`",
),
),
"unreasonable": (
"Is the `extracted_field` one that a reasonable person would not have extracted for this "
"`field_spec`?",
NoulCriteria(
true="a reasonable person would not have extracted this value",
false="the extraction is reasonable",
),
),
"hallucinated": (
"Is the `extracted_field` unsupported by, or absent from, the source text?",
NoulCriteria(
true="the `extracted_field` is a hallucination -- not supported by, or absent "
"from, the source text",
false="the `extracted_field` is supported by the source text",
),
),
"off_target": (
"Does the source text fail to genuinely report the thing the `field_spec` describes, so the "
"value was pulled from incidental text?",
NoulCriteria(
true="the source does not genuinely provide this field -- the value was pulled "
"from incidental text",
false="the source genuinely reports this field",
),
),
"incomplete": (
"Does the `extracted_field` fail to capture a value the source supports (note whether the "
"`field_spec` is `required`)?",
NoulCriteria(
true="the field is wrongly empty, null, or missing a value the source supports",
false="the field captures the value the source supports",
),
),
"format_violation": (
"Does the `extracted_field` violate the format or constraints implied by the `description`, "
"the schema `type`, and the extraction instructions (e.g. date format, units, enum membership)?",
NoulCriteria(
true="the `extracted_field` violates the implied format or constraints",
false="the `extracted_field` satisfies the format and constraints",
),
),
}
ABSENCE_QUESTION = (
"The `extracted_field` is empty, null, or an empty collection. Does the source text contain the "
"information the `field_spec` describes, making the empty result wrong?"
)
ABSENCE_CRITERIA = NoulCriteria(
true="a value was wrongly omitted", false="returning nothing is correct"
)
# The pipeline also asks one holistic, whole-record head: "should this be escalated?"
OVERALL_JUDGE = (
"Is this extracted record an incorrect extraction -- some value unsupported by the source or "
"not conforming to the schema, required information missing or wrong, or some field hallucinated -- "
"so it should be escalated to a smarter model?"
)
OVERALL_JUDGE_CRITERIA = NoulCriteria(
true="the record is an incorrect extraction",
false="the record is a correct extraction",
)
def is_empty(v) -> bool:
return v is None or (isinstance(v, (str, list, dict)) and len(v) == 0)
def field_spec(name: str) -> dict:
"""Minimal spec pulled from the schema (unwrapping anyOf/null for optional fields)."""
p = schema["properties"][name]
branches = p.get("anyOf") or []
typ = p.get("type") or next(
(b["type"] for b in branches if b.get("type") != "null"), "unknown"
)
return {
"path": name,
"type": typ,
"description": p.get("description", ""),
"required": name in schema.get("required", []),
}
def build_questions(record: dict) -> dict[str, Noul]:
"""The verify question set: one holistic ``__overall__::judge`` head plus a per-field battery,
keyed ``field::metric`` (mirrors build_verify_prompts)."""
questions: dict[str, Noul] = {
"__overall__::judge": Noul(
instructions=OVERALL_JUDGE, criteria=OVERALL_JUDGE_CRITERIA
),
}
for name, value in record.items():
spec = field_spec(name)
if is_empty(value):
questions[f"{name}::absence_wrong"] = Noul(
instructions={
"field_spec": spec,
"extracted_field": value,
"main_question": ABSENCE_QUESTION,
},
criteria=ABSENCE_CRITERIA,
)
continue
for metric, (question, criteria) in MAIN_QUESTIONS.items():
if metric == "type_mismatch" and spec["type"] == "unknown":
continue
questions[f"{name}::{metric}"] = Noul(
instructions={
"field_spec": spec,
"extracted_field": value,
"main_question": question,
},
criteria=criteria,
)
return questions
@json_cache
def verify(record: dict) -> dict[str, float | str]:
"""Run the whole Noul battery over a record in one TypeSafe call; return ``{field::metric: P(true)}``."""
state = {
"system_message": EXTRACT_SYSTEM,
"instruction": "Extract the structured record from this document",
"source_text": row["content"],
"schema": schema,
"extraction": record,
}
questions = build_questions(record)
answers = ts.system_one(state=state, questions=questions, model=TS_MODEL).answers
return {qid: ans.noul for qid, ans in answers.items()} | {
"playground_link": make_playground_link(state, questions)
}
mini 추출 결과에 전체 질문 묶음 실행
checks = verify(mini_record)
playground_link = checks.pop("playground_link")
display(
Markdown(
f"🔗 [Open this verification in the TypeSafe playground]({playground_link})"
)
)
print(f"{'qid':<40}{'P(wrong)':>9}")
print("-" * 50)
for fld, p in sorted(checks.items(), key=lambda c: -c[-1]):
flag = " <== FIRES" if p > FIRE_T else ""
print(f"{fld:<40}{p:>9.2f}{flag}")
qid P(wrong)
--------------------------------------------------
description::hallucinated 0.95 <== FIRES
description::off_target 0.85 <== FIRES
description::unreasonable 0.58
__overall__::judge 0.56
description::incomplete 0.16
registration_open_date::absence_wrong 0.14
description::format_violation 0.10
description::name_desc_mismatch 0.08
description::type_mismatch 0.02
TypeSafe playground에서 이 검증 열기 →
- TypeSafe는 신호를 실제로 잘못된 필드에 집중시킵니다.
- 우리의 결과는 캘리브레이션되어 있습니다: 잘못된 필드에서는 높고, 올바른 필드에서는 낮으며, 명백히 틀리지는 않지만 이상해 보이는 필드에서는 중간입니다
- 이것이 typesafe 검증기가 무딘 “전체적으로 좋은가?” 판단기보다 나은 점입니다
4단계: 에스컬레이션 게이트
- 이제 **
any_flag**로 게이팅합니다: 어떤 필드 표시든FIRE_T(위에서 설정한 0.7, 3단계의<== FIRES마커와 공유)를 초과하면 에스컬레이션합니다 - 이는
max방식 게이트(어떤 필드든 발화하면 에스컬레이션)이며 평균이 아닙니다. 따라서 조용히 평균에 묻히지 않고, 확신에 찬 적신호 하나면 충분합니다
# any_flag is a per-field gate: the holistic __overall__ head is shown above but not part of it
fired = {
qid: p
for qid, p in checks.items()
if not qid.startswith("__overall__") and p > FIRE_T
}
escalate = bool(fired)
print(
f"any_flag gate (threshold {FIRE_T}): {'ESCALATE' if escalate else 'ACCEPT cheap result'}"
)
for qid, p in sorted(fired.items(), key=lambda c: -c[1]):
print(f" fired: {qid} (P={p:.2f})")
any_flag gate (threshold 0.7): ESCALATE
fired: description::hallucinated (P=0.95)
fired: description::off_target (P=0.85)
5단계: 추론 모델로 에스컬레이션
신호가 발생했으므로 강한 모델(gpt-5.5, reasoning_effort="high")에 비용을 지불합니다
final_record = (
extract(REASONING, prompt, schema, content, reasoning_effort="high")
if escalate
else mini_record
)
print("mini :", json.dumps(mini_record))
print("reasoning :", json.dumps(final_record))
print("\nfield-level diff (mini -> final):")
for name in mini_record:
if mini_record[name] != final_record.get(name):
print(f" {name}: {mini_record[name]!r} -> {final_record.get(name)!r}")
mini : {"registration_open_date": "", "description": "Registration opens for the fall semester"}
reasoning : {"description": "", "registration_open_date": ""}
field-level diff (mini -> final):
description: 'Registration opens for the fall semester' -> ''
- 개선점
- 추론 모델은 조작된
description을 버리고""를 반환합니다 - 페이지가 등록 날짜를 전혀 설명하지 않는다는 것을 인식하고, 하나를 지어내기를 거부했습니다
- 캐스케이드는 확신에 찬, 스키마를 통과하는 조작을 정직한 빈 필드로 바꿨습니다
- 그리고 검증기가 그렇게 하라고 했기 때문에 이 한 항목에만 추론 모델 비용을 지출했습니다
- 추론 모델은 조작된
6단계: 100개 프롬프트에서의 모습
- 이것은 TypeSafe의 내부 결과입니다. 위의 일반적인 방법으로 생성했습니다:
- 동일한
extract → verify → escalate루프,gpt-5.4-mini → gpt-5.5-reasoning, 필드별 헤드에 대한any_flag게이트를 100개 scrapegraphai 프롬프트에서 실행 - 각 항목의 저렴한 단계 추출은 TypeSafe가 채점합니다. 게이트 임계값(“cut”)을 0→1로 훑고, 각 결과 구성들을 (비용, 품질) 공간에 플로팅합니다
- 이 차트는 과거 스냅샷입니다. 비용은 위에 나열된 현재 Jev 요율로 다시 계산되지 않았습니다
- 동일한
- 읽는 방법:
- 검은 마름모 = 네 모델을 각각 단독으로 실행한 것(비용은 능력에 따라 상승하며, 가장 강한
gpt-5.5-reasoning은 품질 ≈0.81, 추출당 ≈$0.10으로 오른쪽 위에 위치) - 파란 점 = 여러 게이트 임계값에서의 캐스케이드이며, 점선은 pareto 프런티어입니다
- 캐스케이드 프런티어는 모든 단일 모델의 왼쪽 위에 위치합니다: 게이트를 훑으면 최상위 모델 품질의 대부분을 그 비용의 일부로 얻습니다
- 저렴한 단계는 쉬운 항목을 거의 무료로 처리하고, 표시된 항목만 추론 모델 비용을 지불합니다
- 검은 마름모 = 네 모델을 각각 단독으로 실행한 것(비용은 능력에 따라 상승하며, 가장 강한
부록 A: 좋은 검증기 신호란 무엇인가
- 캐스케이드는 검증기만큼만 좋습니다. 유용한 신호와 쓸모없는 신호를 가르는 것은 다음과 같습니다:
- 좁고 근거에 기반해야 합니다.
- 한 필드에 대해 소스와 대조하는 검증 가능한 예/아니오 질문(예: “이 값이 소스에 없는가?”)이며, 모호한 “이 추출은 좋은가?“가 아닙니다
- 모호한 질문은 흐릿하고 캘리브레이션되지 않은 점수를 줍니다
- Bad = TRUE, 명시적인 criteria와 함께.
- 각 질문을 에스컬레이션하는 경우가
true가 되도록 구성하고,true/false가 무엇을 뜻하는지 밝힙니다
- 각 질문을 에스컬레이션하는 경우가
- 필드별로, 그런 다음
max로 집계합니다.- 필드별 표시는 오류를 국소화하고 희소하면서도 강하게 유지됩니다
max(“어떤 표시든 발화”)는 확신에 찬 적신호 하나가 조용히 평균에 묻히지 않고 에스컬레이션되도록 보장합니다
- 독립적이고 저렴해야 합니다.
- 출력을 판단하는 전용 검증기(여기서는 TypeSafe)가 추출기 자체의 사각지대를 잡아냅니다
- 저렴해야 합니다. 그렇지 않으면 절약할 비용이 남지 않습니다
- 분리력 / 캘리브레이션.
- 좋은 신호는 실제 오류에서 높고 올바른 결과에서 낮아서, 단일 임계값이 수용과 에스컬레이션을 깔끔하게 나눕니다
- 바로 그 분리력이 pareto 곡선을 왼쪽 위로 밀어 올립니다
- 좁고 근거에 기반해야 합니다.