병렬 질문
GDPR 위키백과 문서에 대해 13개 질문으로 규제 브리핑을 수행하며, 모든 질문을 하나의 TypeSafe 호출에 묶으면 답변의 변화 없이 12.2배 저렴하고 10.0배 빠르다는 것을 보여줍니다.
문서 하나와 그에 대한 N개의 질문이 있습니다. N개를 모두 담은 요청 하나를 보낼 수도 있고, 질문 하나씩 담은 N개의 요청을 보낼 수도 있습니다. TypeSafe에서는 어느 쪽이든 답이 같게 나옵니다. 각 질문은 문서에 대해 개별적으로 채점되므로, 그 답은 요청에 다른 무엇이 들어 있든 영향을 받지 않습니다.
이를 확인하기 위해 이 cookbook은 각 질문을 두 방식으로 여러 번 던집니다. N개를 한 요청에 담는 방식과 요청마다 질문 하나를 담는 방식입니다. 그리고 회차 간 표준편차를 비교합니다. 즉, 한 답이 반복마다 얼마나 움직이는지를 봅니다. 어떤 질문에 잡음이 있더라도 두 배칭 전략 모두에서 같은 크기로 나타납니다. 배칭이 잡음을 더하지는 않습니다. 대부분의 답은 5회 반복 내내 어느 방식에서든 동일하게 돌아왔고, 매 호출마다 같은 값이었으며 표준편차는 정확히 0.0이었습니다.
비용과 속도는 달라집니다. 모든 요청에서 문서가 대부분을 차지합니다. 질문 하나짜리 호출 N번은 문서 값을 N번, N회 왕복으로 치릅니다. 배칭 호출은 한 번만 치릅니다. 문서가 클수록 그 절약은 N배에 가까워집니다.
여기서 다루는 사례는 규제 브리핑입니다. 문서는 GDPR에 관한 위키백과 문서이며(~54,000자, 모든 요청의 대부분을 문서가 차지하는 문서 중심 워크로드입니다), 규정 준수 팀은 13가지를 확인하려 합니다. Noul 질문 8개, Choice 질문 2개, Score 질문 3개입니다.
설정
pip install ipython 'cooksafe>=0.2.0,<0.3.0'
그런 다음 TYPESAFE_API_KEY를 설정합니다.
import json
import os
import urllib.request
from pathlib import Path
from statistics import mean, stdev
from time import perf_counter
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, ChoiceAnswer, Noul, NoulAnswer, Score, TypeSafeClient
TYPESAFE_MODEL = "jev-1.12"
PRICE = (
0.042,
0.00,
) # $ per 1M tokens (input, output); TypeSafe jev-1.12 as of 2026-09, see README
RUNS = 5 # repeats per batching strategy, to estimate each answer's run-to-run std dev
client = TypeSafeClient(api_key=os.environ["TYPESAFE_API_KEY"], timeout=120.0)
json_cache = JsonCache(Path("json_cache.json"))
문서: GDPR 위키백과 문서
문서의 고정된 리비전에서 일반 텍스트로 가져와 API 호출 옆의 json_cache.json에 캐시합니다. 따라서 실시간 문서가 편집되어도 이 문서와 그 수치는 고정된 상태로 유지됩니다.
WIKIPEDIA_REVISION = 1363040264 # "General Data Protection Regulation", as of 2026-07
@json_cache
def fetch_article(revision_id: int) -> str:
url = (
"https://en.wikipedia.org/w/api.php?action=query&format=json"
f"&prop=extracts&explaintext=1&revids={revision_id}"
)
request = urllib.request.Request(
url, headers={"User-Agent": "typesafe-cookbook/1.0"}
)
with urllib.request.urlopen(request) as response:
pages = json.loads(response.read())["query"]["pages"]
return next(iter(pages.values()))["extract"]
DOCUMENT = {
"source": f"https://en.wikipedia.org/?oldid={WIKIPEDIA_REVISION}",
"text": fetch_article(WIKIPEDIA_REVISION),
}
print(f"{len(DOCUMENT['text']):,} characters")
display(Markdown(f"📄 [Read the pinned Wikipedia revision]({DOCUMENT['source']})"))
53,777 characters
질문: noul 8개 + choice 2개 + score 3개
유형별로 답변마다 추적하는 숫자 하나가 있습니다.
Noul: “예”일 확률입니다.Choice: 최대 확률, 즉 선택된 레이블에 실린 확률입니다.criteria는 각 레이블을 그 의미에 대응시킵니다.Score: 0-1로 정규화한 점수, 즉 점수를 최상위 레벨로 나눈 값입니다.criteria는 레벨 0부터 올라가는 레벨 설명을 나열합니다.
QUESTIONS = {
"breach_72h": Noul(
instructions="Must a personal data breach be reported to the supervisory authority within 72 hours?"
),
"applies_non_eu": Noul(
instructions="Does the regulation apply to organisations established outside the EU that offer goods or services to people in the EU?"
),
"dpo_all_orgs": Noul(
instructions="Must every organisation appoint a Data Protection Officer, regardless of what data it processes?"
),
"pre_ticked_consent": Noul(
instructions="Can valid consent be obtained through pre-ticked boxes or inactivity?"
),
"right_erasure": Noul(
instructions="Does the regulation grant individuals a right to erasure of their personal data?"
),
"data_portability": Noul(
instructions="Does the regulation include a right to data portability?"
),
"us_federal_law": Noul(instructions="Is the GDPR a United States federal law?"),
"criminal_penalties": Noul(
instructions="Does the GDPR itself impose criminal penalties such as imprisonment?"
),
"instrument_type": Choice(
instructions="What kind of EU legal instrument is the GDPR?",
criteria={
"Regulation": "Directly binding law in all member states, no national implementation needed.",
"Directive": "Sets goals that member states implement through national law.",
"Treaty": "An international treaty between states.",
"Recommendation": "Non-binding guidance.",
},
),
"max_fine": Choice(
instructions="What is the maximum administrative fine for the most serious infringements?",
criteria={
"TwentyM_or_4pct": "Up to EUR 20 million or 4% of annual worldwide turnover, whichever is greater.",
"TenM_or_2pct": "Up to EUR 10 million or 2% of annual worldwide turnover, whichever is greater.",
"FixedCap": "A fixed amount not tied to turnover.",
"NoFines": "The GDPR provides no administrative fines.",
},
),
"individual_rights": Score(
instructions="How strong are the rights the GDPR grants to individuals over their data?",
criteria=[
"None: individuals get no rights over their data.",
"Weak: a right to be informed, but little control.",
"Moderate: access and correction rights, but limited means to act on them.",
"Strong: access, erasure, portability, and objection rights, with enforcement behind them.",
],
),
"penalty_severity": Score(
instructions="How severe are the penalties the GDPR provides for non-compliance?",
criteria=[
"None: no penalties of any kind.",
"Symbolic: small fixed fines unlikely to change behavior.",
"Substantial: fines large enough to matter to most companies.",
"Severe: fines scaled to global revenue, material even to the largest companies.",
],
),
"compliance_burden": Score(
instructions="How heavy is the compliance burden the GDPR places on organisations?",
criteria=[
"Negligible: no meaningful obligations.",
"Light: a few notices and disclosures.",
"Moderate: documented processes and some dedicated roles for larger processors.",
"Heavy: records, impact assessments, officers, and breach procedures for many organisations.",
"Extreme: obligations so demanding that ordinary organisations cannot fully comply.",
],
),
}
N = len(QUESTIONS)
METRIC = { # question type -> the one number we track per answer
Noul: "p(yes)",
Choice: "max prob",
Score: "normalized score",
}
두 방식으로 각 5번 질문하기
ask()는 질문의 임의 부분집합을 문서와 함께 보내고, 각 답변을 추적하는 숫자 하나로 환원합니다. 문서는 모든 호출에서 바이트 단위로 동일합니다.
두 배칭 전략을 각각 RUNS = 5번 실행하여, 전략마다 질문 하나에 5개의 답변을 얻습니다. 평균(두 방식이 일치하는가?)과 표준편차(배칭이 잡음을 더하는가?)를 비교하기에 충분합니다. 호출은 cookbook과 함께 제공되는 json_cache.json에 캐시되므로 다시 렌더링하는 데 비용이 들지 않습니다. 파일을 삭제하면 실시간으로 다시 실행합니다.
@json_cache
def ask(keys: tuple[str, ...], run: int):
"""One TypeSafe call -> ({key: tracked metric}, input_tokens, output_tokens, latency_s);
``run`` only forces a distinct live call per repeat."""
started = perf_counter()
response = client.system_one(
state={"article": DOCUMENT},
questions={key: QUESTIONS[key] for key in keys},
model=TYPESAFE_MODEL,
)
values = {}
for key in keys:
answer = response.answers[key]
if isinstance(answer, NoulAnswer):
values[key] = answer.noul
elif isinstance(answer, ChoiceAnswer):
values[key] = max(answer.probabilities.values())
else:
values[key] = answer.score / (len(QUESTIONS[key].criteria) - 1)
return (
values,
response.usage.input_tokens,
response.usage.output_tokens,
perf_counter() - started,
)
def priced(result):
"""({key: metric}, in_tokens, out_tokens, latency) -> ({key: metric}, cost_usd, latency)."""
values, input_tokens, output_tokens, latency = result
return values, input_tokens / 1e6 * PRICE[0] + output_tokens / 1e6 * PRICE[1], latency
# Price after cache retrieval, so a price change needs no new calls.
batched = [
priced(ask(tuple(QUESTIONS), run)) for run in range(RUNS)
] # all N in one call, x RUNS
singles = [
{key: priced(ask((key,), run)) for key in QUESTIONS} for run in range(RUNS)
] # N x 1, x RUNS
배칭은 답을 바꾸지 않습니다
질문마다, 각 배칭 전략에서 5회 실행에 걸친 추적 숫자의 평균과 표준편차입니다. 배칭이 답을 바꿨다면 배칭 열이 단일 열과 달라졌을 것입니다. 평균이 이동한 것은 편향이고, 표준편차가 더 큰 것은 잡음입니다.
print(
f"{'question':<22}{'metric':<18}{'batched mean':>13}{'single mean':>12}"
f"{'batched std':>13}{'single std':>12}"
)
for key, question in QUESTIONS.items():
batched_values = [values[key] for values, _cost, _latency in batched]
single_values = [singles[run][key][0][key] for run in range(RUNS)]
print(
f"{key:<22}{METRIC[type(question)]:<18}{mean(batched_values):>13.3f}"
f"{mean(single_values):>12.3f}{stdev(batched_values):>13.4f}{stdev(single_values):>12.4f}"
)
question metric batched mean single mean batched std single std
breach_72h p(yes) 0.804 0.814 0.0055 0.0055
applies_non_eu p(yes) 0.990 0.990 0.0000 0.0000
dpo_all_orgs p(yes) 0.030 0.030 0.0000 0.0000
pre_ticked_consent p(yes) 0.040 0.040 0.0000 0.0000
right_erasure p(yes) 0.990 0.990 0.0000 0.0000
data_portability p(yes) 0.990 0.990 0.0000 0.0000
us_federal_law p(yes) 0.010 0.010 0.0000 0.0000
criminal_penalties p(yes) 0.108 0.108 0.0045 0.0084
instrument_type max prob 1.000 1.000 0.0000 0.0000
max_fine max prob 1.000 1.000 0.0000 0.0000
individual_rights normalized score 1.000 1.000 0.0000 0.0000
penalty_severity normalized score 1.000 1.000 0.0000 0.0000
compliance_burden normalized score 0.750 0.750 0.0000 0.0000
질문 유형별로 표를 읽으면 다음과 같습니다.
- choice, score, 그리고 8개 noul 중 6개는 5회 반복 내내 동일하게 돌아옵니다. 두 배칭 전략 모두에서 표준편차가 정확히 0.0이고, 모든 배칭 호출과 단일 호출이 같은 숫자를 돌려줍니다. 질문 N개를 담은 호출 하나는 질문 하나짜리 호출 N개와 같은 답을 줍니다.
breach_72h와criminal_penalties에는 회차 간 표집 잡음이 조금 있으며, 그 크기는 두 배칭 전략 모두에서 같고 평균은 그 잡음 범위 안에서 일치합니다. 잡음은 질문의 속성이지 배칭 방식의 속성이 아닙니다. 배칭은 답을 이동시키지도, 분산을 더하지도 않습니다.
어느 쪽이든 배칭 효과는 없습니다. 어떤 질문의 답도 같은 요청을 공유하는 나머지 12개 질문에 의존하지 않습니다.
유일한 차이: 비용과 속도
같은 답, 다른 청구서입니다. ~54,000자 문서가 모든 요청을 지배하므로 다음과 같습니다.
- 비용: 질문 하나짜리 호출 13번은 문서를 13번 다시 보냅니다. 배칭 호출은 한 번만 보냅니다. 이 절약은 호출을 어떻게 발사하든 유지됩니다.
- 속도: 이 수치는 단일 호출 13번의 지연 시간을 합한 것이므로, 그것들이 차례로 실행된다고 가정합니다. 동시에 발사하면 격차는 줄지만, 13배의 토큰 비용은 그대로입니다.
토큰 수와 지연 시간은 답변과 함께 캐시됩니다. 비용은 나중에 적용하며, 둘 다 5회 실행에 걸쳐 평균합니다.
batched_cost = mean(cost for _values, cost, _latency in batched)
batched_latency = mean(latency for _values, _cost, latency in batched)
singles_cost = mean(
sum(singles[run][key][1] for key in QUESTIONS) for run in range(RUNS)
)
singles_latency = mean(
sum(singles[run][key][2] for key in QUESTIONS) for run in range(RUNS)
)
print(f"{'batching':<24}{'calls':>6}{'cost':>12}{'total time':>12}")
print(
f"{f'one call, all {N}':<24}{1:>6}{'$' + format(batched_cost, '.6f'):>12}{format(batched_latency, '.2f') + 's':>12}"
)
print(
f"{f'{N} calls, one each':<24}{N:>6}{'$' + format(singles_cost, '.6f'):>12}{format(singles_latency, '.2f') + 's':>12}"
)
print(
f"\nbatching: {singles_cost / batched_cost:.1f}x cheaper, {singles_latency / batched_latency:.1f}x faster"
)
batching calls cost total time
one call, all 13 1 $0.000497 0.27s
13 calls, one each 13 $0.006090 2.71s
batching: 12.2x cheaper, 10.0x faster
TypeSafe playground에서 열기
같은 문서와 같은 13개 질문을 공유 링크에 담았습니다. 열면 브리핑을 실시간으로 다시 실행하며, 같은 숫자가 돌아옵니다.
playground_link = make_playground_link(
{"article": DOCUMENT}, QUESTIONS, models=[TYPESAFE_MODEL]
)
display(
Markdown(
f"🔗 [Open this article + questions in the TypeSafe playground]({playground_link})"
)
)
TypeSafe playground에서 이 문서 + 질문 열기 →