質問の並列実行
質問の並列実行
GDPR の Wikipedia 記事に対して 13 問の規制ブリーフィングを実行し、すべての質問を 1 回の TypeSafe 呼び出しにまとめると、答えを変えずに 12.2 倍安く、10.0 倍速くなることを示します。
1 つの文書と、それについての N 個の質問があるとします。N 個すべての質問を 1 回のリクエストで送ることも、 質問ごとに 1 回、N 回のリクエストで送ることもできます。TypeSafe ではどちらでも答えは同じに なります。各質問は文書に対して単独でスコアリングされるため、その答えは リクエストに他に何が含まれるかに依存しません。
これを確かめるため、この cookbook は各質問を両方の方法で何度も尋ねます。N 個すべてを 1 回の リクエストで、そして 1 リクエストにつき 1 問で。そして実行間の標準偏差、つまりある答えが 繰り返しごとにどれだけ動くかを比較します。ある質問が持つノイズは、 どちらのバッチ戦略でも同じ大きさです。バッチ処理がノイズを加えることはありません。ほとんどの答えは どちらの方法でも 5 回の繰り返しすべてで同一に返り、毎回同じ値で、標準偏差はちょうど 0.0 でした。
コストと速度は変わります。文書がすべてのリクエストの大部分を占めます。N 回の単一質問呼び出しは その文書を N 回分、N 回の往復で支払います。バッチ呼び出しは 1 回だけ支払います。文書が大きいほど、 この節約は完全な N 倍に近づきます。
ここでのケースは規制のブリーフィングです。文書は GDPR についての Wikipedia 記事
(約 54,000 文字。文書がすべてのリクエストの大部分を占める、文書支配型のワークロードです)で、
コンプライアンスチームは 13 項目の確認を望んでいます:8 問の Noul 質問、2 問の Choice
質問、そして 3 問の Score 質問です。
セットアップ
pip install ipython 'cooksafe>=0.2.0,<0.3.0'
次に TYPESAFE_API_KEY を設定します。
import json
import os
import urllib.request
from pathlib import Path
from statistics import mean, stdev
from time import perf_counter
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, ChoiceAnswer, Noul, NoulAnswer, Score, TypeSafeClient
TYPESAFE_MODEL = "jev-1.12"
PRICE = (
0.042,
0.00,
) # $ per 1M tokens (input, output); TypeSafe jev-1.12 as of 2026-09, see README
RUNS = 5 # repeats per batching strategy, to estimate each answer's run-to-run std dev
client = TypeSafeClient(api_key=os.environ["TYPESAFE_API_KEY"], timeout=120.0)
json_cache = JsonCache(Path("json_cache.json"))
文書:GDPR についての Wikipedia 記事
記事の固定したリビジョンからプレーンテキストとして取得し、API 呼び出しの隣の json_cache.json
にキャッシュします。こうすることで、公開中の記事が編集されても、文書とその数値は
固定されたままになります。
WIKIPEDIA_REVISION = 1363040264 # "General Data Protection Regulation", as of 2026-07
@json_cache
def fetch_article(revision_id: int) -> str:
url = (
"https://en.wikipedia.org/w/api.php?action=query&format=json"
f"&prop=extracts&explaintext=1&revids={revision_id}"
)
request = urllib.request.Request(
url, headers={"User-Agent": "typesafe-cookbook/1.0"}
)
with urllib.request.urlopen(request) as response:
pages = json.loads(response.read())["query"]["pages"]
return next(iter(pages.values()))["extract"]
DOCUMENT = {
"source": f"https://en.wikipedia.org/?oldid={WIKIPEDIA_REVISION}",
"text": fetch_article(WIKIPEDIA_REVISION),
}
print(f"{len(DOCUMENT['text']):,} characters")
display(Markdown(f"📄 [Read the pinned Wikipedia revision]({DOCUMENT['source']})"))
53,777 characters
質問:8 問の noul + 2 問の choice + 3 問の score
型ごとに、答えごとに追跡する数値は 1 つです。
Noul:「はい」の確率。Choice:最大確率、つまり選ばれたラベルに付いた確率。criteriaは各 ラベルをその意味に対応付けます。Score:0-1 に正規化したスコアで、スコアを最上位レベルで割ったもの。criteriaはレベル 0 から順にレベル説明を並べます。
QUESTIONS = {
"breach_72h": Noul(
instructions="Must a personal data breach be reported to the supervisory authority within 72 hours?"
),
"applies_non_eu": Noul(
instructions="Does the regulation apply to organisations established outside the EU that offer goods or services to people in the EU?"
),
"dpo_all_orgs": Noul(
instructions="Must every organisation appoint a Data Protection Officer, regardless of what data it processes?"
),
"pre_ticked_consent": Noul(
instructions="Can valid consent be obtained through pre-ticked boxes or inactivity?"
),
"right_erasure": Noul(
instructions="Does the regulation grant individuals a right to erasure of their personal data?"
),
"data_portability": Noul(
instructions="Does the regulation include a right to data portability?"
),
"us_federal_law": Noul(instructions="Is the GDPR a United States federal law?"),
"criminal_penalties": Noul(
instructions="Does the GDPR itself impose criminal penalties such as imprisonment?"
),
"instrument_type": Choice(
instructions="What kind of EU legal instrument is the GDPR?",
criteria={
"Regulation": "Directly binding law in all member states, no national implementation needed.",
"Directive": "Sets goals that member states implement through national law.",
"Treaty": "An international treaty between states.",
"Recommendation": "Non-binding guidance.",
},
),
"max_fine": Choice(
instructions="What is the maximum administrative fine for the most serious infringements?",
criteria={
"TwentyM_or_4pct": "Up to EUR 20 million or 4% of annual worldwide turnover, whichever is greater.",
"TenM_or_2pct": "Up to EUR 10 million or 2% of annual worldwide turnover, whichever is greater.",
"FixedCap": "A fixed amount not tied to turnover.",
"NoFines": "The GDPR provides no administrative fines.",
},
),
"individual_rights": Score(
instructions="How strong are the rights the GDPR grants to individuals over their data?",
criteria=[
"None: individuals get no rights over their data.",
"Weak: a right to be informed, but little control.",
"Moderate: access and correction rights, but limited means to act on them.",
"Strong: access, erasure, portability, and objection rights, with enforcement behind them.",
],
),
"penalty_severity": Score(
instructions="How severe are the penalties the GDPR provides for non-compliance?",
criteria=[
"None: no penalties of any kind.",
"Symbolic: small fixed fines unlikely to change behavior.",
"Substantial: fines large enough to matter to most companies.",
"Severe: fines scaled to global revenue, material even to the largest companies.",
],
),
"compliance_burden": Score(
instructions="How heavy is the compliance burden the GDPR places on organisations?",
criteria=[
"Negligible: no meaningful obligations.",
"Light: a few notices and disclosures.",
"Moderate: documented processes and some dedicated roles for larger processors.",
"Heavy: records, impact assessments, officers, and breach procedures for many organisations.",
"Extreme: obligations so demanding that ordinary organisations cannot fully comply.",
],
),
}
N = len(QUESTIONS)
METRIC = { # question type -> the one number we track per answer
Noul: "p(yes)",
Choice: "max prob",
Score: "normalized score",
}
2 通りの方法で、それぞれ 5 回尋ねる
ask() は、文書とともに質問の任意の部分集合を送り、各答えを追跡する
1 つの数値に還元します。文書はどの呼び出しでもバイト単位で同一です。
どちらのバッチ戦略も RUNS = 5 回実行し、各質問に戦略ごとに 5 つの答えが得られます。
平均(両者は一致するか)と標準偏差(バッチ処理がノイズを加えるか)を比較するのに十分です。
呼び出しは cookbook に同梱された json_cache.json にキャッシュされるので、再レンダリングは
無料です。ライブで再実行するには削除してください。
@json_cache
def ask(keys: tuple[str, ...], run: int):
"""One TypeSafe call -> ({key: tracked metric}, input_tokens, output_tokens, latency_s);
``run`` only forces a distinct live call per repeat."""
started = perf_counter()
response = client.system_one(
state={"article": DOCUMENT},
questions={key: QUESTIONS[key] for key in keys},
model=TYPESAFE_MODEL,
)
values = {}
for key in keys:
answer = response.answers[key]
if isinstance(answer, NoulAnswer):
values[key] = answer.noul
elif isinstance(answer, ChoiceAnswer):
values[key] = max(answer.probabilities.values())
else:
values[key] = answer.score / (len(QUESTIONS[key].criteria) - 1)
return (
values,
response.usage.input_tokens,
response.usage.output_tokens,
perf_counter() - started,
)
def priced(result):
"""({key: metric}, in_tokens, out_tokens, latency) -> ({key: metric}, cost_usd, latency)."""
values, input_tokens, output_tokens, latency = result
return values, input_tokens / 1e6 * PRICE[0] + output_tokens / 1e6 * PRICE[1], latency
# Price after cache retrieval, so a price change needs no new calls.
batched = [
priced(ask(tuple(QUESTIONS), run)) for run in range(RUNS)
] # all N in one call, x RUNS
singles = [
{key: priced(ask((key,), run)) for key in QUESTIONS} for run in range(RUNS)
] # N x 1, x RUNS
バッチ処理は答えを変えない
質問ごとに、追跡する数値の 5 回の実行にわたる平均と標準偏差を、各バッチ戦略で 示します。バッチ処理が答えを変えるなら、バッチ列は単一列と異なるはずです。 平均がずれればバイアス、標準偏差が大きければノイズです。
print(
f"{'question':<22}{'metric':<18}{'batched mean':>13}{'single mean':>12}"
f"{'batched std':>13}{'single std':>12}"
)
for key, question in QUESTIONS.items():
batched_values = [values[key] for values, _cost, _latency in batched]
single_values = [singles[run][key][0][key] for run in range(RUNS)]
print(
f"{key:<22}{METRIC[type(question)]:<18}{mean(batched_values):>13.3f}"
f"{mean(single_values):>12.3f}{stdev(batched_values):>13.4f}{stdev(single_values):>12.4f}"
)
question metric batched mean single mean batched std single std
breach_72h p(yes) 0.804 0.814 0.0055 0.0055
applies_non_eu p(yes) 0.990 0.990 0.0000 0.0000
dpo_all_orgs p(yes) 0.030 0.030 0.0000 0.0000
pre_ticked_consent p(yes) 0.040 0.040 0.0000 0.0000
right_erasure p(yes) 0.990 0.990 0.0000 0.0000
data_portability p(yes) 0.990 0.990 0.0000 0.0000
us_federal_law p(yes) 0.010 0.010 0.0000 0.0000
criminal_penalties p(yes) 0.108 0.108 0.0045 0.0084
instrument_type max prob 1.000 1.000 0.0000 0.0000
max_fine max prob 1.000 1.000 0.0000 0.0000
individual_rights normalized score 1.000 1.000 0.0000 0.0000
penalty_severity normalized score 1.000 1.000 0.0000 0.0000
compliance_burden normalized score 0.750 0.750 0.0000 0.0000
質問タイプごとに表を読むと:
- choice、score、そして 8 つの noul のうち 6 つは、5 回の繰り返しを通じて同一に返りました。 どちらのバッチ戦略でも標準偏差はちょうど 0.0 で、バッチ呼び出しも単一呼び出しも 毎回同じ数値を返します。N 問を含む 1 回の呼び出しは、1 問ずつの N 回の呼び出しと 同じ答えを与えます。
breach_72hとcriminal_penaltiesは実行間のサンプリングノイズを少し持ちますが、それは どちらのバッチ戦略でも同じ大きさで、平均はそのノイズの範囲内で一致します。 ノイズはバッチの仕方ではなく質問の性質です。バッチ処理は答えをずらすことも 分散を加えることもありません。
いずれにせよ、バッチの効果はありません。どの質問の答えも、リクエストを共有する 他の 12 問に依存しません。
唯一の違い:コストと速度
答えは同じでも、請求は違います。約 54,000 文字の記事がすべてのリクエストの大部分を占めるので:
- コスト:13 回の単一質問呼び出しは記事を 13 回送り直します。バッチ呼び出しは 記事を 1 回送るだけです。この節約は呼び出し方を問わず成り立ちます。
- 速度:この数値は 13 回の単一呼び出しのレイテンシを合計したもので、それらが順番に 実行される前提です。並行に実行すれば差は縮まりますが、13 倍のトークンコストは残ります。
トークン数とレイテンシは答えと一緒にキャッシュされます。コストは後から適用し、 どちらも 5 回の実行で平均します。
batched_cost = mean(cost for _values, cost, _latency in batched)
batched_latency = mean(latency for _values, _cost, latency in batched)
singles_cost = mean(
sum(singles[run][key][1] for key in QUESTIONS) for run in range(RUNS)
)
singles_latency = mean(
sum(singles[run][key][2] for key in QUESTIONS) for run in range(RUNS)
)
print(f"{'batching':<24}{'calls':>6}{'cost':>12}{'total time':>12}")
print(
f"{f'one call, all {N}':<24}{1:>6}{'$' + format(batched_cost, '.6f'):>12}{format(batched_latency, '.2f') + 's':>12}"
)
print(
f"{f'{N} calls, one each':<24}{N:>6}{'$' + format(singles_cost, '.6f'):>12}{format(singles_latency, '.2f') + 's':>12}"
)
print(
f"\nbatching: {singles_cost / batched_cost:.1f}x cheaper, {singles_latency / batched_latency:.1f}x faster"
)
batching calls cost total time
one call, all 13 1 $0.000497 0.27s
13 calls, one each 13 $0.006090 2.71s
batching: 12.2x cheaper, 10.0x faster
TypeSafe プレイグラウンドで開く
同じ記事と同じ 13 問を共有リンクにまとめたものです。開くとブリーフィングを ライブで再実行でき、同じ数値が返ります。
playground_link = make_playground_link(
{"article": DOCUMENT}, QUESTIONS, models=[TYPESAFE_MODEL]
)
display(
Markdown(
f"🔗 [Open this article + questions in the TypeSafe playground]({playground_link})"
)
)
TypeSafe プレイグラウンドでこの記事と質問を開く →