ドキュメント

信頼度を使った分類

信頼度を使った分類

SEC の年次報告書を、それぞれ 1 つの Choice で 75 の業種グループに分類し、その答え自身の信頼度を読んで、そのグループを報告するか、その上位のより広い部門を報告するかを決めます。

SEC に年次報告書を提出するすべての企業は、その中で自社の事業を説明しています。私たちは その説明を標準産業分類(SIC)の下で分類します。75 の業種グループがあり、書類ごとに 1 つの Choice 質問を使います。

ほとんどの提出書類は簡単です。地域銀行は地域銀行です。そうでないものもあります。2 つの セグメントの一方をたった今売却した企業や、実際に営んでいる事業ではなく参入予定の事業を 説明しているスタートアップなどです。モデルはいずれにせよグループを選ばねばならず、 難しいケースの答えは、簡単なケースの答えと見た目が変わりません。難しいケースと簡単な ケースを見分けることに、通常コストがかかります。2 つ目のモデル、追加の呼び出し、人手レビューです。

Choice はすでに教えてくれます。当選した選択肢とともに confidence を返し、確率のほぼ すべてが 1 つの選択肢に載れば高く、複数に散れば低くなります。その 1 つの数値が、信頼できる 答えとそうでない答えを分けます。

信頼できない答えをどう扱うかは、ラベル次第です。SIC のラベルは階層をなします。業種 グループはより広い部門にまとまります。おかげで 1 回の応答がほぼただで済みます。モデルが グループに自信を持てないときは、それが属する部門を報告します。広いラベルは狭いラベルから 導けるので、2 回目の呼び出しは要りません。

60 件の提出書類を通じて、信頼度のしきい値 0.9 がそれらを半分に分けます。確信のある 半分は 90% の確率で正しく、もう半分は 40% です。1 つ上のレベルで報告すると、その 40% が 70% になります。最後に、書類ごとに 1 回のリクエストで、ラベルとその具体度を返す classify() 関数を用意します。

flowchart LR
    doc["Item 1 'Business'<br/>from one 10-K"]

    subgraph request["one request"]
        q["Choice<br/>75 industry groups"]
    end

    sure{"confidence<br/>&ge; 0.9?"}
    grp["report the industry group<br/><i>e.g. 28</i>"]
    div["report its division<br/><i>e.g. manufacturing</i>"]

    doc --> request --> sure
    %% both branches leave the test, so they share a rank and stack on their own
    sure -- "yes" --> grp
    sure -- "no" --> div

セットアップ

pip install ipython matplotlib 'cooksafe>=0.2.0,<0.3.0'

次に TYPESAFE_API_KEY を設定します。すべての API 呼び出しは json_cache.json に キャッシュされ、このファイルは cookbook に同梱されているので、再レンダリング時は API を 呼ばずに公開済みの数値を再生します。そのファイルを削除すれば、すべてをライブで再実行できます。

以下の数値は 2026-08-12 の jev-1.12 によるものです。

import json
from collections import defaultdict
from pathlib import Path

import matplotlib
import matplotlib.pyplot as plt
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, TypeSafeClient

matplotlib.use("Agg")  # headless render

import os  # noqa: E402

TYPESAFE_MODEL = "jev-1.12"
CONFIDENT = 0.9  # above this the group is reported; below it, the division

client = TypeSafeClient(
    api_key=os.environ.get(
        "TYPESAFE_API_KEY", "cache-only"
    ),  # keyless kernels replay the cache
    base_url=os.environ.get("TYPESAFE_ENDPOINT"),
    timeout=120.0,
)
json_cache = JsonCache(Path("json_cache.json"))

分類体系の 2 つのレベルを構築する

sic_codes.tsv は、提出者が自分のコードを選ぶために SEC が公開している業種リストで、 2026-08-10 に取得したものです。444 個の 4 桁コードがあり、それぞれに業種名が付いています。 桁は階層をなします。最初の 2 桁が大グループ(ここでは 75 個あり、01 農業生産から 99 分類不能まで)で、大グループの固定された範囲が 10 の部門を構成します。これは SIC で最も広い区分です。

どちらのレベルも、モデルを介さずその 1 つのファイルから得られます。コードを最初の 2 桁で グループ化し、その桁を部門に対応付けます。

DIVISIONS = [
    (1, 9, "agriculture, forestry and fishing"),
    (10, 14, "mining"),
    (15, 17, "construction"),
    (20, 39, "manufacturing"),
    (40, 49, "transportation, communications and utilities"),
    (50, 51, "wholesale trade"),
    (52, 59, "retail trade"),
    (60, 67, "finance, insurance and real estate"),
    (70, 89, "services"),
    (91, 99, "public administration"),
]

INDUSTRIES: dict[str, str] = {}
for line in Path("sic_codes.tsv").read_text().splitlines()[1:]:
    code, _office, title = line.split("\t")
    INDUSTRIES[code] = title.lower()

GROUPS: dict[str, list[str]] = defaultdict(list)
for code in sorted(INDUSTRIES):
    GROUPS[code[:2]].append(code)

def division(group: str) -> str:
    number = int(group)
    return next(name for low, high, name in DIVISIONS if low <= number <= high)

print(
    f"{len(INDUSTRIES)} industries -> {len(GROUPS)} major groups -> {len(DIVISIONS)} divisions"
)
print(
    f"  group 35 = {division('35')} / {', '.join(INDUSTRIES[c] for c in GROUPS['35'][:3])} ..."
)
444 industries -> 75 major groups -> 10 divisions
  group 35 = manufacturing / engines & turbines, farm machinery & equipment, lawn & garden tractors & home lawn & gardens equip ...

A Choice question needs something to describe each option, and a group’s own name is not always there: 42 of the 75 carry an umbrella title in the SEC’s list, and the rest carry none. So each group is described by the industries inside it, which is what someone reading the filing would match against anyway.

MAX_NAMED = (
    8  # industries listed per group; enough to characterise it without a wall of text
)

def describe(group: str) -> str:
    umbrella = INDUSTRIES.get(f"{group}00")
    inside = [INDUSTRIES[c] for c in GROUPS[group] if c != f"{group}00"][:MAX_NAMED]
    listed = "; ".join(inside)
    return (
        f"{umbrella} — includes: {listed}"
        if umbrella and listed
        else (umbrella or listed)
    )

print(f"group 20: {describe('20')[:150]}")
print(f"\ngroup 65: {describe('65')[:150]}")
group 20: food and kindred products — includes: meat packing plants; sausages & other prepared meat products; poultry slaughtering and processing; dairy product

group 65: real estate — includes: real estate operators (no developers) & lessors; operators of nonresidential buildings; operators of apartment buildings; less

提出書類

filings.jsonl には 60 件の年次報告書(10-K)が入っており、それぞれが Item 1「Business」 ——企業が事業内容を説明する節で、業種コードが関わる唯一の部分——に絞られています。 1993〜2024 年にわたり、700〜2,200 語です。それぞれに、提出者が選んだ SIC コードと、 EDGAR で照会するための accession number が付いています。

そのラベルがどこから来るかは、どんな精度の数値より先に重要です。これは自己申告です。 提出書類を用意した人が一度選んだもので、企業がコードが指す事業を売却してもコードを そのままにすれば古くなります。この 60 件は、自らの本文が付しているコードを裏付ける 提出書類に絞り込まれているので、ここでの数値は EDGAR のメタデータの状態ではなく、 レシピを測っています。

FILINGS = [json.loads(line) for line in Path("filings.jsonl").read_text().splitlines()]
example = FILINGS[7]
print(
    f"{len(FILINGS)} filings, {sum(f['words'] for f in FILINGS) // len(FILINGS)} words on average"
)
print(f"\n{example['id']} (filed {example['year']}, accession {example['accession']}):")
print(f"  {example['text'][:230]}...")
print(f"  filer's code: {example['sic']} {INDUSTRIES[example['sic']]}")
60 filings, 1438 words on average

1389870_2008 (filed 2008, accession 0001079974-09-000155):
  Item 1. DESCRIPTION OF BUSINESS. NARRATIVE DESCRIPTION OF THE BUSINESS Across America Financial Services, Inc. is a corporation which was formed under the laws of the State of Colorado on December 1, 2005. Until March 23, 2007, we...
  filer's code: 6163 loan brokers

Choice の質問を 1 つ行い、信頼度を読む

選択肢が 75 のグループである Choice 質問を 1 つ使います。分類体系の全体が 1 回の リクエストに収まります。Choice はおよそ 240 選択肢まで確実に機能し、75 は十分に その範囲内です。

答えは、当選したグループである choice、75 個それぞれに載った重みである probabilities、 そしてその散らばりがどれほど集中していたかを示す confidence とともに返ります。この レシピは、当選者の確率そのものではなく confidence を読みます。当選者が 0.45 で次点が 0.44 の場合と、当選者が 0.45 で残りの重みが薄く散っている場合とでは状況が異なり、それを 分けるのが confidence です。

QUESTION = (
    "Which broad industry does this company operate in? Judge the company's own operations "
    "as this filing describes them."
)

def questions() -> dict:
    return {
        "group": Choice(
            instructions=QUESTION,
            criteria={group: describe(group) for group in sorted(GROUPS)},
        )
    }

@json_cache
def ask(filing_id: str, text: str) -> dict:
    response = client.system_one(
        state=text, questions=questions(), model=TYPESAFE_MODEL
    )
    answer = response.answers["group"]
    return {
        "group": answer.choice,
        "confidence": answer.confidence,
        "probabilities": dict(answer.probabilities),
    }

確信があればグループを、なければその部門を返す

下の 4 行がレシピの全体です。信頼度が 0.9 以上なら、答えは業種グループとして報告され、 それ未満なら、同じ答えがそのグループの属する部門として報告されます。

どの提出書類も、依然として使えるラベルとともに返ります。モデルが自信を持って分類でき なかったものは、捨てられたり先送りされたりせず、1 つ上のレベルで返ります。部門が粗すぎて アプリケーションが行動に移せないなら、この分岐で人に渡します。

def classify(filing: dict) -> dict:
    answer = ask(filing["id"], filing["text"])
    sure = answer["confidence"] >= CONFIDENT
    return {
        "level": "group" if sure else "division",
        "label": answer["group"] if sure else division(answer["group"]),
        "confidence": answer["confidence"],
        "group": answer["group"],
    }

def show(filing: dict) -> None:
    result = classify(filing)
    named = describe(result["group"]).split(" — ")[0][:46]
    print(
        f"  {filing['id']:>13}  conf {result['confidence']:.2f}  -> {result['level']:<8} "
        f"{result['label']:<14} (group {result['group']}: {named})"
    )

print("three filings the model was sure about:")
for f in sorted(FILINGS, key=lambda f: -ask(f["id"], f["text"])["confidence"])[:3]:
    show(f)
print("\nthree it was not:")
for f in sorted(FILINGS, key=lambda f: ask(f["id"], f["text"])["confidence"])[:3]:
    show(f)
three filings the model was sure about:
    310158_1996  conf 1.00  -> group    28             (group 28: chemicals & allied products)
     33416_1998  conf 1.00  -> group    63             (group 63: life insurance; accident & health insurance; h)
    352541_1996  conf 1.00  -> group    49             (group 49: electric, gas & sanitary services)

three it was not:
   1372167_2013  conf 0.22  -> division manufacturing  (group 38: search, detection, navagation, guidance, aeron)
   1398633_2009  conf 0.23  -> division wholesale trade (group 50: wholesale-durable goods)
     46653_1999  conf 0.29  -> division services       (group 87: services-engineering, accounting, research, ma)

信頼度は、各提出書類の分類の難しさと符合します。1.00 の 3 件は、製薬会社、生命保険会社、 公益企業です。3 社とも書類上は持株会社ですが、それぞれに書類がはっきり名指しする 支配的な事業が 1 つあります。下位の 3 件は、本文を読めば分かる理由でより難しいものです。 2 件はこれから始めようとする事業を説明している開発段階の企業(Nevaeh は「ソフトウェア 開発者として事業を行う予定」、Barricode は「コンピュータセキュリティソフトウェア業界に 参入するために設立」)で、3 件目は 2 つのセグメントを持ち、提出の数週間前にその一方を 売却していました。この 3 件はグループではなく部門として返ります。

classify() がレシピの全体です。ask() を自分の書類に向け、describe() を自分の 分類体系用に書き換えれば、残りはそのまま使えます。

より広い答えがもたらすもの

60 件すべての提出書類を、各提出者が選んだコードに対して、2 つの方針で採点します。 毎回グループを名指しするか、信頼度が 0.9 を下回るたびに部門を報告するかです。

def correct(filing: dict, result: dict) -> bool:
    gold_group = filing["sic"][:2]
    if result["level"] == "group":
        return result["label"] == gold_group
    return result["label"] == division(gold_group)

results = [(f, classify(f)) for f in FILINGS]
sure = [(f, r) for f, r in results if r["level"] == "group"]
unsure = [(f, r) for f, r in results if r["level"] == "division"]

forced = sum(r["group"] == f["sic"][:2] for f, r in results)
broadened = sum(correct(f, r) for f, r in results)

print(f"forced to name a group every time      {forced}/{len(results)} right")
print(
    f"  of those, the {len(sure)} it was sure about  "
    f"{sum(r['group'] == f['sic'][:2] for f, r in sure)}/{len(sure)} right"
)
print(
    f"  and the {len(unsure)} it was not           "
    f"{sum(r['group'] == f['sic'][:2] for f, r in unsure)}/{len(unsure)} right"
)
print(
    f"\nletting it answer coarsely when unsure  {broadened}/{len(results)} useful answers"
)
forced to name a group every time      39/60 right
  of those, the 30 it was sure about  27/30 right
  and the 30 it was not           12/30 right

letting it answer coarsely when unsure  48/60 useful answers

モデルが確信していたところでは、名指ししたグループは 10 回に 9 回正しいです。確信が なかったところでは、グループを名指しすると 40% で、正しいより誤りのほうが多くなります。 同じ答えを部門として報告すると、70% になります。

グラフは、モデルが確信していたかどうかで分けて、2 つの方針を並べて示します。

labels = ["sure\n(group reported)", "unsure\n(division reported)"]
forced_split = [
    sum(r["group"] == f["sic"][:2] for f, r in sure) / len(sure),
    sum(r["group"] == f["sic"][:2] for f, r in unsure) / len(unsure),
]
broad_split = [
    sum(correct(f, r) for f, r in sure) / len(sure),
    sum(correct(f, r) for f, r in unsure) / len(unsure),
]

fig, ax = plt.subplots(figsize=(7, 3.6))
x = range(len(labels))
ax.bar(
    [i - 0.19 for i in x],
    forced_split,
    0.38,
    label="always name a group",
    color="#c8ccd4",
)
ax.bar(
    [i + 0.19 for i in x],
    broad_split,
    0.38,
    label="answer broadly when unsure",
    color="#3b6ea5",
)
for i, (a, b) in enumerate(zip(forced_split, broad_split)):
    ax.text(i - 0.19, a + 0.02, f"{a:.0%}", ha="center", fontsize=9)
    ax.text(i + 0.19, b + 0.02, f"{b:.0%}", ha="center", fontsize=9)
ax.set_xticks(list(x))
ax.set_xticklabels(
    [f"{lab}\nn={n}" for lab, n in zip(labels, [len(sure), len(unsure)])]
)
ax.set_ylabel("labels that are right")
ax.set_ylim(0, 1.12)
ax.set_title("Where the broader answer helps: the filings it was unsure about")
ax.legend(frameon=False, loc="upper right")
ax.spines[["top", "right"]].set_visible(False)
plt.tight_layout()
display(fig)
output

Playground で開く

この共有リンクには 1 件の提出書類と 75 選択肢の質問が入っているので、コードを一切書かずに 分布とそれが生み出す信頼度を確認できます。

playground_link = make_playground_link(
    example["text"], questions(), models=[TYPESAFE_MODEL]
)
display(
    Markdown(
        f"🔗 [Open the filing + question in the TypeSafe playground]({playground_link})"
    )
)
TypeSafe Playground で提出書類と質問を開く →