文件導航

用置信度做分類

每份 SEC 年報用一個 Choice 歸入 75 個行業組之一,再讀答案自身的置信度,決定是報這個組還是它上面更寬的部門。

每一家向 SEC 提交年報的公司,都會在年報裡描述自己的業務。我們把這些描述歸入標準產業分類(Standard Industrial Classification):75 個行業組,每個檔案問一個 Choice 問題。

大多數檔案的分類很輕鬆。一家地區銀行就是一家地區銀行。有些則不然:比如一家剛賣掉自己兩個業務板塊之一的公司,或者一家描述的是一門打算進入、而非正在經營的業務的初創公司。不論如何,模型都得挑出一個組,而難題的答案和簡單題的答案看起來沒什麼兩樣。要在難題和簡單題之間分個高下,通常就得付出成本:再上一個模型、額外的呼叫、人工複核。

一個 Choice 本身就能告訴你。除了勝出的那個選項,它還會返回 confidence:當幾乎全部機率都落在一個選項上時它高,機率分散在幾個選項上時它低。就這一個數字,把你能信的答案和不能信的答案分開了。

拿一個不可信的答案怎麼辦,取決於你的標籤體系。SIC 的標籤構成一個層級:行業組向上歸併成更寬的部門。這讓同一次響應幾乎免費。當模型對是哪個組沒把握時,就報它所屬的部門。寬標籤可以從窄標籤推出來,所以不用再調一次。

在 60 份檔案上,用 0.9 的置信度切一刀,剛好把它們分成兩半。有把握的那一半,90% 是對的;另一半隻有 40%。把它們往上報一個層級,那 40% 就變成了 70%。最後我們得到一個 classify() 函式,每個檔案一次請求,返回一個標籤以及它有多具體。

flowchart LR
    doc["Item 1 'Business'<br/>from one 10-K"]

    subgraph request["one request"]
        q["Choice<br/>75 industry groups"]
    end

    sure{"confidence<br/>&ge; 0.9?"}
    grp["report the industry group<br/><i>e.g. 28</i>"]
    div["report its division<br/><i>e.g. manufacturing</i>"]

    doc --> request --> sure
    %% both branches leave the test, so they share a rank and stack on their own
    sure -- "yes" --> grp
    sure -- "no" --> div

準備工作

pip install ipython matplotlib 'cooksafe>=0.2.0,<0.3.0'

然後設定 TYPESAFE_API_KEY。每次 API 呼叫都快取到 json_cache.json,該檔案隨 cookbook 一起提供,所以重新渲染會回放已釋出的數字,不會呼叫 API。刪掉該檔案就能全部真實重跑。

下面的數字來自 2026-08-12 的 jev-1.12。

import json
from collections import defaultdict
from pathlib import Path

import matplotlib
import matplotlib.pyplot as plt
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, TypeSafeClient

matplotlib.use("Agg")  # headless render

import os  # noqa: E402

TYPESAFE_MODEL = "jev-1.12"
CONFIDENT = 0.9  # above this the group is reported; below it, the division

client = TypeSafeClient(
    api_key=os.environ.get(
        "TYPESAFE_API_KEY", "cache-only"
    ),  # keyless kernels replay the cache
    base_url=os.environ.get("TYPESAFE_ENDPOINT"),
    timeout=120.0,
)
json_cache = JsonCache(Path("json_cache.json"))

搭出分類法的兩個層級

sic_codes.tsv 是 SEC 釋出、供申報人挑選自用程式碼的行業清單,抓取於 2026-08-10:444 個四位程式碼,每個都配一個行業名稱。這些數字本身就是一個層級。前兩位是大類(這裡是 75 個,從 01 農業生產到 99 無法分類),而大類的固定區間構成了 10 個部門,這是 SIC 最粗的一層劃分。

這兩個層級都從這一個檔案裡算出來,不涉及任何模型:按前兩位給程式碼分組,再把這些兩位數字對映到一個部門。

DIVISIONS = [
    (1, 9, "agriculture, forestry and fishing"),
    (10, 14, "mining"),
    (15, 17, "construction"),
    (20, 39, "manufacturing"),
    (40, 49, "transportation, communications and utilities"),
    (50, 51, "wholesale trade"),
    (52, 59, "retail trade"),
    (60, 67, "finance, insurance and real estate"),
    (70, 89, "services"),
    (91, 99, "public administration"),
]

INDUSTRIES: dict[str, str] = {}
for line in Path("sic_codes.tsv").read_text().splitlines()[1:]:
    code, _office, title = line.split("\t")
    INDUSTRIES[code] = title.lower()

GROUPS: dict[str, list[str]] = defaultdict(list)
for code in sorted(INDUSTRIES):
    GROUPS[code[:2]].append(code)

def division(group: str) -> str:
    number = int(group)
    return next(name for low, high, name in DIVISIONS if low <= number <= high)

print(
    f"{len(INDUSTRIES)} industries -> {len(GROUPS)} major groups -> {len(DIVISIONS)} divisions"
)
print(
    f"  group 35 = {division('35')} / {', '.join(INDUSTRIES[c] for c in GROUPS['35'][:3])} ..."
)
444 industries -> 75 major groups -> 10 divisions
  group 35 = manufacturing / engines & turbines, farm machinery & equipment, lawn & garden tractors & home lawn & gardens equip ...

Choice 問題需要給每個選項配上描述,而組本身的名字並不總是存在:75 個組裡有 42 個在 SEC 的清單裡帶一個總括性的名稱,其餘什麼都沒有。所以每個組都用它包含的行業來描述 —— 讀檔案的人本來也是拿這些來對照的。

MAX_NAMED = (
    8  # industries listed per group; enough to characterise it without a wall of text
)

def describe(group: str) -> str:
    umbrella = INDUSTRIES.get(f"{group}00")
    inside = [INDUSTRIES[c] for c in GROUPS[group] if c != f"{group}00"][:MAX_NAMED]
    listed = "; ".join(inside)
    return (
        f"{umbrella} — includes: {listed}"
        if umbrella and listed
        else (umbrella or listed)
    )

print(f"group 20: {describe('20')[:150]}")
print(f"\ngroup 65: {describe('65')[:150]}")
group 20: food and kindred products — includes: meat packing plants; sausages & other prepared meat products; poultry slaughtering and processing; dairy product

group 65: real estate — includes: real estate operators (no developers) & lessors; operators of nonresidential buildings; operators of apartment buildings; less

這些檔案

filings.jsonl 裝著 60 份年報(10-K),每份都裁剪到 Item 1 “Business” —— 公司描述自己做什麼的那一節,也是行業程式碼唯一涉及的部分。它們跨越 1993–2024 年,篇幅從 700 到 2200 詞不等。每份都帶著申報人自己選的 SIC 程式碼,以及用來在 EDGAR 上查到它的 accession number。

在看任何準確率數字之前,得先說清楚這個標籤是怎麼來的。它是自報的:準備檔案的人當初挑一次就定下了,而一旦公司賣掉了程式碼所指的業務卻還留著這個程式碼,它就過時了。這 60 份被篩選到「自身文本支撐其所帶程式碼」的檔案,所以這裡的數字衡量的是這套做法,而不是 EDGAR 後設資料的現狀。

FILINGS = [json.loads(line) for line in Path("filings.jsonl").read_text().splitlines()]
example = FILINGS[7]
print(
    f"{len(FILINGS)} filings, {sum(f['words'] for f in FILINGS) // len(FILINGS)} words on average"
)
print(f"\n{example['id']} (filed {example['year']}, accession {example['accession']}):")
print(f"  {example['text'][:230]}...")
print(f"  filer's code: {example['sic']} {INDUSTRIES[example['sic']]}")
60 filings, 1438 words on average

1389870_2008 (filed 2008, accession 0001079974-09-000155):
  Item 1. DESCRIPTION OF BUSINESS. NARRATIVE DESCRIPTION OF THE BUSINESS Across America Financial Services, Inc. is a corporation which was formed under the laws of the State of Colorado on December 1, 2005. Until March 23, 2007, we...
  filer's code: 6163 loan brokers

問一個 Choice 問題,讀它的置信度

一個 Choice 問題,選項就是那 75 個組。整棵分類法裝得下一次請求:Choice 在約 240 個選項以內都能穩定工作,75 完全在範圍內。

答案裡帶著 choice,也就是勝出的組;probabilities,75 個選項各自的權重;以及 confidence,它說明這份分佈有多集中。這套做法讀的是 confidence,而不是勝出者自己的機率。一個 0.45 的勝出者配上 0.44 的第二名,和一個 0.45 的勝出者配上稀稀拉拉散開的其餘權重,是兩種不同的情形,而 confidence 正是把它們區分開的東西。

QUESTION = (
    "Which broad industry does this company operate in? Judge the company's own operations "
    "as this filing describes them."
)

def questions() -> dict:
    return {
        "group": Choice(
            instructions=QUESTION,
            criteria={group: describe(group) for group in sorted(GROUPS)},
        )
    }

@json_cache
def ask(filing_id: str, text: str) -> dict:
    response = client.system_one(
        state=text, questions=questions(), model=TYPESAFE_MODEL
    )
    answer = response.answers["group"]
    return {
        "group": answer.choice,
        "confidence": answer.confidence,
        "probabilities": dict(answer.probabilities),
    }

有把握時報組,沒把握時報部門

下面這四行就是整套做法。置信度達到 0.9 及以上,答案就作為行業組報出;低於這個值,同一個答案就作為該組所屬的部門報出。

每份檔案依然會返回一個可用的標籤。模型沒能有把握地分類的那一份,不會是被丟掉或轉走,而是往上報一個層級。如果某個部門粗到你的應用沒法據此行動,這個分支就是你把它交給人的地方。

def classify(filing: dict) -> dict:
    answer = ask(filing["id"], filing["text"])
    sure = answer["confidence"] >= CONFIDENT
    return {
        "level": "group" if sure else "division",
        "label": answer["group"] if sure else division(answer["group"]),
        "confidence": answer["confidence"],
        "group": answer["group"],
    }

def show(filing: dict) -> None:
    result = classify(filing)
    named = describe(result["group"]).split(" — ")[0][:46]
    print(
        f"  {filing['id']:>13}  conf {result['confidence']:.2f}  -> {result['level']:<8} "
        f"{result['label']:<14} (group {result['group']}: {named})"
    )

print("three filings the model was sure about:")
for f in sorted(FILINGS, key=lambda f: -ask(f["id"], f["text"])["confidence"])[:3]:
    show(f)
print("\nthree it was not:")
for f in sorted(FILINGS, key=lambda f: ask(f["id"], f["text"])["confidence"])[:3]:
    show(f)
three filings the model was sure about:
    310158_1996  conf 1.00  -> group    28             (group 28: chemicals & allied products)
     33416_1998  conf 1.00  -> group    63             (group 63: life insurance; accident & health insurance; h)
    352541_1996  conf 1.00  -> group    49             (group 49: electric, gas & sanitary services)

three it was not:
   1372167_2013  conf 0.22  -> division manufacturing  (group 38: search, detection, navagation, guidance, aeron)
   1398633_2009  conf 0.23  -> division wholesale trade (group 50: wholesale-durable goods)
     46653_1999  conf 0.29  -> division services       (group 87: services-engineering, accounting, research, ma)

這些置信度和每份檔案分類起來有多難是吻合的。三個 1.00 的分別是製藥商、壽險公司和公用事業公司;三家在紙面上都是控股公司,但每家都有一個主導業務,檔案裡明明白白地點了出來。墊底的三個難在哪,從文本里能讀出來。兩家是處於開發階段的公司,描述的是它們打算開辦的業務(Nevaeh “intends to operate as a software developer”,Barricode “organized to enter into the computer security software industry”),第三家則有兩個業務板塊,在申報前幾週賣掉了其中一個。這三份最後以部門而非組的形式返回。

classify() 就是整套做法。把 ask() 指向你自己的文件,再為你的分類法重寫 describe(),其餘部分照搬即可。

報得更寬能換來什麼

全部 60 份檔案,都以每個申報人自己選的程式碼為基準打分,兩種策略各跑一遍:每次都報一個組,或者只要置信度落到 0.9 以下就報部門。

def correct(filing: dict, result: dict) -> bool:
    gold_group = filing["sic"][:2]
    if result["level"] == "group":
        return result["label"] == gold_group
    return result["label"] == division(gold_group)

results = [(f, classify(f)) for f in FILINGS]
sure = [(f, r) for f, r in results if r["level"] == "group"]
unsure = [(f, r) for f, r in results if r["level"] == "division"]

forced = sum(r["group"] == f["sic"][:2] for f, r in results)
broadened = sum(correct(f, r) for f, r in results)

print(f"forced to name a group every time      {forced}/{len(results)} right")
print(
    f"  of those, the {len(sure)} it was sure about  "
    f"{sum(r['group'] == f['sic'][:2] for f, r in sure)}/{len(sure)} right"
)
print(
    f"  and the {len(unsure)} it was not           "
    f"{sum(r['group'] == f['sic'][:2] for f, r in unsure)}/{len(unsure)} right"
)
print(
    f"\nletting it answer coarsely when unsure  {broadened}/{len(results)} useful answers"
)
forced to name a group every time      39/60 right
  of those, the 30 it was sure about  27/30 right
  and the 30 it was not           12/30 right

letting it answer coarsely when unsure  48/60 useful answers

模型有把握的地方,它報出的組十次裡有九次是對的。沒把握的地方,報組是錯多於對,只有 40%。把這些同樣的答案改報成部門,就升到了 70%。

下圖把兩種策略並排放在一起,並按模型是否有把握做了拆分。

labels = ["sure\n(group reported)", "unsure\n(division reported)"]
forced_split = [
    sum(r["group"] == f["sic"][:2] for f, r in sure) / len(sure),
    sum(r["group"] == f["sic"][:2] for f, r in unsure) / len(unsure),
]
broad_split = [
    sum(correct(f, r) for f, r in sure) / len(sure),
    sum(correct(f, r) for f, r in unsure) / len(unsure),
]

fig, ax = plt.subplots(figsize=(7, 3.6))
x = range(len(labels))
ax.bar(
    [i - 0.19 for i in x],
    forced_split,
    0.38,
    label="always name a group",
    color="#c8ccd4",
)
ax.bar(
    [i + 0.19 for i in x],
    broad_split,
    0.38,
    label="answer broadly when unsure",
    color="#3b6ea5",
)
for i, (a, b) in enumerate(zip(forced_split, broad_split)):
    ax.text(i - 0.19, a + 0.02, f"{a:.0%}", ha="center", fontsize=9)
    ax.text(i + 0.19, b + 0.02, f"{b:.0%}", ha="center", fontsize=9)
ax.set_xticks(list(x))
ax.set_xticklabels(
    [f"{lab}\nn={n}" for lab, n in zip(labels, [len(sure), len(unsure)])]
)
ax.set_ylabel("labels that are right")
ax.set_ylim(0, 1.12)
ax.set_title("Where the broader answer helps: the filings it was unsure about")
ax.legend(frameon=False, loc="upper right")
ax.spines[["top", "right"]].set_visible(False)
plt.tight_layout()
display(fig)
output

在 Playground 裡開啟

這個分享連結裡裝著一份檔案和那個 75 選項的問題,不用寫任何程式碼,就能看到它給出的分佈和置信度。

playground_link = make_playground_link(
    example["text"], questions(), models=[TYPESAFE_MODEL]
)
display(
    Markdown(
        f"🔗 [Open the filing + question in the TypeSafe playground]({playground_link})"
    )
)
在 TypeSafe Playground 裡開啟這份檔案 + 問題 →