문서

신뢰도를 사용한 분류

SEC 연차 보고서를 75개 산업 그룹으로 각각 하나의 Choice로 분류한 다음, 답변 자체의 신뢰도를 읽어 그 그룹을 보고할지 그 위의 더 넓은 부문을 보고할지 결정합니다.

SEC에 연차 보고서를 제출하는 모든 기업은 그 안에서 자기 사업을 설명합니다. 우리는 그 설명을 표준산업분류(Standard Industrial Classification)에 따라 분류합니다: 75개 산업 그룹, 문서마다 Choice 질문 하나입니다.

대부분의 공시는 쉽습니다. 지역 은행은 지역 은행입니다. 일부는 그렇지 않습니다: 두 부문 중 하나를 막 매각한 기업, 또는 운영하는 사업이 아니라 진출할 계획인 사업을 설명하는 스타트업이 그렇습니다. 모델은 어찌 됐든 그룹 하나를 골라야 하고, 어려운 사례의 답은 쉬운 사례의 답과 다르게 보이지 않습니다. 어려운 사례를 쉬운 사례와 구분하는 것이 보통 비용이 드는 곳입니다: 두 번째 모델, 추가 호출, 사람 검토입니다.

Choice는 이미 알려줍니다. 이긴 선택지와 함께 confidence를 반환하는데, 거의 모든 확률이 한 선택지에 쏠렸으면 높고 여러 곳에 퍼졌으면 낮습니다. 이 숫자 하나가 믿을 수 있는 답과 그렇지 않은 답을 구분합니다.

믿을 수 없는 답을 어떻게 할지는 레이블에 달려 있습니다. SIC 레이블은 계층을 이룹니다: 산업 그룹이 더 넓은 부문으로 묶입니다. 그래서 한 번의 응답이 거의 무료가 됩니다. 모델이 그룹을 확신하지 못하면, 그 그룹이 속한 부문을 보고합니다. 넓은 레이블이 좁은 레이블에서 따라 나오므로 두 번째 호출이 없습니다.

60개 공시에서 신뢰도 컷오프 0.9가 이들을 절반으로 나눕니다. 확신하는 절반은 90% 맞습니다. 나머지 절반은 40%입니다. 한 단계 위로 보고하면 그 40%가 70%가 됩니다. 마지막으로 문서마다 요청 하나로 레이블과 그 레이블이 얼마나 구체적인지를 반환하는 classify() 함수로 끝맺습니다.

flowchart LR
    doc["Item 1 'Business'<br/>from one 10-K"]

    subgraph request["one request"]
        q["Choice<br/>75 industry groups"]
    end

    sure{"confidence<br/>&ge; 0.9?"}
    grp["report the industry group<br/><i>e.g. 28</i>"]
    div["report its division<br/><i>e.g. manufacturing</i>"]

    doc --> request --> sure
    %% both branches leave the test, so they share a rank and stack on their own
    sure -- "yes" --> grp
    sure -- "no" --> div

설정

pip install ipython matplotlib 'cooksafe>=0.2.0,<0.3.0'

그런 다음 TYPESAFE_API_KEY를 설정합니다. 모든 API 호출은 쿡북과 함께 제공되는 json_cache.json에 캐시되므로, 다시 렌더링하면 API를 호출하지 않고 공개된 숫자를 재생합니다. 그 파일을 삭제하면 모든 것을 실제로 다시 실행합니다.

아래 숫자는 2026-08-12에 jev-1.12에서 나온 것입니다.

import json
from collections import defaultdict
from pathlib import Path

import matplotlib
import matplotlib.pyplot as plt
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, TypeSafeClient

matplotlib.use("Agg")  # headless render

import os  # noqa: E402

TYPESAFE_MODEL = "jev-1.12"
CONFIDENT = 0.9  # above this the group is reported; below it, the division

client = TypeSafeClient(
    api_key=os.environ.get(
        "TYPESAFE_API_KEY", "cache-only"
    ),  # keyless kernels replay the cache
    base_url=os.environ.get("TYPESAFE_ENDPOINT"),
    timeout=120.0,
)
json_cache = JsonCache(Path("json_cache.json"))

분류 체계의 두 계층 만들기

sic_codes.tsv는 SEC가 제출자가 자기 코드를 고르도록 공개한 산업 목록으로, 2026-08-10에 가져왔습니다: 산업 제목이 붙은 네 자리 코드 444개입니다. 자릿수는 계층입니다. 처음 두 자리는 주요 그룹(여기서는 75개, 01 농업 생산부터 99 분류 불가까지)이고, 주요 그룹의 고정된 범위가 SIC에서 가장 넓은 구분인 열 개 부문을 이룹니다.

두 계층 모두 모델 없이 그 파일 하나에서 나옵니다: 코드를 처음 두 자리로 묶고, 그 자릿수를 부문에 매핑합니다.

DIVISIONS = [
    (1, 9, "agriculture, forestry and fishing"),
    (10, 14, "mining"),
    (15, 17, "construction"),
    (20, 39, "manufacturing"),
    (40, 49, "transportation, communications and utilities"),
    (50, 51, "wholesale trade"),
    (52, 59, "retail trade"),
    (60, 67, "finance, insurance and real estate"),
    (70, 89, "services"),
    (91, 99, "public administration"),
]

INDUSTRIES: dict[str, str] = {}
for line in Path("sic_codes.tsv").read_text().splitlines()[1:]:
    code, _office, title = line.split("\t")
    INDUSTRIES[code] = title.lower()

GROUPS: dict[str, list[str]] = defaultdict(list)
for code in sorted(INDUSTRIES):
    GROUPS[code[:2]].append(code)

def division(group: str) -> str:
    number = int(group)
    return next(name for low, high, name in DIVISIONS if low <= number <= high)

print(
    f"{len(INDUSTRIES)} industries -> {len(GROUPS)} major groups -> {len(DIVISIONS)} divisions"
)
print(
    f"  group 35 = {division('35')} / {', '.join(INDUSTRIES[c] for c in GROUPS['35'][:3])} ..."
)
444 industries -> 75 major groups -> 10 divisions
  group 35 = manufacturing / engines & turbines, farm machinery & equipment, lawn & garden tractors & home lawn & gardens equip ...

Choice 질문은 각 선택지를 설명할 무언가가 필요하고, 그룹 자체의 이름이 항상 있지는 않습니다: 75개 중 42개는 SEC 목록에서 포괄 제목을 가지며 나머지는 없습니다. 그래서 각 그룹은 그 안의 산업들로 설명되며, 이는 공시를 읽는 사람이 어차피 대조할 대상입니다.

MAX_NAMED = (
    8  # industries listed per group; enough to characterise it without a wall of text
)

def describe(group: str) -> str:
    umbrella = INDUSTRIES.get(f"{group}00")
    inside = [INDUSTRIES[c] for c in GROUPS[group] if c != f"{group}00"][:MAX_NAMED]
    listed = "; ".join(inside)
    return (
        f"{umbrella} — includes: {listed}"
        if umbrella and listed
        else (umbrella or listed)
    )

print(f"group 20: {describe('20')[:150]}")
print(f"\ngroup 65: {describe('65')[:150]}")
group 20: food and kindred products — includes: meat packing plants; sausages & other prepared meat products; poultry slaughtering and processing; dairy product

group 65: real estate — includes: real estate operators (no developers) & lessors; operators of nonresidential buildings; operators of apartment buildings; less

공시

filings.jsonl에는 60개 연차 보고서(10-K)가 들어 있고, 각각은 기업이 하는 일을 설명하는 항목 1 “Business”로 잘려 있습니다. 산업 코드가 다루는 부분은 바로 그 부분뿐입니다. 1993–2024년에 걸쳐 있으며 700에서 2,200 단어입니다. 각각은 제출자가 고른 SIC 코드와 EDGAR에서 찾아볼 접수 번호를 담고 있습니다.

그 레이블이 어디서 오는지는 어떤 정확도 수치보다 먼저 중요합니다. 그것은 자기 신고입니다: 공시를 준비한 사람이 한 번 골랐고, 기업이 그 코드가 가리키는 사업을 매각하고도 코드를 유지하면 낡아집니다. 이 60개는 자체 텍스트가 담고 있는 코드를 뒷받침하는 공시로 걸러졌으므로, 여기 숫자는 EDGAR 메타데이터의 상태가 아니라 그 레시피를 측정합니다.

FILINGS = [json.loads(line) for line in Path("filings.jsonl").read_text().splitlines()]
example = FILINGS[7]
print(
    f"{len(FILINGS)} filings, {sum(f['words'] for f in FILINGS) // len(FILINGS)} words on average"
)
print(f"\n{example['id']} (filed {example['year']}, accession {example['accession']}):")
print(f"  {example['text'][:230]}...")
print(f"  filer's code: {example['sic']} {INDUSTRIES[example['sic']]}")
60 filings, 1438 words on average

1389870_2008 (filed 2008, accession 0001079974-09-000155):
  Item 1. DESCRIPTION OF BUSINESS. NARRATIVE DESCRIPTION OF THE BUSINESS Across America Financial Services, Inc. is a corporation which was formed under the laws of the State of Colorado on December 1, 2005. Until March 23, 2007, we...
  filer's code: 6163 loan brokers

Choice 질문 하나를 던지고 신뢰도를 읽기

선택지가 75개 그룹인 Choice 질문 하나입니다. 분류 체계 전체가 요청 하나에 들어갑니다: Choice는 대략 240개 선택지까지 안정적으로 동작하며, 75는 그 안에 넉넉히 들어갑니다.

답은 choice(이긴 그룹), probabilities(75개 각각의 가중치), 그리고 그 분포가 얼마나 집중됐는지 말해주는 confidence와 함께 돌아옵니다. 이 레시피는 이긴 선택지 자체의 확률이 아니라 confidence를 읽습니다. 1위가 0.45이고 2위가 0.44인 경우와, 1위가 0.45인데 나머지 가중치가 얇게 흩어진 경우는 다른 상황이며, confidence가 이 둘을 구분합니다.

QUESTION = (
    "Which broad industry does this company operate in? Judge the company's own operations "
    "as this filing describes them."
)

def questions() -> dict:
    return {
        "group": Choice(
            instructions=QUESTION,
            criteria={group: describe(group) for group in sorted(GROUPS)},
        )
    }

@json_cache
def ask(filing_id: str, text: str) -> dict:
    response = client.system_one(
        state=text, questions=questions(), model=TYPESAFE_MODEL
    )
    answer = response.answers["group"]
    return {
        "group": answer.choice,
        "confidence": answer.confidence,
        "probabilities": dict(answer.probabilities),
    }

확신하면 그룹을, 아니면 그 부문을 반환

아래 네 줄이 레시피 전체입니다. 신뢰도가 0.9 이상이면 답을 산업 그룹으로 보고합니다. 그 아래면 같은 답을 그 그룹이 속한 부문으로 보고합니다.

모든 공시는 여전히 사용 가능한 레이블과 함께 돌아옵니다. 모델이 자신 있게 분류하지 못한 것은 버려지거나 넘겨지는 대신 한 단계 위로 돌아옵니다. 부문이 너무 거칠어 애플리케이션이 행동하기 어렵다면, 이 분기가 그것을 사람에게 넘기는 지점입니다.

def classify(filing: dict) -> dict:
    answer = ask(filing["id"], filing["text"])
    sure = answer["confidence"] >= CONFIDENT
    return {
        "level": "group" if sure else "division",
        "label": answer["group"] if sure else division(answer["group"]),
        "confidence": answer["confidence"],
        "group": answer["group"],
    }

def show(filing: dict) -> None:
    result = classify(filing)
    named = describe(result["group"]).split(" — ")[0][:46]
    print(
        f"  {filing['id']:>13}  conf {result['confidence']:.2f}  -> {result['level']:<8} "
        f"{result['label']:<14} (group {result['group']}: {named})"
    )

print("three filings the model was sure about:")
for f in sorted(FILINGS, key=lambda f: -ask(f["id"], f["text"])["confidence"])[:3]:
    show(f)
print("\nthree it was not:")
for f in sorted(FILINGS, key=lambda f: ask(f["id"], f["text"])["confidence"])[:3]:
    show(f)
three filings the model was sure about:
    310158_1996  conf 1.00  -> group    28             (group 28: chemicals & allied products)
     33416_1998  conf 1.00  -> group    63             (group 63: life insurance; accident & health insurance; h)
    352541_1996  conf 1.00  -> group    49             (group 49: electric, gas & sanitary services)

three it was not:
   1372167_2013  conf 0.22  -> division manufacturing  (group 38: search, detection, navagation, guidance, aeron)
   1398633_2009  conf 0.23  -> division wholesale trade (group 50: wholesale-durable goods)
     46653_1999  conf 0.29  -> division services       (group 87: services-engineering, accounting, research, ma)

신뢰도는 각 공시가 분류하기 얼마나 어려운지와 맞아떨어집니다. 1.00인 세 개는 제약 회사, 생명 보험사, 그리고 유틸리티입니다. 셋 다 서류상으로는 지주회사지만, 각각 공시가 대놓고 밝히는 하나의 지배적 사업을 가지고 있습니다. 맨 아래 세 개는 텍스트에서 읽어낼 수 있는 이유로 더 어렵습니다. 둘은 시작하려는 사업을 설명하는 개발 단계 기업이고(Nevaeh는 “intends to operate as a software developer”, Barricode는 “organized to enter into the computer security software industry”), 세 번째는 두 부문을 가지고 있었는데 공시 몇 주 전에 하나를 매각했습니다. 그 셋은 그룹이 아니라 부문으로 돌아옵니다.

classify()가 레시피 전체입니다. ask()를 자신의 문서로 향하게 하고 describe()를 자신의 분류 체계에 맞게 다시 쓰면 나머지는 그대로 적용됩니다.

더 넓은 답이 사 주는 것

60개 공시 전부를 각 제출자가 고른 코드에 대조해 두 정책 아래에서 채점합니다: 매번 그룹을 명명하거나, 신뢰도가 0.9 아래로 떨어질 때마다 부문을 보고합니다.

def correct(filing: dict, result: dict) -> bool:
    gold_group = filing["sic"][:2]
    if result["level"] == "group":
        return result["label"] == gold_group
    return result["label"] == division(gold_group)

results = [(f, classify(f)) for f in FILINGS]
sure = [(f, r) for f, r in results if r["level"] == "group"]
unsure = [(f, r) for f, r in results if r["level"] == "division"]

forced = sum(r["group"] == f["sic"][:2] for f, r in results)
broadened = sum(correct(f, r) for f, r in results)

print(f"forced to name a group every time      {forced}/{len(results)} right")
print(
    f"  of those, the {len(sure)} it was sure about  "
    f"{sum(r['group'] == f['sic'][:2] for f, r in sure)}/{len(sure)} right"
)
print(
    f"  and the {len(unsure)} it was not           "
    f"{sum(r['group'] == f['sic'][:2] for f, r in unsure)}/{len(unsure)} right"
)
print(
    f"\nletting it answer coarsely when unsure  {broadened}/{len(results)} useful answers"
)
forced to name a group every time      39/60 right
  of those, the 30 it was sure about  27/30 right
  and the 30 it was not           12/30 right

letting it answer coarsely when unsure  48/60 useful answers

모델이 확신한 곳에서는, 명명한 그룹이 열 번 중 아홉 번 맞습니다. 확신하지 못한 곳에서는 그룹을 명명하는 것이 맞는 경우보다 틀린 경우가 더 많았고, 40%였습니다. 같은 답을 부문으로 보고하면 70%가 됩니다.

이 차트는 두 정책을 나란히 놓고, 모델이 확신했는지에 따라 나눕니다.

labels = ["sure\n(group reported)", "unsure\n(division reported)"]
forced_split = [
    sum(r["group"] == f["sic"][:2] for f, r in sure) / len(sure),
    sum(r["group"] == f["sic"][:2] for f, r in unsure) / len(unsure),
]
broad_split = [
    sum(correct(f, r) for f, r in sure) / len(sure),
    sum(correct(f, r) for f, r in unsure) / len(unsure),
]

fig, ax = plt.subplots(figsize=(7, 3.6))
x = range(len(labels))
ax.bar(
    [i - 0.19 for i in x],
    forced_split,
    0.38,
    label="always name a group",
    color="#c8ccd4",
)
ax.bar(
    [i + 0.19 for i in x],
    broad_split,
    0.38,
    label="answer broadly when unsure",
    color="#3b6ea5",
)
for i, (a, b) in enumerate(zip(forced_split, broad_split)):
    ax.text(i - 0.19, a + 0.02, f"{a:.0%}", ha="center", fontsize=9)
    ax.text(i + 0.19, b + 0.02, f"{b:.0%}", ha="center", fontsize=9)
ax.set_xticks(list(x))
ax.set_xticklabels(
    [f"{lab}\nn={n}" for lab, n in zip(labels, [len(sure), len(unsure)])]
)
ax.set_ylabel("labels that are right")
ax.set_ylim(0, 1.12)
ax.set_title("Where the broader answer helps: the filings it was unsure about")
ax.legend(frameon=False, loc="upper right")
ax.spines[["top", "right"]].set_visible(False)
plt.tight_layout()
display(fig)
출력

playground에서 열기

이 공유 링크는 공시 하나와 75개 선택지 질문을 담고 있어, 코드를 한 줄도 쓰지 않고 분포와 그것이 만들어내는 신뢰도를 볼 수 있습니다.

playground_link = make_playground_link(
    example["text"], questions(), models=[TYPESAFE_MODEL]
)
display(
    Markdown(
        f"🔗 [Open the filing + question in the TypeSafe playground]({playground_link})"
    )
)
TypeSafe playground에서 공시 + 질문 열기 →