구조 복원
형식을 잃은 순수 텍스트에서 두 번의 요청으로 Markdown을 재구성합니다. 한 번은 문장 중간에서 끊긴 줄을 다시 이어 붙이고, 다른 한 번은 모든 블록(제목, 목록, 코드, 콜아웃)을 분류합니다.
이 cookbook은 마크업이 벗겨진 순수 텍스트(문장 중간에서 하드 랩된 줄, 제목 표시 없음, 목록 불릿 없음)를 받아 Markdown 구조로 재구성합니다: 제목, 문단, 목록, 인용, 코드, 콜아웃. 입력은 정확히 그런 상태의 팀 메모입니다.
텍스트 생성 모델이 텍스트를 Markdown으로 다시 쓸 수도 있지만, 다시 쓰기는 단어까지 바꿀 수 있습니다. 여기서 모델은 텍스트를 생성하지 않습니다. 문서에 대한 좁은 질문(이 줄은 문장 중간에서 이어지는가? 이 블록은 어떤 종류의 내용인가?)에만 답하고, 렌더링은 코드가 하므로 출력의 모든 문자는 입력에서 나오고, 모든 판단은 확률을 지닙니다.
전체 파이프라인은 문서당 두 번의 API 요청으로, 순서대로 실행됩니다:
- 1단계, 이어 붙이기: 인접한 줄 쌍마다
Noul질문(예/아니오 질문으로, 답은 ’예’가 맞을 확률)을 하나씩 던져, 이 줄바꿈이 하나의 문장을 두 줄로 갈랐는지 묻습니다. 모든 쌍은 하나의 요청으로 보내지고, 갈라진 문장을 이어 가는 줄은 다시 블록으로 병합됩니다. - 2단계, 분류: 병합된 블록마다
Choice질문(목록에서 옵션 하나를 고르며, 모든 옵션에 확률이 있음)을 하나씩 던져, 제목·문단·목록 항목·인용·코드·콜아웃(본문과 구분되는 note, tip, warning) 중에서 고릅니다. 블록은 1단계가 답한 뒤에야 존재하므로 이것은 두 번째 요청입니다. 이 요청은 또한 모든 블록에 대한 동반 질문(제목 레벨, 단계 순서, 콜아웃 종류)을 함께 보내며, 그 답은 블록의 유형이 관련 있을 때만 읽힙니다. - 직접적인 증거는 코드에 남습니다. 빈 줄과 명시적 표시(
-,1.,#)는 코드에서 읽히며, 결코 모델에게 다시 판단하도록 보내지지 않습니다. 이 메모는 빈 줄은 유지했지만 모든 표시를 잃었습니다. 모델은 코드가 텍스트에서 답할 수 없는 질문만 받습니다.
모든 동작은 2단계 질문 criteria에 명시되어 있습니다: 한 줄 설명으로 된 세 개의 딕셔너리, 그리고 classify_questions 안에 있는 단계 질문의 true/false criteria입니다. 나머지 코드는 이들을 둘러싼 배관입니다. 비용과 지연 시간 수치는 부록에 있습니다: 이 메모의 경우 두 번의 왕복, 10,211 토큰, 0.8초, $0.0015입니다.
환경 준비
pip install ipython 'cooksafe>=0.2.0,<0.3.0'
그다음 TYPESAFE_API_KEY를 설정합니다. 모든 API 호출은 cookbook과 함께 제공되는 json_cache.json에 캐시되므로, 다시 렌더링하면 API를 호출하지 않고 게시된 수치를 재생합니다. 그 파일을 삭제하면 모든 것을 실시간으로 다시 실행할 수 있습니다.
import os
import re
import urllib.request
from pathlib import Path
from time import perf_counter
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, Noul, NoulCriteria, TypeSafeClient
TYPESAFE_MODEL = "jev-1.12"
PRICE = (0.042, 0.00) # $ per 1M tokens (input, output); TypeSafe jev-1.12 as of 2026-09
client = TypeSafeClient(api_key=os.environ["TYPESAFE_API_KEY"], timeout=120.0)
json_cache = JsonCache(Path("json_cache.json"))
문서: 형식을 잃은 팀 메모
테스트 문서는 빌드 시스템 이관에 관한 메모로, 순수 텍스트 받은 편지함에 도착한 그대로의 상태입니다: 문단은 문장 중간에서 하드 랩되어 있고, 셸 명령이 아무 표시 없이 한 줄을 차지하며, 두 개의 목록은 불릿도 번호도 없고, 경고는 그것이 경고임을 알리는 아무 표시도 없습니다. 텍스트는 고정된 gist에서 가져오므로 cookbook의 수치가 재현 가능하게 유지됩니다.
GIST = (
"https://gist.githubusercontent.com/eugene-shvarts/6df7daf97233bf92bcdd6b386a0fa561"
"/raw/5da03690611fb6ddcbaabdb91fb9f91d9751b113/build-memo.txt"
)
@json_cache
def fetch_document(url: str) -> str:
request = urllib.request.Request(url, headers={"User-Agent": "typesafe-cookbook/1.0"})
with urllib.request.urlopen(request) as response:
return response.read().decode()
RAW = fetch_document(GIST)
print(RAW[:560])
Migration to the new build system
Hi everyone, quick heads up about the build system migration that is
happening next week. We have been running the new pipeline in shadow
mode for three weeks and the results look solid, so it is time to
make the switch for real.
What changes for you
The old make targets keep working until the end of the month. The new
entrypoint is a single command that wraps everything, including the
docs build that used to be separate.
bun run build
Generated artifacts no longer need to be committed. The new pipeline
uploads them
줄 분할, 빈 줄 추적, id 태깅은 모두 코드에서 이루어집니다. 모델은 관여하지 않습니다.
각 줄은 짧은 id(L014| )를 받습니다. id는 모델이 상태의 일부로 읽는 평범한 텍스트이며, 질문과 답은 이 id로 줄을 가리킵니다(시맨틱 검색 cookbook과 같은 방식입니다).
def to_lines(text: str) -> list[dict]:
lines, gap = [], False
for raw in text.split("\n"):
stripped = re.sub(r"[\t ]+", " ", raw).strip()
if not stripped:
gap = bool(lines) # a leading blank is not a break
continue
lines.append({"text": stripped, "gap": gap})
gap = False
return lines
def tag(items: list[dict], prefix: str) -> str:
return "\n".join(
f"{chr(10) if item['gap'] else ''}{prefix}{i:03d}| {item['text']}"
for i, item in enumerate(items)
)
def line_id(i: int) -> str:
return f"L{i:03d}"
def block_id(i: int) -> str:
return f"B{i:03d}"
LINES = to_lines(RAW)
print(f"{len(LINES)} non-blank lines. The model sees, e.g.:")
print("\n".join(tag(LINES, "L").splitlines()[19:24]))
28 non-blank lines. The model sees, e.g.:
L013| The cutover touches three teams, so check whether you are on this
L014| list before you plan anything for Monday:
L015| The platform team
L016| The web client team
L017| Whoever still owns the release tooling
1단계: 분리된 문장 이어 붙이기
인접한 줄 쌍마다 Noul 질문 하나, 모두 하나의 요청에 담습니다. 빈 줄로 나뉜 쌍은 건너뜁니다. 질문은 의도적으로 좁게(“이 줄은 문장 중간에서 이어지는가?”) 만들어졌으며, 이는 텍스트에 관한 거의 객관적인 사실에 가깝습니다. 부록에서 이 표현 선택과 병합 임계값을 어떻게 도출했는지 모두 다룹니다.
def join_question(i: int) -> Noul:
return Noul(
instructions=f"Does line {line_id(i)} pick up mid-sentence, continuing a sentence left unfinished at the end of line {line_id(i - 1)}?",
criteria=NoulCriteria(
true="The line starts in the middle of a sentence that began on the previous line - the line break tore the sentence apart",
false="The line begins a new sentence, item, heading, or thought of its own",
),
)
@json_cache
def stitch(wording: str = "mid-sentence") -> dict:
make = join_question if wording == "mid-sentence" else naive_join_question
questions = {line_id(i): make(i) for i in range(1, len(LINES)) if not LINES[i]["gap"]}
started = perf_counter()
response = client.system_one(
state=tag(LINES, "L"), questions=questions, model=TYPESAFE_MODEL
)
return {
"joins": [
response.answers[line_id(i)].noul if line_id(i) in response.answers else 0.0
for i in range(len(LINES))
],
"seconds": round(perf_counter() - started, 2),
"usage": [response.usage.input_tokens, response.usage.output_tokens],
}
result = stitch()
print(f"{sum(1 for l in LINES if not l['gap']) - 1} pair questions, one request, "
f"{result['seconds']}s")
16 pair questions, one request, 0.32s
병합 기준값은 이전 줄이 어떻게 끝나는지에 따라 달라집니다. 매달린 줄(문장 종결 부호가 없는 줄) 뒤에서는 병합 확률이 0.2 이상이면 그 쌍을 병합합니다. 종결 부호(. ! ? : ;) 뒤에서는 기준값이 0.5로 올라갑니다. 부록에서 이 두 숫자 뒤에 있는 확률을 짚어 봅니다.
JOIN_AFTER_DANGLING, JOIN_AFTER_TERMINAL = 0.2, 0.5
def ends_terminal(text: str) -> bool:
return re.search(r'[.!?:;…]["\')\]]*$', text) is not None
def merge(joins: list[float]) -> list[dict]:
blocks = []
for i, line in enumerate(LINES):
bar = (
JOIN_AFTER_TERMINAL
if i and ends_terminal(LINES[i - 1]["text"])
else JOIN_AFTER_DANGLING
)
if blocks and not line["gap"] and joins[i] >= bar:
blocks[-1]["text"] += " " + line["text"]
blocks[-1]["lines"].append(i)
else:
blocks.append({"text": line["text"], "lines": [i], "gap": line["gap"]})
return blocks
blocks = merge(result["joins"])
healed = len(LINES) - len(blocks)
print(f"{len(LINES)} lines -> {len(blocks)} blocks ({healed} line breaks healed)")
for i, block in enumerate(blocks):
n = len(block["lines"])
print(f"{block_id(i)} {n} line{'s' if n > 1 else ' '} {block['text'][:62]}")
28 lines -> 17 blocks (11 line breaks healed)
B000 1 line Migration to the new build system
B001 4 lines Hi everyone, quick heads up about the build system migration t
B002 1 line What changes for you
B003 3 lines The old make targets keep working until the end of the month.
B004 1 line bun run build
B005 3 lines Generated artifacts no longer need to be committed. The new pi
B006 2 lines The cutover touches three teams, so check whether you are on t
B007 1 line The platform team
B008 1 line The web client team
B009 1 line Whoever still owns the release tooling
B010 1 line Things to do before Monday
B011 1 line Update your local toolchain to version 2.4 or later
B012 1 line Delete the old build cache directory
B013 1 line Run the doctor script and fix anything it flags
B014 3 lines If the doctor script reports a red result on the toolchain che
B015 2 lines As Dana put it in the kickoff, "a migration nobody notices is
B016 1 line Thanks, and shout if anything looks off.
2단계: 블록 분류
이어 붙인 각 블록은 Choice 질문을 받습니다: 이것은 어떤 종류의 내용인가? 아래 세 개의 딕셔너리와, classify_questions 안의 단계 질문 true/false criteria가 분류기의 전체 명세입니다. 다른 로직은 없습니다. 파이프라인을 여러분의 문서에 맞추려면 이 설명들을 수정하면 됩니다.
TYPE_CRITERIA = {
"heading": "A short label or title that names the document or the section that follows it - not a full sentence of content",
"paragraph": "Running prose: one or more complete sentences of explanatory or narrative text",
"list_item": "One entry in a list of parallel items - an ingredient, a feature, a task, an attendee; reads as one of several sibling entries",
"quote": "Words attributed to a person or source - quoted speech, a citation, an excerpt someone else wrote",
"code": "Computer code, a shell command, terminal output, or a config snippet meant to be read verbatim",
"callout": "A warning, tip, or important note that interrupts the flow to flag something the reader must not miss",
}
HLEVEL_CRITERIA = {
"title": "The title of the whole document",
"section": "A major section heading within the document",
"subsection": "A minor heading nested under a section",
}
CALLOUT_CRITERIA = {
"note": "Neutral extra information the reader should be aware of",
"tip": "A helpful suggestion or shortcut that makes things easier",
"warning": "A caution about something that can go wrong or cause harm",
}
아래의 모든 것은 배관입니다: 질문을 만들고, 요청 하나를 보내고, 답을 다시 읽습니다. 유형이 heading으로 돌아오면 렌더러는 제목 레벨이 필요하고, list_item이면 순서가 중요한지, callout이면 어느 종류인지가 필요합니다. 유형은 아직 알 수 없고, 그것을 기다리면 세 번째 왕복이 되므로, 동반 질문을 같은 요청에서 미리 묻습니다. 이 답들의 대부분은 결코 읽히지 않습니다: 문단의 단계 확률은 아무 의미가 없어 그냥 무시됩니다. 질문을 하나 더 추가하는 비용은 거의 없는데, 상태가 토큰의 대부분이고 어느 쪽이든 한 번만 전송되기 때문이며, 반면 왕복을 한 번 더 추가하면 요청 하나만큼의 지연 시간이 늘어납니다.
HEADING_MAX_CHARS = 90 # longer blocks can't render as headings, so don't ask
def classify_questions(texts: list[str]) -> dict:
questions = {}
for i, text in enumerate(texts):
bid = block_id(i)
questions[f"type_{bid}"] = Choice(
instructions=f"What kind of content is block {bid}?", criteria=TYPE_CRITERIA
)
if len(text) <= HEADING_MAX_CHARS:
questions[f"hlevel_{bid}"] = Choice(
instructions=f"As a heading, what level would block {bid} occupy in this document's structure?",
criteria=HLEVEL_CRITERIA,
)
questions[f"step_{bid}"] = Noul(
instructions=f"Is block {bid} an instruction in a sequence where the order of the items matters?",
criteria=NoulCriteria(
true="It is one step of a procedure - the items around it must happen in order",
false="Order is irrelevant - it is a loose collection, or not a list item at all",
),
)
questions[f"callout_{bid}"] = Choice(
instructions=f"What kind of aside is block {bid}?", criteria=CALLOUT_CRITERIA
)
return questions
@json_cache
def classify(texts: list[str], gaps: list[bool]) -> dict:
tagged = tag([{"text": t, "gap": g} for t, g in zip(texts, gaps)], "B")
questions = classify_questions(texts)
started = perf_counter()
response = client.system_one(state=tagged, questions=questions, model=TYPESAFE_MODEL)
judgments = []
for i in range(len(texts)):
bid = block_id(i)
type_answer = response.answers[f"type_{bid}"]
hlevel = response.answers.get(f"hlevel_{bid}")
judgments.append(
{
"type": type_answer.choice,
"confidence": type_answer.confidence,
"probabilities": type_answer.probabilities,
"hlevel": hlevel.choice if hlevel else "section",
"step": response.answers[f"step_{bid}"].noul,
"callout": response.answers[f"callout_{bid}"].choice,
}
)
return {
"judgments": judgments,
"n_questions": len(questions),
"seconds": round(perf_counter() - started, 2),
"usage": [response.usage.input_tokens, response.usage.output_tokens],
}
classified = classify([b["text"] for b in blocks], [b["gap"] for b in blocks])
for block, judgment in zip(blocks, classified["judgments"]):
block.update(judgment)
print(f"{classified['n_questions']} questions about {len(blocks)} blocks, one request, "
f"{classified['seconds']}s\n")
print(f"{'block':<6}{'type':<11}{'conf':<6}{'companion used':<18}text")
for i, b in enumerate(blocks):
companion = {
"heading": f"level={b['hlevel']}",
"list_item": f"step={b['step']:.2f}",
"callout": f"kind={b['callout']}",
}.get(b["type"], "-")
print(f"{block_id(i):<6}{b['type']:<11}{b['confidence']:.2f} {companion:<18}"
f"{b['text'][:46]}")
62 questions about 17 blocks, one request, 0.51s
block type conf companion used text
B000 heading 0.99 level=title Migration to the new build system
B001 paragraph 0.98 - Hi everyone, quick heads up about the build sy
B002 heading 0.75 level=section What changes for you
B003 paragraph 0.89 - The old make targets keep working until the en
B004 code 1.00 - bun run build
B005 paragraph 0.90 - Generated artifacts no longer need to be commi
B006 paragraph 0.43 - The cutover touches three teams, so check whet
B007 list_item 0.99 step=0.15 The platform team
B008 list_item 1.00 step=0.16 The web client team
B009 list_item 0.99 step=0.12 Whoever still owns the release tooling
B010 heading 0.96 level=section Things to do before Monday
B011 list_item 0.98 step=0.86 Update your local toolchain to version 2.4 or
B012 list_item 0.99 step=0.87 Delete the old build cache directory
B013 list_item 0.92 step=0.90 Run the doctor script and fix anything it flag
B014 callout 0.65 kind=warning If the doctor script reports a red result on t
B015 quote 0.99 - As Dana put it in the kickoff, "a migration no
B016 paragraph 0.92 - Thanks, and shout if anything looks off.
모든 블록의 판단이 그 표에 있고, companion 열은 미리 물어둔 답이 어떻게 쓰이는지 보여 줍니다: “Things to do before Monday” 세 줄은 단계 확률이 0.9에 가깝고(번호가 매겨진 목록으로 렌더링됩니다), 팀 세 줄은 0.1 근처에 있으며(불릿), doctor 스크립트에 관한 표시 없는 경고는 종류가 warning인 콜아웃으로 분류되었습니다. 부록에서는 모델이 확신하지 못한 그 블록 하나를 살펴봅니다.
렌더링
코드가 판단들로부터 페이지를 조립합니다. 연속된 목록 항목은 하나의 목록이 되며, 항목들의 단계 확률 평균이 0.5 이상이면 번호가 매겨집니다. 그 임계값은 어떤 단일 질문도 직접 묻지 않은 그룹 수준의 결정입니다.
STEP_THRESHOLD = 0.5
HEADING_MARK = {"title": "#", "section": "##", "subsection": "###"}
CALLOUT_MARK = {"note": "NOTE", "tip": "TIP", "warning": "WARNING"}
def to_markdown(blocks: list[dict]) -> str:
groups = []
for b in blocks:
if b["type"] in ("list_item", "code") and groups and groups[-1][0] == b["type"]:
groups[-1][1].append(b)
else:
groups.append((b["type"], [b]))
parts = []
for kind, items in groups:
if kind == "list_item":
ordered = sum(b["step"] for b in items) / len(items) >= STEP_THRESHOLD
parts.append("\n".join(
f"{n + 1}. {b['text']}" if ordered else f"- {b['text']}"
for n, b in enumerate(items)
))
elif kind == "code":
parts.append("```\n" + "\n".join(b["text"] for b in items) + "\n```")
elif kind == "heading":
parts.append(f"{HEADING_MARK[items[0]['hlevel']]} {items[0]['text']}")
elif kind == "quote":
parts.append(f"> {items[0]['text']}")
elif kind == "callout":
parts.append(f"> [!{CALLOUT_MARK[items[0]['callout']]}]\n> {items[0]['text']}")
else:
parts.append(items[0]["text"])
return "\n\n".join(parts) + "\n"
markdown = to_markdown(blocks)
print(markdown)
# Migration to the new build system
Hi everyone, quick heads up about the build system migration that is happening next week. We have been running the new pipeline in shadow mode for three weeks and the results look solid, so it is time to make the switch for real.
## What changes for you
The old make targets keep working until the end of the month. The new entrypoint is a single command that wraps everything, including the docs build that used to be separate.
```
bun run build
```
Generated artifacts no longer need to be committed. The new pipeline uploads them to the registry automatically, and checking them in just creates merge conflicts.
The cutover touches three teams, so check whether you are on this list before you plan anything for Monday:
- The platform team
- The web client team
- Whoever still owns the release tooling
## Things to do before Monday
1. Update your local toolchain to version 2.4 or later
2. Delete the old build cache directory
3. Run the doctor script and fix anything it flags
> [!WARNING]
> If the doctor script reports a red result on the toolchain check, do not proceed with the migration. Ping the infra channel first and we will sort it out together.
> As Dana put it in the kickoff, "a migration nobody notices is the only kind worth shipping."
Thanks, and shout if anything looks off.
위의 모든 단어는 입력에서 온 것입니다. 파이프라인은 경계, 유형, 마크업만 선택했습니다.
Playground에서 열기
이 공유 링크에는 이어 붙인 블록과 전체 2단계 질문 세트가 담겨 있습니다. 열면 분류를 실시간으로 다시 실행할 수 있습니다.
playground_link = make_playground_link(
tag(blocks, "B"),
classify_questions([b["text"] for b in blocks]),
models=[TYPESAFE_MODEL],
)
display(Markdown(f"🔗 [Open the stitched memo + questions in the TypeSafe playground]({playground_link})"))
TypeSafe playground에서 이어 붙인 메모와 질문을 엽니다 →
부록
비용과 지연 시간
tokens = [result["usage"], classified["usage"]]
total_in, total_out = sum(t[0] for t in tokens), sum(t[1] for t in tokens)
cost = total_in / 1e6 * PRICE[0] + total_out / 1e6 * PRICE[1]
n_joins = sum(1 for l in LINES if not l["gap"]) - 1
print(f"pass 1 {n_joins} questions {result['seconds']}s")
print(f"pass 2 {classified['n_questions']} questions {classified['seconds']}s")
print(f"total {total_in + total_out:,} tokens "
f"{result['seconds'] + classified['seconds']:.1f}s ${cost:.4f}")
pass 1 16 questions 0.32s
pass 2 62 questions 0.51s
total 10,211 tokens 0.8s $0.0003
두 번의 왕복, 10,211 토큰, 0.8초, $0.0015.
병합 임계값의 출처
1단계에서 얻은 줄별 병합 확률:
print("join line")
for i, line in enumerate(LINES[:18]):
join = " " if i == 0 or line["gap"] else f"{result['joins'][i]:.2f}"
print(f"{join} {line_id(i)}| {line['text'][:66]}")
join line
L000| Migration to the new build system
L001| Hi everyone, quick heads up about the build system migration that
0.77 L002| happening next week. We have been running the new pipeline in shad
0.62 L003| mode for three weeks and the results look solid, so it is time to
0.39 L004| make the switch for real.
L005| What changes for you
L006| The old make targets keep working until the end of the month. The
0.42 L007| entrypoint is a single command that wraps everything, including th
0.59 L008| docs build that used to be separate.
L009| bun run build
L010| Generated artifacts no longer need to be committed. The new pipeli
0.48 L011| uploads them to the registry automatically, and checking them in
0.40 L012| just creates merge conflicts.
L013| The cutover touches three teams, so check whether you are on this
0.50 L014| list before you plan anything for Monday:
0.22 L015| The platform team
0.11 L016| The web client team
0.12 L017| Whoever still owns the release tooling
확률은 두 개의 뚜렷한 띠에 자리 잡습니다: 문장을 가른 줄바꿈은 0.39 이상을 얻고, 작성자가 의도한 줄바꿈은 거의 0에 가깝습니다. 하지만 두 띠 사이 어디에 기준값을 둘지는 이전 줄이 어떻게 끝나는지에 달려 있으며, 이는 코드가 직접 읽을 수 있는 사실입니다:
- 매달린 줄(문장 종결 부호가 없는 줄) 뒤에서는 0.2 이상이면 무엇이든 이어짐으로 간주합니다. 여기서 진짜 이어짐은 0.39까지 낮게 나오므로(
L004| make the switch for real.), 신중하게 0.5로 단일 기준값을 두면 멀쩡한 문단을 갈라 놓게 됩니다. - 종결 부호(문장이나 절을 끝내는 문자:
.!?:;) 뒤에서는 기준값이 0.5로 올라갑니다. 메모의 팀 목록이 그 이유를 보여 줍니다:L015| The platform team은 콜론 뒤에 오는데 0.22를 얻습니다. 이것은 낮지만 0이 아닌 “이 문장이 계속된다”는 신호이며, 0.2 기준값을 넘어 목록을 그것을 이끄는 문장에 병합해 버립니다. 두 경우 모두에 통하는 단일 임계값은 없습니다. 코드가 부호를 먼저 확인하면 두 띠가 분리됩니다.
질문이 “문장 중간”이 아니라 “같은 문단”인 이유
이 파이프라인의 첫 버전은 뻔한 질문을 던졌습니다: “이 두 줄은 같은 문단에 속하는가?” 그것은 특정한 방식으로 실패했습니다. 제목 아래에 이어지는 짧은 줄들(불릿 없이 입력한 목록)은 넓은 의미에서 문단입니다: 줄들이 함께 붙어 있고 주제를 공유합니다. 문단에 대해 물으면 모델은 모든 쌍에 ’예’라고 답하고, 이어 붙이기 단계는 목록 전체를 하나의 긴 블록으로 병합합니다.
같은 문서, 같은 요청 형태, 오직 표현만 바뀌었습니다:
def naive_join_question(i: int) -> Noul:
return Noul(
instructions=f"Are lines {line_id(i - 1)} and {line_id(i)} part of the same paragraph?",
criteria=NoulCriteria(
true="The two lines belong to the same paragraph of running text",
false="The two lines belong to different paragraphs or different pieces of content",
),
)
naive = stitch("same-paragraph")
print(f"{'':14}{'mid-sentence':>13}{'same paragraph':>16}")
for i in (15, 16, 17, 20, 21):
print(f"{line_id(i)}{'':2}{LINES[i]['text'][:36]:<38}"
f"{result['joins'][i]:>7.2f}{naive['joins'][i]:>13.2f}")
print(f"\nblocks after merge: {len(blocks)} (mid-sentence) vs "
f"{len(merge(naive['joins']))} (same paragraph)")
mid-sentence same paragraph
L015 The platform team 0.22 0.77
L016 The web client team 0.11 0.81
L017 Whoever still owns the release tooli 0.12 0.78
L020 Delete the old build cache directory 0.08 0.88
L021 Run the doctor script and fix anythi 0.05 0.91
blocks after merge: 17 (mid-sentence) vs 12 (same paragraph)
문단이라는 표현을 쓰면 표시 없는 모든 목록 항목이 0.75를 넘고 두 목록이 모두 무너집니다. 메모는 몇 개의 질질 끄는 블록으로 병합됩니다. “같은 문단”은 모델에게 주제가 이어지는지를 판단하라고 요구하는데, 목록 항목들 사이에서는 주제가 이어집니다. “문장 중간에서 이어짐”은 텍스트 자체에 대해 묻습니다. 어떤 판단이 임계값에 입력될 때, 질문은 그것을 결정하는 가장 좁은 사실을 지목해야 합니다. 여기서는 표현이 17개 블록과 12개 블록의 차이입니다.
신뢰도가 가장 낮은 블록
uncertain = min(blocks, key=lambda b: b["confidence"])
print(f'"{uncertain["text"]}"')
print(f"confidence {uncertain['confidence']:.2f}: ", end="")
print(", ".join(f"{k} {v:.2f}" for k, v in
sorted(uncertain["probabilities"].items(), key=lambda kv: -kv[1])[:3]))
"The cutover touches three teams, so check whether you are on this list before you plan anything for Monday:"
confidence 0.43: paragraph 0.53, list_item 0.24, callout 0.19
팀 목록을 이끄는 그 문장은 정말로 모호합니다. 뒤에 오는 내용을 지목하고(제목처럼), 완전한 문장이며(문단처럼), 콜아웃이 놓일 자리에 있습니다. 확률도 그에 따라 퍼집니다(paragraph 0.53, list_item 0.24, callout 0.19). UI는 이를 드러낼 수 있습니다. 예를 들어 유형 신뢰도(승리한 선택 뒤의 확률)가 0.55 미만인 블록에는 검토용 밑줄을 긋는 것입니다.