ドキュメント

構造の復元

構造の復元

書式を失ったプレーンテキストから 2 回のリクエストで Markdown を再構築します。1 つは折り返された行をつなぎ、1 つは各ブロックを分類します。

この cookbook は、マークアップが取り除かれたプレーンテキスト(文の途中で折り返された 行、見出しマーカーなし、リストの箇条書き記号なし)を受け取り、その構造を Markdown として 再構築します。見出し、段落、リスト、引用、コード、コールアウトです。入力は、まさにその 状態のチーム向けメモです。

テキスト生成モデルならテキストを Markdown に書き換えられますが、書き換えは語句まで 変えてしまうことがあります。ここではモデルはテキストを一切生成しません。ドキュメントに ついての狭い質問(この行は文の途中から続いているか?このブロックはどんな種類の 内容か?)に答えるだけで、描画はコードが行います。そのため出力のすべての文字が入力に 由来し、すべての判断が確率を持ちます。

パイプライン全体は、ドキュメントごとに 2 回の API リクエストを順に実行するだけです。

  • パス 1、つなぎ合わせ: 隣り合う行の各ペアに Noul の質問(「はい」が正しい確率を 答えとする真偽の質問)を 1 つずつ、行の折り返しが文を 2 行にまたがって分割したかを 尋ねます。すべてのペアは 1 回のリクエストに入り、分割された文を継続する行は ブロックに再結合されます。
  • パス 2、分類: 結合後の各ブロックに Choice の質問(リストから選択肢を 1 つ選び、 各選択肢に確率が付く)を 1 つずつ、見出し・段落・リスト項目・引用・コード・コールアウト (本文から切り離された note・tip・warning)の中から選びます。ブロックはパス 1 が 答えて初めて存在するので、これは 2 回目のリクエストになります。同時に各ブロックの 付随質問(見出しレベル、ステップの順序、コールアウトの種類)も運び、その答えは ブロックの type がそれを関連付けるときだけ読み取ります。
  • 直接的な証拠はコードに残す。 空行と明示的なマーカー(- 、1.、#)は コードで読み取り、モデルに見直しを求めません。このメモは空行を保っていましたが、 マーカーはすべて失っていました。モデルには、コードがテキストから答えられない質問 だけを渡します。

すべての挙動はパス 2 の質問の criteria で規定されています。1 行ずつの説明からなる 3 つの dict と、classify_questions 内にあるステップ質問の true/false の criteria です。残りの コードはそれらを囲む配管です。コストとレイテンシの数値は付録にあります。2 回の 往復、10,211 トークン、0.8 秒、このメモで $0.0015 です。

準備

pip install ipython 'cooksafe>=0.2.0,<0.3.0'

次に TYPESAFE_API_KEY を設定します。すべての API 呼び出しは cookbook に同梱される json_cache.json にキャッシュされるので、再描画すると API を呼ばずに公開済みの数値を 再生します。すべてを実際に再実行するには、このファイルを削除します。

import os
import re
import urllib.request
from pathlib import Path
from time import perf_counter

from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, Noul, NoulCriteria, TypeSafeClient

TYPESAFE_MODEL = "jev-1.12"
PRICE = (0.042, 0.00)  # $ per 1M tokens (input, output); TypeSafe jev-1.12 as of 2026-09
client = TypeSafeClient(api_key=os.environ["TYPESAFE_API_KEY"], timeout=120.0)
json_cache = JsonCache(Path("json_cache.json"))

ドキュメント:書式を失ったチーム向けメモ

テスト用ドキュメントは、ビルドシステム移行についてのメモで、プレーンテキストの受信 トレイに届くままの状態です。段落は文の途中で折り返され、シェルコマンドは裸の行に 置かれ、2 つのリストには箇条書き記号も番号もなく、警告にはそれを警告と示すものが 何もありません。cookbook の数値が再現可能であるように、テキストは固定した gist から 取得します。

GIST = (
    "https://gist.githubusercontent.com/eugene-shvarts/6df7daf97233bf92bcdd6b386a0fa561"
    "/raw/5da03690611fb6ddcbaabdb91fb9f91d9751b113/build-memo.txt"
)

@json_cache
def fetch_document(url: str) -> str:
    request = urllib.request.Request(url, headers={"User-Agent": "typesafe-cookbook/1.0"})
    with urllib.request.urlopen(request) as response:
        return response.read().decode()

RAW = fetch_document(GIST)
print(RAW[:560])
Migration to the new build system

Hi everyone, quick heads up about the build system migration that is
happening next week. We have been running the new pipeline in shadow
mode for three weeks and the results look solid, so it is time to
make the switch for real.

What changes for you

The old make targets keep working until the end of the month. The new
entrypoint is a single command that wraps everything, including the
docs build that used to be separate.

bun run build

Generated artifacts no longer need to be committed. The new pipeline
uploads them

行の分割、空行の追跡、id の付与はすべてコードで行い、モデルは関与しません。 各行には短い id(L014| )が付きます。id はモデルが state の一部として読む普通の テキストで、質問と答えはこの id で行を参照します(セマンティック検索の cookbook と同じ方式です)。

def to_lines(text: str) -> list[dict]:
    lines, gap = [], False
    for raw in text.split("\n"):
        stripped = re.sub(r"[\t ]+", " ", raw).strip()
        if not stripped:
            gap = bool(lines)  # a leading blank is not a break
            continue
        lines.append({"text": stripped, "gap": gap})
        gap = False
    return lines

def tag(items: list[dict], prefix: str) -> str:
    return "\n".join(
        f"{chr(10) if item['gap'] else ''}{prefix}{i:03d}| {item['text']}"
        for i, item in enumerate(items)
    )

def line_id(i: int) -> str:
    return f"L{i:03d}"

def block_id(i: int) -> str:
    return f"B{i:03d}"

LINES = to_lines(RAW)
print(f"{len(LINES)} non-blank lines. The model sees, e.g.:")
print("\n".join(tag(LINES, "L").splitlines()[19:24]))
28 non-blank lines. The model sees, e.g.:
L013| The cutover touches three teams, so check whether you are on this
L014| list before you plan anything for Monday:
L015| The platform team
L016| The web client team
L017| Whoever still owns the release tooling

パス 1:分割された文のつなぎ合わせ

隣り合う行の各ペアに Noul の質問を 1 つずつ、すべて 1 回のリクエストにまとめます。 空行で区切られたペアはスキップします。質問は意図的に狭く(「この行は文の途中から 続いているか?」)、テキストについての客観的な事実に近づけています。文言の選び方と、 結合のしきい値をどう導いたかは付録で扱います。

def join_question(i: int) -> Noul:
    return Noul(
        instructions=f"Does line {line_id(i)} pick up mid-sentence, continuing a sentence left unfinished at the end of line {line_id(i - 1)}?",
        criteria=NoulCriteria(
            true="The line starts in the middle of a sentence that began on the previous line - the line break tore the sentence apart",
            false="The line begins a new sentence, item, heading, or thought of its own",
        ),
    )

@json_cache
def stitch(wording: str = "mid-sentence") -> dict:
    make = join_question if wording == "mid-sentence" else naive_join_question
    questions = {line_id(i): make(i) for i in range(1, len(LINES)) if not LINES[i]["gap"]}
    started = perf_counter()
    response = client.system_one(
        state=tag(LINES, "L"), questions=questions, model=TYPESAFE_MODEL
    )
    return {
        "joins": [
            response.answers[line_id(i)].noul if line_id(i) in response.answers else 0.0
            for i in range(len(LINES))
        ],
        "seconds": round(perf_counter() - started, 2),
        "usage": [response.usage.input_tokens, response.usage.output_tokens],
    }

result = stitch()
print(f"{sum(1 for l in LINES if not l['gap']) - 1} pair questions, one request, "
      f"{result['seconds']}s")
16 pair questions, one request, 0.32s

結合するかどうかのカットオフは、前の行がどう終わるかによります。宙ぶらりんな行 (文末の句読点がない行)の後では、結合確率が 0.2 以上ならペアを結合します。終端の 句読点(. ! ? : ;)の後では、カットオフは 0.5 に上がります。この 2 つの 数値の背後にある確率は付録でたどります。

JOIN_AFTER_DANGLING, JOIN_AFTER_TERMINAL = 0.2, 0.5

def ends_terminal(text: str) -> bool:
    return re.search(r'[.!?:;…]["\')\]]*$', text) is not None

def merge(joins: list[float]) -> list[dict]:
    blocks = []
    for i, line in enumerate(LINES):
        bar = (
            JOIN_AFTER_TERMINAL
            if i and ends_terminal(LINES[i - 1]["text"])
            else JOIN_AFTER_DANGLING
        )
        if blocks and not line["gap"] and joins[i] >= bar:
            blocks[-1]["text"] += " " + line["text"]
            blocks[-1]["lines"].append(i)
        else:
            blocks.append({"text": line["text"], "lines": [i], "gap": line["gap"]})
    return blocks

blocks = merge(result["joins"])
healed = len(LINES) - len(blocks)
print(f"{len(LINES)} lines -> {len(blocks)} blocks ({healed} line breaks healed)")
for i, block in enumerate(blocks):
    n = len(block["lines"])
    print(f"{block_id(i)}  {n} line{'s' if n > 1 else ' '}  {block['text'][:62]}")
28 lines -> 17 blocks (11 line breaks healed)
B000  1 line   Migration to the new build system
B001  4 lines  Hi everyone, quick heads up about the build system migration t
B002  1 line   What changes for you
B003  3 lines  The old make targets keep working until the end of the month.
B004  1 line   bun run build
B005  3 lines  Generated artifacts no longer need to be committed. The new pi
B006  2 lines  The cutover touches three teams, so check whether you are on t
B007  1 line   The platform team
B008  1 line   The web client team
B009  1 line   Whoever still owns the release tooling
B010  1 line   Things to do before Monday
B011  1 line   Update your local toolchain to version 2.4 or later
B012  1 line   Delete the old build cache directory
B013  1 line   Run the doctor script and fix anything it flags
B014  3 lines  If the doctor script reports a red result on the toolchain che
B015  2 lines  As Dana put it in the kickoff, "a migration nobody notices is
B016  1 line   Thanks, and shout if anything looks off.

パス 2:ブロックの分類

結合後の各ブロックには Choice の質問を 1 つ与えます。これはどんな種類の内容か? 以下の 3 つの dict と、下の classify_questions 内にあるステップ質問の true/false の criteria が、分類器の仕様のすべてです。他のロジックはありません。パイプラインを自分の ドキュメントに合わせるには、これらの説明を編集します。

TYPE_CRITERIA = {
    "heading": "A short label or title that names the document or the section that follows it - not a full sentence of content",
    "paragraph": "Running prose: one or more complete sentences of explanatory or narrative text",
    "list_item": "One entry in a list of parallel items - an ingredient, a feature, a task, an attendee; reads as one of several sibling entries",
    "quote": "Words attributed to a person or source - quoted speech, a citation, an excerpt someone else wrote",
    "code": "Computer code, a shell command, terminal output, or a config snippet meant to be read verbatim",
    "callout": "A warning, tip, or important note that interrupts the flow to flag something the reader must not miss",
}
HLEVEL_CRITERIA = {
    "title": "The title of the whole document",
    "section": "A major section heading within the document",
    "subsection": "A minor heading nested under a section",
}
CALLOUT_CRITERIA = {
    "note": "Neutral extra information the reader should be aware of",
    "tip": "A helpful suggestion or shortcut that makes things easier",
    "warning": "A caution about something that can go wrong or cause harm",
}

以下はすべて配管です。質問を組み立て、1 回のリクエストを送り、答えを読み戻します。 type が heading で返ってくれば、描画側は見出しレベルを必要とします。list_item なら 順序が重要かどうか、callout ならどの種類かです。type はまだ分かっておらず、それを 待つと 3 回目の往復になるので、付随質問は同じリクエストで先回りして尋ねます。これらの 答えのほとんどは決して読まれません。段落のステップ確率は何の意味も持たず、単に無視 されます。余分な質問はほとんどコストを増やしません。state がトークンの大半で、どちらに しても一度だけ送られるからです。一方、往復が 1 回増えるとリクエスト 1 回分のレイテンシが 加わります。

HEADING_MAX_CHARS = 90  # longer blocks can't render as headings, so don't ask

def classify_questions(texts: list[str]) -> dict:
    questions = {}
    for i, text in enumerate(texts):
        bid = block_id(i)
        questions[f"type_{bid}"] = Choice(
            instructions=f"What kind of content is block {bid}?", criteria=TYPE_CRITERIA
        )
        if len(text) <= HEADING_MAX_CHARS:
            questions[f"hlevel_{bid}"] = Choice(
                instructions=f"As a heading, what level would block {bid} occupy in this document's structure?",
                criteria=HLEVEL_CRITERIA,
            )
        questions[f"step_{bid}"] = Noul(
            instructions=f"Is block {bid} an instruction in a sequence where the order of the items matters?",
            criteria=NoulCriteria(
                true="It is one step of a procedure - the items around it must happen in order",
                false="Order is irrelevant - it is a loose collection, or not a list item at all",
            ),
        )
        questions[f"callout_{bid}"] = Choice(
            instructions=f"What kind of aside is block {bid}?", criteria=CALLOUT_CRITERIA
        )
    return questions

@json_cache
def classify(texts: list[str], gaps: list[bool]) -> dict:
    tagged = tag([{"text": t, "gap": g} for t, g in zip(texts, gaps)], "B")
    questions = classify_questions(texts)
    started = perf_counter()
    response = client.system_one(state=tagged, questions=questions, model=TYPESAFE_MODEL)
    judgments = []
    for i in range(len(texts)):
        bid = block_id(i)
        type_answer = response.answers[f"type_{bid}"]
        hlevel = response.answers.get(f"hlevel_{bid}")
        judgments.append(
            {
                "type": type_answer.choice,
                "confidence": type_answer.confidence,
                "probabilities": type_answer.probabilities,
                "hlevel": hlevel.choice if hlevel else "section",
                "step": response.answers[f"step_{bid}"].noul,
                "callout": response.answers[f"callout_{bid}"].choice,
            }
        )
    return {
        "judgments": judgments,
        "n_questions": len(questions),
        "seconds": round(perf_counter() - started, 2),
        "usage": [response.usage.input_tokens, response.usage.output_tokens],
    }

classified = classify([b["text"] for b in blocks], [b["gap"] for b in blocks])
for block, judgment in zip(blocks, classified["judgments"]):
    block.update(judgment)
print(f"{classified['n_questions']} questions about {len(blocks)} blocks, one request, "
      f"{classified['seconds']}s\n")
print(f"{'block':<6}{'type':<11}{'conf':<6}{'companion used':<18}text")
for i, b in enumerate(blocks):
    companion = {
        "heading": f"level={b['hlevel']}",
        "list_item": f"step={b['step']:.2f}",
        "callout": f"kind={b['callout']}",
    }.get(b["type"], "-")
    print(f"{block_id(i):<6}{b['type']:<11}{b['confidence']:.2f}  {companion:<18}"
          f"{b['text'][:46]}")
62 questions about 17 blocks, one request, 0.51s

block type       conf  companion used    text
B000  heading    0.99  level=title       Migration to the new build system
B001  paragraph  0.98  -                 Hi everyone, quick heads up about the build sy
B002  heading    0.75  level=section     What changes for you
B003  paragraph  0.89  -                 The old make targets keep working until the en
B004  code       1.00  -                 bun run build
B005  paragraph  0.90  -                 Generated artifacts no longer need to be commi
B006  paragraph  0.43  -                 The cutover touches three teams, so check whet
B007  list_item  0.99  step=0.15         The platform team
B008  list_item  1.00  step=0.16         The web client team
B009  list_item  0.99  step=0.12         Whoever still owns the release tooling
B010  heading    0.96  level=section     Things to do before Monday
B011  list_item  0.98  step=0.86         Update your local toolchain to version 2.4 or
B012  list_item  0.99  step=0.87         Delete the old build cache directory
B013  list_item  0.92  step=0.90         Run the doctor script and fix anything it flag
B014  callout    0.65  kind=warning      If the doctor script reports a red result on t
B015  quote      0.99  -                 As Dana put it in the kickoff, "a migration no
B016  paragraph  0.92  -                 Thanks, and shout if anything looks off.

すべてのブロックの判定はこの表にあり、companion 列は先回りした答えが活用されている 様子を示します。「Things to do before Monday」の 3 行はステップ確率が 0.9 付近で (番号付きリストとして描画されます)、3 つのチームの行は 0.1 付近で(箇条書き)、 医者スクリプトに関するマークのない警告は kind が warning のコールアウトに分類され ました。付録では、モデルが確信を持てなかった 1 つのブロックを見ます。

描画

コードが判定からページを組み立てます。連続するリスト項目は 1 つのリストになり、項目の ステップ確率の平均が 0.5 以上なら番号付きになります。このしきい値は、どの単一の 質問も直接尋ねていないグループ単位の判断です。

STEP_THRESHOLD = 0.5
HEADING_MARK = {"title": "#", "section": "##", "subsection": "###"}
CALLOUT_MARK = {"note": "NOTE", "tip": "TIP", "warning": "WARNING"}

def to_markdown(blocks: list[dict]) -> str:
    groups = []
    for b in blocks:
        if b["type"] in ("list_item", "code") and groups and groups[-1][0] == b["type"]:
            groups[-1][1].append(b)
        else:
            groups.append((b["type"], [b]))
    parts = []
    for kind, items in groups:
        if kind == "list_item":
            ordered = sum(b["step"] for b in items) / len(items) >= STEP_THRESHOLD
            parts.append("\n".join(
                f"{n + 1}. {b['text']}" if ordered else f"- {b['text']}"
                for n, b in enumerate(items)
            ))
        elif kind == "code":
            parts.append("```\n" + "\n".join(b["text"] for b in items) + "\n```")
        elif kind == "heading":
            parts.append(f"{HEADING_MARK[items[0]['hlevel']]} {items[0]['text']}")
        elif kind == "quote":
            parts.append(f"> {items[0]['text']}")
        elif kind == "callout":
            parts.append(f"> [!{CALLOUT_MARK[items[0]['callout']]}]\n> {items[0]['text']}")
        else:
            parts.append(items[0]["text"])
    return "\n\n".join(parts) + "\n"

markdown = to_markdown(blocks)
print(markdown)
# Migration to the new build system

Hi everyone, quick heads up about the build system migration that is happening next week. We have been running the new pipeline in shadow mode for three weeks and the results look solid, so it is time to make the switch for real.

## What changes for you

The old make targets keep working until the end of the month. The new entrypoint is a single command that wraps everything, including the docs build that used to be separate.

```
bun run build
```

Generated artifacts no longer need to be committed. The new pipeline uploads them to the registry automatically, and checking them in just creates merge conflicts.

The cutover touches three teams, so check whether you are on this list before you plan anything for Monday:

- The platform team
- The web client team
- Whoever still owns the release tooling

## Things to do before Monday

1. Update your local toolchain to version 2.4 or later
2. Delete the old build cache directory
3. Run the doctor script and fix anything it flags

> [!WARNING]
> If the doctor script reports a red result on the toolchain check, do not proceed with the migration. Ping the infra channel first and we will sort it out together.

> As Dana put it in the kickoff, "a migration nobody notices is the only kind worth shipping."

Thanks, and shout if anything looks off.

上のすべての語は入力に由来します。パイプラインが選んだのは境界、type、マークアップ だけです。

playground で開く

この共有リンクには、結合後のブロックとパス 2 の質問一式が入っています。開くと分類を 実際に再実行できます。

playground_link = make_playground_link(
    tag(blocks, "B"),
    classify_questions([b["text"] for b in blocks]),
    models=[TYPESAFE_MODEL],
)
display(Markdown(f"🔗 [Open the stitched memo + questions in the TypeSafe playground]({playground_link})"))
TypeSafe playground で結合済みのメモと質問を開く →

付録

コストとレイテンシ

tokens = [result["usage"], classified["usage"]]
total_in, total_out = sum(t[0] for t in tokens), sum(t[1] for t in tokens)
cost = total_in / 1e6 * PRICE[0] + total_out / 1e6 * PRICE[1]
n_joins = sum(1 for l in LINES if not l["gap"]) - 1
print(f"pass 1  {n_joins} questions  {result['seconds']}s")
print(f"pass 2  {classified['n_questions']} questions  {classified['seconds']}s")
print(f"total   {total_in + total_out:,} tokens  "
      f"{result['seconds'] + classified['seconds']:.1f}s  ${cost:.4f}")
pass 1  16 questions  0.32s
pass 2  62 questions  0.51s
total   10,211 tokens  0.8s  $0.0003

2 回の往復、10,211 トークン、0.8 秒、$0.0015 です。

結合しきい値の出どころ

パス 1 の行ごとの結合確率です。

print("join  line")
for i, line in enumerate(LINES[:18]):
    join = "    " if i == 0 or line["gap"] else f"{result['joins'][i]:.2f}"
    print(f"{join}  {line_id(i)}| {line['text'][:66]}")
join  line
      L000| Migration to the new build system
      L001| Hi everyone, quick heads up about the build system migration that
0.77  L002| happening next week. We have been running the new pipeline in shad
0.62  L003| mode for three weeks and the results look solid, so it is time to
0.39  L004| make the switch for real.
      L005| What changes for you
      L006| The old make targets keep working until the end of the month. The
0.42  L007| entrypoint is a single command that wraps everything, including th
0.59  L008| docs build that used to be separate.
      L009| bun run build
      L010| Generated artifacts no longer need to be committed. The new pipeli
0.48  L011| uploads them to the registry automatically, and checking them in
0.40  L012| just creates merge conflicts.
      L013| The cutover touches three teams, so check whether you are on this
0.50  L014| list before you plan anything for Monday:
0.22  L015| The platform team
0.11  L016| The web client team
0.12  L017| Whoever still owns the release tooling

確率は 2 つの帯に分かれます。文を分割した行の折り返しは 0.39 以上、作者が意図した 折り返しはほぼゼロです。しかし、帯の間にカットオフをどこに置くかは、前の行がどう 終わるかによって変わります。これはコードが直接読み取れる事実です。

  • 宙ぶらりん な行(文末の句読点がない行)の後では、0.2 以上なら継続とみなします。 本当の継続はここでは 0.39 という低いスコアです(L004| make the switch for real.)。そのため 0.5 という単一の慎重なカットオフでは、健全な 段落まで分割してしまいます。
  • 終端 の句読点(文や節を終わらせる文字。. ! ? : ;)の後では、カットオフは 0.5 に上がります。メモのチームリストがその理由を 示します。L015| The platform team はコロンの後に続き、0.22 を付けます。これは低いものの非ゼロの「この文は 続いている」というシグナルで、0.2 のカットオフを超え、リストを導入する 文 に結合してしまいます。どちらのケースにも効く単一のしきい値はありません。コードが 最初に句読点を確認すれば、2 つの帯は分かれます。

なぜ質問は「same paragraph」ではなく「mid-sentence」なのか

このパイプラインの最初のバージョンは、誰もが思いつく質問をしていました。「この 2 行は 同じ段落の一部か?」です。これは特定の仕方で失敗しました。見出しの下に並ぶ短い行の 連なり(箇条書き記号なしで入力されたリスト)は、広い意味では段落 です。行がまとまって いてトピックを共有しているからです。段落について尋ねると、モデルはどのペアにも「はい」と 答え、つなぎ合わせのパスがリスト全体を 1 つの長いブロックに結合してしまいます。

同じドキュメント、同じリクエストの形で、文言だけを変えた場合:

def naive_join_question(i: int) -> Noul:
    return Noul(
        instructions=f"Are lines {line_id(i - 1)} and {line_id(i)} part of the same paragraph?",
        criteria=NoulCriteria(
            true="The two lines belong to the same paragraph of running text",
            false="The two lines belong to different paragraphs or different pieces of content",
        ),
    )

naive = stitch("same-paragraph")
print(f"{'':14}{'mid-sentence':>13}{'same paragraph':>16}")
for i in (15, 16, 17, 20, 21):
    print(f"{line_id(i)}{'':2}{LINES[i]['text'][:36]:<38}"
          f"{result['joins'][i]:>7.2f}{naive['joins'][i]:>13.2f}")
print(f"\nblocks after merge: {len(blocks)} (mid-sentence) vs "
      f"{len(merge(naive['joins']))} (same paragraph)")
               mid-sentence  same paragraph
L015  The platform team                        0.22         0.77
L016  The web client team                      0.11         0.81
L017  Whoever still owns the release tooli     0.12         0.78
L020  Delete the old build cache directory     0.08         0.88
L021  Run the doctor script and fix anythi     0.05         0.91

blocks after merge: 17 (mid-sentence) vs 12 (same paragraph)

段落という文言では、マークのないリスト項目がすべて 0.75 を超え、両方のリストが つぶれてしまいます。メモはいくつかの長ったらしいブロックに結合されます。「Same paragraph」 はトピックが引き継がれるかどうかをモデルに判断させますが、リスト項目の間では 引き継がれてしまいます。「Picks up mid-sentence」はテキストそのものについて尋ねます。 判断がしきい値に供給されるとき、質問はそれを決める最も狭い事実を名指しすべきです。 ここでは、文言が 17 ブロックと 12 ブロックの違いになっています。

最も信頼度の低いブロック

uncertain = min(blocks, key=lambda b: b["confidence"])
print(f'"{uncertain["text"]}"')
print(f"confidence {uncertain['confidence']:.2f}: ", end="")
print(", ".join(f"{k} {v:.2f}" for k, v in
                sorted(uncertain["probabilities"].items(), key=lambda kv: -kv[1])[:3]))
"The cutover touches three teams, so check whether you are on this list before you plan anything for Monday:"
confidence 0.43: paragraph 0.53, list_item 0.24, callout 0.19

チームリストを導入する文は本当に曖昧です。後続を示し(見出し寄り)、完全な文であり (段落寄り)、コールアウトが来そうな位置にあります。確率はそれに応じて散らばり (paragraph 0.53、list_item 0.24、callout 0.19)、UI はそれを可視化できます。たとえば、 type の信頼度(勝った選択肢の背後の確率)が 0.55 未満のブロックに、レビュー用の下線を 引くなどです。