ドキュメント

行単位の検索

行単位の検索

GitHub の利用規約で行単位のセマンティック検索を構築します。1 回のリクエストで Choice が 218 個の行 id を採点し、Noul が答えの有無を確認します。

GitHub の利用規約と、それについての平易な質問があります。必要なのは、質問に答える行と、 ドキュメントに答えがないときにそれを検知する手段です。同梱のクエリは、直接の答えがある行を 上位に並べます。exists のしきい値が、残りのケースを「欠落」か「部分」に分類します。 最終的に得られるのは find() です。これは exists の確率と、行ごとの関連度スコアを返します。

クエリがドキュメントを走査し、一致する行に付いた答えを明らかにする

検索バックエンドは 3 つの部分から組み立てられます。

  1. TypeSafe が行を指せるように、各行に ID を付けます。
  2. Choice の質問を使い、各行 ID がクエリにどれだけ答えているかで順位を付けます。Choice の質問の確率は常に合計 1 になるため、どの行もクエリに答えていなくても、必ず 1 行が 1 位になります。
  3. 同じリクエストで Noul の質問を使い、ドキュメントにそもそも答えがあるかどうかを確認します。

準備

TypeSafe API キーを取得する

TypeSafe コンソールでキーを作成し、エクスポートします。

export TYPESAFE_API_KEY="your-key-here"

依存関係をインストールする

pip install 'cooksafe>=0.2.0,<0.3.0'

JsonCache は同梱の API レスポンスを再生するので、以下の手順は API キーなし・費用なしで 実行できます。代わりにリクエストを実際に送るには、TYPESAFE_API_KEY を設定して json_cache.json を削除します。

スクリプトを作成する

semantic_search.py を、インポートとクライアントから始めます。

import os
import urllib.request
from pathlib import Path

from cooksafe import JsonCache
from typesafe_sdk import Choice, Noul, NoulCriteria, TypeSafeClient

TYPESAFE_MODEL = "jev-1.12"

client = TypeSafeClient(
    api_key=os.environ.get("TYPESAFE_API_KEY", "cache-only"), timeout=120.0
)
json_cache = JsonCache(Path("json_cache.json"))

ステップ 1:各行に ID を付ける

テスト用ドキュメントは GitHub の利用規約で、218 個の条項に分割されているため、各検索 結果は引用可能な 1 行を指します。

semantic_search.py に追記します。

GIST = (
    "https://gist.githubusercontent.com/eugene-shvarts/900632789a24983d5678ffd508dd01f6"
    "/raw/cf9c2ab422d568deade949ef0a06bed6896964b9/github-tos.txt"
)

@json_cache
def fetch_document(url: str) -> str:
    request = urllib.request.Request(
        url, headers={"User-Agent": "typesafe-cookbook/1.0"}
    )
    with urllib.request.urlopen(request) as response:
        return response.read().decode()

LINES = fetch_document(GIST).splitlines()

キャッシュが再ダウンロードを防ぎ、splitlines() は 218 個の文字列からなるリストを残します。

次に各行の先頭に短い ID を付け、行を 1 つのドキュメントに再結合します。モデルはこの ID を使って答えを指します。

def line_id(i: int) -> str:
    return f"L{i:03d}"

DOCUMENT = "\n".join(f"{line_id(i)}| {line}" for i, line in enumerate(LINES))

DOCUMENT はこのようになります。

L052| You own Your Content. If you post Content you did not create, you are responsible for...
L053| You grant us and other Users the licenses in Sections D.4–D.8. These licenses apply...
L054| 4. License Grant to Us

ステップ 2:答えがどこにあるかを尋ねる

Choice の質問は、各選択肢に対して確率を返します。行 ID を選択肢にすれば、 「選択肢を 1 つ選ぶ」が「ある行を指す」になります。

def where_question(query: str) -> Choice:
    return Choice(
        instructions=f'Which line of the document contains the answer to: "{query}"?',
        criteria={line_id(i): None for i in range(len(LINES))},
    )

選択肢の説明が None なのは、各 ID に対応するテキストがドキュメントにすでに含まれて いるためです。クエリは instructions に入れます。state は検索の間ずっと変わりません。

ステップ 3:答えが存在するかを確認する

Choice の確率は常に合計 1 になるため、ドキュメントが質問に答えていなくても、ある行が 1 位になります。順位付けだけでは、本当の答えと最も近い無関係な行を区別できません。

そこで、同じリクエストでもう 1 つ質問します。

def exists_question(query: str) -> Noul:
    return Noul(
        instructions=f'Does any line of the document address or answer: "{query}"?',
        criteria=NoulCriteria(
            true="At least one line of the document states or directly implies the answer",
            false="No line of the document addresses this",
        ),
    )

Choice の確率と違い、Noul の確率は他の選択肢に依存しないため、ドキュメントに答えが ないときはほぼゼロまで下がります。

ステップ 4:2 つの質問を 1 回のリクエストで送る

system_one メソッドは 1 回の処理で両方の質問に答えます。state は一度だけ送られるので、 存在チェックを追加しても増える出力はわずかです。

タグ付きのドキュメントとユーザーの質問が 1 つの TypeSafe リクエストに入ります。Choice の質問が
すべての行を採点し、Noul の質問が答えが存在するかを確認します。その後、ローカルコードが
行を順位付けし、ドキュメントの判定を適用します。
@json_cache
def _find(
    model: str,
    state: str,
    where: Choice,
    exists: Noul,
) -> dict:
    response = client.system_one(
        state=state,
        questions={"where": where, "exists": exists},
        model=model,
    )
    probabilities = response.answers["where"].probabilities
    return {
        "exists": response.answers["exists"].noul,
        "relevance": [probabilities.get(line_id(i), 0.0) for i in range(len(LINES))],
    }

def find(query: str) -> dict:
    return _find(
        TYPESAFE_MODEL,
        DOCUMENT,
        where_question(query),
        exists_question(query),
    )

relevance リストは、ドキュメント順に 1 行につき 1 つのスコアを保持します。

ステップ 5:結果を読む

仕上げはローカルコードの 2 つの部分です。verdict() は生の exists 確率を 3 つの 状態に変換し、部分的な答え用に中間の状態を用意します。show() は relevance を 棒グラフとして描画し、ターミナルでも順位を読み取れるようにします。

FOUND, ABSENT = 0.7, 0.35  # present answers typically read >=0.9, absent <=0.05

def verdict(exists: float) -> str:
    if exists >= FOUND:
        return "answered in this document"
    return "not in this document" if exists < ABSENT else "partially addressed"

def show(query: str, top: int = 4) -> dict:
    result = find(query)
    print(f'"{query}"')
    print(f"  exists {result['exists']:.2f} -> {verdict(result['exists'])}")
    ranked = sorted(
        range(len(LINES)), key=lambda i: result["relevance"][i], reverse=True
    )
    for i in ranked[:top]:
        bar = "#" * max(1, round(result["relevance"][i] * 12))
        preview = LINES[i][:58].rstrip()
        print(f"  {line_id(i)}  {result['relevance'][i]:.2f}  {bar:<12}  {preview}")
    return result

これらのしきい値は以下の例を分離しますが、本番で使う前に自分のドキュメントで 調整してください。

ステップ 6:検索を実行する

直接の答えがある質問を 2 つ、答えがない質問を 1 つ、部分的な答えがある質問を 1 つ、 計 4 つ尋ねます。

print(f"{len(LINES)} lines, {len(DOCUMENT):,} characters\n")
show("who owns the code I upload?")
print()
show("can GitHub kick me off the platform without warning?")
print()
show("do I have to take disputes to arbitration?", top=2)
print()
show("can minors use GitHub with parental permission?", top=2)
218 lines, 43,980 characters

"who owns the code I upload?"
  exists 0.98 -> answered in this document
  L052  0.95  ###########   You own Your Content. If you post Content you did not crea
  L046  0.02  #             Short version: You own content you create, but you allow u
  L051  0.02  #             3. Ownership and License Grants
  L217  0.01  #             Questions about the Terms of Service? Contact us through t

"can GitHub kick me off the platform without warning?"
  exists 0.97 -> answered in this document
  L168  0.97  ############  GitHub has the right to suspend or terminate your access t
  L167  0.03  #             3. GitHub May Terminate
  L000  0.00  #             Effective date: April 27, 2026 · A. Definitions
  L001  0.00  #             Short version: We use these basic terms throughout the agr

"do I have to take disputes to arbitration?"
  exists 0.14 -> not in this document
  L205  0.86  ##########    Except to the extent applicable law provides otherwise, th
  L168  0.02  #             GitHub has the right to suspend or terminate your access t

"can minors use GitHub with parental permission?"
  exists 0.46 -> partially addressed
  L029  0.90  ###########   You must be age 13 or older. While we are thrilled to see
  L012  0.07  #             “User,” “You,” and “Your” refer to the individual person,

スコアが意味するもの

最初の 2 つのクエリは、直接の答えと、それを検証するのに必要な原文の行を返します。

残りの 2 つは、存在チェックがなぜ重要かを示します。

  • 仲裁: 順位付けは最も近い行に 0.86 のスコアを与えますが、exists は 0.14 しか ありません。答えはドキュメントの中にはありません。
  • 保護者の許可: 年齢のルールが 1 位になりますが、保護者の許可がそのルールを変えるか どうかには答えていません。結果は部分的に該当です。

順位付けはどこを見るべきかを教え、exists スコアはその結果が質問に答えているかを 教えます。

自分のドキュメントで試す

タグ付きの契約書を TypeSafe playground で開く と、同じテキストに対して 質問を編集できます。自分のドキュメントを検索するには、fetch_document() の URL を 差し替えます。スクリプトの他の行はすべて LINES を基に動作します。