行単位の検索
行単位の検索
GitHub の利用規約で行単位のセマンティック検索を構築します。1 回のリクエストで Choice が 218 個の行 id を採点し、Noul が答えの有無を確認します。
GitHub の利用規約と、それについての平易な質問があります。必要なのは、質問に答える行と、
ドキュメントに答えがないときにそれを検知する手段です。同梱のクエリは、直接の答えがある行を
上位に並べます。exists のしきい値が、残りのケースを「欠落」か「部分」に分類します。
最終的に得られるのは find() です。これは exists の確率と、行ごとの関連度スコアを返します。
検索バックエンドは 3 つの部分から組み立てられます。
- TypeSafe が行を指せるように、各行に ID を付けます。
Choiceの質問を使い、各行 ID がクエリにどれだけ答えているかで順位を付けます。Choice の質問の確率は常に合計 1 になるため、どの行もクエリに答えていなくても、必ず 1 行が 1 位になります。- 同じリクエストで
Noulの質問を使い、ドキュメントにそもそも答えがあるかどうかを確認します。
準備
TypeSafe API キーを取得する
TypeSafe コンソールでキーを作成し、エクスポートします。
export TYPESAFE_API_KEY="your-key-here"
依存関係をインストールする
pip install 'cooksafe>=0.2.0,<0.3.0'
JsonCache は同梱の API レスポンスを再生するので、以下の手順は API キーなし・費用なしで
実行できます。代わりにリクエストを実際に送るには、TYPESAFE_API_KEY を設定して
json_cache.json を削除します。
スクリプトを作成する
semantic_search.py を、インポートとクライアントから始めます。
import os
import urllib.request
from pathlib import Path
from cooksafe import JsonCache
from typesafe_sdk import Choice, Noul, NoulCriteria, TypeSafeClient
TYPESAFE_MODEL = "jev-1.12"
client = TypeSafeClient(
api_key=os.environ.get("TYPESAFE_API_KEY", "cache-only"), timeout=120.0
)
json_cache = JsonCache(Path("json_cache.json"))
ステップ 1:各行に ID を付ける
テスト用ドキュメントは GitHub の利用規約で、218 個の条項に分割されているため、各検索 結果は引用可能な 1 行を指します。
semantic_search.py に追記します。
GIST = (
"https://gist.githubusercontent.com/eugene-shvarts/900632789a24983d5678ffd508dd01f6"
"/raw/cf9c2ab422d568deade949ef0a06bed6896964b9/github-tos.txt"
)
@json_cache
def fetch_document(url: str) -> str:
request = urllib.request.Request(
url, headers={"User-Agent": "typesafe-cookbook/1.0"}
)
with urllib.request.urlopen(request) as response:
return response.read().decode()
LINES = fetch_document(GIST).splitlines()
キャッシュが再ダウンロードを防ぎ、splitlines() は 218 個の文字列からなるリストを残します。
次に各行の先頭に短い ID を付け、行を 1 つのドキュメントに再結合します。モデルはこの ID を使って答えを指します。
def line_id(i: int) -> str:
return f"L{i:03d}"
DOCUMENT = "\n".join(f"{line_id(i)}| {line}" for i, line in enumerate(LINES))
DOCUMENT はこのようになります。
L052| You own Your Content. If you post Content you did not create, you are responsible for...
L053| You grant us and other Users the licenses in Sections D.4–D.8. These licenses apply...
L054| 4. License Grant to Us
ステップ 2:答えがどこにあるかを尋ねる
Choice の質問は、各選択肢に対して確率を返します。行 ID を選択肢にすれば、
「選択肢を 1 つ選ぶ」が「ある行を指す」になります。
def where_question(query: str) -> Choice:
return Choice(
instructions=f'Which line of the document contains the answer to: "{query}"?',
criteria={line_id(i): None for i in range(len(LINES))},
)
選択肢の説明が None なのは、各 ID に対応するテキストがドキュメントにすでに含まれて
いるためです。クエリは instructions に入れます。state は検索の間ずっと変わりません。
ステップ 3:答えが存在するかを確認する
Choice の確率は常に合計 1 になるため、ドキュメントが質問に答えていなくても、ある行が 1 位になります。順位付けだけでは、本当の答えと最も近い無関係な行を区別できません。
そこで、同じリクエストでもう 1 つ質問します。
def exists_question(query: str) -> Noul:
return Noul(
instructions=f'Does any line of the document address or answer: "{query}"?',
criteria=NoulCriteria(
true="At least one line of the document states or directly implies the answer",
false="No line of the document addresses this",
),
)
Choice の確率と違い、Noul の確率は他の選択肢に依存しないため、ドキュメントに答えが ないときはほぼゼロまで下がります。
ステップ 4:2 つの質問を 1 回のリクエストで送る
system_one メソッドは 1 回の処理で両方の質問に答えます。state は一度だけ送られるので、
存在チェックを追加しても増える出力はわずかです。
@json_cache
def _find(
model: str,
state: str,
where: Choice,
exists: Noul,
) -> dict:
response = client.system_one(
state=state,
questions={"where": where, "exists": exists},
model=model,
)
probabilities = response.answers["where"].probabilities
return {
"exists": response.answers["exists"].noul,
"relevance": [probabilities.get(line_id(i), 0.0) for i in range(len(LINES))],
}
def find(query: str) -> dict:
return _find(
TYPESAFE_MODEL,
DOCUMENT,
where_question(query),
exists_question(query),
)
relevance リストは、ドキュメント順に 1 行につき 1 つのスコアを保持します。
ステップ 5:結果を読む
仕上げはローカルコードの 2 つの部分です。verdict() は生の exists 確率を 3 つの
状態に変換し、部分的な答え用に中間の状態を用意します。show() は relevance を
棒グラフとして描画し、ターミナルでも順位を読み取れるようにします。
FOUND, ABSENT = 0.7, 0.35 # present answers typically read >=0.9, absent <=0.05
def verdict(exists: float) -> str:
if exists >= FOUND:
return "answered in this document"
return "not in this document" if exists < ABSENT else "partially addressed"
def show(query: str, top: int = 4) -> dict:
result = find(query)
print(f'"{query}"')
print(f" exists {result['exists']:.2f} -> {verdict(result['exists'])}")
ranked = sorted(
range(len(LINES)), key=lambda i: result["relevance"][i], reverse=True
)
for i in ranked[:top]:
bar = "#" * max(1, round(result["relevance"][i] * 12))
preview = LINES[i][:58].rstrip()
print(f" {line_id(i)} {result['relevance'][i]:.2f} {bar:<12} {preview}")
return result
これらのしきい値は以下の例を分離しますが、本番で使う前に自分のドキュメントで 調整してください。
ステップ 6:検索を実行する
直接の答えがある質問を 2 つ、答えがない質問を 1 つ、部分的な答えがある質問を 1 つ、 計 4 つ尋ねます。
print(f"{len(LINES)} lines, {len(DOCUMENT):,} characters\n")
show("who owns the code I upload?")
print()
show("can GitHub kick me off the platform without warning?")
print()
show("do I have to take disputes to arbitration?", top=2)
print()
show("can minors use GitHub with parental permission?", top=2)
218 lines, 43,980 characters
"who owns the code I upload?"
exists 0.98 -> answered in this document
L052 0.95 ########### You own Your Content. If you post Content you did not crea
L046 0.02 # Short version: You own content you create, but you allow u
L051 0.02 # 3. Ownership and License Grants
L217 0.01 # Questions about the Terms of Service? Contact us through t
"can GitHub kick me off the platform without warning?"
exists 0.97 -> answered in this document
L168 0.97 ############ GitHub has the right to suspend or terminate your access t
L167 0.03 # 3. GitHub May Terminate
L000 0.00 # Effective date: April 27, 2026 · A. Definitions
L001 0.00 # Short version: We use these basic terms throughout the agr
"do I have to take disputes to arbitration?"
exists 0.14 -> not in this document
L205 0.86 ########## Except to the extent applicable law provides otherwise, th
L168 0.02 # GitHub has the right to suspend or terminate your access t
"can minors use GitHub with parental permission?"
exists 0.46 -> partially addressed
L029 0.90 ########### You must be age 13 or older. While we are thrilled to see
L012 0.07 # “User,” “You,” and “Your” refer to the individual person,
スコアが意味するもの
最初の 2 つのクエリは、直接の答えと、それを検証するのに必要な原文の行を返します。
残りの 2 つは、存在チェックがなぜ重要かを示します。
- 仲裁: 順位付けは最も近い行に 0.86 のスコアを与えますが、
existsは 0.14 しか ありません。答えはドキュメントの中にはありません。 - 保護者の許可: 年齢のルールが 1 位になりますが、保護者の許可がそのルールを変えるか どうかには答えていません。結果は部分的に該当です。
順位付けはどこを見るべきかを教え、exists スコアはその結果が質問に答えているかを
教えます。
自分のドキュメントで試す
タグ付きの契約書を TypeSafe playground で開く と、同じテキストに対して
質問を編集できます。自分のドキュメントを検索するには、fetch_document() の URL を
差し替えます。スクリプトの他の行はすべて LINES を基に動作します。