逐行搜尋
為 GitHub 的服務條款構建語義搜尋。一次請求裡,用一個 Choice 問題給 218 個行 id 相對一句大白話查詢打分,再用一個 Noul 問題檢查文件裡到底有沒有答案。
你手上有 GitHub 的服務條款,還有一個關於它的大白話問題。你需要的是能回答這個問題的行,
以及一種在文件裡沒有答案時能察覺出來的辦法。示例裡的查詢把有直接答案的行排在前面。
exists 閾值把其餘情況分成「缺失」或「部分」。最終你會得到一個 find(),它返回
exists 機率,以及每一行的一個相關性分數。
搜尋後端由三部分拼起來:
- 給每一行打上一個 ID,這樣 TypeSafe 才能指向它。
- 用一個
Choice問題,按各行回答該查詢的好壞給這些行 ID 排序。Choice 問題的機率總和恆為 1,所以即使沒有任何一行回答了查詢,也總會有一行排在第一。 - 在同一次請求裡,用一個
Noul問題檢查文件裡到底有沒有答案。
準備工作
獲取 TypeSafe API key
在 TypeSafe 控制台建立一個 key 並匯出:
export TYPESAFE_API_KEY="your-key-here"
安裝依賴
pip install 'cooksafe>=0.2.0,<0.3.0'
JsonCache 會回放隨附的 API 響應,所以下面的步驟不需要 API key
也不花一分錢。若想讓請求真正發出去,就設定 TYPESAFE_API_KEY 並刪掉
json_cache.json。
建立指令碼
semantic_search.py 從匯入和客戶端開始:
import os
import urllib.request
from pathlib import Path
from cooksafe import JsonCache
from typesafe_sdk import Choice, Noul, NoulCriteria, TypeSafeClient
TYPESAFE_MODEL = "jev-1.12"
client = TypeSafeClient(
api_key=os.environ.get("TYPESAFE_API_KEY", "cache-only"), timeout=120.0
)
json_cache = JsonCache(Path("json_cache.json"))
第 1 步:給每一行打上 ID
測試文件是 GitHub 的服務條款,切成 218 條,這樣每條搜尋結果都指向一行可以引用的文字。
往 semantic_search.py 里加上:
GIST = (
"https://gist.githubusercontent.com/eugene-shvarts/900632789a24983d5678ffd508dd01f6"
"/raw/cf9c2ab422d568deade949ef0a06bed6896964b9/github-tos.txt"
)
@json_cache
def fetch_document(url: str) -> str:
request = urllib.request.Request(
url, headers={"User-Agent": "typesafe-cookbook/1.0"}
)
with urllib.request.urlopen(request) as response:
return response.read().decode()
LINES = fetch_document(GIST).splitlines()
快取避免了重複下載,splitlines() 留下一個含 218 個字串的列表。
現在給每一行加一個短 ID 字首,再把這些行拼回一個文件。模型用這些 ID 指向它的答案。
def line_id(i: int) -> str:
return f"L{i:03d}"
DOCUMENT = "\n".join(f"{line_id(i)}| {line}" for i, line in enumerate(LINES))
DOCUMENT 現在長這樣:
L052| You own Your Content. If you post Content you did not create, you are responsible for...
L053| You grant us and other Users the licenses in Sections D.4–D.8. These licenses apply...
L054| 4. License Grant to Us
第 2 步:問答案在哪
一個 Choice 問題會為每個選項返回一個機率。把行 ID 當作選
項,“挑一個選項” 就變成了 “指向某一行”。
def where_question(query: str) -> Choice:
return Choice(
instructions=f'Which line of the document contains the answer to: "{query}"?',
criteria={line_id(i): None for i in range(len(LINES))},
)
選項的描述都是 None,因為文件裡已經含有每個
ID 對應的文本了。查詢放在 instructions 裡;狀態在多次搜尋之間保持不變。
第 3 步:檢查是否存在答案
Choice 的機率總和恆為 1,所以即使文件 回答不了這個問題,也總有某一行排在第一。光靠排序,分不清真答案和離得最近的 不相關行。
所以在同一次請求裡再問第二個問題:
def exists_question(query: str) -> Noul:
return Noul(
instructions=f'Does any line of the document address or answer: "{query}"?',
criteria=NoulCriteria(
true="At least one line of the document states or directly implies the answer",
false="No line of the document addresses this",
),
)
與 Choice 的機率不同,Noul 的機率不取決於其它選項, 所以當文件裡沒有答案時,它可以掉到接近零。
第 4 步:把兩個問題放進一次請求
system_one 方法一趟就回答這兩個問題。狀態只發送一次,所以
加上存在性檢查只需要多出一點點輸出。
@json_cache
def _find(
model: str,
state: str,
where: Choice,
exists: Noul,
) -> dict:
response = client.system_one(
state=state,
questions={"where": where, "exists": exists},
model=model,
)
probabilities = response.answers["where"].probabilities
return {
"exists": response.answers["exists"].noul,
"relevance": [probabilities.get(line_id(i), 0.0) for i in range(len(LINES))],
}
def find(query: str) -> dict:
return _find(
TYPESAFE_MODEL,
DOCUMENT,
where_question(query),
exists_question(query),
)
relevance 列表按文件順序為每一行保留一個分數。
第 5 步:讀結果
兩段原生代碼收尾:verdict() 把原始的 exists 機率
變成三種狀態,其中給部分回答留了一箇中間態;show() 把 relevance
渲染成柱狀圖,讓排序在終端裡也能讀。
FOUND, ABSENT = 0.7, 0.35 # present answers typically read >=0.9, absent <=0.05
def verdict(exists: float) -> str:
if exists >= FOUND:
return "answered in this document"
return "not in this document" if exists < ABSENT else "partially addressed"
def show(query: str, top: int = 4) -> dict:
result = find(query)
print(f'"{query}"')
print(f" exists {result['exists']:.2f} -> {verdict(result['exists'])}")
ranked = sorted(
range(len(LINES)), key=lambda i: result["relevance"][i], reverse=True
)
for i in ranked[:top]:
bar = "#" * max(1, round(result["relevance"][i] * 12))
preview = LINES[i][:58].rstrip()
print(f" {line_id(i)} {result['relevance'][i]:.2f} {bar:<12} {preview}")
return result
這些閾值能區分下面的例子,但在生產環境使用之前, 請拿你自己的文件調一調。
第 6 步:執行搜尋
問兩個有直接答案的問題,一個沒有答案的,一個只有 部分答案的,一共四個。
print(f"{len(LINES)} lines, {len(DOCUMENT):,} characters\n")
show("who owns the code I upload?")
print()
show("can GitHub kick me off the platform without warning?")
print()
show("do I have to take disputes to arbitration?", top=2)
print()
show("can minors use GitHub with parental permission?", top=2)
218 lines, 43,980 characters
"who owns the code I upload?"
exists 0.98 -> answered in this document
L052 0.95 ########### You own Your Content. If you post Content you did not crea
L046 0.02 # Short version: You own content you create, but you allow u
L051 0.02 # 3. Ownership and License Grants
L217 0.01 # Questions about the Terms of Service? Contact us through t
"can GitHub kick me off the platform without warning?"
exists 0.97 -> answered in this document
L168 0.97 ############ GitHub has the right to suspend or terminate your access t
L167 0.03 # 3. GitHub May Terminate
L000 0.00 # Effective date: April 27, 2026 · A. Definitions
L001 0.00 # Short version: We use these basic terms throughout the agr
"do I have to take disputes to arbitration?"
exists 0.14 -> not in this document
L205 0.86 ########## Except to the extent applicable law provides otherwise, th
L168 0.02 # GitHub has the right to suspend or terminate your access t
"can minors use GitHub with parental permission?"
exists 0.46 -> partially addressed
L029 0.90 ########### You must be age 13 or older. While we are thrilled to see
L012 0.07 # “User,” “You,” and “Your” refer to the individual person,
這些分數意味著什麼
前兩個查詢返回直接答案,以及核對它們所需的源文本行。
另外兩個則說明存在性檢查為什麼重要:
- 仲裁: 排序把最接近的一行給了 0.86 分,但
exists只有 0.14。答案並不在文件裡。 - 家長許可: 年齡規定排在第一位,但它沒有回答家長許可是否會改變這條規定。結果是部分涉及。
排序告訴你該去哪兒找;exists 分數告訴你這個結果
有沒有回答問題。
在你自己的文件上試試
在 TypeSafe playground 中開啟這份打好標籤的合同,對著同一段文本編輯問題。想搜你自己的文件,就替換 fetch_document() 裡的 URL;指令碼其餘
每一行都基於 LINES 工作。