줄 단위 검색
GitHub 서비스 약관을 위한 시맨틱 검색을 구축합니다. 한 번의 요청에서 Choice 질문으로 218개 줄 id를 자연어 질의에 대해 채점하고, Noul 질문으로 문서에 답이 있는지 검사합니다.
여러분에게는 GitHub 서비스 약관과, 그것에 관한 자연어 질문이 있습니다. 필요한 것은 그
질문에 답하는 줄들과, 문서에 답이 없을 때를 감지하는 방법입니다. 예시에 포함된 질의는
직접적인 답이 있는 줄을 앞에 오도록 순위를 매깁니다. exists 임계값은 나머지 경우를
누락 또는 부분으로 분류합니다. 최종적으로 find()가 나오는데, 이 함수는 exists
확률과 줄마다 하나의 관련도 점수를 반환합니다.
검색 백엔드는 세 부분으로 구성됩니다.
- TypeSafe가 가리킬 수 있도록 각 줄에 ID를 붙입니다.
Choice질문을 사용해, 질의에 얼마나 잘 답하는지를 기준으로 그 줄 ID들의 순위를 매깁니다. Choice 질문의 확률은 항상 합이 1이므로, 질의에 답하는 줄이 하나도 없어도 어떤 줄이 반드시 1위가 됩니다.- 같은 요청에서
Noul질문을 사용해 문서에 답이 애초에 들어 있는지 검사합니다.
준비
TypeSafe API 키 발급
TypeSafe 콘솔에서 키를 만들고 내보냅니다.
export TYPESAFE_API_KEY="your-key-here"
의존성 설치
pip install 'cooksafe>=0.2.0,<0.3.0'
JsonCache는 함께 실린 API 응답을 재생하므로, 아래 단계는 API 키 없이, 비용도 들이지
않고 실행됩니다. 대신 요청을 실제로 보내려면 TYPESAFE_API_KEY를 설정하고
json_cache.json을 삭제하십시오.
스크립트 작성
semantic_search.py를 임포트와 클라이언트로 시작합니다.
import os
import urllib.request
from pathlib import Path
from cooksafe import JsonCache
from typesafe_sdk import Choice, Noul, NoulCriteria, TypeSafeClient
TYPESAFE_MODEL = "jev-1.12"
client = TypeSafeClient(
api_key=os.environ.get("TYPESAFE_API_KEY", "cache-only"), timeout=120.0
)
json_cache = JsonCache(Path("json_cache.json"))
1단계: 모든 줄에 ID 붙이기
테스트 문서는 GitHub 서비스 약관으로, 218개 조항으로 나뉘어 있어 모든 검색 결과가 인용할 수 있는 한 줄을 가리킵니다.
semantic_search.py에 다음을 추가합니다.
GIST = (
"https://gist.githubusercontent.com/eugene-shvarts/900632789a24983d5678ffd508dd01f6"
"/raw/cf9c2ab422d568deade949ef0a06bed6896964b9/github-tos.txt"
)
@json_cache
def fetch_document(url: str) -> str:
request = urllib.request.Request(
url, headers={"User-Agent": "typesafe-cookbook/1.0"}
)
with urllib.request.urlopen(request) as response:
return response.read().decode()
LINES = fetch_document(GIST).splitlines()
캐시는 반복 다운로드를 막고, splitlines()는 218개 문자열의 리스트를 남깁니다.
이제 각 줄 앞에 짧은 ID를 붙이고, 그 줄들을 다시 하나의 문서로 이어 붙입니다. 모델은 이 ID를 사용해 자기 답을 가리킵니다.
def line_id(i: int) -> str:
return f"L{i:03d}"
DOCUMENT = "\n".join(f"{line_id(i)}| {line}" for i, line in enumerate(LINES))
DOCUMENT는 이제 다음과 같습니다.
L052| You own Your Content. If you post Content you did not create, you are responsible for...
L053| You grant us and other Users the licenses in Sections D.4–D.8. These licenses apply...
L054| 4. License Grant to Us
2단계: 답이 어디 있는지 묻기
Choice 질문은 선택지마다 확률을 반환합니다. 줄 ID를 선택지로 사용하면 “선택지를
고른다”가 “줄을 가리킨다”가 됩니다.
def where_question(query: str) -> Choice:
return Choice(
instructions=f'Which line of the document contains the answer to: "{query}"?',
criteria={line_id(i): None for i in range(len(LINES))},
)
선택지 설명이 None인 이유는 문서가 이미 각 ID에 해당하는 텍스트를 담고 있기
때문입니다. 질의는 instructions에 들어가고, 상태는 검색 사이에 변하지 않습니다.
3단계: 답이 존재하는지 검사하기
Choice 확률은 항상 합이 1이므로, 문서가 질문에 답하지 못할 때도 어떤 줄이 1위가 됩니다. 순위만으로는 진짜 답과 가장 가까운 무관한 줄을 구별할 수 없습니다.
그래서 같은 요청에서 두 번째 질문을 던집니다.
def exists_question(query: str) -> Noul:
return Noul(
instructions=f'Does any line of the document address or answer: "{query}"?',
criteria=NoulCriteria(
true="At least one line of the document states or directly implies the answer",
false="No line of the document addresses this",
),
)
Choice 확률과 달리 Noul 확률은 다른 선택지에 의존하지 않으므로, 문서에 답이 없을 때 0에 가까이 떨어질 수 있습니다.
4단계: 두 질문을 한 번의 요청으로 보내기
system_one 메서드는 두 질문에 한 번에 답합니다. 상태는 한 번만 전송되므로, 존재
여부 검사를 추가하는 데 필요한 추가 출력은 아주 적습니다.
@json_cache
def _find(
model: str,
state: str,
where: Choice,
exists: Noul,
) -> dict:
response = client.system_one(
state=state,
questions={"where": where, "exists": exists},
model=model,
)
probabilities = response.answers["where"].probabilities
return {
"exists": response.answers["exists"].noul,
"relevance": [probabilities.get(line_id(i), 0.0) for i in range(len(LINES))],
}
def find(query: str) -> dict:
return _find(
TYPESAFE_MODEL,
DOCUMENT,
where_question(query),
exists_question(query),
)
relevance 리스트는 문서 순서대로 줄마다 점수 하나를 유지합니다.
5단계: 결과 읽기
로컬 코드 두 조각이 마무리합니다. verdict()는 원시 exists 확률을 세 가지 상태로
바꾸는데, 부분 답변을 위한 중간 상태가 하나 있습니다. show()는 relevance를 막대
그래프로 렌더링하여 순위를 터미널에서도 읽을 수 있게 합니다.
FOUND, ABSENT = 0.7, 0.35 # present answers typically read >=0.9, absent <=0.05
def verdict(exists: float) -> str:
if exists >= FOUND:
return "answered in this document"
return "not in this document" if exists < ABSENT else "partially addressed"
def show(query: str, top: int = 4) -> dict:
result = find(query)
print(f'"{query}"')
print(f" exists {result['exists']:.2f} -> {verdict(result['exists'])}")
ranked = sorted(
range(len(LINES)), key=lambda i: result["relevance"][i], reverse=True
)
for i in ranked[:top]:
bar = "#" * max(1, round(result["relevance"][i] * 12))
preview = LINES[i][:58].rstrip()
print(f" {line_id(i)} {result['relevance'][i]:.2f} {bar:<12} {preview}")
return result
이 임계값들은 아래 예시들을 구분해 주지만, 프로덕션에서 사용하기 전에 여러분의 문서에 맞게 조정하십시오.
6단계: 검색 실행하기
직접적인 답이 있는 질문 둘, 답이 없는 질문 하나, 부분적인 답이 있는 질문 하나, 이렇게 네 개를 던집니다.
print(f"{len(LINES)} lines, {len(DOCUMENT):,} characters\n")
show("who owns the code I upload?")
print()
show("can GitHub kick me off the platform without warning?")
print()
show("do I have to take disputes to arbitration?", top=2)
print()
show("can minors use GitHub with parental permission?", top=2)
218 lines, 43,980 characters
"who owns the code I upload?"
exists 0.98 -> answered in this document
L052 0.95 ########### You own Your Content. If you post Content you did not crea
L046 0.02 # Short version: You own content you create, but you allow u
L051 0.02 # 3. Ownership and License Grants
L217 0.01 # Questions about the Terms of Service? Contact us through t
"can GitHub kick me off the platform without warning?"
exists 0.97 -> answered in this document
L168 0.97 ############ GitHub has the right to suspend or terminate your access t
L167 0.03 # 3. GitHub May Terminate
L000 0.00 # Effective date: April 27, 2026 · A. Definitions
L001 0.00 # Short version: We use these basic terms throughout the agr
"do I have to take disputes to arbitration?"
exists 0.14 -> not in this document
L205 0.86 ########## Except to the extent applicable law provides otherwise, th
L168 0.02 # GitHub has the right to suspend or terminate your access t
"can minors use GitHub with parental permission?"
exists 0.46 -> partially addressed
L029 0.90 ########### You must be age 13 or older. While we are thrilled to see
L012 0.07 # “User,” “You,” and “Your” refer to the individual person,
점수가 의미하는 것
처음 두 질의는 직접적인 답과 그것을 검증하는 데 필요한 원문 줄을 반환합니다.
나머지 둘은 존재 여부 검사가 왜 중요한지 보여줍니다.
- 중재: 순위는 가장 가까운 줄에 0.86점을 주지만,
exists는 0.14에 불과합니다. 답은 문서에 없습니다. - 부모 동의: 연령 규정이 1위로 올라오지만, 부모 동의가 이 규정을 바꾸는지는 답하지 않습니다. 결과는 부분적으로 다뤄진 상태입니다.
순위는 어디를 봐야 하는지 알려주고, exists 점수는 그 결과가 질문에 답하는지
알려줍니다.
여러분의 문서에서 시험해 보기
TypeSafe playground에서 태그가 붙은 계약을 열어 같은 텍스트에 대해
질문을 편집합니다. 여러분의 문서를 검색하려면 fetch_document()의 URL을 바꾸십시오.
스크립트의 나머지 모든 줄은 LINES를 기반으로 동작합니다.