리랭킹
40개의 CLERC 법률 쿼리에 대해 30개 구절 BM25 후보 목록을 만든 뒤, 쿼리-후보 쌍마다 TypeSafe 질문 하나로 top-1 정확도를 5%에서 18%로, top-10 정확도를 38%에서 62%로 끌어올립니다.
수천 개의 문서가 있고, 특정 질문에 답하는 하나를 찾아야 한다고 합시다. 그럼 어떻게 찾을까요?
먼저 키워드 매칭 같은 빠른 방법으로 그 수천 개 후보를 그럴듯한 것들의 후보 목록으로 줄입니다. 이것을 빠른 검색(fast search)이라 부릅니다.
빠른 검색은 그것을 잘하지만, 후보 목록의 어느 후보가 정답인지는 알려 주지 못합니다. 여기서 리랭킹이 등장합니다. 리랭킹은 후보 목록의 모든 후보를 쿼리에 직접 점수 매기고, 가장 좋은 것을 맨 앞에 놓습니다.
두 단계 모두 아래에서 CLERC 데이터셋의 법원 의견 구절 3,565개에 대해 실행됩니다. BM25가 40개 쿼리 각각에 대해 30개 후보의 빠른 검색 후보 목록을 만들고, TypeSafe가 각 후보 목록을 리랭킹합니다. 리랭킹을 하면 정답 구절이 쿼리의 18%에서 1위에 오르며, 이는 빠른 검색만 썼을 때의 5%에서 오른 것입니다.
과정에서 여러분은 다음을 배웁니다:
- 빠른 검색이 하는 일, 그리고 그것이 답 전체가 아닌 이유
- 리랭킹이 무엇이며, 빠른 검색 단계 뒤에 어떻게 들어맞는지
- TypeSafe가 한 후보를 쿼리에 어떻게 점수 매기고, 그것이 결과를 얼마나 개선하는지
직접 해 보기
TypeSafe Playground에서 쿼리, 후보, 리랭킹 질문 열기
수천 개 중에서 문서 하나를 어떻게 찾을까요?
문서 더미와 쿼리, 즉 여러분이 찾는 것을 기술하는 텍스트 한 조각이 있습니다. 더미 어딘가에 그것에 답하는 문서 하나가 있습니다.
모든 문서를 쿼리와 하나씩 대조하는 것은 문서당 비교 한 번으로 동작합니다. 문서가 수백만 개면 질의당 비교가 수백만 번입니다. 두 단계 접근으로 성능을 개선할 수 있습니다.
- 전체 더미에 실행할 만큼 빠른 방법으로, 더미를 가능성 있는 후보들의 짧은 목록으로 줄입니다.
- 그 짧은 목록에 더 정확한 단계를 적용하여, 정확히 맞는 답을 찾습니다.
이 cookbook은 그 구성을 법원 의견 데이터셋에 대해 테스트하며, 아래 리랭킹 예제에서 다룹니다.
빠른 검색이란?
빠른 검색은 큰 코퍼스의 모든 문서에 대해 쿼리를 비교하고 빠르게 순위가 매겨진 후보 목록을 반환할 수 있는 모든 방법입니다. 흔한 방법으로는 BM25 같은 키워드 검색과, 의미로 구절을 비교하는 밀집 임베딩이 있습니다. 시스템은 종종 두 방법을 결합합니다.
여기서 첫 단계는 BM25이고 그뿐입니다. BM25는 공유 단어로 구절의 순위를 매깁니다. 이 단계를 단순하게 유지하면 관심이 리랭킹에 남고, 그것이 이 cookbook의 요점입니다. 빠른 검색 방법의 선택은 부차적인 문제입니다. 리랭킹은 후보 목록에 들어온 구절만 보기 때문입니다.
리랭킹이란?
리랭킹은 빠른 검색이 이미 만들어 낸 후보 목록을 가져다 더 나은 순서로 놓습니다. 쿼리를 코퍼스 전체와 한꺼번에 비교하는 대신, 쿼리를 후보 목록의 각 후보와 개별적으로 비교하고, 그 점수로 후보 목록을 정렬합니다.
점수는 언어 모델에서 나올 수 있습니다. 모델에게 쿼리와 후보 하나를 함께 주고 그 후보가 쿼리에 얼마나 잘 답하는지 묻습니다. 그러면 리랭킹은 후보의 표현이 쿼리와 다르더라도 후보 목록에서 가장 잘 맞는 것을 찾습니다.
TypeSafe로 리랭킹하기
리랭커는 모든 쿼리-후보 쌍에 대해 비교 가능한 점수가 필요합니다. 범용 언어 모델이 이 점수들을 생성하거나, 후보 목록 전체를 직접 순위 매길 수 있습니다. 하지만 독립적인 쌍 점수 매기기에서는 점수 척도를 정의하고 모델에게 모든 후보에 같은 기준을 적용하도록 프롬프트해야 합니다. 반복 호출은 같은 쌍에 대해 여전히 다른 점수를 낼 수 있고, 범용 생성은 숫자 하나만 필요한 작업에 시간과 비용을 더합니다.
TypeSafe가 반환하는 것
TypeSafe에서는 점수 매기기 요청이 예/아니오 질문으로 남을 수 있습니다.
Could this candidate passage be from the cited precedent?
단순한 예 또는 아니오로는 30개 후보의 순위를 매기기에 충분하지 않습니다. 대신 Noul은 0과 1 사이의 숫자를 반환하며, 이를 noul이라 합니다. noul은 답이 예일 가능성에 대한 TypeSafe의 추정입니다.
이 질문의 기준이 참과 거짓을 무엇으로 볼지 정의합니다. TypeSafe는 그것을 모든 쿼리-후보 쌍에 적용하고 noul을 그대로 반환합니다. 그 noul이 애플리케이션이 정렬에 쓰는 점수입니다. 범용 모델을 위해 점수 척도를 만들어 낼 필요가 없으며, TypeSafe는 이 반복 점수 매기기를 더 빠르고, 더 저렴하고, 더 일관되게 하도록 만들어졌습니다.
단순화한 의사코드로, TypeSafe 점수 매기기 호출 하나는 이렇게 생겼습니다:
question = Noul(
instructions="Is this candidate the cited case?",
criteria=NoulCriteria(
true="The candidate states the specific rule the query cites.",
false="The candidate is only on a similar topic.",
),
)
response = client.system_one(state={...}, questions={"is_cited_source": question})
response.answers["is_cited_source"].noul # -> 0.87
TypeSafe는 쿼리와 후보 하나를 그 질문에 비추어 함께 읽고, noul을 반환합니다.
이것을 사용하면 후보 목록의 모든 후보에 같은 질문을 실행한 뒤, 각 호출이 돌려주는 noul로 후보 목록을 정렬하여(가장 높은 것이 먼저) 후보 목록을 리랭킹할 수 있습니다.
nouls = {candidate: ask_typesafe(query, candidate) for candidate in shortlist}
reranked = sorted(shortlist, key=lambda c: nouls[c], reverse=True) # highest noul first
아래 다이어그램은 후보마다 요청 하나가 후보 목록을 재정렬하는 데 쓰이는 점수를 어떻게 만들어 내는지 보여줍니다.
flowchart LR
q["query excerpt<br/><i>one opinion passage,<br/>citation removed</i>"]
sl["shortlist from fast search<br/><i>30 candidate passages</i>"]
quest["<b>one Noul</b><br/>could this candidate be<br/>from the cited precedent?<br/><i>criteria fix true and false</i>"]
%% direction LR inside an LR chart keeps each state beside its noul, two columns,
%% so the fan-out is four rows tall instead of eight
subgraph fan["one request per candidate · no request sees another"]
direction LR
d1["state<br/>{query, candidate 1}"] --> n1["noul<br/>0.87"]
d2["state<br/>{query, candidate 2}"] --> n2["noul<br/>0.41"]
dx["⋮"] --> nx["⋮"]
d30["state<br/>{query, candidate 30}"] --> n30["noul<br/>0.12"]
end
sort["sort by noul,<br/>highest first"]
out["re-ranked shortlist<br/><i>same 30, better order</i>"]
q --> fan
sl --> fan
quest --> fan
fan --> sort --> out
%% the elision is not a node - drop its box so it reads as "and so on"
classDef elide fill:none,stroke:none
class dx,nx elide
linkStyle 2 stroke:none
리랭킹 예제
이제 빠른 검색과 리랭킹을 법률 검색 데이터셋인 CLERC에서 실행합니다. 이 예제는 법원 의견 구절 3,565개와 쿼리 40개를 사용합니다.
준비
첫 단계는 이 워크스루가 의존하는 패키지를 설치합니다.
bm25s와datasets가 빠른 검색 후보 목록을 만듭니다.typesafe-sdk와cooksafe가 리랭킹과 API 캐싱을 처리합니다.matplotlib이 결과 차트를 그립니다.
pip install bm25s datasets matplotlib 'cooksafe>=0.2.0,<0.3.0'
다음 블록은 TypeSafe 클라이언트와 워크스루의 나머지가 사용하는 상수를 설정합니다. 어떤 TypeSafe 모델을 호출할지, 빠른 검색이 리랭커에게 넘기는 후보 목록이 얼마나 큰지 등입니다. TypeSafe를 호출하려면 TYPESAFE_API_KEY가 필요합니다.
import hashlib
import json
import os
import random
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
from cooksafe import JsonCache
from IPython.display import display
from typesafe_sdk import Noul, NoulCriteria, TypeSafeClient
TYPESAFE_MODEL = "jev-1.12"
PRICE = (
0.042,
0.00,
) # $ per 1M tokens (input, output); TypeSafe jev-1.12 as of 2026-08
N_ROWS = 170 # CLERC rows pooled into the shared corpus
N_QUERIES = 40 # rows we evaluate
TOP_K = 30 # candidates the shortlist hands to the re-ranker, per query
client = TypeSafeClient(
api_key=os.environ.get(
"TYPESAFE_API_KEY", "cache-only"
), # keyless kernels replay the cache
base_url=os.environ.get("TYPESAFE_ENDPOINT"),
timeout=120.0,
)
json_cache = JsonCache(Path("json_cache.json"))
빠른 검색으로 구절 순위 매기기
여기 사용된 데이터셋은 미국 법원 의견 코퍼스로, 170개 행을 모은 것입니다. 각 행은 이렇게 구성됩니다.
- Query: 인용이 제거된 의견 발췌.
- Gold: 제거된 인용이 가리키던 구절, 즉 쿼리에 대한 유일한 정답.
- Candidates: 코퍼스의 나머지 모든 구절로, 각각 쿼리가 실수로 매칭될 수 있는 것.
170개 행 중 40개를 쿼리로 평가하기 위해 고릅니다. 나머지 130개는 후보로만 등장합니다.
다음 셀은 위에서 설명한 기법으로 후보 목록을 만듭니다:
- 코퍼스를 불러옵니다.
- BM25로 모든 쿼리에 대해 순위를 매깁니다.
여기에는 아직 TypeSafe가 없습니다. 이것은 빠른 검색 단계일 뿐입니다.
CLERC_FILE = (
"https://huggingface.co/datasets/jhu-clsp/CLERC/resolve/main/"
"teva_train_dir/train_data.jsonl.gz"
)
def cid(text: str) -> str:
"""Corpus id: a content hash, so passages shared across queries dedupe."""
return hashlib.sha1(text.encode("utf-8")).hexdigest()[:16]
@json_cache
def build_slice(n_rows: int, n_queries: int, seed: int) -> dict:
"""Stream CLERC rows, pool ``n_rows`` of them into a corpus, pick ``n_queries`` to evaluate."""
from datasets import load_dataset # heavy import, keep local
stream = load_dataset("json", data_files=CLERC_FILE, streaming=True, split="train")
rows = []
for row in stream:
if (
row.get("positive_passages")
and len(row.get("negative_passages") or []) == 20
):
rows.append(row)
if len(rows) >= 1000:
break
rng = random.Random(seed)
picked = rng.sample(rows, n_rows)
corpus, pool = {}, []
for row in picked:
gold = row["positive_passages"][0]["text"]
corpus[cid(gold)] = gold
for neg in row["negative_passages"]:
corpus[cid(neg["text"])] = neg["text"]
pool.append(
{"qid": str(row["query_id"]), "query": row["query"], "gold": cid(gold)}
)
# hold out the first 20 pooled rows; evaluate on the rest
queries = rng.sample(pool[20:], n_queries)
# sort the corpus by id so every run — live or cache replay — iterates it identically
return {"queries": queries, "corpus": dict(sorted(corpus.items()))}
def bm25_rankings(corpus: dict[str, str], queries: dict[str, str], k: int = 100):
"""Rank every passage in the corpus by word overlap with each query."""
import bm25s
cids = list(corpus)
retriever = bm25s.BM25()
retriever.index(bm25s.tokenize([corpus[c] for c in cids], stopwords="en"))
qids = list(queries)
idxs, _ = retriever.retrieve(
bm25s.tokenize([queries[q] for q in qids], stopwords="en"), k=min(k, len(cids))
)
return {q: [cids[i] for i in idxs[row]] for row, q in enumerate(qids)}
def gold_rank(ranked: list[str], gold: str) -> int | None:
"""1-based rank of the gold id, or None if it isn't in the list."""
return ranked.index(gold) + 1 if gold in ranked else None
SURFACE, INK, INK2, MUTED = "#f8f8f2", "#34342f", "#34342f", "#7c7c77"
GRID, AXIS, BLUE, GREEN = "#d8d8cf", "#d8d8cf", "#5d76a2", "#6f9b52"
def bar_chart(labels: list[str], shares: list[float], title: str) -> None:
"""A small single-series bar chart of shares (0-1, shown as percentages)."""
import matplotlib.pyplot as plt
fig, ax = plt.subplots(figsize=(5, 3.2), facecolor=SURFACE)
ax.set_facecolor(SURFACE)
for side in ("top", "right"):
ax.spines[side].set_visible(False)
for side in ("left", "bottom"):
ax.spines[side].set_color(AXIS)
ax.tick_params(colors=MUTED, labelcolor=INK2, labelsize=9)
ax.set_axisbelow(True)
ax.grid(axis="y", color=GRID, linewidth=0.8)
bars = ax.bar(labels, shares, width=0.55, color=[BLUE, GREEN][: len(labels)])
ax.bar_label(
bars,
labels=[f"{s * 100:.0f}%" for s in shares],
padding=4,
color=INK,
fontsize=11,
)
ax.set_ylim(0, 1.1)
ax.set_yticks([0, 0.25, 0.5, 0.75, 1.0])
ax.set_yticklabels(["0%", "25%", "50%", "75%", "100%"])
ax.set_ylabel(f"share of {len(queries)} queries", color=INK2, fontsize=9)
ax.set_title(title, loc="left", color=INK, fontsize=11)
plt.tight_layout()
display(fig)
plt.close(fig)
ds = build_slice(N_ROWS, N_QUERIES, seed=0)
corpus: dict[str, str] = ds["corpus"]
queries = {q["qid"]: q["query"] for q in ds["queries"]}
golds = {q["qid"]: q["gold"] for q in ds["queries"]}
candidates = {q: ranked[:TOP_K] for q, ranked in bm25_rankings(corpus, queries).items()}
in_top_k = sum(golds[q] in candidates[q] for q in queries)
at_rank_1 = sum(candidates[q][0] == golds[q] for q in queries)
bar_chart(
[f"In top {TOP_K}", "At rank 1"],
[in_top_k / len(queries), at_rank_1 / len(queries)],
f"Where the correct passage lands, {len(queries)} queries against {len(corpus):,} candidates",
)
빠른 검색이 정답 구절을 1위로 올릴 가능성은 낮습니다
차트는 3,565개 후보 중에서 빠른 검색이 정답 구절을 어디에 놓는지 보여줍니다.
빠른 검색은 코퍼스를 정답을 포함하는 후보 목록으로 안정적으로 좁힙니다. 40개 쿼리 100%에서 정답을 포함합니다. 하지만 그 구절이 후보 목록에서 1위인 경우는 드물어, 5%에 불과합니다.
아래의 리랭킹은 이미 후보 목록에 있는 상위 30개 후보의 순서만 바꿉니다. 빠른 검색이 선택하지 않은 구절을 추가할 수는 없습니다. 여기서는 후보 목록이 40개 쿼리 모두의 정답 구절을 포함하므로, 리랭킹은 각각을 더 나은 위치에 놓는 데 집중할 수 있습니다.
TypeSafe로 리랭킹하기
리랭킹은 후보 목록의 모든 후보를 그 쿼리에 대해 점수 매기고, 그 점수로 정렬합니다. TypeSafe가 각 쌍에 대해 묻는 질문은 그 후보가 쿼리의 제거된 인용이 가리키는 구절일 수 있는지입니다.
다음 셀은 다음을 합니다:
- 그 질문을 정의합니다.
- 모든 후보 목록의 후보마다 한 번씩, 즉 쿼리 40개 곱하기 후보 30개, 총 1,200번 호출을 차례로가 아니라 동시에 실행합니다.
- TypeSafe가 반환한 점수로 각 후보 목록을 정렬하여 리랭킹된 결과를 만듭니다.
is_cited_source = Noul(
instructions=(
"The query excerpt comes from a US federal court opinion and was written "
"immediately around a citation to a precedent; the citation itself has been "
"removed. Could the candidate passage be from that cited precedent — does it "
"establish the specific legal proposition the query excerpt invokes at its "
"citation point?"
),
criteria=NoulCriteria(
true=(
"The candidate passage states or establishes the specific rule, standard, "
"holding, or fact pattern that the query excerpt attributes to its removed "
"citation."
),
false=(
"The candidate passage is merely on a similar topic or doctrine; it does not "
"supply the specific proposition the query excerpt relies on."
),
),
)
@json_cache
def score_candidate(model: str, query: str, candidate: str, question_json: str) -> dict:
"""One TypeSafe call about one (query, candidate) pair: a noul, plus token usage."""
# the SDK takes a question as its JSON dict, so the cached string decodes straight in
question = json.loads(question_json)
response = client.system_one(
state={"query_excerpt": query, "candidate_passage": candidate},
questions={"is_cited_source": question},
model=model,
)
return {
"noul": response.answers["is_cited_source"].noul,
"input_tokens": response.usage.input_tokens or 0,
"output_tokens": response.usage.output_tokens or 0,
}
# Each of the 40 queries has 30 candidates, so re-ranking every shortlist means 1,200 independent
# calls — cheap enough to fire all at once with a thread pool instead of one after another.
pair_list = [(q, c) for q in queries for c in candidates[q]]
question_json = is_cited_source.model_dump_json(exclude_none=True)
with ThreadPoolExecutor(max_workers=12) as pool:
results = pool.map(
lambda p: score_candidate(
TYPESAFE_MODEL, queries[p[0]], corpus[p[1]], question_json
),
pair_list,
)
pair_scores = {q: {} for q in queries}
for (q, c), result in zip(pair_list, results):
pair_scores[q][c] = result
reranked = {
q: sorted(candidates[q], key=lambda c: -pair_scores[q][c]["noul"]) for q in queries
}
def chart_before_after(
runs: dict[str, dict[str, list[str]]], thresholds: list[int]
) -> None:
"""Grouped bar chart: how often the correct passage lands in the top N, for each run."""
import numpy as np
import matplotlib.pyplot as plt
labels = list(runs)
colors = [BLUE, GREEN]
def share_in_top(rankings, k):
return sum(
gold_rank(rankings[q], golds[q]) in range(1, k + 1) for q in queries
) / len(queries)
fig, ax = plt.subplots(figsize=(6.5, 3.6), facecolor=SURFACE)
ax.set_facecolor(SURFACE)
for side in ("top", "right"):
ax.spines[side].set_visible(False)
for side in ("left", "bottom"):
ax.spines[side].set_color(AXIS)
ax.tick_params(colors=MUTED, labelcolor=INK2, labelsize=9)
ax.set_axisbelow(True)
ax.grid(axis="y", color=GRID, linewidth=0.8)
x = np.arange(len(thresholds))
width = 0.35
for i, (label, rankings) in enumerate(runs.items()):
shares = [share_in_top(rankings, k) for k in thresholds]
offset = (i - (len(labels) - 1) / 2) * width
bars = ax.bar(x + offset, shares, width * 0.92, color=colors[i], label=label)
ax.bar_label(
bars,
labels=[f"{s * 100:.0f}%" for s in shares],
padding=3,
color=INK2,
fontsize=8.5,
)
ax.set_xticks(x, [f"top {k}" for k in thresholds])
ax.set_ylim(0, 1)
ax.set_yticks([0, 0.25, 0.5, 0.75, 1.0])
ax.set_yticklabels(["0%", "25%", "50%", "75%", "100%"])
ax.set_ylabel(f"share of {len(queries)} queries", color=INK2, fontsize=9)
ax.set_title(
"How often the correct passage lands near the top",
loc="left",
color=INK,
fontsize=11,
)
ax.legend(frameon=False, labelcolor=INK2, fontsize=9, loc="upper left")
plt.tight_layout()
display(fig)
plt.close(fig)
chart_before_after(
{"Fast search": candidates, "+ TypeSafe re-rank": reranked}, [1, 5, 10]
)
calls = [pair_scores[q][c] for q in queries for c in pair_scores[q]]
input_tokens = sum(call["input_tokens"] for call in calls)
output_tokens = sum(call["output_tokens"] for call in calls)
cost = input_tokens / 1_000_000 * PRICE[0] + output_tokens / 1_000_000 * PRICE[1]
print(
f"{len(calls)} TypeSafe calls used {input_tokens:,} input and "
f"{output_tokens:,} output tokens, costing ${cost:.4f}."
)
1200 TypeSafe calls used 1,536,002 input and 25,200 output tokens, costing $0.0645.
리랭킹이 정답을 위쪽으로 옮깁니다
차트는 빠른 검색과, 빠른 검색 더하기 리랭킹을 세 임계값에서 비교합니다. 리랭킹은 모든 임계값에서 정답 구절을 위쪽에 더 가깝게 옮깁니다:
- Top 1 — 5% → 18%
- Top 5 — 15% → 35%
- Top 10 — 38% → 62%
보고된 토큰 수와 비용은 40개 후보 목록을 리랭킹하는 데 쓰인 1,200번의 TypeSafe 호출 전체를 포함합니다.
각 CLERC 행은 정답 구절 하나와 부정 구절 20개를 담습니다. 이 워크스루는 170개 행의 구절을 하나의 공유 코퍼스로 모읍니다. 40개 평가 쿼리 각각에 대해, BM25는 그 행과 함께 제공된 20개 부정 구절만이 아니라 전체 코퍼스에서 30개 후보를 선택합니다. 그런 다음 TypeSafe가 각 선택된 후보에 대해 쿼리를 읽고 그 30개 구절을 리랭킹합니다.
이 워크스루는 명료함을 위해 쌍마다 질문 하나를 물었습니다. 실제 애플리케이션은 같은 쌍에 대해 여러 질문을 한 번의 호출로 물을 것입니다. 방법은 병렬 질문 cookbook과 Speculative Fan-Out 패턴을 보십시오.
다음은 무엇인가
같은 구성 요소가 TypeSafe 문서의 다른 곳에도 나타납니다:
- Noul, TypeSafe가 예/아니오 질문을 점수로 바꾸는 방법.
- Speculative Fan-Out, 한 문서에 대해 여러 질문을 한 번의 호출로 묻는 방법.
- Line-by-line Search, 키워드가 아니라 의미로 코퍼스를 검색하는 또 다른 방법.