知識圖譜實體對齊
用一個 Score 問題加三個配套 Noul,判斷兩份啤酒目錄中的 450 個候選對裡哪些描述的是同一款產品,並呈現是哪些欄位不一致。
知識圖譜裡的一個關鍵問題是:判斷一個外來實體是否與已有實體重複,
尤其是在手頭只有來自不同來源的自然語言時。給定若干可能是重複的候選對,一個 TypeSafe Score
就能判斷每一對是不是重複,或者是否值得讓策展人再細看。
假設有兩個資料來源,描述的是同一批事物中相互重疊的部分,而你需要知道 一側的哪一條與另一側的哪一條是同一個東西。知識圖譜把這些條目稱作實體, 並儲存記錄下來的關於每個實體的事實。某個便宜但粗糙的初篩已經比對過兩個來源, 挑出了 450 個值得細看的候選對。剩下的就是對每一對做出判斷。
錯誤地合併兩個實體代價更大,因為此後關於任一實體的事實都變成了在描述合併後的實體, 任何與二者之一相連的東西也會一併跟過來。事後要撤銷,就得弄清哪條事實來自哪裡。 漏掉一個匹配只是留下一條重複,所以這個判斷需要一個第三選項: 既不安全到可以合併、也不安全到可以丟棄的那些對。
這個判斷是一個 Score 問題,三個結果各對應一個檔位:
- 不同產品 —— 兩個實體不建立關聯
- 相關,但可能並非同一個 —— 交給策展人決定
- 同一產品 —— 合併它們
我們用 Score 問題,是因為我們想把一個語義標籤——也就是 score 的 criteria—— 直接貼在每個結果上,中間那個結果也不例外。換成 Noul 問題,就只能靠對它的輸出取閾值 來間接做到;換成 Choice 問題,則會丟掉這三個結果之間的順序關係。
接下來,對我們想考察的實體的每個欄位,都可以在同一次請求裡順帶問一個關於該
欄位是否匹配的 Noul 問題。如果 score 既沒落在 “same product” 也沒落在
“different product” 檔位,這些 noul 就能給策展人提供更詳細的資訊。
最終你會得到一個 route(),它接收一個候選對,返回三個結果之一,
不需要你針對自己的資料去擬合任何閾值。
flowchart LR
PAIR["one candidate pair<br/><i>both entities, one state</i>"] --> CALL
subgraph CALL["one request, four questions"]
direction TB
S["<b>Score:</b> how do the two relate?<br/>· different product<br/>· related, but possibly not the same<br/>· same product"]
N["<b>Nouls:</b> one per compared field<br/>· same name?<br/>· same brewery?<br/>· same style?"]
%% invisible link: without an edge these two share a rank, which in a TB
%% subgraph puts them side by side instead of stacked
S ~~~ N
end
S --> R{"round to the<br/>nearest level"}
R -->|"different"| DROP["leave unlinked"]
R -->|"same"| M["assert sameAs"]
%% the queue is last so the dotted edge below reaches it without crossing
%% the arrow into `assert sameAs`
R -->|"related"| Q["curator queue"]
N -.->|"which field<br/>they disagree on"| Q
準備工作
pip install matplotlib ipython 'cooksafe>=0.2.0,<0.3.0'
然後設定 TYPESAFE_API_KEY。每次呼叫都會快取到隨 cookbook 附帶的 json_cache.json,
所以重新渲染時會回放已釋出的數字,不呼叫 API。刪掉該檔案即可把所有內容重新即時跑一遍。
下面的數字來自 2026-08-11 的 jev-1.12。
import json
import os
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
import matplotlib
import matplotlib.pyplot as plt
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Noul, Score, TypeSafeClient
matplotlib.use("Agg") # headless render
TYPESAFE_MODEL = "jev-1.12"
MAX_WORKERS = 6 # small pool; the public endpoint rate-limits above roughly eight
client = TypeSafeClient(
api_key=os.environ.get(
"TYPESAFE_API_KEY", "cache-only"
), # keyless kernels replay the cache
base_url=os.environ.get("TYPESAFE_ENDPOINT"),
timeout=120.0,
)
json_cache = JsonCache(Path("json_cache.json"))
載入候選對
這些候選對來自一個已釋出的基準資料集——Magellan 合集裡的 Beer 資料:
兩份從不同網站抓取的啤酒目錄,已經被前面那趟粗篩削減到 450 對。每個實體帶四個欄位:
name、brewery、style 和酒精含量。每一對還帶有一個 known_same_as,
也就是基準自己的答案。
文本原樣保留,不做預處理:有些 HTML 實體從未還原成字元,撇號被拆成了單獨的 詞,還有幾個字元解碼錯了。
每一對發出一次請求,所以你的花銷取決於拿到多少對, 而不是任一來源的大小。
PAIRS = json.loads(Path("candidate_pairs.json").read_text(encoding="utf-8"))
BY_ID = {pair["id"]: pair for pair in PAIRS}
print(f"{len(PAIRS)} candidate pairs. The first one, as the model will see it:")
print(json.dumps({k: PAIRS[0][k] for k in ("entity_a", "entity_b")}, indent=2)[:420])
450 candidate pairs. The first one, as the model will see it:
{
"entity_a": {
"name": "C N Red Imperial Red Ale",
"brewery": "Redwood Lodge",
"style": "American Amber / Red Ale",
"abv": "8.10 %"
},
"entity_b": {
"name": "Kinetic Infrared Imperial Red Ale",
"brewery": "Kinetic Brewing Company",
"style": "American Strong Ale",
"abv": "9.30 %"
}
}
每個候選對問一個 Score 問題和三個 Noul 問題
兩個實體作為 entity_a 和 entity_b 放進同一個狀態,所以問題是關於這一對的,
而不是單獨關於任何一側。四個問題都在同一次請求裡。
下面這三條檔位描述就是整個決策:每個檔位對應一個結果。 這個檔案裡任何地方都沒有閾值常量。而且你可以在看到任何一個 score 之前就把這些描述寫好, 這一點是那些需要你去擬合的數字做不到的。
中間那個檔位值得仔細寫。在這裡,它覆蓋了變體、特別版,以及說得通、可能指任一產品 的名字,於是這些會被送到策展人那裡,而不是被合併或丟棄。
OUTCOME 給三個結果起了名字。合併那個結果叫 assert sameAs,因為
sameAs 是記錄「兩個實體是同一個東西」的標準方式,而寫出這樣一條,正是合併真正發生的方式。
四個欄位裡有三個得到 Noul 問題:name、brewery 和 style。酒精含量沒有,
因為比較兩個數字屬於算術;需要的話在程式碼裡算就行。
要把它用到另一類資料上,你只要改寫 QUESTIONS 和 LEVELS。
此外只有列印結果的那兩個函式知道「啤酒」這件事,因為裡面點了這些欄位名。
LEVELS = [
"They describe two different products.",
"They describe closely related products that may or may not be the same one: "
"a variant, a special edition, or a name that could plausibly refer to either.",
"They describe one and the same product.",
]
OUTCOME = {0: "leave unlinked", 1: "curator queue", 2: "assert sameAs"}
QUESTIONS = {
"link_state": Score(
instructions="How do the two entity descriptions relate as products?",
criteria=LEVELS,
),
"same_name": Noul(
instructions="Do the two entities state the same beer name?",
),
"same_brewery": Noul(
instructions="Are the two entities from the same brewery?",
),
"same_style": Noul(
instructions="Do the two entities describe the same beer style?",
),
}
@json_cache
def score(pair_id: str) -> dict:
"""One request about one candidate pair -> the score plus the three noul answers."""
pair = BY_ID[pair_id]
response = client.system_one(
state={"entity_a": pair["entity_a"], "entity_b": pair["entity_b"]},
questions=QUESTIONS,
model=TYPESAFE_MODEL,
)
link = response.answers["link_state"]
return {
"score": link.score,
"probabilities": link.probabilities,
"confidence": link.confidence,
"properties": {
k: response.answers[k].noul for k in QUESTIONS if k != "link_state"
},
# tokens and requests are the durable units; don't cache a derived cost
"input_tokens": response.usage.input_tokens or 0,
"output_tokens": response.usage.output_tokens or 0,
}
def route(score_value: float) -> str:
"""The whole decision rule: the nearest level names the outcome."""
return OUTCOME[min(int(score_value + 0.5), len(LEVELS) - 1)]
def show(pair_id: str) -> None:
pair, result = BY_ID[pair_id], score(pair_id)
print(
f"{pair_id} score {result['score']:.2f} confidence {result['confidence']:.2f}"
f" -> {route(result['score'])}"
)
for side in ("entity_a", "entity_b"):
e = pair[side]
print(f" {e['name'][:44]:<46}{e['brewery'][:30]:<32}{e['style'][:22]}")
nouls = result["properties"]
print(
f" name {nouls['same_name']:.2f} brewery {nouls['same_brewery']:.2f} "
f"style {nouls['same_style']:.2f}"
)
四對。c446 是同一款產品,c427 是兩款。另外兩對因為不同的原因落在中間檔位:
c100 名字和酒廠都相同,但兩個來源對它的風格說法不同;c428 則把一款啤酒和它的
水果加啤酒花變體配成了一對。
for pair_id in ("c446", "c427", "c100", "c428"):
show(pair_id)
print()
c446 score 1.94 confidence 0.92 -> assert sameAs
Thomas Hooker Old Marley Barleywine Thomas Hooker Brewing Company American Barleywine
Thomas Hooker Old Marley Barleywine Thomas Hooker Brewing Company Barley Wine
name 0.97 brewery 0.99 style 0.81
c427 score 0.03 confidence 0.95 -> leave unlinked
Frost Quake Bourbon Barrel Aged Barley Wine Wellington County Brewery American Barleywine
Lompoc Bourbon Barrel Aged Proletariat Red A Lompoc Brewing Amber Ale
name 0.02 brewery 0.09 style 0.08
c100 score 1.30 confidence 0.27 -> curator queue
Belle Gueule Rousse Brasseurs R.J. American Amber / Red A
Belle Gueule Rousse Brasseurs RJ Amber Lager/Vienna
name 0.95 brewery 0.94 style 0.35
c428 score 1.10 confidence 0.77 -> curator queue
Ambleside Amber Ale Bridge Brewing Company American Amber / Red A
Bridge Ambleside Amber Ale - Pomegranate & G Bridge Brewing Company Amber Ale
name 0.63 brewery 0.98 style 0.74
給每個候選對路由
# 450 candidate pairs, one request each; a small pool keeps a live run to a few minutes.
with ThreadPoolExecutor(max_workers=MAX_WORKERS) as pool:
scored = list(pool.map(lambda pair: score(pair["id"]), PAIRS))
scores = [result["score"] for result in scored]
by_outcome: dict[str, list[str]] = {name: [] for name in OUTCOME.values()}
for pair, s in zip(PAIRS, scores):
by_outcome[route(s)].append(pair["id"])
SURFACE, INK, INK2, MUTED = "#fcfcfb", "#0b0b0b", "#52514e", "#898781"
GRID, AXIS, BLUE, ORANGE = "#e1e0d9", "#c3c2b7", "#2a78d6", "#eb6834"
BINS, TOP = 20, len(LEVELS) - 1
counts = [0] * BINS
for s in scores:
counts[min(int(s / TOP * BINS), BINS - 1)] += 1
centers = [(i + 0.5) / BINS * TOP for i in range(BINS)]
queued = [c if route(x) == "curator queue" else 0 for c, x in zip(counts, centers)]
settled = [c if route(x) != "curator queue" else 0 for c, x in zip(counts, centers)]
fig, ax = plt.subplots(figsize=(7.2, 3.6), facecolor=SURFACE)
ax.set_facecolor(SURFACE)
for side in ("top", "right"):
ax.spines[side].set_visible(False)
for side in ("left", "bottom"):
ax.spines[side].set_color(AXIS)
ax.tick_params(colors=MUTED, labelcolor=INK2, labelsize=9)
ax.set_axisbelow(True)
ax.grid(axis="y", color=GRID, linewidth=0.8)
ax.bar(
centers, settled, width=TOP / BINS * 0.9, color=BLUE, label="settled automatically"
)
ax.bar(
centers, queued, width=TOP / BINS * 0.9, color=ORANGE, label="sent to the curator"
)
for edge in (0.5, 1.5):
ax.axvline(edge, color=INK2, linewidth=1, linestyle="--")
ax.set_xticks([0, 0.5, 1, 1.5, 2])
ax.set_xticklabels(["0\ndifferent", "0.5", "1\nrelated", "1.5", "2\nsame"])
ax.set_xlabel("score for the pair", color=INK2, fontsize=9)
ax.set_ylabel("candidate pairs", color=INK2, fontsize=9)
ax.set_title(
f"{len(PAIRS)} candidate pairs, scored once each",
loc="left",
color=INK,
fontsize=11,
)
ax.legend(frameon=False, labelcolor=INK2, fontsize=9)
display(fig)
plt.close(fig)
for name in ("assert sameAs", "curator queue", "leave unlinked"):
n = len(by_outcome[name])
print(f"{name:<16}{n:>5} ({n / len(PAIRS):>5.1%})")
assert sameAs 40 ( 8.9%)
curator queue 50 (11.1%)
leave unlinked 360 (80.0%)
route() 改變答案的那兩個 score 值就是切分點。大多數候選對都能定下來:
360 對低於下切分點,40 對高於上切分點,剩下 50 對交給策展人。
在這個資料集上,score 並不整齊地落在整數上。大多數落在 0.25 附近。兩款毫無 共同點的啤酒可能仍共用一個風格名,酒廠名也可能看著相似,於是模型會把一些機率分給 中間檔位,而不是一點不給。決定一對結果的是它落在切分點的哪一側, 它離某個檔位有多近並不參與其中。
兩個切分點周圍的擁擠程度並不一樣。有九對落在上切分點 1.5 的 0.1 範圍內, 這個點決定的是哪些會被合併進圖譜。有四十七對落在下切分點 0.5 的 0.1 範圍內, 而這個點只決定策展人要不要看到這一對。這兩個數字都不是你調出來的。 兩者都取決於你如何措辭這些檔位,而中間檔位的措辭,正是決定候選對是進策展人手裡、 還是留在未關聯一邊的關鍵。
在 playground 中開啟
下面的 playground 連結開啟的是 c428,它得分 1.10,去了策展人那裡。
它把 Ambleside Amber Ale 和 Bridge Ambleside Amber Ale - Pomegranate & Galena
Hops 配成一對:同一家酒廠,同樣的酒精含量。四個問題都隨連結一起帶上。
playground_link = make_playground_link(
{"entity_a": BY_ID["c428"]["entity_a"], "entity_b": BY_ID["c428"]["entity_b"]},
QUESTIONS,
models=[TYPESAFE_MODEL],
)
display(
Markdown(
f"🔗 [Open this pair + questions in the TypeSafe playground]({playground_link})"
)
)
在 TypeSafe playground 中開啟這一對和這些問題 →