文档导航

知识图谱实体对齐

知识图谱实体对齐

用一个 Score 问题加三个配套 Noul,判断两份啤酒目录中的 450 个候选对里哪些描述的是同一款产品,并呈现是哪些字段不一致。

知识图谱里的一个关键问题是:判断一个外来实体是否与已有实体重复, 尤其是在手头只有来自不同来源的自然语言时。给定若干可能是重复的候选对,一个 TypeSafe Score 就能判断每一对是不是重复,或者是否值得让策展人再细看。

假设有两个数据源,描述的是同一批事物中相互重叠的部分,而你需要知道 一侧的哪一条与另一侧的哪一条是同一个东西。知识图谱把这些条目称作实体, 并保存记录下来的关于每个实体的事实。某个便宜但粗糙的初筛已经比对过两个来源, 挑出了 450 个值得细看的候选对。剩下的就是对每一对做出判断。

错误地合并两个实体代价更大,因为此后关于任一实体的事实都变成了在描述合并后的实体, 任何与二者之一相连的东西也会一并跟过来。事后要撤销,就得弄清哪条事实来自哪里。 漏掉一个匹配只是留下一条重复,所以这个判断需要一个第三选项: 既不安全到可以合并、也不安全到可以丢弃的那些对。

这个判断是一个 Score 问题,三个结果各对应一个档位:

  • 不同产品 —— 两个实体不建立关联
  • 相关,但可能并非同一个 —— 交给策展人决定
  • 同一产品 —— 合并它们

我们用 Score 问题,是因为我们想把一个语义标签——也就是 score 的 criteria—— 直接贴在每个结果上,中间那个结果也不例外。换成 Noul 问题,就只能靠对它的输出取阈值 来间接做到;换成 Choice 问题,则会丢掉这三个结果之间的顺序关系。

接下来,对我们想考察的实体的每个字段,都可以在同一次请求里顺带问一个关于该 字段是否匹配的 Noul 问题。如果 score 既没落在 “same product” 也没落在 “different product” 档位,这些 noul 就能给策展人提供更详细的信息。

最终你会得到一个 route(),它接收一个候选对,返回三个结果之一, 不需要你针对自己的数据去拟合任何阈值。

flowchart LR
    PAIR["one candidate pair<br/><i>both entities, one state</i>"] --> CALL

    subgraph CALL["one request, four questions"]
        direction TB
        S["<b>Score:</b> how do the two relate?<br/>· different product<br/>· related, but possibly not the same<br/>· same product"]
        N["<b>Nouls:</b> one per compared field<br/>· same name?<br/>· same brewery?<br/>· same style?"]
        %% invisible link: without an edge these two share a rank, which in a TB
        %% subgraph puts them side by side instead of stacked
        S ~~~ N
    end

    S --> R{"round to the<br/>nearest level"}
    R -->|"different"| DROP["leave unlinked"]
    R -->|"same"| M["assert sameAs"]
    %% the queue is last so the dotted edge below reaches it without crossing
    %% the arrow into `assert sameAs`
    R -->|"related"| Q["curator queue"]
    N -.->|"which field<br/>they disagree on"| Q

准备工作

pip install matplotlib ipython 'cooksafe>=0.2.0,<0.3.0'

然后设置 TYPESAFE_API_KEY。每次调用都会缓存到随 cookbook 附带的 json_cache.json, 所以重新渲染时会回放已发布的数字,不调用 API。删掉该文件即可把所有内容重新实时跑一遍。

下面的数字来自 2026-08-11 的 jev-1.12。

import json
import os
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path

import matplotlib
import matplotlib.pyplot as plt
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Noul, Score, TypeSafeClient

matplotlib.use("Agg")  # headless render

TYPESAFE_MODEL = "jev-1.12"
MAX_WORKERS = 6  # small pool; the public endpoint rate-limits above roughly eight

client = TypeSafeClient(
    api_key=os.environ.get(
        "TYPESAFE_API_KEY", "cache-only"
    ),  # keyless kernels replay the cache
    base_url=os.environ.get("TYPESAFE_ENDPOINT"),
    timeout=120.0,
)
json_cache = JsonCache(Path("json_cache.json"))

加载候选对

这些候选对来自一个已发布的基准数据集——Magellan 合集里的 Beer 数据: 两份从不同网站抓取的啤酒目录,已经被前面那趟粗筛削减到 450 对。每个实体带四个字段: name、brewery、style 和酒精含量。每一对还带有一个 known_same_as, 也就是基准自己的答案。

文本原样保留,不做预处理:有些 HTML 实体从未还原成字符,撇号被拆成了单独的 词,还有几个字符解码错了。

每一对发出一次请求,所以你的花销取决于拿到多少对, 而不是任一来源的大小。

PAIRS = json.loads(Path("candidate_pairs.json").read_text(encoding="utf-8"))
BY_ID = {pair["id"]: pair for pair in PAIRS}

print(f"{len(PAIRS)} candidate pairs. The first one, as the model will see it:")
print(json.dumps({k: PAIRS[0][k] for k in ("entity_a", "entity_b")}, indent=2)[:420])
450 candidate pairs. The first one, as the model will see it:
{
  "entity_a": {
    "name": "C N Red Imperial Red Ale",
    "brewery": "Redwood Lodge",
    "style": "American Amber / Red Ale",
    "abv": "8.10 %"
  },
  "entity_b": {
    "name": "Kinetic Infrared Imperial Red Ale",
    "brewery": "Kinetic Brewing Company",
    "style": "American Strong Ale",
    "abv": "9.30 %"
  }
}

每个候选对问一个 Score 问题和三个 Noul 问题

两个实体作为 entity_a 和 entity_b 放进同一个状态,所以问题是关于这一对的, 而不是单独关于任何一侧。四个问题都在同一次请求里。

下面这三条档位描述就是整个决策:每个档位对应一个结果。 这个文件里任何地方都没有阈值常量。而且你可以在看到任何一个 score 之前就把这些描述写好, 这一点是那些需要你去拟合的数字做不到的。

中间那个档位值得仔细写。在这里,它覆盖了变体、特别版,以及说得通、可能指任一产品 的名字,于是这些会被送到策展人那里,而不是被合并或丢弃。

OUTCOME 给三个结果起了名字。合并那个结果叫 assert sameAs,因为 sameAs 是记录「两个实体是同一个东西」的标准方式,而写出这样一条,正是合并真正发生的方式。

四个字段里有三个得到 Noul 问题:name、brewery 和 style。酒精含量没有, 因为比较两个数字属于算术;需要的话在代码里算就行。 要把它用到另一类数据上,你只要改写 QUESTIONS 和 LEVELS。 此外只有打印结果的那两个函数知道「啤酒」这件事,因为里面点了这些字段名。

LEVELS = [
    "They describe two different products.",
    "They describe closely related products that may or may not be the same one: "
    "a variant, a special edition, or a name that could plausibly refer to either.",
    "They describe one and the same product.",
]
OUTCOME = {0: "leave unlinked", 1: "curator queue", 2: "assert sameAs"}

QUESTIONS = {
    "link_state": Score(
        instructions="How do the two entity descriptions relate as products?",
        criteria=LEVELS,
    ),
    "same_name": Noul(
        instructions="Do the two entities state the same beer name?",
    ),
    "same_brewery": Noul(
        instructions="Are the two entities from the same brewery?",
    ),
    "same_style": Noul(
        instructions="Do the two entities describe the same beer style?",
    ),
}

@json_cache
def score(pair_id: str) -> dict:
    """One request about one candidate pair -> the score plus the three noul answers."""
    pair = BY_ID[pair_id]
    response = client.system_one(
        state={"entity_a": pair["entity_a"], "entity_b": pair["entity_b"]},
        questions=QUESTIONS,
        model=TYPESAFE_MODEL,
    )
    link = response.answers["link_state"]
    return {
        "score": link.score,
        "probabilities": link.probabilities,
        "confidence": link.confidence,
        "properties": {
            k: response.answers[k].noul for k in QUESTIONS if k != "link_state"
        },
        # tokens and requests are the durable units; don't cache a derived cost
        "input_tokens": response.usage.input_tokens or 0,
        "output_tokens": response.usage.output_tokens or 0,
    }

def route(score_value: float) -> str:
    """The whole decision rule: the nearest level names the outcome."""
    return OUTCOME[min(int(score_value + 0.5), len(LEVELS) - 1)]

def show(pair_id: str) -> None:
    pair, result = BY_ID[pair_id], score(pair_id)
    print(
        f"{pair_id}  score {result['score']:.2f}  confidence {result['confidence']:.2f}"
        f"  ->  {route(result['score'])}"
    )
    for side in ("entity_a", "entity_b"):
        e = pair[side]
        print(f"    {e['name'][:44]:<46}{e['brewery'][:30]:<32}{e['style'][:22]}")
    nouls = result["properties"]
    print(
        f"    name {nouls['same_name']:.2f}   brewery {nouls['same_brewery']:.2f}   "
        f"style {nouls['same_style']:.2f}"
    )

四对。c446 是同一款产品,c427 是两款。另外两对因为不同的原因落在中间档位: c100 名字和酒厂都相同,但两个来源对它的风格说法不同;c428 则把一款啤酒和它的 水果加啤酒花变体配成了一对。

for pair_id in ("c446", "c427", "c100", "c428"):
    show(pair_id)
    print()
c446  score 1.94  confidence 0.92  ->  assert sameAs
    Thomas Hooker Old Marley Barleywine           Thomas Hooker Brewing Company   American Barleywine
    Thomas Hooker Old Marley Barleywine           Thomas Hooker Brewing Company   Barley Wine
    name 0.97   brewery 0.99   style 0.81

c427  score 0.03  confidence 0.95  ->  leave unlinked
    Frost Quake Bourbon Barrel Aged Barley Wine   Wellington County Brewery       American Barleywine
    Lompoc Bourbon Barrel Aged Proletariat Red A  Lompoc Brewing                  Amber Ale
    name 0.02   brewery 0.09   style 0.08

c100  score 1.30  confidence 0.27  ->  curator queue
    Belle Gueule Rousse                           Brasseurs R.J.                  American Amber / Red A
    Belle Gueule Rousse                           Brasseurs RJ                    Amber Lager/Vienna
    name 0.95   brewery 0.94   style 0.35

c428  score 1.10  confidence 0.77  ->  curator queue
    Ambleside Amber Ale                           Bridge Brewing Company          American Amber / Red A
    Bridge Ambleside Amber Ale - Pomegranate & G  Bridge Brewing Company          Amber Ale
    name 0.63   brewery 0.98   style 0.74

给每个候选对路由

# 450 candidate pairs, one request each; a small pool keeps a live run to a few minutes.
with ThreadPoolExecutor(max_workers=MAX_WORKERS) as pool:
    scored = list(pool.map(lambda pair: score(pair["id"]), PAIRS))

scores = [result["score"] for result in scored]
by_outcome: dict[str, list[str]] = {name: [] for name in OUTCOME.values()}
for pair, s in zip(PAIRS, scores):
    by_outcome[route(s)].append(pair["id"])

SURFACE, INK, INK2, MUTED = "#fcfcfb", "#0b0b0b", "#52514e", "#898781"
GRID, AXIS, BLUE, ORANGE = "#e1e0d9", "#c3c2b7", "#2a78d6", "#eb6834"

BINS, TOP = 20, len(LEVELS) - 1
counts = [0] * BINS
for s in scores:
    counts[min(int(s / TOP * BINS), BINS - 1)] += 1
centers = [(i + 0.5) / BINS * TOP for i in range(BINS)]
queued = [c if route(x) == "curator queue" else 0 for c, x in zip(counts, centers)]
settled = [c if route(x) != "curator queue" else 0 for c, x in zip(counts, centers)]

fig, ax = plt.subplots(figsize=(7.2, 3.6), facecolor=SURFACE)
ax.set_facecolor(SURFACE)
for side in ("top", "right"):
    ax.spines[side].set_visible(False)
for side in ("left", "bottom"):
    ax.spines[side].set_color(AXIS)
ax.tick_params(colors=MUTED, labelcolor=INK2, labelsize=9)
ax.set_axisbelow(True)
ax.grid(axis="y", color=GRID, linewidth=0.8)
ax.bar(
    centers, settled, width=TOP / BINS * 0.9, color=BLUE, label="settled automatically"
)
ax.bar(
    centers, queued, width=TOP / BINS * 0.9, color=ORANGE, label="sent to the curator"
)
for edge in (0.5, 1.5):
    ax.axvline(edge, color=INK2, linewidth=1, linestyle="--")
ax.set_xticks([0, 0.5, 1, 1.5, 2])
ax.set_xticklabels(["0\ndifferent", "0.5", "1\nrelated", "1.5", "2\nsame"])
ax.set_xlabel("score for the pair", color=INK2, fontsize=9)
ax.set_ylabel("candidate pairs", color=INK2, fontsize=9)
ax.set_title(
    f"{len(PAIRS)} candidate pairs, scored once each",
    loc="left",
    color=INK,
    fontsize=11,
)
ax.legend(frameon=False, labelcolor=INK2, fontsize=9)
display(fig)
plt.close(fig)

for name in ("assert sameAs", "curator queue", "leave unlinked"):
    n = len(by_outcome[name])
    print(f"{name:<16}{n:>5}  ({n / len(PAIRS):>5.1%})")
assert sameAs      40  ( 8.9%)
curator queue      50  (11.1%)
leave unlinked    360  (80.0%)
output

route() 改变答案的那两个 score 值就是切分点。大多数候选对都能定下来: 360 对低于下切分点,40 对高于上切分点,剩下 50 对交给策展人。

在这个数据集上,score 并不整齐地落在整数上。大多数落在 0.25 附近。两款毫无 共同点的啤酒可能仍共用一个风格名,酒厂名也可能看着相似,于是模型会把一些概率分给 中间档位,而不是一点不给。决定一对结果的是它落在切分点的哪一侧, 它离某个档位有多近并不参与其中。

两个切分点周围的拥挤程度并不一样。有九对落在上切分点 1.5 的 0.1 范围内, 这个点决定的是哪些会被合并进图谱。有四十七对落在下切分点 0.5 的 0.1 范围内, 而这个点只决定策展人要不要看到这一对。这两个数字都不是你调出来的。 两者都取决于你如何措辞这些档位,而中间档位的措辞,正是决定候选对是进策展人手里、 还是留在未关联一边的关键。

在 playground 中打开

下面的 playground 链接打开的是 c428,它得分 1.10,去了策展人那里。 它把 Ambleside Amber Ale 和 Bridge Ambleside Amber Ale - Pomegranate & Galena Hops 配成一对:同一家酒厂,同样的酒精含量。四个问题都随链接一起带上。

playground_link = make_playground_link(
    {"entity_a": BY_ID["c428"]["entity_a"], "entity_b": BY_ID["c428"]["entity_b"]},
    QUESTIONS,
    models=[TYPESAFE_MODEL],
)
display(
    Markdown(
        f"🔗 [Open this pair + questions in the TypeSafe playground]({playground_link})"
    )
)
在 TypeSafe playground 中打开这一对和这些问题 →