ドキュメント

ナレッジグラフのエンティティアラインメント

ナレッジグラフのエンティティアラインメント

2 つのビールカタログの 450 組の候補ペアから同一製品を判定します。1 つの Score と、どのフィールドが食い違うかを示す 3 つの Noul を使います。

ナレッジグラフの重要な問題の 1 つは、入ってくるエンティティが既存のものを重複しているかを 判定することです。特に、異なるソースからの自然言語しか手がかりがない場合にそうです。 重複の可能性があるペアが与えられたとき、1 つの TypeSafe の Score が、各ペアが重複か どうか、あるいはキュレーターによる詳しい確認に値するかどうかを判定します。

2 つのデータソースが同じものの重なり合う集合を記述していて、一方のどのエントリが他方の どのエントリと同じものかを知る必要があるとします。ナレッジグラフはそうしたエントリを エンティティ と呼び、それぞれについて記録された事実を保持します。安価ですが粗い 最初のパスがすでに 2 つのソースを比較し、詳しく見る価値のある 450 組のペアを選び出して います。残るのは、各ペアについて判断を下すことです。

2 つのエンティティを不適切にマージする方が高くつく間違いです。どちらのエンティティに ついての事実も、マージ後のものを記述することになり、どちらかにリンクされたものも すべて付いてきます。後から取り消すには、どの事実がどこから来たかを突き止める必要が あります。一致を見逃しても重複が残るだけなので、判断には 3 つ目の選択肢が必要です。 マージするのも安全でなく、捨てるのも安全でないペアです。

判定は Score の質問で、3 つの結果それぞれに 1 つのレベルがあります。

  • 別の製品 — 2 つのエンティティをリンクしないままにする
  • 関連するが、同じではないかもしれない — キュレーターに判断を委ねる
  • 同じ製品 — マージする

Score の質問を使うのは、セマンティックなラベル(score の criteria)を、中間の結果を 含む各結果に直接結び付けたいからです。Noul の質問ならその出力にしきい値を設けることで 間接的に実現できますが、Choice の質問では 3 つの結果の順序関係が失われます。

次に、考慮したいエンティティの各フィールドについて、それらのフィールドが一致するか どうかを尋ねる Noul の質問を同じリクエストに同乗させられます。これらの noul は、 score が「same product」レベルにも「different product」レベルにも落ちなかった場合に、 キュレーターにより詳しい情報を提供します。

最終的には、1 組の候補ペアを受け取り 3 つの結果のいずれかを返す route() が得られます。 自分のデータに合わせて調整しなければならないしきい値はありません。

flowchart LR
    PAIR["one candidate pair<br/><i>both entities, one state</i>"] --> CALL

    subgraph CALL["one request, four questions"]
        direction TB
        S["<b>Score:</b> how do the two relate?<br/>· different product<br/>· related, but possibly not the same<br/>· same product"]
        N["<b>Nouls:</b> one per compared field<br/>· same name?<br/>· same brewery?<br/>· same style?"]
        %% invisible link: without an edge these two share a rank, which in a TB
        %% subgraph puts them side by side instead of stacked
        S ~~~ N
    end

    S --> R{"round to the<br/>nearest level"}
    R -->|"different"| DROP["leave unlinked"]
    R -->|"same"| M["assert sameAs"]
    %% the queue is last so the dotted edge below reaches it without crossing
    %% the arrow into `assert sameAs`
    R -->|"related"| Q["curator queue"]
    N -.->|"which field<br/>they disagree on"| Q

準備

pip install matplotlib ipython 'cooksafe>=0.2.0,<0.3.0'

次に TYPESAFE_API_KEY を設定します。すべての呼び出しは cookbook に同梱される json_cache.json にキャッシュされるので、再描画すると API を呼ばずに公開済みの数値を 再生します。すべてを実際に再実行するには、このファイルを削除します。

以下の数値は 2026-08-11 の jev-1.12 によるものです。

import json
import os
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path

import matplotlib
import matplotlib.pyplot as plt
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Noul, Score, TypeSafeClient

matplotlib.use("Agg")  # headless render

TYPESAFE_MODEL = "jev-1.12"
MAX_WORKERS = 6  # small pool; the public endpoint rate-limits above roughly eight

client = TypeSafeClient(
    api_key=os.environ.get(
        "TYPESAFE_API_KEY", "cache-only"
    ),  # keyless kernels replay the cache
    base_url=os.environ.get("TYPESAFE_ENDPOINT"),
    timeout=120.0,
)
json_cache = JsonCache(Path("json_cache.json"))

候補ペアを読み込む

ペアは公開ベンチマークセットである Magellan コレクションの Beer データから来ています。 異なるウェブサイトからスクレイピングされた 2 つのビールカタログで、最初の粗いパスに よってすでに 450 組に絞り込まれています。各エンティティは 4 つのフィールドを持ちます。 name、brewery、style、アルコール度数です。各ペアはベンチマーク自身の答えである known_same_as も持ちます。

テキストは公開されたまま、前処理なしで残されています。文字に戻されなかった HTML エンティティ、別々の語として分かれてしまったアポストロフィ、誤ってデコードされた いくつかの文字です。

ペアごとに 1 回のリクエストが送られるので、コストはどちらのソースの大きさではなく、 渡されたペアの数に従います。

PAIRS = json.loads(Path("candidate_pairs.json").read_text(encoding="utf-8"))
BY_ID = {pair["id"]: pair for pair in PAIRS}

print(f"{len(PAIRS)} candidate pairs. The first one, as the model will see it:")
print(json.dumps({k: PAIRS[0][k] for k in ("entity_a", "entity_b")}, indent=2)[:420])
450 candidate pairs. The first one, as the model will see it:
{
  "entity_a": {
    "name": "C N Red Imperial Red Ale",
    "brewery": "Redwood Lodge",
    "style": "American Amber / Red Ale",
    "abv": "8.10 %"
  },
  "entity_b": {
    "name": "Kinetic Infrared Imperial Red Ale",
    "brewery": "Kinetic Brewing Company",
    "style": "American Strong Ale",
    "abv": "9.30 %"
  }
}

候補ペアごとに 1 つの Score の質問と 3 つの Noul の質問をする

両方のエンティティが entity_a と entity_b として 1 つの state に入るので、質問は どちらか片方ではなく ペア についてのものです。4 つすべてが 1 回のリクエストに 同乗します。

以下の 3 つのレベルの説明が判断のすべてです。各レベルが 1 つの結果です。このファイルの どこにもしきい値の定数はありません。これらの説明は、1 つの score も見る前に書くことも できます。これは、調整しなければならない数値には当てはまりません。

注意して書く価値があるのは中間のレベルです。ここでは、変種、特別版、どちらの製品を 指していてもおかしくない名前をカバーするので、それらはマージも破棄もされず キュレーターに届きます。

OUTCOME は 3 つの結果に名前を付けます。マージの結果が assert sameAs と呼ばれるのは、 sameAs が 2 つのエンティティが同じものであることを記録する標準的な方法であり、それを 書くことが実際のマージだからです。

4 つのフィールドのうち 3 つに Noul の質問が付きます。name、brewery、style です。 アルコール度数には付きません。2 つの数値の比較は算術だからです。必要ならコードで計算 します。これを別の種類のデータに使うには、QUESTIONS と LEVELS を書き換えます。 他にビールを知っているコードは、フィールド名を述べる、結果を出力する 2 つの関数だけ です。

LEVELS = [
    "They describe two different products.",
    "They describe closely related products that may or may not be the same one: "
    "a variant, a special edition, or a name that could plausibly refer to either.",
    "They describe one and the same product.",
]
OUTCOME = {0: "leave unlinked", 1: "curator queue", 2: "assert sameAs"}

QUESTIONS = {
    "link_state": Score(
        instructions="How do the two entity descriptions relate as products?",
        criteria=LEVELS,
    ),
    "same_name": Noul(
        instructions="Do the two entities state the same beer name?",
    ),
    "same_brewery": Noul(
        instructions="Are the two entities from the same brewery?",
    ),
    "same_style": Noul(
        instructions="Do the two entities describe the same beer style?",
    ),
}

@json_cache
def score(pair_id: str) -> dict:
    """One request about one candidate pair -> the score plus the three noul answers."""
    pair = BY_ID[pair_id]
    response = client.system_one(
        state={"entity_a": pair["entity_a"], "entity_b": pair["entity_b"]},
        questions=QUESTIONS,
        model=TYPESAFE_MODEL,
    )
    link = response.answers["link_state"]
    return {
        "score": link.score,
        "probabilities": link.probabilities,
        "confidence": link.confidence,
        "properties": {
            k: response.answers[k].noul for k in QUESTIONS if k != "link_state"
        },
        # tokens and requests are the durable units; don't cache a derived cost
        "input_tokens": response.usage.input_tokens or 0,
        "output_tokens": response.usage.output_tokens or 0,
    }

def route(score_value: float) -> str:
    """The whole decision rule: the nearest level names the outcome."""
    return OUTCOME[min(int(score_value + 0.5), len(LEVELS) - 1)]

def show(pair_id: str) -> None:
    pair, result = BY_ID[pair_id], score(pair_id)
    print(
        f"{pair_id}  score {result['score']:.2f}  confidence {result['confidence']:.2f}"
        f"  ->  {route(result['score'])}"
    )
    for side in ("entity_a", "entity_b"):
        e = pair[side]
        print(f"    {e['name'][:44]:<46}{e['brewery'][:30]:<32}{e['style'][:22]}")
    nouls = result["properties"]
    print(
        f"    name {nouls['same_name']:.2f}   brewery {nouls['same_brewery']:.2f}   "
        f"style {nouls['same_style']:.2f}"
    )

4 組のペアです。c446 は 1 つの製品で、c427 は 2 つです。残りの 2 つは別々の理由で 中間のレベルに落ちます。c100 は名前と brewery が同じですが、ソースが style を違う 言葉で書いており、c428 はビールをそのフルーツとホップの変種と組み合わせています。

for pair_id in ("c446", "c427", "c100", "c428"):
    show(pair_id)
    print()
c446  score 1.94  confidence 0.92  ->  assert sameAs
    Thomas Hooker Old Marley Barleywine           Thomas Hooker Brewing Company   American Barleywine
    Thomas Hooker Old Marley Barleywine           Thomas Hooker Brewing Company   Barley Wine
    name 0.97   brewery 0.99   style 0.81

c427  score 0.03  confidence 0.95  ->  leave unlinked
    Frost Quake Bourbon Barrel Aged Barley Wine   Wellington County Brewery       American Barleywine
    Lompoc Bourbon Barrel Aged Proletariat Red A  Lompoc Brewing                  Amber Ale
    name 0.02   brewery 0.09   style 0.08

c100  score 1.30  confidence 0.27  ->  curator queue
    Belle Gueule Rousse                           Brasseurs R.J.                  American Amber / Red A
    Belle Gueule Rousse                           Brasseurs RJ                    Amber Lager/Vienna
    name 0.95   brewery 0.94   style 0.35

c428  score 1.10  confidence 0.77  ->  curator queue
    Ambleside Amber Ale                           Bridge Brewing Company          American Amber / Red A
    Bridge Ambleside Amber Ale - Pomegranate & G  Bridge Brewing Company          Amber Ale
    name 0.63   brewery 0.98   style 0.74

すべての候補ペアをルーティングする

# 450 candidate pairs, one request each; a small pool keeps a live run to a few minutes.
with ThreadPoolExecutor(max_workers=MAX_WORKERS) as pool:
    scored = list(pool.map(lambda pair: score(pair["id"]), PAIRS))

scores = [result["score"] for result in scored]
by_outcome: dict[str, list[str]] = {name: [] for name in OUTCOME.values()}
for pair, s in zip(PAIRS, scores):
    by_outcome[route(s)].append(pair["id"])

SURFACE, INK, INK2, MUTED = "#fcfcfb", "#0b0b0b", "#52514e", "#898781"
GRID, AXIS, BLUE, ORANGE = "#e1e0d9", "#c3c2b7", "#2a78d6", "#eb6834"

BINS, TOP = 20, len(LEVELS) - 1
counts = [0] * BINS
for s in scores:
    counts[min(int(s / TOP * BINS), BINS - 1)] += 1
centers = [(i + 0.5) / BINS * TOP for i in range(BINS)]
queued = [c if route(x) == "curator queue" else 0 for c, x in zip(counts, centers)]
settled = [c if route(x) != "curator queue" else 0 for c, x in zip(counts, centers)]

fig, ax = plt.subplots(figsize=(7.2, 3.6), facecolor=SURFACE)
ax.set_facecolor(SURFACE)
for side in ("top", "right"):
    ax.spines[side].set_visible(False)
for side in ("left", "bottom"):
    ax.spines[side].set_color(AXIS)
ax.tick_params(colors=MUTED, labelcolor=INK2, labelsize=9)
ax.set_axisbelow(True)
ax.grid(axis="y", color=GRID, linewidth=0.8)
ax.bar(
    centers, settled, width=TOP / BINS * 0.9, color=BLUE, label="settled automatically"
)
ax.bar(
    centers, queued, width=TOP / BINS * 0.9, color=ORANGE, label="sent to the curator"
)
for edge in (0.5, 1.5):
    ax.axvline(edge, color=INK2, linewidth=1, linestyle="--")
ax.set_xticks([0, 0.5, 1, 1.5, 2])
ax.set_xticklabels(["0\ndifferent", "0.5", "1\nrelated", "1.5", "2\nsame"])
ax.set_xlabel("score for the pair", color=INK2, fontsize=9)
ax.set_ylabel("candidate pairs", color=INK2, fontsize=9)
ax.set_title(
    f"{len(PAIRS)} candidate pairs, scored once each",
    loc="left",
    color=INK,
    fontsize=11,
)
ax.legend(frameon=False, labelcolor=INK2, fontsize=9)
display(fig)
plt.close(fig)

for name in ("assert sameAs", "curator queue", "leave unlinked"):
    n = len(by_outcome[name])
    print(f"{name:<16}{n:>5}  ({n / len(PAIRS):>5.1%})")
assert sameAs      40  ( 8.9%)
curator queue      50  (11.1%)
leave unlinked    360  (80.0%)
出力

route() が答えを変える 2 つの score 値がカットポイントです。ほとんどのペアは確定 します。360 組は下のカットポイントを下回り、40 組は上のカットポイントを上回り、 50 組がキュレーターに残ります。

このセットでは、score は綺麗に整数に乗りません。ほとんどが 0.25 付近に落ちます。 共通点のない 2 つのビールでも style 名を共有していたり、brewery 名が似て見えたりする ことがあるので、モデルは中間のレベルにゼロではなくいくらかの確率を与えます。ペアを 決めるのは、カットポイントのどちら側に落ちるかです。レベルにどれだけ近いかは関係 ありません。

2 つのカットポイントの混み具合は同じではありません。9 組が上のカットポイント 1.5 の 0.1 以内にあり、これはグラフにマージされるものを決める方です。47 組が下の カットポイント 0.5 のそれだけ近くにあり、これはキュレーターがペアを見るかどうかだけを 決めます。どちらの数値も調整するものではありません。どちらもレベルの書き方から 従うもので、中間レベルの文言が、キュレーター行きとリンクされないままのペアの間で ペアを動かします。

playground で開く

下の playground リンクは c428 を開きます。これは 1.10 を付け、キュレーターに送られ ました。Ambleside Amber Ale を Bridge Ambleside Amber Ale - Pomegranate & Galena Hops と組にしています。brewery が同じで、アルコール度数も同じです。4 つの質問すべてが 付いてきます。

playground_link = make_playground_link(
    {"entity_a": BY_ID["c428"]["entity_a"], "entity_b": BY_ID["c428"]["entity_b"]},
    QUESTIONS,
    models=[TYPESAFE_MODEL],
)
display(
    Markdown(
        f"🔗 [Open this pair + questions in the TypeSafe playground]({playground_link})"
    )
)
TypeSafe playground でこのペアと質問を開く →