문서

지식 그래프 개체 정렬

두 맥주 카탈로그에서 나온 450개 후보 쌍 중 어느 것이 같은 제품을 나타내는지, Score 질문 하나와 어느 필드가 어긋나는지 드러내는 세 개의 Noul로 판단합니다.

지식 그래프에서 핵심 문제는 들어오는 개체가 기존 개체를 중복하는지 판단하는 것이며, 특히 서로 다른 출처의 자연어만 있을 때 그렇습니다. 중복 후보 쌍들이 주어지면, TypeSafe Score 하나로 각 쌍이 중복인지, 아니면 큐레이터의 면밀한 검토가 필요한지 판단합니다.

두 데이터 소스가 같은 대상들의 겹치는 집합을 기술하고 있고, 한쪽의 어느 항목이 다른 쪽의 어느 항목과 같은 것인지 알아야 한다고 합시다. 지식 그래프는 그런 항목을 *개체(entity)*라 부르고, 각 개체에 대해 기록된 사실을 담습니다. 값싸지만 거친 1차 패스가 이미 두 소스를 비교하여 면밀히 볼 가치가 있는 450개 쌍을 골라냈다고 합시다. 남은 일은 각 쌍에 대해 판단을 내리는 것입니다.

두 개체를 부적절하게 병합하는 것이 더 비싼 실수입니다. 어느 한쪽 개체에 대한 모든 사실이 이제 병합된 개체를 기술하게 되고, 어느 한쪽에 연결된 모든 것이 함께 딸려 오기 때문입니다. 나중에 이를 되돌리려면 어느 사실이 어디서 왔는지 따져야 합니다. 매칭을 놓치는 것은 중복을 남길 뿐이므로, 이 판단에는 세 번째 선택지가 필요합니다. 병합해도 안전하지 않고 버려도 안전하지 않은 쌍입니다.

그 판단은 세 가지 결과 각각에 대해 레벨이 하나씩 있는 Score 질문입니다.

  • different product — 두 개체를 연결하지 않고 둡니다
  • related, but possibly not the same — 큐레이터에게 넘겨 결정하게 합니다
  • same product — 병합합니다

Score 질문을 쓰는 이유는 의미 레이블, 즉 score 기준을 각 결과에 직접 붙이고 싶기 때문이며, 중간 결과도 포함해서입니다. Noul 질문은 그 대신 출력에 임계값을 적용하는 방식으로 이를 간접적으로 달성할 수 있고, Choice 질문은 세 결과의 순서 관계를 잃게 됩니다.

다음으로, 고려하려는 개체의 각 필드에 대해 그 필드들이 일치하는지 묻는 Noul 질문을 같은 요청에 함께 실어 보낼 수 있습니다. 이 noul들은 score가 “same product”에도 “different product” 레벨에도 속하지 않을 때 큐레이터에게 더 자세한 정보를 제공합니다.

결국 후보 쌍 하나를 받아 세 결과 중 하나를 반환하는 route()가 만들어지며, 여러분의 데이터에 맞춰 조정해야 하는 임계값은 없습니다.

flowchart LR
    PAIR["one candidate pair<br/><i>both entities, one state</i>"] --> CALL

    subgraph CALL["one request, four questions"]
        direction TB
        S["<b>Score:</b> how do the two relate?<br/>· different product<br/>· related, but possibly not the same<br/>· same product"]
        N["<b>Nouls:</b> one per compared field<br/>· same name?<br/>· same brewery?<br/>· same style?"]
        %% invisible link: without an edge these two share a rank, which in a TB
        %% subgraph puts them side by side instead of stacked
        S ~~~ N
    end

    S --> R{"round to the<br/>nearest level"}
    R -->|"different"| DROP["leave unlinked"]
    R -->|"same"| M["assert sameAs"]
    %% the queue is last so the dotted edge below reaches it without crossing
    %% the arrow into `assert sameAs`
    R -->|"related"| Q["curator queue"]
    N -.->|"which field<br/>they disagree on"| Q

준비

pip install matplotlib ipython 'cooksafe>=0.2.0,<0.3.0'

그다음 TYPESAFE_API_KEY를 설정하십시오. 모든 호출은 cookbook과 함께 제공되는 json_cache.json에 캐시되므로, 다시 렌더링하면 API를 호출하지 않고 게시된 숫자를 재생합니다. 그 파일을 삭제하면 전부 실제로 다시 실행합니다.

아래 숫자는 2026-08-11의 jev-1.12에서 나왔습니다.

import json
import os
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path

import matplotlib
import matplotlib.pyplot as plt
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Noul, Score, TypeSafeClient

matplotlib.use("Agg")  # headless render

TYPESAFE_MODEL = "jev-1.12"
MAX_WORKERS = 6  # small pool; the public endpoint rate-limits above roughly eight

client = TypeSafeClient(
    api_key=os.environ.get(
        "TYPESAFE_API_KEY", "cache-only"
    ),  # keyless kernels replay the cache
    base_url=os.environ.get("TYPESAFE_ENDPOINT"),
    timeout=120.0,
)
json_cache = JsonCache(Path("json_cache.json"))

후보 쌍 불러오기

쌍은 공개된 벤치마크 세트인 Magellan 컬렉션의 Beer 데이터에서 나옵니다. 서로 다른 웹사이트에서 스크레이핑한 두 맥주 카탈로그이며, 그 거친 1차 패스로 이미 450개 쌍으로 줄여 놓았습니다. 각 개체는 네 개 필드를 가집니다. 이름, 양조장, 스타일, 알코올 도수입니다. 각 쌍은 벤치마크 자체의 정답인 known_same_as도 함께 가집니다.

텍스트는 게시된 그대로 두고 전처리하지 않습니다. 문자로 되돌려지지 않은 HTML 엔티티, 별도 단어로 쪼개진 아포스트로피, 잘못 디코딩된 몇몇 문자 등입니다.

쌍마다 요청 하나가 나가므로, 드는 비용은 두 소스 중 어느 쪽의 크기가 아니라 여러분이 받은 쌍의 수를 따릅니다.

PAIRS = json.loads(Path("candidate_pairs.json").read_text(encoding="utf-8"))
BY_ID = {pair["id"]: pair for pair in PAIRS}

print(f"{len(PAIRS)} candidate pairs. The first one, as the model will see it:")
print(json.dumps({k: PAIRS[0][k] for k in ("entity_a", "entity_b")}, indent=2)[:420])
450 candidate pairs. The first one, as the model will see it:
{
  "entity_a": {
    "name": "C N Red Imperial Red Ale",
    "brewery": "Redwood Lodge",
    "style": "American Amber / Red Ale",
    "abv": "8.10 %"
  },
  "entity_b": {
    "name": "Kinetic Infrared Imperial Red Ale",
    "brewery": "Kinetic Brewing Company",
    "style": "American Strong Ale",
    "abv": "9.30 %"
  }
}

후보 쌍마다 Score 질문 하나와 Noul 질문 세 개 묻기

두 개체는 entity_a와 entity_b로 하나의 state에 들어가므로, 질문은 어느 한쪽이 아니라 쌍에 관한 것입니다. 네 질문 모두 하나의 요청에 실려 갑니다.

아래 세 개의 레벨 설명이 곧 결정 전체입니다. 각 레벨이 하나의 결과입니다. 이 파일 어디에도 임계값 상수는 없습니다. 또한 이 설명들은 score를 하나도 보기 전에 작성할 수 있는데, 맞춰 넣어야 하는 숫자라면 그렇지 않습니다.

중간 레벨이 신중하게 써야 할 레벨입니다. 여기서는 변형, 스페셜 에디션, 그리고 둘 중 어느 제품을 가리킬 수도 있는 이름을 포괄하므로, 그런 것들은 병합되거나 버려지는 대신 큐레이터에게 도달합니다.

OUTCOME은 세 결과에 이름을 붙입니다. 병합 결과를 assert sameAs라고 부르는 이유는, sameAs가 두 개체가 같은 것임을 기록하는 표준 방식이고, 그것을 하나 작성하는 것이 곧 병합이 실제로 일어나는 방식이기 때문입니다.

네 필드 중 세 개가 Noul 질문을 받습니다. 이름, 양조장, 스타일입니다. 알코올 도수는 받지 않는데, 두 숫자를 비교하는 것은 산술이기 때문입니다. 필요하면 코드에서 계산하십시오. 이 코드를 다른 종류의 데이터에 쓰려면 QUESTIONS와 LEVELS를 다시 작성하면 됩니다. 맥주에 대해 아는 다른 코드는 결과를 출력하는 두 함수뿐이며, 그것들은 필드 이름을 붙입니다.

LEVELS = [
    "They describe two different products.",
    "They describe closely related products that may or may not be the same one: "
    "a variant, a special edition, or a name that could plausibly refer to either.",
    "They describe one and the same product.",
]
OUTCOME = {0: "leave unlinked", 1: "curator queue", 2: "assert sameAs"}

QUESTIONS = {
    "link_state": Score(
        instructions="How do the two entity descriptions relate as products?",
        criteria=LEVELS,
    ),
    "same_name": Noul(
        instructions="Do the two entities state the same beer name?",
    ),
    "same_brewery": Noul(
        instructions="Are the two entities from the same brewery?",
    ),
    "same_style": Noul(
        instructions="Do the two entities describe the same beer style?",
    ),
}

@json_cache
def score(pair_id: str) -> dict:
    """One request about one candidate pair -> the score plus the three noul answers."""
    pair = BY_ID[pair_id]
    response = client.system_one(
        state={"entity_a": pair["entity_a"], "entity_b": pair["entity_b"]},
        questions=QUESTIONS,
        model=TYPESAFE_MODEL,
    )
    link = response.answers["link_state"]
    return {
        "score": link.score,
        "probabilities": link.probabilities,
        "confidence": link.confidence,
        "properties": {
            k: response.answers[k].noul for k in QUESTIONS if k != "link_state"
        },
        # tokens and requests are the durable units; don't cache a derived cost
        "input_tokens": response.usage.input_tokens or 0,
        "output_tokens": response.usage.output_tokens or 0,
    }

def route(score_value: float) -> str:
    """The whole decision rule: the nearest level names the outcome."""
    return OUTCOME[min(int(score_value + 0.5), len(LEVELS) - 1)]

def show(pair_id: str) -> None:
    pair, result = BY_ID[pair_id], score(pair_id)
    print(
        f"{pair_id}  score {result['score']:.2f}  confidence {result['confidence']:.2f}"
        f"  ->  {route(result['score'])}"
    )
    for side in ("entity_a", "entity_b"):
        e = pair[side]
        print(f"    {e['name'][:44]:<46}{e['brewery'][:30]:<32}{e['style'][:22]}")
    nouls = result["properties"]
    print(
        f"    name {nouls['same_name']:.2f}   brewery {nouls['same_brewery']:.2f}   "
        f"style {nouls['same_style']:.2f}"
    )

네 쌍입니다. c446은 하나의 제품이고 c427은 두 개입니다. 나머지 둘은 서로 다른 이유로 중간 레벨에 들어갑니다. c100은 이름과 양조장이 같지만 소스가 스타일을 다르게 표현하고, c428은 어떤 맥주를 그것의 과일-홉 변형과 짝지었습니다.

for pair_id in ("c446", "c427", "c100", "c428"):
    show(pair_id)
    print()
c446  score 1.94  confidence 0.92  ->  assert sameAs
    Thomas Hooker Old Marley Barleywine           Thomas Hooker Brewing Company   American Barleywine
    Thomas Hooker Old Marley Barleywine           Thomas Hooker Brewing Company   Barley Wine
    name 0.97   brewery 0.99   style 0.81

c427  score 0.03  confidence 0.95  ->  leave unlinked
    Frost Quake Bourbon Barrel Aged Barley Wine   Wellington County Brewery       American Barleywine
    Lompoc Bourbon Barrel Aged Proletariat Red A  Lompoc Brewing                  Amber Ale
    name 0.02   brewery 0.09   style 0.08

c100  score 1.30  confidence 0.27  ->  curator queue
    Belle Gueule Rousse                           Brasseurs R.J.                  American Amber / Red A
    Belle Gueule Rousse                           Brasseurs RJ                    Amber Lager/Vienna
    name 0.95   brewery 0.94   style 0.35

c428  score 1.10  confidence 0.77  ->  curator queue
    Ambleside Amber Ale                           Bridge Brewing Company          American Amber / Red A
    Bridge Ambleside Amber Ale - Pomegranate & G  Bridge Brewing Company          Amber Ale
    name 0.63   brewery 0.98   style 0.74

후보 쌍마다 라우팅하기

# 450 candidate pairs, one request each; a small pool keeps a live run to a few minutes.
with ThreadPoolExecutor(max_workers=MAX_WORKERS) as pool:
    scored = list(pool.map(lambda pair: score(pair["id"]), PAIRS))

scores = [result["score"] for result in scored]
by_outcome: dict[str, list[str]] = {name: [] for name in OUTCOME.values()}
for pair, s in zip(PAIRS, scores):
    by_outcome[route(s)].append(pair["id"])

SURFACE, INK, INK2, MUTED = "#fcfcfb", "#0b0b0b", "#52514e", "#898781"
GRID, AXIS, BLUE, ORANGE = "#e1e0d9", "#c3c2b7", "#2a78d6", "#eb6834"

BINS, TOP = 20, len(LEVELS) - 1
counts = [0] * BINS
for s in scores:
    counts[min(int(s / TOP * BINS), BINS - 1)] += 1
centers = [(i + 0.5) / BINS * TOP for i in range(BINS)]
queued = [c if route(x) == "curator queue" else 0 for c, x in zip(counts, centers)]
settled = [c if route(x) != "curator queue" else 0 for c, x in zip(counts, centers)]

fig, ax = plt.subplots(figsize=(7.2, 3.6), facecolor=SURFACE)
ax.set_facecolor(SURFACE)
for side in ("top", "right"):
    ax.spines[side].set_visible(False)
for side in ("left", "bottom"):
    ax.spines[side].set_color(AXIS)
ax.tick_params(colors=MUTED, labelcolor=INK2, labelsize=9)
ax.set_axisbelow(True)
ax.grid(axis="y", color=GRID, linewidth=0.8)
ax.bar(
    centers, settled, width=TOP / BINS * 0.9, color=BLUE, label="settled automatically"
)
ax.bar(
    centers, queued, width=TOP / BINS * 0.9, color=ORANGE, label="sent to the curator"
)
for edge in (0.5, 1.5):
    ax.axvline(edge, color=INK2, linewidth=1, linestyle="--")
ax.set_xticks([0, 0.5, 1, 1.5, 2])
ax.set_xticklabels(["0\ndifferent", "0.5", "1\nrelated", "1.5", "2\nsame"])
ax.set_xlabel("score for the pair", color=INK2, fontsize=9)
ax.set_ylabel("candidate pairs", color=INK2, fontsize=9)
ax.set_title(
    f"{len(PAIRS)} candidate pairs, scored once each",
    loc="left",
    color=INK,
    fontsize=11,
)
ax.legend(frameon=False, labelcolor=INK2, fontsize=9)
display(fig)
plt.close(fig)

for name in ("assert sameAs", "curator queue", "leave unlinked"):
    n = len(by_outcome[name])
    print(f"{name:<16}{n:>5}  ({n / len(PAIRS):>5.1%})")
assert sameAs      40  ( 8.9%)
curator queue      50  (11.1%)
leave unlinked    360  (80.0%)
output

route()가 답을 바꾸는 두 score 값이 절단점입니다. 대부분의 쌍은 정착합니다. 360개는 아래쪽 절단점 아래로, 40개는 위쪽 절단점 위로 점수가 나오며, 50개가 큐레이터에게 남습니다.

이 세트에서 score는 정수에 깔끔하게 놓이지 않습니다. 대부분 0.25 근처에 위치합니다. 공통점이 전혀 없는 두 맥주도 스타일 이름을 공유할 수 있고, 양조장 이름이 비슷해 보일 수 있으므로, 모델은 중간 레벨에 확률을 전혀 주지 않고 일부를 줍니다. 한 쌍을 결정하는 것은 그것이 절단점의 어느 쪽에 놓이는지입니다. 레벨에 얼마나 가까이 놓이는지는 고려되지 않습니다.

두 절단점은 똑같이 붐비지 않습니다. 9개 쌍이 위쪽 절단점인 1.5의 0.1 이내에 놓이며, 이것이 그래프에 무엇이 병합될지를 결정하는 절단점입니다. 47개가 아래쪽 절단점인 0.5에 그만큼 가까이 놓이며, 이것은 큐레이터가 그 쌍을 보는지 여부만 결정합니다. 어느 숫자도 여러분이 조정하는 값이 아닙니다. 둘 다 레벨을 어떻게 표현했는지에서 나오며, 중간 레벨의 문구가 쌍을 큐레이터와 연결되지 않은 채 남겨진 쌍 사이에서 이동시키는 요인입니다.

playground에서 열기

아래 playground 링크는 score 1.10으로 큐레이터에게 간 c428을 엽니다. 이 쌍은 Ambleside Amber Ale과 Bridge Ambleside Amber Ale - Pomegranate & Galena Hops를 짝짓습니다. 같은 양조장, 같은 알코올 도수입니다. 네 질문이 모두 함께 제공됩니다.

playground_link = make_playground_link(
    {"entity_a": BY_ID["c428"]["entity_a"], "entity_b": BY_ID["c428"]["entity_b"]},
    QUESTIONS,
    models=[TYPESAFE_MODEL],
)
display(
    Markdown(
        f"🔗 [Open this pair + questions in the TypeSafe playground]({playground_link})"
    )
)
TypeSafe playground에서 이 쌍 + 질문 열기 →