事前パース済みの値の抽出
事前パース済みの値の抽出
正規表現でメール、電話番号、金額の候補を見つけ、TypeSafe に要求された span を選ばせて、コードが逐語的な値を正規化できるようにします。
正規表現が候補の値を見つけ、TypeSafe が質問の求めているものを選び、 コードがそれを逐語的にコピーします。
ここで紹介する find と pick の組は、自分のドキュメントに向けられるものです。3 つの実例でその使い方を示します:送信者が領収書の送付先として希望するアドレス、+14155550177 という電話番号、そして請求としてフラグが立てられた 1315.50 USD という請求書の合計額です。
TypeSafe は渡された選択肢の中から 1 つを選ぶので、まず候補を見つけておく必要があります。正規表現が見つけ、TypeSafe が 1 つ選び、コードがその選択をコピーする、という 3 つのステップです:
- 正規表現がテキスト内の候補の値を見つけます。過剰に見つけるように調整します。
- TypeSafe が質問の求めている候補を選び、コードが後段で必要とする属性(通貨、国、金額がクレジットか請求か)を読み取ります。
- コードが選ばれた値をコピーし、正規化します。
TypeSafe は正規表現が見つけた span の中からしか選ばないので、返ってくる値はそれらの span の 1 つをそのままコピーしたものです。値をでっち上げたり、桁を入れ替えたりすることはできません。
正規表現がドキュメント内の候補の値を見つけ、TypeSafe が 1 つ選び、後段のコードがそれを正規化して処理します。
準備
pip install ipython phonenumbers 'cooksafe>=0.2.0,<0.3.0'
次に TYPESAFE_API_KEY を設定します。
import os
import re
from decimal import Decimal
from pathlib import Path
import phonenumbers
from cooksafe import JsonCache, make_playground_link
from IPython.display import Markdown, display
from typesafe_sdk import Choice, Noul, TypeSafeClient
TYPESAFE_MODEL = "jev-1.12"
NONE = "none" # the escape hatch on every selection: "none of the candidates fits"
# base_url defaults to https://api.typesafe.ai/ ; the env override points at another deployment.
ts = TypeSafeClient(
api_key=os.environ.get(
"TYPESAFE_API_KEY", "cache-only"
), # cached re-renders need no key
base_url=os.environ.get("TYPESAFE_BASE_URL"),
timeout=30.0,
)
json_cache = JsonCache(Path("json_cache.json"))
ヘルパー
find は過剰に見つけるように調整した正規表現を実行し、一致を重複排除します。pick は、find が返す span を選択肢とする Choice の質問です。したがってその答えは、それらの span の 1 つを正確にコピーしたものか、どの候補も当てはまらない場合は none です。classify は固定のラベル集合に対する Choice の質問で、ここでは通貨と国に使います。
is_true は Noul で、ここでは金額がクレジットかどうかを尋ねるのに使います。
すべての呼び出しは json_cache.json にキャッシュされるので、再レンダリングしても API 呼び出しは発生しません。
EMAIL_RE = re.compile(r"[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}")
PHONE_RE = re.compile(r"\(?\+?\d[\d\s()\-.]{6,}\d")
MONEY_RE = re.compile(r"[$€£¥]\s?\d[\d,]*(?:\.\d{2})?")
def find(pattern: re.Pattern, text: str) -> list[str]:
"""Code-side candidate finder: recall-tuned regex, deduped, in document order."""
seen: set[str] = set()
out: list[str] = []
for match in pattern.findall(text):
span = match.strip()
if span and span not in seen:
seen.add(span)
out.append(span)
return out
@json_cache
def pick(document: str, candidates: list[str], question: str) -> dict:
"""TypeSafe selects which found span plays the role. Returns {choice, confidence}.
The options ARE the candidate spans, so ``choice`` is a verbatim copy of one of them (or the
``none`` hatch) - the model chooses, code owns the string."""
criteria = {c: None for c in candidates} | {
NONE: "None of these is the requested value."
}
answer = ts.system_one(
state=document,
questions={"pick": Choice(instructions=question, criteria=criteria)},
model=TYPESAFE_MODEL,
).answers["pick"]
return {"choice": answer.choice, "confidence": answer.confidence}
@json_cache
def classify(document: str, question: str, options: list[str]) -> dict:
"""A small Choice over a fixed label set (currency, country, ...). Returns {choice, confidence}."""
answer = ts.system_one(
state=document,
questions={
"q": Choice(instructions=question, criteria={o: None for o in options})
},
model=TYPESAFE_MODEL,
).answers["q"]
return {"choice": answer.choice, "confidence": answer.confidence}
@json_cache
def is_true(document: str, question: str) -> float:
"""A yes/no Noul. Returns P(yes)."""
return (
ts.system_one(
state=document,
questions={"q": Noul(instructions=question)},
model=TYPESAFE_MODEL,
)
.answers["q"]
.noul
)
メール:役割に応じて正しいアドレスを選ぶ
ヘッダーに 4 つのアドレスがあります。本文では、領収書を To: の請求用エイリアスではなく個人のアドレスに送るよう求めているので、答えは本文を読むかどうかにかかっています。ここでの質問は 2 つ:どのアドレスが領収書を受け取るか、そしてどれがメッセージを送ったかです。
EMAIL_DOC = """From: Dana Whit <dana.whit@acme-corp.com>
To: billing@acme-corp.com
Cc: orders@acme-corp.com
Reply-To: dana.personal@gmail.com
Hi team - please don't use the billing alias for this one. Send my receipt to my
personal address instead. Thanks, Dana."""
emails = find(EMAIL_RE, EMAIL_DOC)
receipt = pick(
EMAIL_DOC, emails, "Which email address does the sender want their receipt sent to?"
)
sender = pick(
EMAIL_DOC, emails, "Which email address did this message come from (the From line)?"
)
print("candidates :", emails)
# code copies the picked value verbatim and normalizes (lowercase); it never re-types it
print(
f"receipt -> : {receipt['choice'].lower():<28} (conf {receipt['confidence']:.2f})"
)
print(f"sender -> : {sender['choice'].lower():<28} (conf {sender['confidence']:.2f})")
candidates : ['dana.whit@acme-corp.com', 'billing@acme-corp.com', 'orders@acme-corp.com', 'dana.personal@gmail.com']
receipt -> : dana.personal@gmail.com (conf 0.98)
sender -> : dana.whit@acme-corp.com (conf 1.00)
receipt は Reply-To: 行にある個人の Gmail アドレスで、本文が求めているものです。sender は From 行にあるものです。どちらも正規表現の一致をコピーし、コードで小文字化したものです。
電話:携帯番号を選び、E.164 に正規化する
3 つの番号があり、いずれも国コードを含んでいません。TypeSafe が携帯番号を選び、テキストから国を読み取ります。phonenumbers がその 2 つの答えを組み合わせて E.164、つまり + と国コードで始まる国際形式にします。
PHONE_DOC = """Reach our San Francisco office at these numbers: main desk (415) 555-0199,
billing fax (415) 555-0142, and my direct cell (415) 555-0177. Call the cell if it's urgent."""
phones = find(PHONE_RE, PHONE_DOC)
mobile = pick(PHONE_DOC, phones, "Which of these is the direct mobile / cell number?")
region = classify(
PHONE_DOC,
"In what country is this office located?",
["US", "GB", "DE", "FR", "CA", "AU"],
)
# code copies the picked value and normalizes it with the model-supplied country
parsed = phonenumbers.parse(mobile["choice"], region["choice"])
e164 = phonenumbers.format_number(parsed, phonenumbers.PhoneNumberFormat.E164)
print("candidates :", phones)
print(f"mobile -> : {mobile['choice']} (conf {mobile['confidence']:.2f})")
print(f"country -> : {region['choice']} (conf {region['confidence']:.2f})")
print(f"E.164 -> : {e164}")
candidates : ['(415) 555-0199', '(415) 555-0142', '(415) 555-0177']
mobile -> : (415) 555-0177 (conf 1.00)
country -> : US (conf 0.90)
E.164 -> : +14155550177
数字そのものは、どれが携帯番号か、どの国にあるかを示していません。それを示すのは周囲の言葉です。TypeSafe がその言葉を読み、phonenumbers が選ばれた番号を +14155550177 としてフォーマットします。
金額:金額を選び、通貨を分類し、クレジットか請求かを判定する
4 つの金額が載った請求書です。TypeSafe が支払総額とクレジットを選び、通貨を読み取り、選ばれた各金額に請求かクレジットかのフラグを立てます。コードは選ばれた各文字列をコピーし、Decimal にパースします。
MONEY_DOC = """Invoice INV-2087.
Subtotal: $1,200.00
Sales tax: $115.50
Total due: $1,315.50
A $50.00 courtesy credit from last month has already been applied."""
amounts = find(MONEY_RE, MONEY_DOC)
currency = classify(
MONEY_DOC,
"What currency are these amounts in?",
["USD", "EUR", "GBP", "JPY", "CAD"],
)
total = pick(MONEY_DOC, amounts, "Which amount is the total the customer must pay?")
credit = pick(
MONEY_DOC, amounts, "Which amount is the courtesy credit that was applied?"
)
def to_decimal(value: str) -> Decimal:
"""Copy the picked value and parse the number in code (US grouping/decimal here)."""
return Decimal(re.sub(r"[^\d.]", "", value))
for label, chosen in [("total due", total), ("credit", credit)]:
is_credit = is_true(
MONEY_DOC,
f"Is the amount {chosen['choice']} a credit or refund to the customer, not a charge?",
)
kind = "credit" if is_credit > 0.5 else "charge"
print(
f"{label:<10}: {chosen['choice']:<10} -> {to_decimal(chosen['choice'])} {currency['choice']} "
f"({kind}, P(credit)={is_credit:.2f})"
)
print("\ncandidates :", amounts)
total due : $1,315.50 -> 1315.50 USD (charge, P(credit)=0.01)
credit : $50.00 -> 50.00 USD (credit, P(credit)=0.99)
candidates : ['$1,200.00', '$115.50', '$1,315.50', '$50.00']
支払総額は $1,315.50、クレジットは $50.00 で、どちらも USD です。クレジットか請求かを判定する Noul は、総額に対して 0.01、クレジットに対して 0.99 を返すので、コードはパースする各 Decimal の符号を把握できます。
to_decimalは、カンマが桁区切りでドットが小数点だと仮定しています。これは$1,315.50には当てはまりますが、€1.315,50では逆になります。ドキュメントがどちらの慣習を使っているかをNoulの質問で尋ね、コードでそれに応じて分岐してください。
TypeSafe playground で開く
ブラウザでメールスレッドを開く共有リンクです。領収書の質問が含まれ、正規表現が見つけた 4 つのアドレスがその選択肢に含まれています。
receipt_criteria = {e: None for e in emails} | {
NONE: "None of these is the requested value."
}
playground_link = make_playground_link(
EMAIL_DOC,
{
"receipt": Choice(
instructions="Which email address does the sender want their receipt sent to?",
criteria=receipt_criteria,
)
},
models=[TYPESAFE_MODEL],
)
display(
Markdown(
f"🔗 [Open this thread + selection in the TypeSafe playground]({playground_link})"
)
)
TypeSafe playground でこのスレッドと選択を開く →
2 つの制約
Choiceの質問は最大 255 個の選択肢を扱えます。候補がそれより多い場合は、2 段階で絞り込みます:まずセクションを選び、次にその中の span を選びます。- 手間がかかるのは候補を見つける部分です。メール、電話番号、金額にはそれらをカバーする正規表現がありますが、名前にはありません。そのため名前の候補は、すでに持っている名簿から取るか、固有表現認識器か、候補を提案する LLM から取る必要があります。その後 TypeSafe が、質問の求めているものを選びます。