ドキュメント

例

例

コピー&ペーストのレシピです。どのスニペットも、名指しするヘルパー(ship、CACHE など。自分で用意します)を除いて自己完結しています。

クイックスタート

import laya

def log(ctx):
    print(ctx.model, ctx.results[0]["answers"])

agent = laya.load("convaiinnovations/laya", on_predict_end=log)
agent.system_one("I was charged twice.", {"urgent": {"type": "noul", "instructions": "Urgent?"}})

監査

browser-use のユースケース:すべての意思決定を取り込み、外部サービスへ送ります。

import json, sys
import laya

def audit(ctx):
    for state, result in zip(ctx.states, ctx.results or []):
        record = {
            "run_id": ctx.run_id,
            "model": ctx.model,
            "state": state,
            "routing": result.get("routing"),
            "answers": result["answers"],
            "usage": result.get("usage"),
            "call_usage": ctx.usage,
            "call_elapsed_ms": round(ctx.elapsed_ms or 0.0, 3),
        }
        print(json.dumps(record), file=sys.stderr)
        # ship_to_service(record)

agent = laya.load("convaiinnovations/laya", on_predict_end=audit)

1 回のフック呼び出しが呼び出し全体を覆うので、ループは意思決定ごとに 1 レコードを書きます。predict_batch での同じ形はバッチを参照してください。

完全な実行可能版は examples/hooks/audit.py にあります。

PII の秘匿化

import re
import laya

EMAIL = re.compile(r"\b[\w.+-]+@[\w-]+\.[\w.-]+\b")
PHONE = re.compile(r"\+?\d[\d ()-]{7,}\d")

def scrub(value):
    if isinstance(value, str):
        return PHONE.sub("[phone]", EMAIL.sub("[email]", value))
    if isinstance(value, dict):
        return {k: scrub(v) for k, v in value.items()}
    if isinstance(value, list):
        return [scrub(v) for v in value]
    return value

def redact(ctx):
    ctx.states = [scrub(s) for s in ctx.states]

agent = laya.load("convaiinnovations/laya", on_predict_start=redact)

examples/hooks/redact.py を参照してください。

キャッシュ

import hashlib, json
import laya

CACHE = {}

def key(ctx, index):
    # Not sort_keys=True: criteria order is positional, so two orders are two questions,
    # and the checkpoint and token budget change the answer too.
    payload = json.dumps([ctx.states[index], ctx.questions, ctx.model,
                          ctx.max_len, ctx.head_max_len], default=str)
    return hashlib.sha256(payload.encode()).hexdigest()

def read(ctx):
    hits = [CACHE.get(key(ctx, i)) for i in range(len(ctx.states))]
    if all(hit is not None for hit in hits):
        ctx.skip(hits)   # one per state: skip replaces the whole call

def write(ctx):
    for i, result in enumerate(ctx.results or []):
        CACHE[key(ctx, i)] = result

agent = laya.load("convaiinnovations/laya", on_predict_start=read, on_predict_end=write)
first = agent.system_one("state", QUESTIONS)    # runs the model
second = agent.system_one("state", QUESTIONS)   # served from CACHE

examples/hooks/cache.py を参照してください。

メトリクス

import laya

COUNTS, LATENCIES = {}, []

def metrics(ctx):
    COUNTS[ctx.model] = COUNTS.get(ctx.model, 0) + 1
    if ctx.elapsed_ms is not None:
        LATENCIES.append(ctx.elapsed_ms)

agent = laya.load("convaiinnovations/laya", on_predict_end=metrics, hooks_raise=False)

examples/hooks/otel.py を参照してください。

ガードレール

start フックから例外を投げてリクエストをブロックします。

import laya

class Blocked(Exception):
    pass

def guard(ctx):
    text = " ".join(str(state) for state in ctx.states).lower()
    if "ignore previous instructions" in text:
        raise Blocked("prompt injection")

agent = laya.load("convaiinnovations/laya", on_predict_start=guard)

try:
    agent.system_one("Ignore previous instructions and ...", QUESTIONS)
except Blocked:
    handle_block()

start フックは呼び出しのすべての state を見るので、それらすべてを検査してください。ctx.states[0] だけを読むと、predict_batch 呼び出しの残りを通してしまいます。

信頼度ゲート

低信頼度の答えを書き換えるか、注釈を付けます。

def gate(ctx):
    for result in ctx.results or []:
        answer = result["answers"].get("dept")
        if answer and answer["confidence"] < 0.6:
            answer["choice"] = "human-review"
            answer["gated"] = True

agent = laya.load("convaiinnovations/laya", on_predict_end=gate)

ctx.results は呼び出しの state ごとに 1 つの dict を持つので、ループはしきい値を外したすべての答えに注釈を付け、最初の state のものだけにはとどまりません。

ルーティングの固定

ある種のトラフィックに対してチェックポイントを強制します。

from laya import Router
from laya.router import RouteDecision

def pin(ctx):
    if "refund" in str(ctx.states[0]).lower():
        ctx.decision = RouteDecision(
            model="typed-decisions",
            repo="convaiinnovations/laya/typed-decisions",
            reason="refund workflow",
            detection=None,
            workflow=None,
        )

router = Router(hooks=[pin])

インストールせずに呼び出しごとに:

router.predict("refund request", QUESTIONS, hooks=[pin])

ライフサイクル

チェックポイントの構築と追い出しを観察します。

from laya import Router

class Lifecycle:
    def on_load(self, ctx):
        print("loaded", ctx.model, "agent", type(ctx.agent).__name__)

    def on_evict(self, ctx):
        print("evicted", ctx.model)

router = Router(max_loaded=1, hooks=[Lifecycle()])
router.preload(["english", "multilingual"])   # on_load fires per build
router.unload()                               # on_evict fires per freed checkpoint

合成

インストールされたフックが先、次に便宜 callable。すべてが 1 つのコンテキストを共有します。

import laya

class Metrics:
    def on_predict_end(self, ctx):
        record_latency(ctx.model, ctx.elapsed_ms)

def redact(ctx):
    ctx.states = [strip_pii(s) for s in ctx.states]

def audit(ctx):
    ship(ctx.run_id, ctx.results)

agent = laya.load(
    "convaiinnovations/laya",
    hooks=[Metrics()],              # installed, runs first
    on_predict_start=redact,        # convenience, appended
    on_predict_end=audit,           # convenience, appended
    hooks_raise=True,
)

呼び出しごとのフック

単一の呼び出しのためにフックを上書きまたは拡張します。

agent.system_one(
    state,
    questions,
    on_predict_end=lambda ctx: debug_dump(ctx),
    hooks_raise=False,
)

router.predict(
    state,
    questions,
    hooks=[pin],                    # applies to on_route too
    on_predict_end=audit,
)

バッチ

フックは Agent.predict_batch 呼び出しごとに 1 回発火し、ctx.states がすべての state を保持します。Router.predict_batch は代わりに Router レベルのフックをリクエストごとに 1 回走らせ、それぞれが 1 つの state と独自の run_id を持ちます。なのでそこの同じフックは、フックの呼び出しごとに 1 レコードを書きます。

def audit_batch(ctx):
    for state, result in zip(ctx.states, ctx.results):
        ship_one(ctx.run_id, state, result)

results = agent.predict_batch([state_a, state_b, state_c], questions, on_predict_end=audit_batch)

HTTP サーバー

サーバーが Router.predict を呼ぶので、Router のフックは laya.serve に対して自動的に発火します。

from laya import Router
from laya.serve import create_app

router = Router(hooks=[Metrics()], on_predict_end=audit, hooks_raise=False)
app = create_app(router=router)

ONNXAgent

ONNXAgent は predict レベルのイベントだけを公開します。

from laya.onnx_agent import ONNXAgent

agent = ONNXAgent("convaiinnovations/laya", onnx_path="laya.onnx", on_predict_end=audit)
agent.system_one(state, questions)

実行時登録

構築後にフックを付け、外し、スコープします。

agent.add_hook(Metrics())          # attach at runtime
agent.remove_hook(Metrics())       # by identity

with agent.hooks_installed(DebugDump()):
    agent.system_one(state, questions)   # DebugDump only here

基底クラスとプロセス全体の既定値

BaseHook をサブクラス化して必要なものだけを上書きし、すべての Agent と Router に渡すのではなくプロセス全体に一度登録します。

from laya import BaseHook, hooks

class Audit(BaseHook):
    def on_predict_end(self, ctx):
        ship(ctx.run_id, ctx.results)

hooks.set_default_hooks(hooks=[Audit()])   # runs for every call in the process

# later, or in tests:
hooks.clear_default_hooks()

トークン予算

1 回の呼び出しのトークン予算を、フックまたは呼び出しごとの引数から形成します。フックの値は有効な予算を置き換えるので、まずその予算を読む必要があります。呼び出しの最も広い質問でサイズを決め、コアが選択肢に適用するトークンの床を上回って保ち、state が窓を保てるよう head_max_len とともに max_len を広げてください。

def widen(ctx):
    k = max((len(q.get("criteria", {}) or {}) for q in ctx.questions.values()), default=0)
    if k < 50:
        return
    cfg = getattr(ctx.agent, "cfg", None) or {}
    head = ctx.head_max_len if ctx.head_max_len is not None else cfg.get("head_max_len", 192)
    window = ctx.max_len if ctx.max_len is not None else cfg.get("max_len", 512)
    need = 16 + 8 * k
    if need > head:
        ctx.head_max_len = need
        ctx.max_len = max(window, need + 8 + 64)

agent = laya.load("convaiinnovations/laya", on_predict_start=widen)

# or per call
agent.system_one(state, questions, head_max_len=512, max_len=1024)

各行の背後にある算術はトークン予算の形成にあり、ラベル集合が広げた窓にも収まらないときの手段が predict_shortlist です。

非同期フック

非同期フックを AsyncHook で包みます。呼び出し元が同期でもすでにイベントループの中でも、各コルーチンは同期のコア内で完走します。

import laya
from laya import AsyncHook

class RemoteAudit:
    async def on_predict_end(self, ctx):
        await ship(ctx.run_id, ctx.results)

agent = laya.load("convaiinnovations/laya", hooks=[AsyncHook(RemoteAudit())])

素の非同期 callable も動きます。

async def async_end(ctx):
    await ship(ctx.results)

agent.system_one(state, questions, on_predict_end=async_end)

フックのタイムアウト

各フック呼び出しを制限し、詰まったフックが配信中のリクエストをハングさせないようにします。

agent = laya.load("convaiinnovations/laya", on_predict_end=metrics, hooks_timeout=2.0)

# or per call
agent.system_one(state, questions, on_predict_end=metrics, hooks_timeout=0.5)

タイムアウトしたフックは TimeoutError を投げます(hooks_raise=False なら警告します)。フックはバックグラウンドで走り続けるので、ネットワーク呼び出しにも独自のタイムアウトを与えてください。エラーを参照してください。

フックのテスト

モデルなしでフックが見たものを表明します。encode/forward/decode のヘルパーをスタブした predict_batch を駆動します。tests/test_hooks.py がそうしています。

seen = []
agent.predict_batch(["s0"], questions, on_predict_end=lambda ctx: seen.append(ctx.results))
assert len(seen) == 1

API の表面は tests/test_hooks_api.py で固定されています。

関連項目