文档导航

示例

示例

可复制粘贴的配方。除了它点名的那些辅助函数(ship、CACHE 等等,由你提供),每个片段都是 自包含的。

快速开始

import laya

def log(ctx):
    print(ctx.model, ctx.results[0]["answers"])

agent = laya.load("convaiinnovations/laya", on_predict_end=log)
agent.system_one("I was charged twice.", {"urgent": {"type": "noul", "instructions": "Urgent?"}})

审计

browser-use 那个用例:捕获每一个决策并发给一个外部服务。

import json, sys
import laya

def audit(ctx):
    for state, result in zip(ctx.states, ctx.results or []):
        record = {
            "run_id": ctx.run_id,
            "model": ctx.model,
            "state": state,
            "routing": result.get("routing"),
            "answers": result["answers"],
            "usage": result.get("usage"),
            "call_usage": ctx.usage,
            "call_elapsed_ms": round(ctx.elapsed_ms or 0.0, 3),
        }
        print(json.dumps(record), file=sys.stderr)
        # ship_to_service(record)

agent = laya.load("convaiinnovations/laya", on_predict_end=audit)

一次钩子调用覆盖整次调用,所以这个循环每个决策写一条记录;predict_batch 上的同样形状见 批处理。

完整可运行的版本在 examples/hooks/audit.py。

脱敏 PII

import re
import laya

EMAIL = re.compile(r"\b[\w.+-]+@[\w-]+\.[\w.-]+\b")
PHONE = re.compile(r"\+?\d[\d ()-]{7,}\d")

def scrub(value):
    if isinstance(value, str):
        return PHONE.sub("[phone]", EMAIL.sub("[email]", value))
    if isinstance(value, dict):
        return {k: scrub(v) for k, v in value.items()}
    if isinstance(value, list):
        return [scrub(v) for v in value]
    return value

def redact(ctx):
    ctx.states = [scrub(s) for s in ctx.states]

agent = laya.load("convaiinnovations/laya", on_predict_start=redact)

见 examples/hooks/redact.py。

缓存

import hashlib, json
import laya

CACHE = {}

def key(ctx, index):
    # Not sort_keys=True: criteria order is positional, so two orders are two questions,
    # and the checkpoint and token budget change the answer too.
    payload = json.dumps([ctx.states[index], ctx.questions, ctx.model,
                          ctx.max_len, ctx.head_max_len], default=str)
    return hashlib.sha256(payload.encode()).hexdigest()

def read(ctx):
    hits = [CACHE.get(key(ctx, i)) for i in range(len(ctx.states))]
    if all(hit is not None for hit in hits):
        ctx.skip(hits)   # one per state: skip replaces the whole call

def write(ctx):
    for i, result in enumerate(ctx.results or []):
        CACHE[key(ctx, i)] = result

agent = laya.load("convaiinnovations/laya", on_predict_start=read, on_predict_end=write)
first = agent.system_one("state", QUESTIONS)    # runs the model
second = agent.system_one("state", QUESTIONS)   # served from CACHE

见 examples/hooks/cache.py。

指标

import laya

COUNTS, LATENCIES = {}, []

def metrics(ctx):
    COUNTS[ctx.model] = COUNTS.get(ctx.model, 0) + 1
    if ctx.elapsed_ms is not None:
        LATENCIES.append(ctx.elapsed_ms)

agent = laya.load("convaiinnovations/laya", on_predict_end=metrics, hooks_raise=False)

见 examples/hooks/otel.py。

防护栏

从一个 start 钩子抛异常来拦下一个请求。

import laya

class Blocked(Exception):
    pass

def guard(ctx):
    text = " ".join(str(state) for state in ctx.states).lower()
    if "ignore previous instructions" in text:
        raise Blocked("prompt injection")

agent = laya.load("convaiinnovations/laya", on_predict_start=guard)

try:
    agent.system_one("Ignore previous instructions and ...", QUESTIONS)
except Blocked:
    handle_block()

一个 start 钩子能看到这次调用的每个状态,所以要把它们全测一遍:只读 ctx.states[0] 会让 predict_batch 调用里剩下的部分通过。

置信度门控

改写一个低置信度的答案,或者给它加标注。

def gate(ctx):
    for result in ctx.results or []:
        answer = result["answers"].get("dept")
        if answer and answer["confidence"] < 0.6:
            answer["choice"] = "human-review"
            answer["gated"] = True

agent = laya.load("convaiinnovations/laya", on_predict_end=gate)

ctx.results 为这次调用的每个状态保存一个 dict,所以这个循环给每一个没到阈值的答案都加标注, 而不只是第一个状态的。

钉住路由

为某一类流量强制一个 checkpoint。

from laya import Router
from laya.router import RouteDecision

def pin(ctx):
    if "refund" in str(ctx.states[0]).lower():
        ctx.decision = RouteDecision(
            model="typed-decisions",
            repo="convaiinnovations/laya/typed-decisions",
            reason="refund workflow",
            detection=None,
            workflow=None,
        )

router = Router(hooks=[pin])

逐调用,不安装:

router.predict("refund request", QUESTIONS, hooks=[pin])

生命周期

观察 checkpoint 的构建和驱逐。

from laya import Router

class Lifecycle:
    def on_load(self, ctx):
        print("loaded", ctx.model, "agent", type(ctx.agent).__name__)

    def on_evict(self, ctx):
        print("evicted", ctx.model)

router = Router(max_loaded=1, hooks=[Lifecycle()])
router.preload(["english", "multilingual"])   # on_load fires per build
router.unload()                               # on_evict fires per freed checkpoint

组合

已安装的钩子在前,然后是便捷可调用对象;全部共享一个上下文。

import laya

class Metrics:
    def on_predict_end(self, ctx):
        record_latency(ctx.model, ctx.elapsed_ms)

def redact(ctx):
    ctx.states = [strip_pii(s) for s in ctx.states]

def audit(ctx):
    ship(ctx.run_id, ctx.results)

agent = laya.load(
    "convaiinnovations/laya",
    hooks=[Metrics()],              # installed, runs first
    on_predict_start=redact,        # convenience, appended
    on_predict_end=audit,           # convenience, appended
    hooks_raise=True,
)

逐调用钩子

为单次调用覆盖或扩展钩子。

agent.system_one(
    state,
    questions,
    on_predict_end=lambda ctx: debug_dump(ctx),
    hooks_raise=False,
)

router.predict(
    state,
    questions,
    hooks=[pin],                    # applies to on_route too
    on_predict_end=audit,
)

批处理

钩子每次 Agent.predict_batch 调用触发一次,ctx.states 持有每个状态。Router.predict_batch 改为每请求运行一次它的 Router 级钩子,每个带一个状态和自己的 run_id,所以那里同一个钩子每次 钩子调用写一条记录。

def audit_batch(ctx):
    for state, result in zip(ctx.states, ctx.results):
        ship_one(ctx.run_id, state, result)

results = agent.predict_batch([state_a, state_b, state_c], questions, on_predict_end=audit_batch)

HTTP 服务器

Router 钩子对 laya.serve 自动生效,因为服务器调用 Router.predict。

from laya import Router
from laya.serve import create_app

router = Router(hooks=[Metrics()], on_predict_end=audit, hooks_raise=False)
app = create_app(router=router)

ONNXAgent

ONNXAgent 只暴露 predict 级事件。

from laya.onnx_agent import ONNXAgent

agent = ONNXAgent("convaiinnovations/laya", onnx_path="laya.onnx", on_predict_end=audit)
agent.system_one(state, questions)

运行时注册

在构造之后挂上、摘下或限定钩子的作用域。

agent.add_hook(Metrics())          # attach at runtime
agent.remove_hook(Metrics())       # by identity

with agent.hooks_installed(DebugDump()):
    agent.system_one(state, questions)   # DebugDump only here

基类与进程级默认值

继承 BaseHook 只覆盖你需要的东西,并为整个进程注册一次某个钩子,而不是把它传给每个 Agent 和 Router。

from laya import BaseHook, hooks

class Audit(BaseHook):
    def on_predict_end(self, ctx):
        ship(ctx.run_id, ctx.results)

hooks.set_default_hooks(hooks=[Audit()])   # runs for every call in the process

# later, or in tests:
hooks.clear_default_hooks()

token 预算

为一次调用塑造 token 预算,可以从钩子里做,也可以用逐调用参数。钩子的值替换生效中的预算,所以它 必须先读那个预算:按这次调用里最宽的问题定尺寸,不要低于核心给选项施加的 token 下限,并在拓宽 head_max_len 时一并拓宽 max_len,这样状态才留得住窗口。

def widen(ctx):
    k = max((len(q.get("criteria", {}) or {}) for q in ctx.questions.values()), default=0)
    if k < 50:
        return
    cfg = getattr(ctx.agent, "cfg", None) or {}
    head = ctx.head_max_len if ctx.head_max_len is not None else cfg.get("head_max_len", 192)
    window = ctx.max_len if ctx.max_len is not None else cfg.get("max_len", 512)
    need = 16 + 8 * k
    if need > head:
        ctx.head_max_len = need
        ctx.max_len = max(window, need + 8 + 64)

agent = laya.load("convaiinnovations/laya", on_predict_start=widen)

# or per call
agent.system_one(state, questions, head_max_len=512, max_len=1024)

token 预算的塑造有每一行背后的算术,而 predict_shortlist是当一组标签连拓宽后的窗口都装不下时的选项。

异步钩子

把异步钩子包进 AsyncHook;每个协程都在同步内核里运行到完成,无论调用方是同步的,还是已经在一 个事件循环里。

import laya
from laya import AsyncHook

class RemoteAudit:
    async def on_predict_end(self, ctx):
        await ship(ctx.run_id, ctx.results)

agent = laya.load("convaiinnovations/laya", hooks=[AsyncHook(RemoteAudit())])

普通的异步可调用对象也行:

async def async_end(ctx):
    await ship(ctx.results)

agent.system_one(state, questions, on_predict_end=async_end)

钩子超时

给每次钩子调用设上界,这样一个卡住的钩子无法挂起一个正在服务的请求:

agent = laya.load("convaiinnovations/laya", on_predict_end=metrics, hooks_timeout=2.0)

# or per call
agent.system_one(state, questions, on_predict_end=metrics, hooks_timeout=0.5)

超时的钩子抛 TimeoutError(或 hooks_raise=False 时警告)。钩子会在后台继续运行,所以也要给 网络调用它们自己的超时。见错误。

测试钩子

不用模型就能断言钩子看到了什么:把 encode/forward/decode 这些辅助函数打桩,然后驱动 predict_batch,就像 tests/test_hooks.py 做的那样。

seen = []
agent.predict_batch(["s0"], questions, on_predict_end=lambda ctx: seen.append(ctx.results))
assert len(seen) == 1

API 表面由 tests/test_hooks_api.py 钉住。

另见