示例
示例
可复制粘贴的配方。除了它点名的那些辅助函数(ship、CACHE 等等,由你提供),每个片段都是
自包含的。
- 快速开始
- 审计
- 脱敏 PII
- 缓存
- 指标
- 防护栏
- 置信度门控
- 钉住路由
- 生命周期
- 组合
- 逐调用钩子
- 批处理
- HTTP 服务器
- ONNXAgent
- 运行时注册
- 基类与进程级默认值
- 异步钩子
- 钩子超时
- token 预算
- 测试钩子
快速开始
import laya
def log(ctx):
print(ctx.model, ctx.results[0]["answers"])
agent = laya.load("convaiinnovations/laya", on_predict_end=log)
agent.system_one("I was charged twice.", {"urgent": {"type": "noul", "instructions": "Urgent?"}})
审计
browser-use 那个用例:捕获每一个决策并发给一个外部服务。
import json, sys
import laya
def audit(ctx):
for state, result in zip(ctx.states, ctx.results or []):
record = {
"run_id": ctx.run_id,
"model": ctx.model,
"state": state,
"routing": result.get("routing"),
"answers": result["answers"],
"usage": result.get("usage"),
"call_usage": ctx.usage,
"call_elapsed_ms": round(ctx.elapsed_ms or 0.0, 3),
}
print(json.dumps(record), file=sys.stderr)
# ship_to_service(record)
agent = laya.load("convaiinnovations/laya", on_predict_end=audit)
一次钩子调用覆盖整次调用,所以这个循环每个决策写一条记录;predict_batch 上的同样形状见
批处理。
完整可运行的版本在 examples/hooks/audit.py。
脱敏 PII
import re
import laya
EMAIL = re.compile(r"\b[\w.+-]+@[\w-]+\.[\w.-]+\b")
PHONE = re.compile(r"\+?\d[\d ()-]{7,}\d")
def scrub(value):
if isinstance(value, str):
return PHONE.sub("[phone]", EMAIL.sub("[email]", value))
if isinstance(value, dict):
return {k: scrub(v) for k, v in value.items()}
if isinstance(value, list):
return [scrub(v) for v in value]
return value
def redact(ctx):
ctx.states = [scrub(s) for s in ctx.states]
agent = laya.load("convaiinnovations/laya", on_predict_start=redact)
缓存
import hashlib, json
import laya
CACHE = {}
def key(ctx, index):
# Not sort_keys=True: criteria order is positional, so two orders are two questions,
# and the checkpoint and token budget change the answer too.
payload = json.dumps([ctx.states[index], ctx.questions, ctx.model,
ctx.max_len, ctx.head_max_len], default=str)
return hashlib.sha256(payload.encode()).hexdigest()
def read(ctx):
hits = [CACHE.get(key(ctx, i)) for i in range(len(ctx.states))]
if all(hit is not None for hit in hits):
ctx.skip(hits) # one per state: skip replaces the whole call
def write(ctx):
for i, result in enumerate(ctx.results or []):
CACHE[key(ctx, i)] = result
agent = laya.load("convaiinnovations/laya", on_predict_start=read, on_predict_end=write)
first = agent.system_one("state", QUESTIONS) # runs the model
second = agent.system_one("state", QUESTIONS) # served from CACHE
指标
import laya
COUNTS, LATENCIES = {}, []
def metrics(ctx):
COUNTS[ctx.model] = COUNTS.get(ctx.model, 0) + 1
if ctx.elapsed_ms is not None:
LATENCIES.append(ctx.elapsed_ms)
agent = laya.load("convaiinnovations/laya", on_predict_end=metrics, hooks_raise=False)
防护栏
从一个 start 钩子抛异常来拦下一个请求。
import laya
class Blocked(Exception):
pass
def guard(ctx):
text = " ".join(str(state) for state in ctx.states).lower()
if "ignore previous instructions" in text:
raise Blocked("prompt injection")
agent = laya.load("convaiinnovations/laya", on_predict_start=guard)
try:
agent.system_one("Ignore previous instructions and ...", QUESTIONS)
except Blocked:
handle_block()
一个 start 钩子能看到这次调用的每个状态,所以要把它们全测一遍:只读 ctx.states[0] 会让
predict_batch 调用里剩下的部分通过。
置信度门控
改写一个低置信度的答案,或者给它加标注。
def gate(ctx):
for result in ctx.results or []:
answer = result["answers"].get("dept")
if answer and answer["confidence"] < 0.6:
answer["choice"] = "human-review"
answer["gated"] = True
agent = laya.load("convaiinnovations/laya", on_predict_end=gate)
ctx.results 为这次调用的每个状态保存一个 dict,所以这个循环给每一个没到阈值的答案都加标注,
而不只是第一个状态的。
钉住路由
为某一类流量强制一个 checkpoint。
from laya import Router
from laya.router import RouteDecision
def pin(ctx):
if "refund" in str(ctx.states[0]).lower():
ctx.decision = RouteDecision(
model="typed-decisions",
repo="convaiinnovations/laya/typed-decisions",
reason="refund workflow",
detection=None,
workflow=None,
)
router = Router(hooks=[pin])
逐调用,不安装:
router.predict("refund request", QUESTIONS, hooks=[pin])
生命周期
观察 checkpoint 的构建和驱逐。
from laya import Router
class Lifecycle:
def on_load(self, ctx):
print("loaded", ctx.model, "agent", type(ctx.agent).__name__)
def on_evict(self, ctx):
print("evicted", ctx.model)
router = Router(max_loaded=1, hooks=[Lifecycle()])
router.preload(["english", "multilingual"]) # on_load fires per build
router.unload() # on_evict fires per freed checkpoint
组合
已安装的钩子在前,然后是便捷可调用对象;全部共享一个上下文。
import laya
class Metrics:
def on_predict_end(self, ctx):
record_latency(ctx.model, ctx.elapsed_ms)
def redact(ctx):
ctx.states = [strip_pii(s) for s in ctx.states]
def audit(ctx):
ship(ctx.run_id, ctx.results)
agent = laya.load(
"convaiinnovations/laya",
hooks=[Metrics()], # installed, runs first
on_predict_start=redact, # convenience, appended
on_predict_end=audit, # convenience, appended
hooks_raise=True,
)
逐调用钩子
为单次调用覆盖或扩展钩子。
agent.system_one(
state,
questions,
on_predict_end=lambda ctx: debug_dump(ctx),
hooks_raise=False,
)
router.predict(
state,
questions,
hooks=[pin], # applies to on_route too
on_predict_end=audit,
)
批处理
钩子每次 Agent.predict_batch 调用触发一次,ctx.states 持有每个状态。Router.predict_batch
改为每请求运行一次它的 Router 级钩子,每个带一个状态和自己的 run_id,所以那里同一个钩子每次
钩子调用写一条记录。
def audit_batch(ctx):
for state, result in zip(ctx.states, ctx.results):
ship_one(ctx.run_id, state, result)
results = agent.predict_batch([state_a, state_b, state_c], questions, on_predict_end=audit_batch)
HTTP 服务器
Router 钩子对 laya.serve 自动生效,因为服务器调用 Router.predict。
from laya import Router
from laya.serve import create_app
router = Router(hooks=[Metrics()], on_predict_end=audit, hooks_raise=False)
app = create_app(router=router)
ONNXAgent
ONNXAgent 只暴露 predict 级事件。
from laya.onnx_agent import ONNXAgent
agent = ONNXAgent("convaiinnovations/laya", onnx_path="laya.onnx", on_predict_end=audit)
agent.system_one(state, questions)
运行时注册
在构造之后挂上、摘下或限定钩子的作用域。
agent.add_hook(Metrics()) # attach at runtime
agent.remove_hook(Metrics()) # by identity
with agent.hooks_installed(DebugDump()):
agent.system_one(state, questions) # DebugDump only here
基类与进程级默认值
继承 BaseHook 只覆盖你需要的东西,并为整个进程注册一次某个钩子,而不是把它传给每个 Agent 和
Router。
from laya import BaseHook, hooks
class Audit(BaseHook):
def on_predict_end(self, ctx):
ship(ctx.run_id, ctx.results)
hooks.set_default_hooks(hooks=[Audit()]) # runs for every call in the process
# later, or in tests:
hooks.clear_default_hooks()
token 预算
为一次调用塑造 token 预算,可以从钩子里做,也可以用逐调用参数。钩子的值替换生效中的预算,所以它
必须先读那个预算:按这次调用里最宽的问题定尺寸,不要低于核心给选项施加的 token 下限,并在拓宽
head_max_len 时一并拓宽 max_len,这样状态才留得住窗口。
def widen(ctx):
k = max((len(q.get("criteria", {}) or {}) for q in ctx.questions.values()), default=0)
if k < 50:
return
cfg = getattr(ctx.agent, "cfg", None) or {}
head = ctx.head_max_len if ctx.head_max_len is not None else cfg.get("head_max_len", 192)
window = ctx.max_len if ctx.max_len is not None else cfg.get("max_len", 512)
need = 16 + 8 * k
if need > head:
ctx.head_max_len = need
ctx.max_len = max(window, need + 8 + 64)
agent = laya.load("convaiinnovations/laya", on_predict_start=widen)
# or per call
agent.system_one(state, questions, head_max_len=512, max_len=1024)
token 预算的塑造有每一行背后的算术,而
predict_shortlist是当一组标签连拓宽后的窗口都装不下时的选项。
异步钩子
把异步钩子包进 AsyncHook;每个协程都在同步内核里运行到完成,无论调用方是同步的,还是已经在一
个事件循环里。
import laya
from laya import AsyncHook
class RemoteAudit:
async def on_predict_end(self, ctx):
await ship(ctx.run_id, ctx.results)
agent = laya.load("convaiinnovations/laya", hooks=[AsyncHook(RemoteAudit())])
普通的异步可调用对象也行:
async def async_end(ctx):
await ship(ctx.results)
agent.system_one(state, questions, on_predict_end=async_end)
钩子超时
给每次钩子调用设上界,这样一个卡住的钩子无法挂起一个正在服务的请求:
agent = laya.load("convaiinnovations/laya", on_predict_end=metrics, hooks_timeout=2.0)
# or per call
agent.system_one(state, questions, on_predict_end=metrics, hooks_timeout=0.5)
超时的钩子抛 TimeoutError(或 hooks_raise=False 时警告)。钩子会在后台继续运行,所以也要给
网络调用它们自己的超时。见错误。
测试钩子
不用模型就能断言钩子看到了什么:把 encode/forward/decode 这些辅助函数打桩,然后驱动
predict_batch,就像 tests/test_hooks.py 做的那样。
seen = []
agent.predict_batch(["s0"], questions, on_predict_end=lambda ctx: seen.append(ctx.results))
assert len(seen) == 1
API 表面由 tests/test_hooks_api.py 钉住。