Документация

Примеры

Рецепты для копирования и вставки. Каждый фрагмент самодостаточен, кроме helpers, которые он называет (ship, CACHE и так далее), — их предоставляете вы.

Быстрый старт

import laya

def log(ctx):
    print(ctx.model, ctx.results[0]["answers"])

agent = laya.load("convaiinnovations/laya", on_predict_end=log)
agent.system_one("I was charged twice.", {"urgent": {"type": "noul", "instructions": "Urgent?"}})

Аудит

Сценарий использования browser-use: захватывать каждое решение и отправлять его во внешний сервис.

import json, sys
import laya

def audit(ctx):
    for state, result in zip(ctx.states, ctx.results or []):
        record = {
            "run_id": ctx.run_id,
            "model": ctx.model,
            "state": state,
            "routing": result.get("routing"),
            "answers": result["answers"],
            "usage": result.get("usage"),
            "call_usage": ctx.usage,
            "call_elapsed_ms": round(ctx.elapsed_ms or 0.0, 3),
        }
        print(json.dumps(record), file=sys.stderr)
        # ship_to_service(record)

agent = laya.load("convaiinnovations/laya", on_predict_end=audit)

Один вызов hook охватывает весь вызов, поэтому цикл пишет одну запись на каждое решение; см. Пакетная обработка для той же формы на predict_batch.

Полная исполняемая версия — в examples/hooks/audit.py.

Сокрытие PII

import re
import laya

EMAIL = re.compile(r"\b[\w.+-]+@[\w-]+\.[\w.-]+\b")
PHONE = re.compile(r"\+?\d[\d ()-]{7,}\d")

def scrub(value):
    if isinstance(value, str):
        return PHONE.sub("[phone]", EMAIL.sub("[email]", value))
    if isinstance(value, dict):
        return {k: scrub(v) for k, v in value.items()}
    if isinstance(value, list):
        return [scrub(v) for v in value]
    return value

def redact(ctx):
    ctx.states = [scrub(s) for s in ctx.states]

agent = laya.load("convaiinnovations/laya", on_predict_start=redact)

См. examples/hooks/redact.py.

Кэш

import hashlib, json
import laya

CACHE = {}

def key(ctx, index):
    # Not sort_keys=True: criteria order is positional, so two orders are two questions,
    # and the checkpoint and token budget change the answer too.
    payload = json.dumps([ctx.states[index], ctx.questions, ctx.model,
                          ctx.max_len, ctx.head_max_len], default=str)
    return hashlib.sha256(payload.encode()).hexdigest()

def read(ctx):
    hits = [CACHE.get(key(ctx, i)) for i in range(len(ctx.states))]
    if all(hit is not None for hit in hits):
        ctx.skip(hits)   # one per state: skip replaces the whole call

def write(ctx):
    for i, result in enumerate(ctx.results or []):
        CACHE[key(ctx, i)] = result

agent = laya.load("convaiinnovations/laya", on_predict_start=read, on_predict_end=write)
first = agent.system_one("state", QUESTIONS)    # runs the model
second = agent.system_one("state", QUESTIONS)   # served from CACHE

См. examples/hooks/cache.py.

Метрики

import laya

COUNTS, LATENCIES = {}, []

def metrics(ctx):
    COUNTS[ctx.model] = COUNTS.get(ctx.model, 0) + 1
    if ctx.elapsed_ms is not None:
        LATENCIES.append(ctx.elapsed_ms)

agent = laya.load("convaiinnovations/laya", on_predict_end=metrics, hooks_raise=False)

См. examples/hooks/otel.py.

Ограничитель

Заблокируйте запрос, вызвав исключение из hook начала.

import laya

class Blocked(Exception):
    pass

def guard(ctx):
    text = " ".join(str(state) for state in ctx.states).lower()
    if "ignore previous instructions" in text:
        raise Blocked("prompt injection")

agent = laya.load("convaiinnovations/laya", on_predict_start=guard)

try:
    agent.system_one("Ignore previous instructions and ...", QUESTIONS)
except Blocked:
    handle_block()

Hook начала видит каждое состояние вызова, поэтому проверяйте их все: чтение только ctx.states[0] пропускает остальные состояния вызова predict_batch.

Порог уверенности

Перепишите ответ с низкой уверенностью или пометьте его.

def gate(ctx):
    for result in ctx.results or []:
        answer = result["answers"].get("dept")
        if answer and answer["confidence"] < 0.6:
            answer["choice"] = "human-review"
            answer["gated"] = True

agent = laya.load("convaiinnovations/laya", on_predict_end=gate)

ctx.results содержит один словарь на каждое состояние вызова, поэтому цикл помечает каждый ответ, не дотягивающий до порога, а не только ответ первого состояния.

Фиксация маршрутизации

Принудительно назначьте чекпойнт для класса трафика.

from laya import Router
from laya.router import RouteDecision

def pin(ctx):
    if "refund" in str(ctx.states[0]).lower():
        ctx.decision = RouteDecision(
            model="typed-decisions",
            repo="convaiinnovations/laya/typed-decisions",
            reason="refund workflow",
            detection=None,
            workflow=None,
        )

router = Router(hooks=[pin])

На вызов, без установки:

router.predict("refund request", QUESTIONS, hooks=[pin])

Жизненный цикл

Наблюдайте за сборкой и вытеснением чекпойнтов.

from laya import Router

class Lifecycle:
    def on_load(self, ctx):
        print("loaded", ctx.model, "agent", type(ctx.agent).__name__)

    def on_evict(self, ctx):
        print("evicted", ctx.model)

router = Router(max_loaded=1, hooks=[Lifecycle()])
router.preload(["english", "multilingual"])   # on_load fires per build
router.unload()                               # on_evict fires per freed checkpoint

Композиция

Сначала установленные hooks, затем удобные callable-объекты; все разделяют один контекст.

import laya

class Metrics:
    def on_predict_end(self, ctx):
        record_latency(ctx.model, ctx.elapsed_ms)

def redact(ctx):
    ctx.states = [strip_pii(s) for s in ctx.states]

def audit(ctx):
    ship(ctx.run_id, ctx.results)

agent = laya.load(
    "convaiinnovations/laya",
    hooks=[Metrics()],              # installed, runs first
    on_predict_start=redact,        # convenience, appended
    on_predict_end=audit,           # convenience, appended
    hooks_raise=True,
)

Hooks на вызов

Переопределите или расширьте hooks для одного вызова.

agent.system_one(
    state,
    questions,
    on_predict_end=lambda ctx: debug_dump(ctx),
    hooks_raise=False,
)

router.predict(
    state,
    questions,
    hooks=[pin],                    # applies to on_route too
    on_predict_end=audit,
)

Пакетная обработка

Hooks срабатывают один раз на вызов Agent.predict_batch, при этом ctx.states содержит каждое состояние. Вместо этого Router.predict_batch выполняет свои hooks уровня Router один раз на запрос, каждый — с одним состоянием и своим run_id, поэтому тот же hook там пишет одну запись на каждый вызов hook.

def audit_batch(ctx):
    for state, result in zip(ctx.states, ctx.results):
        ship_one(ctx.run_id, state, result)

results = agent.predict_batch([state_a, state_b, state_c], questions, on_predict_end=audit_batch)

HTTP-сервер

Hooks Router срабатывают для laya.serve автоматически, потому что сервер вызывает Router.predict.

from laya import Router
from laya.serve import create_app

router = Router(hooks=[Metrics()], on_predict_end=audit, hooks_raise=False)
app = create_app(router=router)

ONNXAgent

ONNXAgent предоставляет только события уровня predict.

from laya.onnx_agent import ONNXAgent

agent = ONNXAgent("convaiinnovations/laya", onnx_path="laya.onnx", on_predict_end=audit)
agent.system_one(state, questions)

Регистрация во время выполнения

Присоединяйте, отсоединяйте hooks или ограничивайте их область после создания.

agent.add_hook(Metrics())          # attach at runtime
agent.remove_hook(Metrics())       # by identity

with agent.hooks_installed(DebugDump()):
    agent.system_one(state, questions)   # DebugDump only here

Базовый класс и значения по умолчанию для всего процесса

Унаследуйте от BaseHook, чтобы переопределить только нужное, и зарегистрируйте что-либо один раз для всего процесса, вместо того чтобы передавать это каждому Agent и Router.

from laya import BaseHook, hooks

class Audit(BaseHook):
    def on_predict_end(self, ctx):
        ship(ctx.run_id, ctx.results)

hooks.set_default_hooks(hooks=[Audit()])   # runs for every call in the process

# later, or in tests:
hooks.clear_default_hooks()

Бюджет токенов

Настройте бюджет токенов для одного вызова — из hook или аргументом на вызов. Значение hook заменяет действующий бюджет, поэтому сначала нужно прочитать этот бюджет: рассчитывайте на самый широкий вопрос вызова, оставайтесь выше нижней границы токенов, которую ядро применяет к вариантам, и расширяйте max_len вместе с head_max_len, чтобы состояние сохранило окно.

def widen(ctx):
    k = max((len(q.get("criteria", {}) or {}) for q in ctx.questions.values()), default=0)
    if k < 50:
        return
    cfg = getattr(ctx.agent, "cfg", None) or {}
    head = ctx.head_max_len if ctx.head_max_len is not None else cfg.get("head_max_len", 192)
    window = ctx.max_len if ctx.max_len is not None else cfg.get("max_len", 512)
    need = 16 + 8 * k
    if need > head:
        ctx.head_max_len = need
        ctx.max_len = max(window, need + 8 + 64)

agent = laya.load("convaiinnovations/laya", on_predict_start=widen)

# or per call
agent.system_one(state, questions, head_max_len=512, max_len=1024)

Настройка бюджета токенов содержит арифметику за каждой строкой, а predict_shortlist — вариант, когда набор меток не помещается даже в расширенное окно.

Асинхронные hooks

Оберните асинхронный hook в AsyncHook; каждая корутина выполняется до завершения в синхронном ядре, независимо от того, синхронный ли вызывающий код или уже внутри цикла событий.

import laya
from laya import AsyncHook

class RemoteAudit:
    async def on_predict_end(self, ctx):
        await ship(ctx.run_id, ctx.results)

agent = laya.load("convaiinnovations/laya", hooks=[AsyncHook(RemoteAudit())])

Обычный асинхронный callable тоже работает:

async def async_end(ctx):
    await ship(ctx.results)

agent.system_one(state, questions, on_predict_end=async_end)

Тайм-аут hook

Ограничьте каждый вызов hook, чтобы зависший hook не подвесил обслуживаемый запрос:

agent = laya.load("convaiinnovations/laya", on_predict_end=metrics, hooks_timeout=2.0)

# or per call
agent.system_one(state, questions, on_predict_end=metrics, hooks_timeout=0.5)

Hook, превысивший тайм-аут, вызывает TimeoutError (или выдаёт предупреждение, когда hooks_raise=False). Hook продолжает выполняться в фоне, поэтому дайте и сетевым вызовам собственный тайм-аут. См. ошибки.

Тестирование hooks

Проверьте, что увидел hook, без модели: прогоните predict_batch с заглушками для helpers encode/forward/decode, как это делает tests/test_hooks.py.

seen = []
agent.predict_batch(["s0"], questions, on_predict_end=lambda ctx: seen.append(ctx.results))
assert len(seen) == 1

Поверхность API зафиксирована в tests/test_hooks_api.py.

См. также