Docs

Router

Router

laya.Router detects the language of each state and sends the request to the matching checkpoint, loading checkpoints on first use.

Names, types, defaults and code stay in English; the rest is translated (entries not translated yet are shown in the original English).

Router

Router(
    models: Optional[Dict[str, str]] = None,
    device: Optional[str] = None,
    token: Optional[str] = None,
    revision: Optional[str] = None,
    revisions: Optional[Dict[str, Optional[str]]] = None,
    max_loaded: int = 2,
    default: str = "english",
    auto_task_detection: bool = False,
    standalone_repos: bool = False,
    preload: bool = False,
    lang_guess: Optional[Any] = None,
    hooks=None,
    on_predict_start=None,
    on_predict_end=None,
    hooks_raise: bool = True,
    hooks_concurrent: bool = True,
    hooks_timeout: Optional[float] = None,
    agent_kwargs: Optional[Dict[str, Any]] = None,
    sha256_digests: Optional[Dict[str, Optional[Dict[str, str]]]] = None,
)

Bases: HookRegistry

Lazily loads Laya checkpoints and sends each request to the right one.

from laya import Router

r = Router()
r.predict({"message": "Mein Konto wurde zweimal belastet"}, questions)   # -> multilingual
r.predict({"message": "I was charged twice"}, questions)                 # -> english
r.predict(state, questions, model="typed-decisions")                     # explicit

Models are downloaded and built on first use. max_loaded caps how many stay resident (least-recently-used is evicted), because all three together are ~1.16B parameters.

The default is 2, because automatic routing only ever chooses between english and multilingual: a cap of one rebuilds the checkpoint it just evicted on every script switch, which is seconds per request on exactly the traffic the Router exists for. Traffic that only ever sees one language never builds the second checkpoint, so the default costs it nothing. Lower it to 1 for a memory-constrained host, and raise it to 3 (or preload) when auto_task_detection, an explicit model= or an explicit task= can reach typed-decisions as well.

For a server or a demo, preload instead: a cold load costs seconds, while detection costs microseconds, so even the default still pays a load the first time a language appears.

r = Router(preload=True)                    # all three resident, routing is free
r = Router(preload=True, device="cuda")
r.preload(["english", "multilingual"])      # or just the two you serve

Hub revisions are opt-in. revision applies one commit to every model; revisions={"english": "...", "multilingual": "..."} overrides that per model, which is useful when standalone repositories were reviewed at different commits. Without either, huggingface_hub's normal default and existing offline cache are used.

Anything else laya.Agent accepts is reachable through agent_kwargs, which is merged into every checkpoint the Router builds:

Router(agent_kwargs={"lang_temperatures": {"de": {"temperature": [1.0, 1.4, 2.0]}}})
Router(agent_kwargs={"expected_sha256": {"model.safetensors": "a3f1..."}})
Router(agent_kwargs={"fast": True})

The names the Router sets for itself -- model_id_or_path, device, token, subfolder, revision and the hook arguments -- are refused here rather than silently shadowed, and the remaining names are checked against Agent.__init__ at construction, so a misspelled option fails on the Router(...) line instead of on the first request.

Artifact digests are opt-in and always per model: sha256_digests={"english": {...}} passes that {path relative to the checkpoint dir: hexdigest} map to the Agent that loads it, so a tampered or substituted weight file is refused before it is parsed. There is no Router-wide equivalent of revision because digests, unlike a commit SHA, are not shareable: the bundled repository ships a separate model.safetensors for each of english, multilingual and typed-decisions, so one flat map can only ever match one of them. A model listed with None or {} is loaded unverified.

The same split is available to a process configured only by environment: when LAYA_SHA256_DIGESTS holds a model-keyed map ({"english": {...}, "multilingual": {...}}) this seeds it per checkpoint, so a server that keeps several resident can pin each with its own digests instead of refusing to start on the second one. A flat LAYA_SHA256_DIGESTS keeps its existing meaning, applied by laya.revisions to every checkpoint the process loads, which is right for a single-checkpoint one. An argument entry wins over the environment for the model it names.

Hooks are opt-in and run at the Router level: on_route sees the routing decision, on_load / on_evict see model lifecycle, and on_predict_start / on_predict_end wrap the whole route+infer call. See laya.hooks.

Parameters

modelsOptional[Dict[str, str]]= None
deviceOptional[str]= None
tokenOptional[str]= None
revisionOptional[str]= None
revisionsOptional[Dict[str, Optional[str]]]= None
max_loadedint= 2
defaultstr= "english"
auto_task_detectionbool= False
standalone_reposbool= False
preloadbool= False
lang_guessOptional[Any]= None
hooks= None
on_predict_start= None
on_predict_end= None
hooks_raisebool= True
hooks_concurrentbool= True
hooks_timeoutOptional[float]= None
agent_kwargsOptional[Dict[str, Any]]= None
sha256_digestsOptional[Dict[str, Optional[Dict[str, str]]]]= None

load

load(name: str)

Return the Agent for name, downloading and building it on first use.

Concurrent callers share a single Agent instead of building duplicates.

Parameters

namestr

attach

attach(name: str, agent: Any)

Register an already-built Agent under name instead of loading a second copy.

Useful when the process has a checkpoint loaded for other reasons: a demo that already built convaiinnovations/laya can hand it to the router rather than pay for -- and hold in memory -- a duplicate 421M parameters.

Parameters

namestr
agentAny

preload

preload(names: Optional[List[str]] = None)

Download and build checkpoints up front so no request ever pays a model load.

A cold load costs seconds; language detection costs microseconds. With every checkpoint resident, routing is effectively free -- which is what you want in a server or a demo. max_loaded is raised to fit both the requested checkpoints and all already-resident agents, so incremental preloading does not evict either.

Parameters

namesOptional[List[str]]= None

unload

unload(name: Optional[str] = None)

Free one model, or all of them.

Parameters

nameOptional[str]= None

loaded_revisions

loaded_revisions: Dict[str, Optional[str]]

Commit SHA each resident agent was loaded from (None for local paths).

route

route(
    state: Union[str, dict, list, None],
    questions: Optional[Dict[str, Any]] = None,
    model: Optional[str] = None,
    task: Optional[str] = None,
    lang: Optional[str] = None,
    lang_guess: Optional[Any] = None,
    hooks=None,
    hooks_raise: Optional[bool] = None,
    hooks_timeout: Optional[float] = None,
) -> RouteDecision

Decide which checkpoint to use, then let on_route hooks observe or replace it.

ctx.decision is the RouteDecision; a hook may replace it (for example to pin a checkpoint) and the replacement is what gets returned and used. hooks are per-call hooks, appended after any installed on the Router.

Parameters

stateUnion[str, dict, list, None]
questionsOptional[Dict[str, Any]]= None
modelOptional[str]= None
taskOptional[str]= None
langOptional[str]= None
lang_guessOptional[Any]= None
hooks= None
hooks_raiseOptional[bool]= None
hooks_timeoutOptional[float]= None

predict

predict(
    state: Union[str, dict, list],
    questions: Dict[str, Any],
    model: Optional[str] = None,
    task: Optional[str] = None,
    lang: Optional[str] = None,
    lang_guess: Optional[Any] = None,
    hooks=None,
    on_predict_start=None,
    on_predict_end=None,
    hooks_raise: Optional[bool] = None,
    hooks_timeout: Optional[float] = None,
    max_len: Optional[int] = None,
    head_max_len: Optional[int] = None,
    min_confidence: Optional[float] = None,
) -> Dict[str, Any]

Route, then answer every question in one forward pass on the chosen checkpoint.

The result is the usual system_one payload plus a routing key recording the decision. Router-level on_predict_start / on_predict_end hooks wrap the whole route+infer call and see ctx.decision; see laya.hooks. max_len / head_max_len override the agent token budget for this call (a start hook may set ctx.max_len / ctx.head_max_len).

Parameters

stateUnion[str, dict, list]
questionsDict[str, Any]
modelOptional[str]= None
taskOptional[str]= None
langOptional[str]= None
lang_guessOptional[Any]= None
hooks= None
on_predict_start= None
on_predict_end= None
hooks_raiseOptional[bool]= None
hooks_timeoutOptional[float]= None
max_lenOptional[int]= None
head_max_lenOptional[int]= None
min_confidenceOptional[float]= None

predict_long

predict_long(
    state: Union[str, dict, list],
    questions: Dict[str, Any],
    model: Optional[str] = None,
    task: Optional[str] = None,
    lang: Optional[str] = None,
    lang_guess: Optional[Any] = None,
    window: Optional[int] = None,
    stride: Optional[int] = None,
    aggregate: str = "auto",
    batch_size: Optional[int] = None,
    hooks=None,
    on_predict_start=None,
    on_predict_end=None,
    hooks_raise: Optional[bool] = None,
    hooks_timeout: Optional[float] = None,
) -> Dict[str, Any]

Route, then scan every window of the state instead of only its first one.

predict scores a state from a single window: anything past max_len is cut off (the first window, or for a conversation list the last) and never reaches the model. This routes exactly as predict does -- the same model/task/lang hints, the same router-level hooks, the same routing key and usage -- and scores the routed state with that agent's predict_long, which splits it into overlapping windows and aggregates per question. The aggregation rules are laya.agent.Agent.predict_long's: noul takes the strongest window, choice/score the most confident one.

Per-call hooks (hooks, on_predict_start, on_predict_end, hooks_raise, hooks_timeout) wrap the whole route+scan exactly as they wrap predict: the scan runs last, so a start hook that answers (ctx.skip(...)) or rewrites the state wins. max_len / head_max_len are not accepted here -- a window is sized by window or the checkpoint budget, and overriding the single-window truncation is what predict_long is for.

Parameters

stateUnion[str, dict, list]
questionsDict[str, Any]
modelOptional[str]= None
taskOptional[str]= None
langOptional[str]= None
lang_guessOptional[Any]= None
windowOptional[int]= None

state tokens per window. Defaults to the routed checkpoint's budget (max_len - head_max_len - 8); a smaller window isolates a localized span.

strideOptional[int]= None

token step between windows; defaults to window // 2 (50% overlap).

aggregatestr= "auto"

"auto" (the per-type rules above) is the only mode.

batch_sizeOptional[int]= None

cap on windows per forward pass, to bound memory on very long states.

hooks= None
on_predict_start= None
on_predict_end= None
hooks_raiseOptional[bool]= None
hooks_timeoutOptional[float]= None

Returns

The usual predict payload, with usage["windows"] counting the windows scored.

Raises

TypeError: the routed agent has no predict_long (an ONNX agent, or one attached by hand), so there is nothing to scan with. Raised under the router's hooks_raise policy, which defaults to raising.

decide

decide(
    state: Union[str, dict, list],
    schema: Any = None,
    questions: Optional[Dict[str, Any]] = None,
    return_details: bool = False,
    min_confidence: Optional[float] = None,
    predict_kwargs,
) -> Any

Answer state against a schema (JSON schema or pydantic model) and return typed values.

See laya.structured. Pass exactly one of schema or questions; extra keyword arguments (for example model=, task=, hooks=) are forwarded to predict.

Parameters

stateUnion[str, dict, list]
schemaAny= None
questionsOptional[Dict[str, Any]]= None
return_detailsbool= False
min_confidenceOptional[float]= None
predict_kwargs

decide_batch

decide_batch(
    states: Sequence[Any],
    schema: Any = None,
    questions: Optional[Dict[str, Any]] = None,
    return_details: bool = False,
    min_confidence: Optional[float] = None,
    predict_kwargs,
) -> List[Any]

Answer many states against one schema (JSON schema or pydantic model) in one batched call.

The throughput form of :meth:decide: the schema is planned once and its questions run over every state through :meth:predict_batch (grouped forward passes, results in input order), then each state's answers are projected as decide does. Extra keyword arguments (batch_size=, model=, hooks=, ...) are forwarded to predict_batch. See laya.structured.

Parameters

statesSequence[Any]
schemaAny= None
questionsOptional[Dict[str, Any]]= None
return_detailsbool= False
min_confidenceOptional[float]= None
predict_kwargs

route_batch

route_batch(
    requests: Sequence[Dict[str, Any]],
    hooks_timeout: Optional[float] = None,
) -> List[RouteDecision]

Route a heterogeneous request batch without loading any checkpoints.

Each request is a mapping with state and questions plus the same optional routing overrides accepted by :meth:route: model, task, lang and lang_guess. The returned decisions preserve input order.

This is intentionally separate from inference so callers can inspect or aggregate routing decisions before paying model-load cost.

Parameters

requestsSequence[Dict[str, Any]]

Sequence of request dictionaries, each requiring state and questions.

hooks_timeoutOptional[float]= None

Override the Router's hooks_timeout for this call, applied to every request's on_route dispatch, as on :meth:route.

predict_batch

predict_batch(
    requests: Sequence[Dict[str, Any]],
    batch_size: Optional[int] = None,
    hooks_timeout: Optional[float] = None,
    min_confidence: Optional[float] = None,
    sort_by_length: bool = False,
) -> List[Dict[str, Any]]

Route and execute a heterogeneous request batch with minimal model churn.

Requests are routed first and grouped by checkpoint. Within each checkpoint, requests that share the same question schema are passed to Agent.predict_batch so their states can share forward passes. Results are then restored to the original request order.

Requests may independently specify model, task, lang, lang_guess, max_len or head_max_len and may use different question schemas. max_len / head_max_len are the per-request form of the token-budget override predict takes as call arguments: they set the checkpoint's state and question-head budgets for that one request, so a wide question can be asked without shrinking the batch's other requests to the same window. Requests that ask for different budgets are split into separate forward passes, since one Agent.predict_batch call carries one budget for all its states. A start hook may still replace either value on ctx.

Router-level predict hooks run per request, as predict runs them: each request gets its own PredictContext, so on_predict_start can replace that request's state, questions or token budget, or ctx.skip(...) it, and on_predict_end sees and may replace its result. Requests are grouped for the forward pass after their start hooks have run, and a checkpoint group's requests end in reverse of the order they started. If a checkpoint group fails, every request of it whose start hook ran fails with the exception, a cache hit included: each gets on_error and then on_predict_end before the exception propagates.

Parameters

requestsSequence[Dict[str, Any]]

Sequence of request dictionaries. Every item requires state and questions and may include model, task, lang or lang_guess routing overrides and max_len / head_max_len token-budget overrides.

batch_sizeOptional[int]= None

Optional maximum number of states per Agent forward-pass batch.

hooks_timeoutOptional[float]= None

Override the Router's hooks_timeout for this call.

min_confidenceOptional[float]= None
sort_by_lengthbool= False

Forwarded to every Agent.predict_batch call, so each question group pads to a shorter maximum; see Agent.predict_batch. Results retain the input order either way. Silently dropped for an attached agent whose predict_batch predates the knob (#294).

Returns

One normal Router prediction result per request, in the same order as the input.

RouteDecision

RouteDecision()

Bases: dict

The routing outcome: which model, why, and what was detected.

Behaves as a dict so it serialises straight into an API response.

DEFAULT_MODELS

DEFAULT_MODELS = {
    "english": (BUNDLE_REPO, None),
    "multilingual": (BUNDLE_REPO, "multilingual"),
    "typed-decisions": (BUNDLE_REPO, "typed-decisions"),
}