Router
Router
laya.Router detects the language of each state and sends the request to the matching
checkpoint, loading checkpoints on first use.
Names, types, defaults and code stay in English; the rest is translated (entries not translated yet are shown in the original English).
Router
Router(
models: Optional[Dict[str, str]] = None,
device: Optional[str] = None,
token: Optional[str] = None,
revision: Optional[str] = None,
revisions: Optional[Dict[str, Optional[str]]] = None,
max_loaded: int = 2,
default: str = "english",
auto_task_detection: bool = False,
standalone_repos: bool = False,
preload: bool = False,
lang_guess: Optional[Any] = None,
hooks=None,
on_predict_start=None,
on_predict_end=None,
hooks_raise: bool = True,
hooks_concurrent: bool = True,
hooks_timeout: Optional[float] = None,
agent_kwargs: Optional[Dict[str, Any]] = None,
sha256_digests: Optional[Dict[str, Optional[Dict[str, str]]]] = None,
)Bases: HookRegistry
Lazily loads Laya checkpoints and sends each request to the right one.
from laya import Router
r = Router()
r.predict({"message": "Mein Konto wurde zweimal belastet"}, questions) # -> multilingual
r.predict({"message": "I was charged twice"}, questions) # -> english
r.predict(state, questions, model="typed-decisions") # explicit
Models are downloaded and built on first use. max_loaded caps how many stay resident
(least-recently-used is evicted), because all three together are ~1.16B parameters.
The default is 2, because automatic routing only ever chooses between english and
multilingual: a cap of one rebuilds the checkpoint it just evicted on every script switch,
which is seconds per request on exactly the traffic the Router exists for. Traffic that only
ever sees one language never builds the second checkpoint, so the default costs it nothing.
Lower it to 1 for a memory-constrained host, and raise it to 3 (or preload) when
auto_task_detection, an explicit model= or an explicit task= can reach
typed-decisions as well.
For a server or a demo, preload instead: a cold load costs seconds, while detection costs microseconds, so even the default still pays a load the first time a language appears.
r = Router(preload=True) # all three resident, routing is free
r = Router(preload=True, device="cuda")
r.preload(["english", "multilingual"]) # or just the two you serve
Hub revisions are opt-in. revision applies one commit to every model;
revisions={"english": "...", "multilingual": "..."} overrides that per model,
which is useful when standalone repositories were reviewed at different commits.
Without either, huggingface_hub's normal default and existing offline cache are used.
Anything else laya.Agent accepts is reachable through agent_kwargs, which is merged into
every checkpoint the Router builds:
Router(agent_kwargs={"lang_temperatures": {"de": {"temperature": [1.0, 1.4, 2.0]}}})
Router(agent_kwargs={"expected_sha256": {"model.safetensors": "a3f1..."}})
Router(agent_kwargs={"fast": True})
The names the Router sets for itself -- model_id_or_path, device, token, subfolder,
revision and the hook arguments -- are refused here rather than silently shadowed, and the
remaining names are checked against Agent.__init__ at construction, so a misspelled option
fails on the Router(...) line instead of on the first request.
Artifact digests are opt-in and always per model: sha256_digests={"english": {...}}
passes that {path relative to the checkpoint dir: hexdigest} map to the Agent that
loads it, so a tampered or substituted weight file is refused before it is parsed. There
is no Router-wide equivalent of revision because digests, unlike a commit SHA, are not
shareable: the bundled repository ships a separate model.safetensors for each of
english, multilingual and typed-decisions, so one flat map can only ever match one
of them. A model listed with None or {} is loaded unverified.
The same split is available to a process configured only by environment: when
LAYA_SHA256_DIGESTS holds a model-keyed map ({"english": {...}, "multilingual": {...}})
this seeds it per checkpoint, so a server that keeps several resident can pin each with its
own digests instead of refusing to start on the second one. A flat LAYA_SHA256_DIGESTS
keeps its existing meaning, applied by laya.revisions to every checkpoint the process
loads, which is right for a single-checkpoint one. An argument entry wins over the
environment for the model it names.
Hooks are opt-in and run at the Router level: on_route sees the routing decision,
on_load / on_evict see model lifecycle, and on_predict_start / on_predict_end
wrap the whole route+infer call. See laya.hooks.
Parameters
modelsOptional[Dict[str, str]]=NonedeviceOptional[str]=NonetokenOptional[str]=NonerevisionOptional[str]=NonerevisionsOptional[Dict[str, Optional[str]]]=Nonemax_loadedint=2defaultstr="english"auto_task_detectionbool=Falsestandalone_reposbool=Falsepreloadbool=Falselang_guessOptional[Any]=Nonehooks=Noneon_predict_start=Noneon_predict_end=Nonehooks_raisebool=Truehooks_concurrentbool=Truehooks_timeoutOptional[float]=Noneagent_kwargsOptional[Dict[str, Any]]=Nonesha256_digestsOptional[Dict[str, Optional[Dict[str, str]]]]=None
load
load(name: str)Return the Agent for name, downloading and building it on first use.
Concurrent callers share a single Agent instead of building duplicates.
Parameters
namestr
attach
attach(name: str, agent: Any)Register an already-built Agent under name instead of loading a second copy.
Useful when the process has a checkpoint loaded for other reasons: a demo that already
built convaiinnovations/laya can hand it to the router rather than pay for -- and hold
in memory -- a duplicate 421M parameters.
Parameters
namestragentAny
preload
preload(names: Optional[List[str]] = None)Download and build checkpoints up front so no request ever pays a model load.
A cold load costs seconds; language detection costs microseconds. With every
checkpoint resident, routing is effectively free -- which is what you want in a
server or a demo. max_loaded is raised to fit both the requested checkpoints and
all already-resident agents, so incremental preloading does not evict either.
Parameters
namesOptional[List[str]]=None
unload
unload(name: Optional[str] = None)Free one model, or all of them.
Parameters
nameOptional[str]=None
loaded_revisions
loaded_revisions: Dict[str, Optional[str]]Commit SHA each resident agent was loaded from (None for local paths).
route
route(
state: Union[str, dict, list, None],
questions: Optional[Dict[str, Any]] = None,
model: Optional[str] = None,
task: Optional[str] = None,
lang: Optional[str] = None,
lang_guess: Optional[Any] = None,
hooks=None,
hooks_raise: Optional[bool] = None,
hooks_timeout: Optional[float] = None,
) -> RouteDecisionDecide which checkpoint to use, then let on_route hooks observe or replace it.
ctx.decision is the RouteDecision; a hook may replace it (for example to pin a
checkpoint) and the replacement is what gets returned and used. hooks are per-call
hooks, appended after any installed on the Router.
Parameters
stateUnion[str, dict, list, None]questionsOptional[Dict[str, Any]]=NonemodelOptional[str]=NonetaskOptional[str]=NonelangOptional[str]=Nonelang_guessOptional[Any]=Nonehooks=Nonehooks_raiseOptional[bool]=Nonehooks_timeoutOptional[float]=None
predict
predict(
state: Union[str, dict, list],
questions: Dict[str, Any],
model: Optional[str] = None,
task: Optional[str] = None,
lang: Optional[str] = None,
lang_guess: Optional[Any] = None,
hooks=None,
on_predict_start=None,
on_predict_end=None,
hooks_raise: Optional[bool] = None,
hooks_timeout: Optional[float] = None,
max_len: Optional[int] = None,
head_max_len: Optional[int] = None,
min_confidence: Optional[float] = None,
) -> Dict[str, Any]Route, then answer every question in one forward pass on the chosen checkpoint.
The result is the usual system_one payload plus a routing key recording the decision.
Router-level on_predict_start / on_predict_end hooks wrap the whole route+infer call
and see ctx.decision; see laya.hooks. max_len / head_max_len override the agent
token budget for this call (a start hook may set ctx.max_len / ctx.head_max_len).
Parameters
stateUnion[str, dict, list]questionsDict[str, Any]modelOptional[str]=NonetaskOptional[str]=NonelangOptional[str]=Nonelang_guessOptional[Any]=Nonehooks=Noneon_predict_start=Noneon_predict_end=Nonehooks_raiseOptional[bool]=Nonehooks_timeoutOptional[float]=Nonemax_lenOptional[int]=Nonehead_max_lenOptional[int]=Nonemin_confidenceOptional[float]=None
predict_long
predict_long(
state: Union[str, dict, list],
questions: Dict[str, Any],
model: Optional[str] = None,
task: Optional[str] = None,
lang: Optional[str] = None,
lang_guess: Optional[Any] = None,
window: Optional[int] = None,
stride: Optional[int] = None,
aggregate: str = "auto",
batch_size: Optional[int] = None,
hooks=None,
on_predict_start=None,
on_predict_end=None,
hooks_raise: Optional[bool] = None,
hooks_timeout: Optional[float] = None,
) -> Dict[str, Any]Route, then scan every window of the state instead of only its first one.
predict scores a state from a single window: anything past max_len is cut off (the
first window, or for a conversation list the last) and never reaches the model. This
routes exactly as predict does -- the same model/task/lang hints, the same
router-level hooks, the same routing key and usage -- and scores the routed state with
that agent's predict_long, which splits it into overlapping windows and aggregates per
question. The aggregation rules are laya.agent.Agent.predict_long's: noul takes the
strongest window, choice/score the most confident one.
Per-call hooks (hooks, on_predict_start, on_predict_end, hooks_raise,
hooks_timeout) wrap the whole route+scan exactly as they wrap predict: the scan runs
last, so a start hook that answers (ctx.skip(...)) or rewrites the state wins.
max_len / head_max_len are not accepted here -- a window is sized by window or the
checkpoint budget, and overriding the single-window truncation is what predict_long is for.
Parameters
stateUnion[str, dict, list]questionsDict[str, Any]modelOptional[str]=NonetaskOptional[str]=NonelangOptional[str]=Nonelang_guessOptional[Any]=NonewindowOptional[int]=Nonestate tokens per window. Defaults to the routed checkpoint's budget (
max_len - head_max_len - 8); a smaller window isolates a localized span.strideOptional[int]=Nonetoken step between windows; defaults to
window // 2(50% overlap).aggregatestr="auto""auto" (the per-type rules above) is the only mode.
batch_sizeOptional[int]=Nonecap on windows per forward pass, to bound memory on very long states.
hooks=Noneon_predict_start=Noneon_predict_end=Nonehooks_raiseOptional[bool]=Nonehooks_timeoutOptional[float]=None
Returns
The usual predict payload, with usage["windows"] counting the windows scored.
Raises
TypeError: the routed agent has no predict_long (an ONNX agent, or one attached by
hand), so there is nothing to scan with. Raised under the router's
hooks_raise policy, which defaults to raising.
decide
decide(
state: Union[str, dict, list],
schema: Any = None,
questions: Optional[Dict[str, Any]] = None,
return_details: bool = False,
min_confidence: Optional[float] = None,
predict_kwargs,
) -> AnyAnswer state against a schema (JSON schema or pydantic model) and return typed values.
See laya.structured. Pass exactly one of schema or questions; extra keyword arguments
(for example model=, task=, hooks=) are forwarded to predict.
Parameters
stateUnion[str, dict, list]schemaAny=NonequestionsOptional[Dict[str, Any]]=Nonereturn_detailsbool=Falsemin_confidenceOptional[float]=Nonepredict_kwargs
decide_batch
decide_batch(
states: Sequence[Any],
schema: Any = None,
questions: Optional[Dict[str, Any]] = None,
return_details: bool = False,
min_confidence: Optional[float] = None,
predict_kwargs,
) -> List[Any]Answer many states against one schema (JSON schema or pydantic model) in one batched call.
The throughput form of :meth:decide: the schema is planned once and its questions
run over every state through :meth:predict_batch (grouped forward passes, results in
input order), then each state's answers are projected as decide does. Extra keyword
arguments (batch_size=, model=, hooks=, ...) are forwarded to
predict_batch. See laya.structured.
Parameters
statesSequence[Any]schemaAny=NonequestionsOptional[Dict[str, Any]]=Nonereturn_detailsbool=Falsemin_confidenceOptional[float]=Nonepredict_kwargs
route_batch
route_batch(
requests: Sequence[Dict[str, Any]],
hooks_timeout: Optional[float] = None,
) -> List[RouteDecision]Route a heterogeneous request batch without loading any checkpoints.
Each request is a mapping with state and questions plus the same optional
routing overrides accepted by :meth:route: model, task, lang and
lang_guess. The returned decisions preserve input order.
This is intentionally separate from inference so callers can inspect or aggregate routing decisions before paying model-load cost.
Parameters
requestsSequence[Dict[str, Any]]Sequence of request dictionaries, each requiring
stateandquestions.hooks_timeoutOptional[float]=NoneOverride the Router's
hooks_timeoutfor this call, applied to every request'son_routedispatch, as on :meth:route.
predict_batch
predict_batch(
requests: Sequence[Dict[str, Any]],
batch_size: Optional[int] = None,
hooks_timeout: Optional[float] = None,
min_confidence: Optional[float] = None,
sort_by_length: bool = False,
) -> List[Dict[str, Any]]Route and execute a heterogeneous request batch with minimal model churn.
Requests are routed first and grouped by checkpoint. Within each checkpoint,
requests that share the same question schema are passed to
Agent.predict_batch so their states can share forward passes. Results are
then restored to the original request order.
Requests may independently specify model, task, lang,
lang_guess, max_len or head_max_len and may use different question schemas.
max_len / head_max_len are the per-request form of the token-budget override
predict takes as call arguments: they set the checkpoint's state and question-head
budgets for that one request, so a wide question can be asked without shrinking the
batch's other requests to the same window. Requests that ask for different budgets are
split into separate forward passes, since one Agent.predict_batch call carries one
budget for all its states. A start hook may still replace either value on ctx.
Router-level predict hooks run per request, as predict runs them: each request
gets its own PredictContext, so on_predict_start can replace that request's
state, questions or token budget, or ctx.skip(...) it, and on_predict_end
sees and may replace its result. Requests are grouped for the forward pass after
their start hooks have run, and a checkpoint group's requests end in reverse of the
order they started. If a checkpoint group fails, every request of it whose start hook
ran fails with the exception, a cache hit included: each gets on_error and then
on_predict_end before the exception propagates.
Parameters
requestsSequence[Dict[str, Any]]Sequence of request dictionaries. Every item requires
stateandquestionsand may includemodel,task,langorlang_guessrouting overrides andmax_len/head_max_lentoken-budget overrides.batch_sizeOptional[int]=NoneOptional maximum number of states per Agent forward-pass batch.
hooks_timeoutOptional[float]=NoneOverride the Router's
hooks_timeoutfor this call.min_confidenceOptional[float]=Nonesort_by_lengthbool=FalseForwarded to every
Agent.predict_batchcall, so each question group pads to a shorter maximum; seeAgent.predict_batch. Results retain the input order either way. Silently dropped for an attached agent whosepredict_batchpredates the knob (#294).
Returns
One normal Router prediction result per request, in the same order as the input.
RouteDecision
RouteDecision()Bases: dict
The routing outcome: which model, why, and what was detected.
Behaves as a dict so it serialises straight into an API response.
DEFAULT_MODELS
DEFAULT_MODELS = {
"english": (BUNDLE_REPO, None),
"multilingual": (BUNDLE_REPO, "multilingual"),
"typed-decisions": (BUNDLE_REPO, "typed-decisions"),
}