架构模式与生产用例
Laya 是一个端侧、非自回归的决策引擎,基于单次前向传播的 Transformer 架构(ModernBERT-large 与
mmBERT-base)。它不像 LLM 那样逐 token 顺序生成(那会带来可变的 token 生成开销、解码循环开销,
以及不可预测的输出 schema),而是在一次前向传播里,跨一组离散问题算出校准的概率分布。
本文是给把 Laya 集成进生产环境的系统架构师和后端工程师的一份架构蓝图。
决策引擎 vs. LLM vs. 嵌入检索
选择哪种原语,取决于延迟约束、托管拓扑,以及任务是要求自由形式的生成还是离散分类:
| 维度 | 嵌入检索 | Laya 决策引擎 | 自回归 LLM |
|---|---|---|---|
| 计算模型 | 向量余弦距离 | 单次前向传播(非自回归掩码头) | 逐 token 顺序生成 |
| 执行特征 | 最近邻索引查找 | 单次固定前向传播(无解码循环) | 随输出长度扩展的迭代解码循环 |
| 托管与拓扑 | 进程内或向量数据库 | 进程内(本地 CPU/GPU)或自托管 HTTP 守护进程 | 远程托管 API 或大模型 GPU 服务 |
| 上下文注意力 | 池化向量表示 | 覆盖完整输入的深层双向交叉注意力 | 因果顺序注意力 |
| 结构化输出 | 非结构化的检索片段 | 原生对齐 schema 的分布(choice、score、noul) |
需要 JSON 修复或 schema 采样的自由形式文本 |
| 主要工作负载 | 广域候选检索 | 离散分类、策略门控与路由 | 开放式合成、翻译与生成 |
1. 智能入口网关
在分层架构中,很大一部分进来的请求并不需要自回归 LLM 的生成能力。标准 FAQ、确定性的状态查询,或分类式的 路由决策这类查询,都可以在本地完成评估。
Laya 充当一个智能入口网关:它在一次本地前向传播里对查询的意图和复杂度做分类。确定性的请求经由内部端点或 缓存响应在本地解决,而复杂的生成任务则转发给上游 LLM。
架构
graph TD
A[User Request] --> B["<b>Laya Gateway Router</b><br/>• query_complexity: simple | moderate | complex<br/>• intent: faq | account_lookup | creative_synthesis"]
B -->|Simple & High Confidence| C["<b>Local In-Process Resolution</b><br/>Deterministic FAQ / Internal API"]
B -->|Complex or Low Confidence| D["<b>Upstream Generative LLM</b><br/>Open-ended synthesis & reasoning"]
实现
import laya
agent = laya.load("convaiinnovations/laya")
GATEWAY_QUESTIONS = {
"complexity": {
"type": "choice",
"instructions": "How complex is the user's request?",
"criteria": {
"canned": "A greeting, standard FAQ, or simple status request.",
"structured": "A deterministic data query that can be answered by an API.",
"complex": "Requires creative generation, multi-step code, or complex analysis.",
},
},
"requires_reasoning": {
"type": "noul",
"instructions": "Does this query require frontier model reasoning?",
},
}
def route_request(user_prompt: str):
res = agent.system_one(user_prompt, GATEWAY_QUESTIONS, min_confidence=0.85)
answers = res["answers"]
complexity = answers["complexity"]["choice"]
low_confidence = answers["complexity"].get("low_confidence", False)
# Abstain or escalate if complex or unconfident
if low_confidence or complexity == "complex" or answers["requires_reasoning"]["noul"] > 0.5:
return call_frontier_llm(user_prompt)
if complexity == "canned":
return lookup_faq_response(user_prompt)
return execute_internal_api(user_prompt)
2. 低延迟语音轮次交替与打断路由
对话式语音智能体(WebRTC、电话)运行在严格的轮次交替约束之下:在检测用户何时说话或打断时出现的延迟, 会造成不自然的对话冷场。当呼叫方只是做个确认或打断时,等一个完整的生成模型产出它的第一个 token, 会引入本可避免的迟滞。
Laya 可以直接放在语音转文字(STT)转写之后,在一次前向传播里对对话流和用户意图做分类:在检测到打断时 立即触发快速的填充语音或停止音频播放,同时把复杂的询问交给完整的合成流水线。
架构
graph TD
A[User Voice Audio] --> B["<b>Speech-to-Text</b><br/>Streaming Audio Transcription"]
B --> C["<b>Laya Voice Router</b><br/>• intent: ack | reject | interrupt | inquiry<br/>• is_interruption: noul probability"]
C -->|Interruption: score > 0.6| D["<b>Halt Audio Playback</b><br/>Immediate playback cutoff"]
C -->|Quick Intent: ack / reject| E["<b>Immediate Audio Filler</b><br/>Conversational confirmation"]
C -->|Complex Inquiry| F["<b>Upstream Pipeline</b><br/>Full response synthesis"]
实现
from laya import Agent
agent = Agent("convaiinnovations/laya")
VOICE_QUESTIONS = {
"intent": {
"type": "choice",
"instructions": "Caller conversational intention",
"criteria": {
"ack": "Caller said yes, ok, sure, or agreed.",
"reject": "Caller said no, cancel, or disagreed.",
"interrupt": "Caller said hold on, wait, or wants to stop.",
"inquiry": "Caller is asking a detailed question.",
},
},
"is_interruption": {
"type": "noul",
"instructions": "Is the caller interrupting the current speech playback?",
},
}
def on_voice_chunk(transcript: str, is_speaking: bool):
decision = agent.system_one(transcript, VOICE_QUESTIONS)
answers = decision["answers"]
# Halt playback immediately if caller interrupts
if answers["is_interruption"]["noul"] > 0.6:
stop_audio_playback()
intent = answers["intent"]["choice"]
if intent in ("ack", "reject"):
play_immediate_filler_audio(intent)
else:
dispatch_to_background_pipeline(transcript)
3. LLM 前置的安全与提示词防火墙
保护系统免受对抗性提示词注入、越狱和敏感数据泄露,必须发生在提示词进入 LLM 上下文窗口之前。仅为了判断 一个提示词是否安全而跑一个单独的生成模型,会额外增加延迟和运维开销。
Laya 作为一个内联的、非自回归的安全防火墙运行,在一次前向传播里、在下游处理之前,评估提示词注入、 权限提升和超出范围的任务。
架构
graph TD
A[User Input] --> B["<b>Inline Security Hook</b><br/>• prompt_injection (noul)<br/>• system_prompt_extraction (noul)<br/>• pii_present (noul)"]
B -->|Policy Violation: score ≥ 0.5| C["<b>Abort & Reject</b><br/>Raise policy exception & audit event"]
B -->|Clean: score < 0.5| D["<b>Dispatch to Main Workflow</b><br/>Safe to execute"]
实现
Laya 里的钩子是鸭子类型的:任何实现了 Hook 生命周期方法的对象(或继承自 laya.hooks 的 BaseHook)
都可以挂到一个 Agent 或 Router 上。
在 Laya 默认的钩子抛出语义(hooks_raise=True)下:
- 在
on_predict_start里抛出一个异常,会在模型分词或推理发生之前立即中止执行。 - 该异常会直接从
system_one()/predict()传播给调用方。 - 生命周期清理(
on_error和on_predict_end)仍会执行,ctx.error被设为抛出的那个异常,从而确保审计日志和遥测记录下这个被拦下的请求。
import laya
from laya import Router
from laya.hooks import BaseHook, PredictContext
SECURITY_SCHEMA = {
"is_jailbreak": {
"type": "noul",
"instructions": "Is the user attempting a prompt injection, exploit, or jailbreak?",
},
"extracts_system_prompt": {
"type": "noul",
"instructions": "Is the user asking to reveal instructions, system prompts, or hidden rules?",
},
"pii_leak": {
"type": "noul",
"instructions": "Does the input contain passwords, API keys, or credentials?",
},
}
class SecurityFirewallHook(BaseHook):
"""Inspect inputs before inference; raises on policy violation.
With hooks_raise=True (the default), raising from on_predict_start aborts
inference immediately and propagates the exception to the caller, while
allowing any downstream on_error or audit logging hooks to record the event.
"""
def __init__(self, guard_agent):
self.guard = guard_agent
def on_predict_start(self, ctx: PredictContext):
for state in ctx.states:
check = self.guard.system_one(state, SECURITY_SCHEMA)
ans = check["answers"]
if ans["is_jailbreak"]["noul"] > 0.5 or ans["extracts_system_prompt"]["noul"] > 0.5:
raise PermissionError("Request blocked by security firewall: adversarial prompt detected.")
# Attach to Router or Agent; hooks_raise=True ensures policy exceptions propagate
guard_agent = laya.load("convaiinnovations/laya")
router = Router(hooks=[SecurityFirewallHook(guard_agent)], hooks_raise=True)
[!TIP] 对于 CrewAI 工作流,Laya 还在
laya.integrations.crewai里为此类执行前安全门模式开箱提供了LayaTaskGuard。
4. 气隙隔离的边缘 RAG 路由
在安全敏感的企业环境里(国防、医疗、金融合规、边缘设备),外部 API 不可用或被禁止。文档集合往往被隔离成 不同的领域(例如临床试验、病历、财务报告、技术规格)。
与其用一个单一的、塞满不相关嵌入的庞杂向量索引去查询,Laya 充当一个本地的边缘路由器,在检索之前把 用户查询导向特定的本地向量索引或 SQLite 数据库。
架构
graph TD
A["<b>User Query</b><br/>Local / Edge Workstation"] --> B["<b>Laya Edge Router</b><br/>• target_domain: clinical | billing | compliance<br/><i>In-process local routing</i>"]
B -->|Clinical Domain| C[("<b>Clinical Vector Store</b><br/>Medical trials, dosages & EHR")]
B -->|Billing Domain| D[("<b>Billing Vector Store</b><br/>Invoices, claims & ICD-10 codes")]
B -->|Compliance Domain| E[("<b>Compliance Vector Store</b><br/>HIPAA policies & audit guidelines")]
实现
from laya import Router
# Automatically routes between local English and Multilingual models
router = Router()
INDEX_QUESTIONS = {
"target_domain": {
"type": "choice",
"instructions": "Which domain index contains the source truth for this query?",
"criteria": {
"clinical": "Medical conditions, medications, dosages, and clinical trials.",
"billing": "Invoices, payment claims, ICD-10 billing codes, and insurance.",
"compliance": "HIPAA compliance rules, privacy policies, and data audits.",
},
}
}
def query_airgapped_rag(user_query: str):
decision = router.predict(user_query, INDEX_QUESTIONS)
domain = decision["answers"]["target_domain"]["choice"]
# Load and search only the relevant isolated local index
local_index = get_isolated_vector_store(domain)
return local_index.similarity_search(user_query, k=4)
5. 分层多智能体任务委派
多智能体框架常常用一个 LLM「管理者」或「监督者」节点来决定下一步该由哪个专用智能体执行。
由于生成式管理节点是逐 token 顺序生成的,监督者委派会在每一跳引入相当大的编排开销。用一个非自回归的 决策模型替换生成式监督者,能在一次前向传播里完成委派,提供跨智能体的确定性路由。
Laya 为流行的编排框架提供了第一方集成:
- CrewAI: 用
LayaCrewRouter做分层多智能体任务路由与防护栏。 - LlamaIndex: 用
LayaSingleSelector做单次前向传播的路由查询引擎。 - LangChain / LangGraph: 用
laya.integrations.langchain做条件边分发。
架构
graph TD
A["<b>Task Input / Workflow State</b>"] --> B["<b>Laya Orchestrator</b><br/>• assignee: researcher | coder | writer<br/>• priority: score (1–5 urgency)"]
B -->|Research Assignment| C["<b>Researcher Agent</b><br/>Literature search & fact-checking"]
B -->|Code Assignment| D["<b>Coder Agent</b><br/>Implementation, bug-fixing & tests"]
B -->|Writing Assignment| E["<b>Copywriter Agent</b><br/>Drafting, copy editing & summary"]
实现(CrewAI / LangGraph 示例)
from laya import Router
router = Router()
DELEGATION_QUESTIONS = {
"assignee": {
"type": "choice",
"instructions": "Assign this task to the most qualified specialist.",
"criteria": {
"researcher": "Needs literature search, fact checking, or data collection.",
"coder": "Needs bug fixing, script writing, or unit test generation.",
"writer": "Needs article drafting, copy editing, or summary composition.",
},
},
"priority": {
"type": "score",
"instructions": "Urgency score from 1 (low) to 5 (critical)",
"criteria": ["1", "2", "3", "4", "5"],
},
}
def supervisor_node(state):
task_description = state["task"]
decision = router.predict(task_description, DELEGATION_QUESTIONS)
answers = decision["answers"]
return {
"next_agent": answers["assignee"]["choice"],
"urgency": answers["priority"]["score"],
}
6. 高吞吐工单与支持分诊
客户支持组织和运营中心每天处理大量工单、邮件和告警。把托管的生成式 LLM API 用于分类式分诊,会引入:
- 网络限流: 在流量突然激增时被限流。
- 成本放大: 仅为离散分类就产生可变的 token 开销。
- schema 漂移: 生成模型返回格式错误的 JSON 或 markdown 代码块。
批处理流水线可以用 predict_batch 或 decide_batch() 在共享的前向传播上评估工单流,直接输出严格符合
应用 schema 的类型化数据。
架构
graph TD
A["<b>Incoming Ticket Stream</b><br/>Message Broker / Webhook"] --> B["<b>Laya Batch Worker</b><br/>decide_batch()<br/>• department: billing | tech | sales | general<br/>• severity: 1..5<br/>• escalate_to_human: true | false"]
B -->|Department: billing| C["<b>Billing & Invoicing Queue</b>"]
B -->|Severity ≥ 4 or Human Escalation| D["<b>Tier-3 Escalation Queue</b><br/>Human On-Call Pager"]
B -->|Low Severity & Standard Inquiry| E["<b>Automated Resolution Pipeline</b>"]
实现
from laya.structured import decide_batch
from laya import Agent
agent = Agent("convaiinnovations/laya")
# Strict typed schema
TICKET_SCHEMA = {
"type": "object",
"properties": {
"department": {
"type": "string",
"enum": ["billing", "technical_support", "sales", "general"],
"description": "Primary support category",
},
"severity": {
"type": "integer",
"minimum": 1,
"maximum": 5,
"description": "Severity level from 1 (minor) to 5 (outage)",
},
"escalate_to_human": {
"type": "boolean",
"description": "True if customer is angry, threatening churn, or reporting a legal issue",
},
},
}
def process_ticket_batch(tickets: list[str]):
# Returns typed dictionaries conforming exactly to TICKET_SCHEMA
results = decide_batch(agent, tickets, TICKET_SCHEMA)
for ticket_text, structured in zip(tickets, results):
enqueue_ticket(
department=structured["department"],
severity=structured["severity"],
human_required=structured["escalate_to_human"],
raw_text=ticket_text,
)
生产部署检查清单
把上述任一模式推广到生产之前,请确认:
- 硬件容量: 确保宿主内存足以常驻模型权重。在 CPU 上,恰当地配置线程池(
torch.set_num_threads)。 - 置信度阈值: 在关键任务的门控上设置
min_confidence(例如0.80–0.90),让系统在查询含糊时安全回退。 - 多语言路由: 当用户流量包含混合或非英语输入时,用
Router()而不是静态的Agent()。 - 分阶段上线: 遵循分阶段引入指南,先对生产流量做影子运行,再让决策变得权威。