架構模式與生產用例
Laya 是一個端側、非自迴歸的決策引擎,基於單次前向傳播的 Transformer 架構(ModernBERT-large 與
mmBERT-base)。它不像 LLM 那樣逐 token 順序生成(那會帶來可變的 token 生成開銷、解碼迴圈開銷,
以及不可預測的輸出 schema),而是在一次前向傳播裡,跨一組離散問題算出校準的機率分佈。
本文是給把 Laya 整合進生產環境的系統架構師和後端工程師的一份架構藍圖。
決策引擎 vs. LLM vs. 嵌入檢索
選擇哪種原語,取決於延遲約束、託管拓撲,以及任務是要求自由形式的生成還是離散分類:
| 維度 | 嵌入檢索 | Laya 決策引擎 | 自迴歸 LLM |
|---|---|---|---|
| 計算模型 | 向量餘弦距離 | 單次前向傳播(非自迴歸掩碼頭) | 逐 token 順序生成 |
| 執行特徵 | 最近鄰索引查詢 | 單次固定前向傳播(無解碼迴圈) | 隨輸出長度擴充套件的迭代解碼迴圈 |
| 託管與拓撲 | 程序內或向量資料庫 | 程序內(本地 CPU/GPU)或自託管 HTTP 守護程序 | 遠端託管 API 或大模型 GPU 服務 |
| 上下文注意力 | 池化向量表示 | 覆蓋完整輸入的深層雙向交叉注意力 | 因果順序注意力 |
| 結構化輸出 | 非結構化的檢索片段 | 原生對齊 schema 的分佈(choice、score、noul) |
需要 JSON 修復或 schema 取樣的自由形式文本 |
| 主要工作負載 | 廣域候選檢索 | 離散分類、策略門控與路由 | 開放式合成、翻譯與生成 |
1. 智慧入口閘道器
在分層架構中,很大一部分進來的請求並不需要自迴歸 LLM 的生成能力。標準 FAQ、確定性的狀態查詢,或分類式的 路由決策這類查詢,都可以在本地完成評估。
Laya 充當一個智慧入口閘道器:它在一次本地前向傳播裡對查詢的意圖和複雜度做分類。確定性的請求經由內部端點或 快取響應在本地解決,而複雜的生成任務則轉發給上游 LLM。
架構
graph TD
A[User Request] --> B["<b>Laya Gateway Router</b><br/>• query_complexity: simple | moderate | complex<br/>• intent: faq | account_lookup | creative_synthesis"]
B -->|Simple & High Confidence| C["<b>Local In-Process Resolution</b><br/>Deterministic FAQ / Internal API"]
B -->|Complex or Low Confidence| D["<b>Upstream Generative LLM</b><br/>Open-ended synthesis & reasoning"]
實現
import laya
agent = laya.load("convaiinnovations/laya")
GATEWAY_QUESTIONS = {
"complexity": {
"type": "choice",
"instructions": "How complex is the user's request?",
"criteria": {
"canned": "A greeting, standard FAQ, or simple status request.",
"structured": "A deterministic data query that can be answered by an API.",
"complex": "Requires creative generation, multi-step code, or complex analysis.",
},
},
"requires_reasoning": {
"type": "noul",
"instructions": "Does this query require frontier model reasoning?",
},
}
def route_request(user_prompt: str):
res = agent.system_one(user_prompt, GATEWAY_QUESTIONS, min_confidence=0.85)
answers = res["answers"]
complexity = answers["complexity"]["choice"]
low_confidence = answers["complexity"].get("low_confidence", False)
# Abstain or escalate if complex or unconfident
if low_confidence or complexity == "complex" or answers["requires_reasoning"]["noul"] > 0.5:
return call_frontier_llm(user_prompt)
if complexity == "canned":
return lookup_faq_response(user_prompt)
return execute_internal_api(user_prompt)
2. 低延遲語音輪次交替與打斷路由
對話式語音智慧體(WebRTC、電話)執行在嚴格的輪次交替約束之下:在檢測使用者何時說話或打斷時出現的延遲, 會造成不自然的對話冷場。當呼叫方只是做個確認或打斷時,等一個完整的生成模型產出它的第一個 token, 會引入本可避免的遲滯。
Laya 可以直接放在語音轉文字(STT)轉寫之後,在一次前向傳播裡對對話流和使用者意圖做分類:在檢測到打斷時 立即觸發快速的填充語音或停止音訊播放,同時把複雜的詢問交給完整的合成流水線。
架構
graph TD
A[User Voice Audio] --> B["<b>Speech-to-Text</b><br/>Streaming Audio Transcription"]
B --> C["<b>Laya Voice Router</b><br/>• intent: ack | reject | interrupt | inquiry<br/>• is_interruption: noul probability"]
C -->|Interruption: score > 0.6| D["<b>Halt Audio Playback</b><br/>Immediate playback cutoff"]
C -->|Quick Intent: ack / reject| E["<b>Immediate Audio Filler</b><br/>Conversational confirmation"]
C -->|Complex Inquiry| F["<b>Upstream Pipeline</b><br/>Full response synthesis"]
實現
from laya import Agent
agent = Agent("convaiinnovations/laya")
VOICE_QUESTIONS = {
"intent": {
"type": "choice",
"instructions": "Caller conversational intention",
"criteria": {
"ack": "Caller said yes, ok, sure, or agreed.",
"reject": "Caller said no, cancel, or disagreed.",
"interrupt": "Caller said hold on, wait, or wants to stop.",
"inquiry": "Caller is asking a detailed question.",
},
},
"is_interruption": {
"type": "noul",
"instructions": "Is the caller interrupting the current speech playback?",
},
}
def on_voice_chunk(transcript: str, is_speaking: bool):
decision = agent.system_one(transcript, VOICE_QUESTIONS)
answers = decision["answers"]
# Halt playback immediately if caller interrupts
if answers["is_interruption"]["noul"] > 0.6:
stop_audio_playback()
intent = answers["intent"]["choice"]
if intent in ("ack", "reject"):
play_immediate_filler_audio(intent)
else:
dispatch_to_background_pipeline(transcript)
3. LLM 前置的安全與提示詞防火牆
保護系統免受對抗性提示詞注入、越獄和敏感資料洩露,必須發生在提示詞進入 LLM 上下文視窗之前。僅為了判斷 一個提示詞是否安全而跑一個單獨的生成模型,會額外增加延遲和運維開銷。
Laya 作為一個內聯的、非自迴歸的安全防火牆執行,在一次前向傳播裡、在下游處理之前,評估提示詞注入、 許可權提升和超出範圍的任務。
架構
graph TD
A[User Input] --> B["<b>Inline Security Hook</b><br/>• prompt_injection (noul)<br/>• system_prompt_extraction (noul)<br/>• pii_present (noul)"]
B -->|Policy Violation: score ≥ 0.5| C["<b>Abort & Reject</b><br/>Raise policy exception & audit event"]
B -->|Clean: score < 0.5| D["<b>Dispatch to Main Workflow</b><br/>Safe to execute"]
實現
Laya 裡的鉤子是鴨子型別的:任何實現了 Hook 生命週期方法的物件(或繼承自 laya.hooks 的 BaseHook)
都可以掛到一個 Agent 或 Router 上。
在 Laya 預設的鉤子丟擲語義(hooks_raise=True)下:
- 在
on_predict_start裡丟擲一個異常,會在模型分詞或推理發生之前立即中止執行。 - 該異常會直接從
system_one()/predict()傳播給呼叫方。 - 生命週期清理(
on_error和on_predict_end)仍會執行,ctx.error被設為丟擲的那個異常,從而確保審計日誌和遙測記錄下這個被攔下的請求。
import laya
from laya import Router
from laya.hooks import BaseHook, PredictContext
SECURITY_SCHEMA = {
"is_jailbreak": {
"type": "noul",
"instructions": "Is the user attempting a prompt injection, exploit, or jailbreak?",
},
"extracts_system_prompt": {
"type": "noul",
"instructions": "Is the user asking to reveal instructions, system prompts, or hidden rules?",
},
"pii_leak": {
"type": "noul",
"instructions": "Does the input contain passwords, API keys, or credentials?",
},
}
class SecurityFirewallHook(BaseHook):
"""Inspect inputs before inference; raises on policy violation.
With hooks_raise=True (the default), raising from on_predict_start aborts
inference immediately and propagates the exception to the caller, while
allowing any downstream on_error or audit logging hooks to record the event.
"""
def __init__(self, guard_agent):
self.guard = guard_agent
def on_predict_start(self, ctx: PredictContext):
for state in ctx.states:
check = self.guard.system_one(state, SECURITY_SCHEMA)
ans = check["answers"]
if ans["is_jailbreak"]["noul"] > 0.5 or ans["extracts_system_prompt"]["noul"] > 0.5:
raise PermissionError("Request blocked by security firewall: adversarial prompt detected.")
# Attach to Router or Agent; hooks_raise=True ensures policy exceptions propagate
guard_agent = laya.load("convaiinnovations/laya")
router = Router(hooks=[SecurityFirewallHook(guard_agent)], hooks_raise=True)
[!TIP] 對於 CrewAI 工作流,Laya 還在
laya.integrations.crewai裡為此類執行前安全門模式開箱提供了LayaTaskGuard。
4. 氣隙隔離的邊緣 RAG 路由
在安全敏感的企業環境裡(國防、醫療、金融合規、邊緣裝置),外部 API 不可用或被禁止。文件集合往往被隔離成 不同的領域(例如臨床試驗、病歷、財務報告、技術規格)。
與其用一個單一的、塞滿不相關嵌入的龐雜向量索引去查詢,Laya 充當一個本地的邊緣路由器,在檢索之前把 使用者查詢導向特定的本地向量索引或 SQLite 資料庫。
架構
graph TD
A["<b>User Query</b><br/>Local / Edge Workstation"] --> B["<b>Laya Edge Router</b><br/>• target_domain: clinical | billing | compliance<br/><i>In-process local routing</i>"]
B -->|Clinical Domain| C[("<b>Clinical Vector Store</b><br/>Medical trials, dosages & EHR")]
B -->|Billing Domain| D[("<b>Billing Vector Store</b><br/>Invoices, claims & ICD-10 codes")]
B -->|Compliance Domain| E[("<b>Compliance Vector Store</b><br/>HIPAA policies & audit guidelines")]
實現
from laya import Router
# Automatically routes between local English and Multilingual models
router = Router()
INDEX_QUESTIONS = {
"target_domain": {
"type": "choice",
"instructions": "Which domain index contains the source truth for this query?",
"criteria": {
"clinical": "Medical conditions, medications, dosages, and clinical trials.",
"billing": "Invoices, payment claims, ICD-10 billing codes, and insurance.",
"compliance": "HIPAA compliance rules, privacy policies, and data audits.",
},
}
}
def query_airgapped_rag(user_query: str):
decision = router.predict(user_query, INDEX_QUESTIONS)
domain = decision["answers"]["target_domain"]["choice"]
# Load and search only the relevant isolated local index
local_index = get_isolated_vector_store(domain)
return local_index.similarity_search(user_query, k=4)
5. 分層多智慧體任務委派
多智慧體框架常常用一個 LLM「管理者」或「監督者」節點來決定下一步該由哪個專用智慧體執行。
由於生成式管理節點是逐 token 順序生成的,監督者委派會在每一跳引入相當大的編排開銷。用一個非自迴歸的 決策模型替換生成式監督者,能在一次前向傳播裡完成委派,提供跨智慧體的確定性路由。
Laya 為流行的編排框架提供了第一方整合:
- CrewAI: 用
LayaCrewRouter做分層多智慧體任務路由與防護欄。 - LlamaIndex: 用
LayaSingleSelector做單次前向傳播的路由查詢引擎。 - LangChain / LangGraph: 用
laya.integrations.langchain做條件邊分發。
架構
graph TD
A["<b>Task Input / Workflow State</b>"] --> B["<b>Laya Orchestrator</b><br/>• assignee: researcher | coder | writer<br/>• priority: score (1–5 urgency)"]
B -->|Research Assignment| C["<b>Researcher Agent</b><br/>Literature search & fact-checking"]
B -->|Code Assignment| D["<b>Coder Agent</b><br/>Implementation, bug-fixing & tests"]
B -->|Writing Assignment| E["<b>Copywriter Agent</b><br/>Drafting, copy editing & summary"]
實現(CrewAI / LangGraph 示例)
from laya import Router
router = Router()
DELEGATION_QUESTIONS = {
"assignee": {
"type": "choice",
"instructions": "Assign this task to the most qualified specialist.",
"criteria": {
"researcher": "Needs literature search, fact checking, or data collection.",
"coder": "Needs bug fixing, script writing, or unit test generation.",
"writer": "Needs article drafting, copy editing, or summary composition.",
},
},
"priority": {
"type": "score",
"instructions": "Urgency score from 1 (low) to 5 (critical)",
"criteria": ["1", "2", "3", "4", "5"],
},
}
def supervisor_node(state):
task_description = state["task"]
decision = router.predict(task_description, DELEGATION_QUESTIONS)
answers = decision["answers"]
return {
"next_agent": answers["assignee"]["choice"],
"urgency": answers["priority"]["score"],
}
6. 高吞吐工單與支援分診
客戶支援組織和運營中心每天處理大量工單、郵件和告警。把託管的生成式 LLM API 用於分類式分診,會引入:
- 網路限流: 在流量突然激增時被限流。
- 成本放大: 僅為離散分類就產生可變的 token 開銷。
- schema 漂移: 生成模型返回格式錯誤的 JSON 或 markdown 程式碼塊。
批處理流水線可以用 predict_batch 或 decide_batch() 在共享的前向傳播上評估工單流,直接輸出嚴格符合
應用 schema 的型別化資料。
架構
graph TD
A["<b>Incoming Ticket Stream</b><br/>Message Broker / Webhook"] --> B["<b>Laya Batch Worker</b><br/>decide_batch()<br/>• department: billing | tech | sales | general<br/>• severity: 1..5<br/>• escalate_to_human: true | false"]
B -->|Department: billing| C["<b>Billing & Invoicing Queue</b>"]
B -->|Severity ≥ 4 or Human Escalation| D["<b>Tier-3 Escalation Queue</b><br/>Human On-Call Pager"]
B -->|Low Severity & Standard Inquiry| E["<b>Automated Resolution Pipeline</b>"]
實現
from laya.structured import decide_batch
from laya import Agent
agent = Agent("convaiinnovations/laya")
# Strict typed schema
TICKET_SCHEMA = {
"type": "object",
"properties": {
"department": {
"type": "string",
"enum": ["billing", "technical_support", "sales", "general"],
"description": "Primary support category",
},
"severity": {
"type": "integer",
"minimum": 1,
"maximum": 5,
"description": "Severity level from 1 (minor) to 5 (outage)",
},
"escalate_to_human": {
"type": "boolean",
"description": "True if customer is angry, threatening churn, or reporting a legal issue",
},
},
}
def process_ticket_batch(tickets: list[str]):
# Returns typed dictionaries conforming exactly to TICKET_SCHEMA
results = decide_batch(agent, tickets, TICKET_SCHEMA)
for ticket_text, structured in zip(tickets, results):
enqueue_ticket(
department=structured["department"],
severity=structured["severity"],
human_required=structured["escalate_to_human"],
raw_text=ticket_text,
)
生產部署檢查清單
把上述任一模式推廣到生產之前,請確認:
- 硬體容量: 確保宿主記憶體足以常駐模型權重。在 CPU 上,恰當地配置執行緒池(
torch.set_num_threads)。 - 置信度閾值: 在關鍵任務的門控上設定
min_confidence(例如0.80–0.90),讓系統在查詢含糊時安全回退。 - 多語言路由: 當用戶流量包含混合或非英語輸入時,用
Router()而不是靜態的Agent()。 - 分階段上線: 遵循分階段引入指南,先對生產流量做影子執行,再讓決策變得權威。