アーキテクチャパターンと本番ユースケース
Laya は、単一フォワードパスの Transformer アーキテクチャ(ModernBERT-large と mmBERT-base)に
基づく、オンデバイスの非自己回帰型意思決定エンジンです。LLM のようにトークンを逐次生成するのでは
なく(それには可変のトークン生成コスト、デコードループのオーバーヘッド、予測できない出力スキーマが
伴います)、Laya は 1 回のフォワードパスで離散的な質問上の較正済み確率分布を計算します。
このドキュメントは、Laya を本番環境に組み込むシステムアーキテクトとバックエンドエンジニア向けの アーキテクチャブループリントです。
意思決定エンジン vs. LLM vs. 埋め込み検索
どのプリミティブを選ぶかは、レイテンシの制約、ホスティングのトポロジー、そしてタスクが自由形式の 生成を要するか離散的な分類を要するかによって決まります:
| 観点 | 埋め込み検索 | Laya 意思決定エンジン | 自己回帰型 LLM |
|---|---|---|---|
| 計算モデル | ベクトルのコサイン距離 | 単一のフォワードパス(非自己回帰のマスク済みヘッド) | トークンごとの逐次生成 |
| 実行プロファイル | 最近傍インデックスの検索 | 単一の固定フォワードパス(デコードループなし) | 出力長に応じてスケールする反復デコードループ |
| ホスティングとトポロジー | プロセス内またはベクトルデータベース | プロセス内(ローカル CPU/GPU)または自己ホストの HTTP デーモン | リモートのホスト型 API または大規模モデルの GPU 配信 |
| コンテキストのアテンション | プールされたベクトル表現 | 入力全体にわたる深い双方向クロスアテンション | 因果的な逐次アテンション |
| 構造化出力 | 非構造化の検索済みチャンク | schema に整合したネイティブの分布(choice、score、noul) |
JSON 修復やスキーマサンプリングを要する自由形式テキスト |
| 主なワークロード | 広範な候補の検索 | 離散的な分類、ポリシーのゲーティングとルーティング | 自由形式の合成、翻訳、生成 |
1. スマートイングレスゲートウェイ
階層化アーキテクチャでは、受信リクエストの大部分が自己回帰型 LLM の生成能力を必要としません。 標準的な FAQ、決定論的な状態照会、カテゴリ別のルーティング判断といったクエリは、ローカルで 評価できます。
Laya はインテリジェントなイングレスゲートウェイとして機能します。1 回のローカルフォワードパスで クエリの意図と複雑さを分類します。決定論的なリクエストは内部エンドポイントやキャッシュされた応答 を介してローカルで解決され、複雑な生成タスクは上流の LLM に転送されます。
アーキテクチャ
graph TD
A[User Request] --> B["<b>Laya Gateway Router</b><br/>• query_complexity: simple | moderate | complex<br/>• intent: faq | account_lookup | creative_synthesis"]
B -->|Simple & High Confidence| C["<b>Local In-Process Resolution</b><br/>Deterministic FAQ / Internal API"]
B -->|Complex or Low Confidence| D["<b>Upstream Generative LLM</b><br/>Open-ended synthesis & reasoning"]
実装
import laya
agent = laya.load("convaiinnovations/laya")
GATEWAY_QUESTIONS = {
"complexity": {
"type": "choice",
"instructions": "How complex is the user's request?",
"criteria": {
"canned": "A greeting, standard FAQ, or simple status request.",
"structured": "A deterministic data query that can be answered by an API.",
"complex": "Requires creative generation, multi-step code, or complex analysis.",
},
},
"requires_reasoning": {
"type": "noul",
"instructions": "Does this query require frontier model reasoning?",
},
}
def route_request(user_prompt: str):
res = agent.system_one(user_prompt, GATEWAY_QUESTIONS, min_confidence=0.85)
answers = res["answers"]
complexity = answers["complexity"]["choice"]
low_confidence = answers["complexity"].get("low_confidence", False)
# Abstain or escalate if complex or unconfident
if low_confidence or complexity == "complex" or answers["requires_reasoning"]["noul"] > 0.5:
return call_frontier_llm(user_prompt)
if complexity == "canned":
return lookup_faq_response(user_prompt)
return execute_internal_api(user_prompt)
2. 低レイテンシの音声ターンテイキングと割り込みルーター
対話型音声エージェント(WebRTC、テレフォニー)は、厳格なターンテイキングの制約下で動作します。 ユーザーがいつ発話し、いつ割り込むかを検出する際の遅延は、不自然な会話の空白を生みます。発信者が 単に相づちを打つだけ、あるいは割り込むだけの場合に、生成モデルが最初のトークンを出力するまで 待つと、避けられるはずの遅延が生じます。
Laya は音声認識(STT)の転写直後に置くことができ、会話の流れとユーザーの意図を 1 回のフォワード パスで分類します。割り込みが検出されたら即座に素早いフィラー音声を鳴らすか音声再生を止め、複雑な 問い合わせは完全な合成パイプラインに委ねます。
アーキテクチャ
graph TD
A[User Voice Audio] --> B["<b>Speech-to-Text</b><br/>Streaming Audio Transcription"]
B --> C["<b>Laya Voice Router</b><br/>• intent: ack | reject | interrupt | inquiry<br/>• is_interruption: noul probability"]
C -->|Interruption: score > 0.6| D["<b>Halt Audio Playback</b><br/>Immediate playback cutoff"]
C -->|Quick Intent: ack / reject| E["<b>Immediate Audio Filler</b><br/>Conversational confirmation"]
C -->|Complex Inquiry| F["<b>Upstream Pipeline</b><br/>Full response synthesis"]
実装
from laya import Agent
agent = Agent("convaiinnovations/laya")
VOICE_QUESTIONS = {
"intent": {
"type": "choice",
"instructions": "Caller conversational intention",
"criteria": {
"ack": "Caller said yes, ok, sure, or agreed.",
"reject": "Caller said no, cancel, or disagreed.",
"interrupt": "Caller said hold on, wait, or wants to stop.",
"inquiry": "Caller is asking a detailed question.",
},
},
"is_interruption": {
"type": "noul",
"instructions": "Is the caller interrupting the current speech playback?",
},
}
def on_voice_chunk(transcript: str, is_speaking: bool):
decision = agent.system_one(transcript, VOICE_QUESTIONS)
answers = decision["answers"]
# Halt playback immediately if caller interrupts
if answers["is_interruption"]["noul"] > 0.6:
stop_audio_playback()
intent = answers["intent"]["choice"]
if intent in ("ack", "reject"):
play_immediate_filler_audio(intent)
else:
dispatch_to_background_pipeline(transcript)
3. LLM 前段のセキュリティとプロンプトファイアウォール
敵対的なプロンプトインジェクション、ジェイルブレイク、機密データの漏えいからシステムを守ることは、 プロンプトが LLM のコンテキストウィンドウに到達する前に行わなければなりません。プロンプトが安全 かどうかを判断するためだけに別の生成モデルを走らせると、冗長なレイテンシと運用オーバーヘッドが 加わります。
Laya はインラインの非自己回帰型セキュリティファイアウォールとして動作し、プロンプトインジェクション、 権限昇格、範囲外のタスクを、下流の処理の前に 1 回のフォワードパスで評価します。
アーキテクチャ
graph TD
A[User Input] --> B["<b>Inline Security Hook</b><br/>• prompt_injection (noul)<br/>• system_prompt_extraction (noul)<br/>• pii_present (noul)"]
B -->|Policy Violation: score ≥ 0.5| C["<b>Abort & Reject</b><br/>Raise policy exception & audit event"]
B -->|Clean: score < 0.5| D["<b>Dispatch to Main Workflow</b><br/>Safe to execute"]
実装
Laya のフックはダックタイピングです。Hook のライフサイクルメソッドを実装した任意のオブジェクト
(または laya.hooks の BaseHook を継承したもの)を Agent や Router に取り付けられます。
Laya の既定のフック送出セマンティクス(hooks_raise=True)では:
on_predict_startの内部で例外を送出すると、モデルのトークン化や推論が行われる前に実行が即座に中止されます。- その例外は
system_one()/predict()から呼び出し元へ直接伝播します。 - ライフサイクルの後処理(
on_errorとon_predict_end)は引き続き実行され、ctx.errorには送出された例外が設定されるため、監査ログとテレメトリにブロックされたリクエストが記録されます。
import laya
from laya import Router
from laya.hooks import BaseHook, PredictContext
SECURITY_SCHEMA = {
"is_jailbreak": {
"type": "noul",
"instructions": "Is the user attempting a prompt injection, exploit, or jailbreak?",
},
"extracts_system_prompt": {
"type": "noul",
"instructions": "Is the user asking to reveal instructions, system prompts, or hidden rules?",
},
"pii_leak": {
"type": "noul",
"instructions": "Does the input contain passwords, API keys, or credentials?",
},
}
class SecurityFirewallHook(BaseHook):
"""Inspect inputs before inference; raises on policy violation.
With hooks_raise=True (the default), raising from on_predict_start aborts
inference immediately and propagates the exception to the caller, while
allowing any downstream on_error or audit logging hooks to record the event.
"""
def __init__(self, guard_agent):
self.guard = guard_agent
def on_predict_start(self, ctx: PredictContext):
for state in ctx.states:
check = self.guard.system_one(state, SECURITY_SCHEMA)
ans = check["answers"]
if ans["is_jailbreak"]["noul"] > 0.5 or ans["extracts_system_prompt"]["noul"] > 0.5:
raise PermissionError("Request blocked by security firewall: adversarial prompt detected.")
# Attach to Router or Agent; hooks_raise=True ensures policy exceptions propagate
guard_agent = laya.load("convaiinnovations/laya")
router = Router(hooks=[SecurityFirewallHook(guard_agent)], hooks_raise=True)
[!TIP] CrewAI のワークフロー向けに、Laya はまさにこの実行前の安全ゲートのパターン向けとして、
LayaTaskGuardをlaya.integrations.crewaiにすぐ使える形で用意しています。
4. エアギャップ環境のエッジ RAG ルーター
セキュアな企業環境(防衛、医療、金融コンプライアンス、エッジアプライアンス)では、外部 API が 利用できなかったり禁止されていたりします。ドキュメントの集合は、しばしば別々のドメイン(例:臨床 試験、患者記録、財務報告、技術仕様)に分離されています。
無関係な埋め込みを持つ単一のモノリシックなベクトルインデックスに問い合わせるのではなく、Laya は ローカルのエッジルーターとして機能し、検索の前にユーザーのクエリを特定のローカルベクトルインデッ クスや SQLite データベースへ振り分けます。
アーキテクチャ
graph TD
A["<b>User Query</b><br/>Local / Edge Workstation"] --> B["<b>Laya Edge Router</b><br/>• target_domain: clinical | billing | compliance<br/><i>In-process local routing</i>"]
B -->|Clinical Domain| C[("<b>Clinical Vector Store</b><br/>Medical trials, dosages & EHR")]
B -->|Billing Domain| D[("<b>Billing Vector Store</b><br/>Invoices, claims & ICD-10 codes")]
B -->|Compliance Domain| E[("<b>Compliance Vector Store</b><br/>HIPAA policies & audit guidelines")]
実装
from laya import Router
# Automatically routes between local English and Multilingual models
router = Router()
INDEX_QUESTIONS = {
"target_domain": {
"type": "choice",
"instructions": "Which domain index contains the source truth for this query?",
"criteria": {
"clinical": "Medical conditions, medications, dosages, and clinical trials.",
"billing": "Invoices, payment claims, ICD-10 billing codes, and insurance.",
"compliance": "HIPAA compliance rules, privacy policies, and data audits.",
},
}
}
def query_airgapped_rag(user_query: str):
decision = router.predict(user_query, INDEX_QUESTIONS)
domain = decision["answers"]["target_domain"]["choice"]
# Load and search only the relevant isolated local index
local_index = get_isolated_vector_store(domain)
return local_index.similarity_search(user_query, k=4)
5. 階層型マルチエージェントのタスク委任
マルチエージェントフレームワークは、次にどの専門エージェントが実行すべきかを決めるために、しばしば LLM の「マネージャー」または「スーパーバイザー」ノードを採用します。
生成型のマネージャーノードはトークンを逐次生成するため、スーパーバイザーによる委任はホップごとに 相当なオーケストレーションのオーバーヘッドを招きえます。生成型のスーパーバイザーを非自己回帰型の 意思決定モデルに置き換えると、委任が 1 回のフォワードパスで行われ、エージェント間で決定論的な ルーティングが得られます。
Laya は、主要なオーケストレーションフレームワーク向けにファーストパーティの統合を提供しています:
- CrewAI: 階層型マルチエージェントのタスクルーティングとガードレールには
LayaCrewRouterを使います。 - LlamaIndex: 1 回のフォワードパスで動くルータークエリエンジンには
LayaSingleSelectorを使います。 - LangChain / LangGraph: 条件付きエッジのディスパッチには
laya.integrations.langchainを使います。
アーキテクチャ
graph TD
A["<b>Task Input / Workflow State</b>"] --> B["<b>Laya Orchestrator</b><br/>• assignee: researcher | coder | writer<br/>• priority: score (1–5 urgency)"]
B -->|Research Assignment| C["<b>Researcher Agent</b><br/>Literature search & fact-checking"]
B -->|Code Assignment| D["<b>Coder Agent</b><br/>Implementation, bug-fixing & tests"]
B -->|Writing Assignment| E["<b>Copywriter Agent</b><br/>Drafting, copy editing & summary"]
実装(CrewAI / LangGraph の例)
from laya import Router
router = Router()
DELEGATION_QUESTIONS = {
"assignee": {
"type": "choice",
"instructions": "Assign this task to the most qualified specialist.",
"criteria": {
"researcher": "Needs literature search, fact checking, or data collection.",
"coder": "Needs bug fixing, script writing, or unit test generation.",
"writer": "Needs article drafting, copy editing, or summary composition.",
},
},
"priority": {
"type": "score",
"instructions": "Urgency score from 1 (low) to 5 (critical)",
"criteria": ["1", "2", "3", "4", "5"],
},
}
def supervisor_node(state):
task_description = state["task"]
decision = router.predict(task_description, DELEGATION_QUESTIONS)
answers = decision["answers"]
return {
"next_agent": answers["assignee"]["choice"],
"urgency": answers["priority"]["score"],
}
6. 高スループットのチケットとサポートのトリアージ
顧客サポート組織やオペレーションセンターは、毎日大量のチケット、メール、アラートを処理します。 カテゴリ別のトリアージにホスト型の生成 LLM API を使うと、次のことが起こりえます:
- ネットワークのレート制限: 急激な量の急増時にスロットリングされる。
- コストの増幅: 離散的な分類のためだけに可変のトークンコストがかかる。
- スキーマのドリフト: 生成モデルが不正な JSON や markdown コードブロックを返す。
バッチ処理パイプラインは、predict_batch や decide_batch() を使って共有のフォワードパス上で
チケットのストリームを評価し、アプリケーションのスキーマに直接準拠する厳密な型付きデータを出力
できます。
アーキテクチャ
graph TD
A["<b>Incoming Ticket Stream</b><br/>Message Broker / Webhook"] --> B["<b>Laya Batch Worker</b><br/>decide_batch()<br/>• department: billing | tech | sales | general<br/>• severity: 1..5<br/>• escalate_to_human: true | false"]
B -->|Department: billing| C["<b>Billing & Invoicing Queue</b>"]
B -->|Severity ≥ 4 or Human Escalation| D["<b>Tier-3 Escalation Queue</b><br/>Human On-Call Pager"]
B -->|Low Severity & Standard Inquiry| E["<b>Automated Resolution Pipeline</b>"]
実装
from laya.structured import decide_batch
from laya import Agent
agent = Agent("convaiinnovations/laya")
# Strict typed schema
TICKET_SCHEMA = {
"type": "object",
"properties": {
"department": {
"type": "string",
"enum": ["billing", "technical_support", "sales", "general"],
"description": "Primary support category",
},
"severity": {
"type": "integer",
"minimum": 1,
"maximum": 5,
"description": "Severity level from 1 (minor) to 5 (outage)",
},
"escalate_to_human": {
"type": "boolean",
"description": "True if customer is angry, threatening churn, or reporting a legal issue",
},
},
}
def process_ticket_batch(tickets: list[str]):
# Returns typed dictionaries conforming exactly to TICKET_SCHEMA
results = decide_batch(agent, tickets, TICKET_SCHEMA)
for ticket_text, structured in zip(tickets, results):
enqueue_ticket(
department=structured["department"],
severity=structured["severity"],
human_required=structured["escalate_to_human"],
raw_text=ticket_text,
)
本番デプロイのチェックリスト
上記のいずれかのパターンを本番に展開する前に、次を確認してください:
- ハードウェアのサイジング: モデル重みを常駐させるのに十分なホストメモリを確保します。CPU では、スレッドプールを適切に構成します(
torch.set_num_threads)。 - 信頼度のしきい値: ミッションクリティカルなゲートに
min_confidence(例:0.80–0.90)を設定し、クエリが曖昧なときにシステムが安全にフォールバックするようにします。 - 多言語ルーティング: ユーザートラフィックに混在した入力や英語以外の入力が含まれる場合は、静的な
Agent()ではなくRouter()を使います。 - 段階的なロールアウト: 段階的な導入ガイドに従い、意思決定を正式なものにする前に本番トラフィックでシャドー運用します。