아키텍처 패턴과 프로덕션 사용 사례
Laya는 단일 포워드 패스 트랜스포머 아키텍처(ModernBERT-large 및 mmBERT-base)를 기반으로 하는
온디바이스 비자기회귀 의사결정 엔진입니다. LLM처럼 토큰을 순차적으로 생성하는 대신(이는 가변적인
토큰 생성 비용, 디코드 루프 오버헤드, 예측 불가능한 출력 스키마를 수반합니다), Laya는 한 번의
포워드 패스로 이산적인 질문에 대한 캘리브레이션된 확률 분포를 계산합니다.
이 문서는 Laya를 프로덕션 환경에 통합하는 시스템 아키텍트와 백엔드 엔지니어를 위한 아키텍처 블루프린트 역할을 합니다.
의사결정 엔진 vs. LLM vs. 임베딩 검색
올바른 프리미티브를 선택하는 일은 지연 시간 제약, 호스팅 토폴로지, 그리고 그 작업이 자유 형식 생성과 이산적 분류 중 무엇을 요구하는지에 달려 있습니다:
| 관점 | 임베딩 검색 | Laya 의사결정 엔진 | 자기회귀 LLM |
|---|---|---|---|
| 계산 모델 | 벡터 코사인 거리 | 단일 포워드 패스(비자기회귀 마스크드 헤드) | 토큰 단위 순차 생성 |
| 실행 프로파일 | 최근접 이웃 인덱스 조회 | 단일 고정 포워드 패스(디코딩 루프 없음) | 출력 길이에 따라 확장되는 반복 디코딩 루프 |
| 호스팅과 토폴로지 | 인프로세스 또는 벡터 데이터베이스 | 인프로세스(로컬 CPU/GPU) 또는 자체 호스팅 HTTP 데몬 | 원격 호스팅 API 또는 대형 모델 GPU 서빙 |
| 컨텍스트 어텐션 | 풀링된 벡터 표현 | 전체 입력에 걸친 깊은 양방향 크로스 어텐션 | 인과적 순차 어텐션 |
| 구조화 출력 | 비구조화된 검색 청크 | 스키마에 정렬된 네이티브 분포(choice, score, noul) |
JSON 복구나 스키마 샘플링이 필요한 자유 형식 텍스트 |
| 주요 워크로드 | 광범위한 후보 검색 | 이산적 분류, 정책 게이팅과 라우팅 | 개방형 합성, 번역, 생성 |
1. 스마트 인그레스 게이트웨이
계층형 아키텍처에서는 들어오는 요청의 상당 부분이 자기회귀 LLM의 생성 능력을 필요로 하지 않습니다. 일반 FAQ, 결정론적 상태 조회, 범주형 라우팅 결정 같은 질의는 로컬에서 평가할 수 있습니다.
Laya는 지능형 인그레스 게이트웨이로 동작합니다. 한 번의 로컬 포워드 패스로 질의의 의도와 복잡도를 분류합니다. 결정론적 요청은 내부 엔드포인트나 캐시된 응답을 통해 로컬에서 해결되고, 복잡한 생성 작업은 업스트림 LLM으로 전달됩니다.
아키텍처
graph TD
A[User Request] --> B["<b>Laya Gateway Router</b><br/>• query_complexity: simple | moderate | complex<br/>• intent: faq | account_lookup | creative_synthesis"]
B -->|Simple & High Confidence| C["<b>Local In-Process Resolution</b><br/>Deterministic FAQ / Internal API"]
B -->|Complex or Low Confidence| D["<b>Upstream Generative LLM</b><br/>Open-ended synthesis & reasoning"]
구현
import laya
agent = laya.load("convaiinnovations/laya")
GATEWAY_QUESTIONS = {
"complexity": {
"type": "choice",
"instructions": "How complex is the user's request?",
"criteria": {
"canned": "A greeting, standard FAQ, or simple status request.",
"structured": "A deterministic data query that can be answered by an API.",
"complex": "Requires creative generation, multi-step code, or complex analysis.",
},
},
"requires_reasoning": {
"type": "noul",
"instructions": "Does this query require frontier model reasoning?",
},
}
def route_request(user_prompt: str):
res = agent.system_one(user_prompt, GATEWAY_QUESTIONS, min_confidence=0.85)
answers = res["answers"]
complexity = answers["complexity"]["choice"]
low_confidence = answers["complexity"].get("low_confidence", False)
# Abstain or escalate if complex or unconfident
if low_confidence or complexity == "complex" or answers["requires_reasoning"]["noul"] > 0.5:
return call_frontier_llm(user_prompt)
if complexity == "canned":
return lookup_faq_response(user_prompt)
return execute_internal_api(user_prompt)
2. 저지연 음성 턴테이킹과 끼어들기 라우터
대화형 음성 에이전트(WebRTC, 전화)는 엄격한 턴테이킹 제약 아래에서 동작합니다. 사용자가 언제 말하고 언제 끼어드는지를 감지할 때의 지연은 부자연스러운 대화 공백을 만듭니다. 발신자가 단지 맞장구를 치거나 끼어들 뿐일 때 생성 모델이 첫 토큰을 낼 때까지 기다리면 피할 수 있는 지연이 생깁니다.
Laya는 음성-텍스트(STT) 변환 직후에 배치해 대화 흐름과 사용자 의도를 한 번의 포워드 패스로 분류할 수 있습니다. 끼어들기가 감지되면 즉시 빠른 필러 오디오를 재생하거나 오디오 재생을 중단하고, 복잡한 문의는 전체 합성 파이프라인으로 위임합니다.
아키텍처
graph TD
A[User Voice Audio] --> B["<b>Speech-to-Text</b><br/>Streaming Audio Transcription"]
B --> C["<b>Laya Voice Router</b><br/>• intent: ack | reject | interrupt | inquiry<br/>• is_interruption: noul probability"]
C -->|Interruption: score > 0.6| D["<b>Halt Audio Playback</b><br/>Immediate playback cutoff"]
C -->|Quick Intent: ack / reject| E["<b>Immediate Audio Filler</b><br/>Conversational confirmation"]
C -->|Complex Inquiry| F["<b>Upstream Pipeline</b><br/>Full response synthesis"]
구현
from laya import Agent
agent = Agent("convaiinnovations/laya")
VOICE_QUESTIONS = {
"intent": {
"type": "choice",
"instructions": "Caller conversational intention",
"criteria": {
"ack": "Caller said yes, ok, sure, or agreed.",
"reject": "Caller said no, cancel, or disagreed.",
"interrupt": "Caller said hold on, wait, or wants to stop.",
"inquiry": "Caller is asking a detailed question.",
},
},
"is_interruption": {
"type": "noul",
"instructions": "Is the caller interrupting the current speech playback?",
},
}
def on_voice_chunk(transcript: str, is_speaking: bool):
decision = agent.system_one(transcript, VOICE_QUESTIONS)
answers = decision["answers"]
# Halt playback immediately if caller interrupts
if answers["is_interruption"]["noul"] > 0.6:
stop_audio_playback()
intent = answers["intent"]["choice"]
if intent in ("ack", "reject"):
play_immediate_filler_audio(intent)
else:
dispatch_to_background_pipeline(transcript)
3. LLM 사전 보안과 프롬프트 방화벽
적대적 프롬프트 인젝션, 탈옥, 민감 데이터 유출로부터 시스템을 보호하는 일은 프롬프트가 LLM의 컨텍스트 창에 도달하기 전에 일어나야 합니다. 프롬프트가 안전한지만 판단하려고 별도의 생성 모델을 돌리면 중복된 지연과 운영 오버헤드가 더해집니다.
Laya는 인라인 비자기회귀 보안 방화벽으로 동작하며, 프롬프트 인젝션, 권한 상승, 범위를 벗어난 작업을 하위 처리 전에 한 번의 포워드 패스로 평가합니다.
아키텍처
graph TD
A[User Input] --> B["<b>Inline Security Hook</b><br/>• prompt_injection (noul)<br/>• system_prompt_extraction (noul)<br/>• pii_present (noul)"]
B -->|Policy Violation: score ≥ 0.5| C["<b>Abort & Reject</b><br/>Raise policy exception & audit event"]
B -->|Clean: score < 0.5| D["<b>Dispatch to Main Workflow</b><br/>Safe to execute"]
구현
Laya의 훅은 덕 타이핑입니다. Hook의 수명 주기 메서드를 구현한 임의의 객체(laya.hooks의 BaseHook
을 상속한 객체 포함)를 Agent나 Router에 붙일 수 있습니다.
Laya의 기본 훅 예외 발생 시맨틱(hooks_raise=True)에서:
on_predict_start안에서 예외를 발생시키면 모델 토큰화나 추론이 일어나기 전에 실행이 즉시 중단됩니다.- 그 예외는
system_one()/predict()에서 호출자에게 직접 전파됩니다. - 수명 주기 정리(
on_error와on_predict_end)는 계속 실행되며,ctx.error는 발생한 예외로 설정되어 감사 로그와 텔레메트리가 차단된 요청을 기록하게 합니다.
import laya
from laya import Router
from laya.hooks import BaseHook, PredictContext
SECURITY_SCHEMA = {
"is_jailbreak": {
"type": "noul",
"instructions": "Is the user attempting a prompt injection, exploit, or jailbreak?",
},
"extracts_system_prompt": {
"type": "noul",
"instructions": "Is the user asking to reveal instructions, system prompts, or hidden rules?",
},
"pii_leak": {
"type": "noul",
"instructions": "Does the input contain passwords, API keys, or credentials?",
},
}
class SecurityFirewallHook(BaseHook):
"""Inspect inputs before inference; raises on policy violation.
With hooks_raise=True (the default), raising from on_predict_start aborts
inference immediately and propagates the exception to the caller, while
allowing any downstream on_error or audit logging hooks to record the event.
"""
def __init__(self, guard_agent):
self.guard = guard_agent
def on_predict_start(self, ctx: PredictContext):
for state in ctx.states:
check = self.guard.system_one(state, SECURITY_SCHEMA)
ans = check["answers"]
if ans["is_jailbreak"]["noul"] > 0.5 or ans["extracts_system_prompt"]["noul"] > 0.5:
raise PermissionError("Request blocked by security firewall: adversarial prompt detected.")
# Attach to Router or Agent; hooks_raise=True ensures policy exceptions propagate
guard_agent = laya.load("convaiinnovations/laya")
router = Router(hooks=[SecurityFirewallHook(guard_agent)], hooks_raise=True)
[!TIP] CrewAI 워크플로의 경우, Laya는 바로 이런 실행 전 안전 게이트 패턴을 위해
laya.integrations.crewai에LayaTaskGuard를 즉시 사용할 수 있는 형태로 제공합니다.
4. 에어갭 엣지 RAG 라우터
보안이 중요한 기업 환경(국방, 의료, 금융 규정 준수, 엣지 어플라이언스)에서는 외부 API를 사용할 수 없거나 금지되어 있습니다. 문서 컬렉션은 흔히 서로 다른 도메인(예: 임상 시험, 환자 기록, 재무 보고서, 기술 사양)으로 분리되어 있습니다.
관련 없는 임베딩으로 단일 모놀리식 벡터 인덱스를 조회하는 대신, Laya는 로컬 엣지 라우터로 동작하여 검색 전에 사용자 질의를 특정 로컬 벡터 인덱스나 SQLite 데이터베이스로 보냅니다.
아키텍처
graph TD
A["<b>User Query</b><br/>Local / Edge Workstation"] --> B["<b>Laya Edge Router</b><br/>• target_domain: clinical | billing | compliance<br/><i>In-process local routing</i>"]
B -->|Clinical Domain| C[("<b>Clinical Vector Store</b><br/>Medical trials, dosages & EHR")]
B -->|Billing Domain| D[("<b>Billing Vector Store</b><br/>Invoices, claims & ICD-10 codes")]
B -->|Compliance Domain| E[("<b>Compliance Vector Store</b><br/>HIPAA policies & audit guidelines")]
구현
from laya import Router
# Automatically routes between local English and Multilingual models
router = Router()
INDEX_QUESTIONS = {
"target_domain": {
"type": "choice",
"instructions": "Which domain index contains the source truth for this query?",
"criteria": {
"clinical": "Medical conditions, medications, dosages, and clinical trials.",
"billing": "Invoices, payment claims, ICD-10 billing codes, and insurance.",
"compliance": "HIPAA compliance rules, privacy policies, and data audits.",
},
}
}
def query_airgapped_rag(user_query: str):
decision = router.predict(user_query, INDEX_QUESTIONS)
domain = decision["answers"]["target_domain"]["choice"]
# Load and search only the relevant isolated local index
local_index = get_isolated_vector_store(domain)
return local_index.similarity_search(user_query, k=4)
5. 계층형 멀티 에이전트 작업 위임
멀티 에이전트 프레임워크는 다음 단계를 어느 전문 에이전트가 실행할지 결정하기 위해 LLM “매니저” 또는 “슈퍼바이저” 노드를 두는 경우가 많습니다.
생성형 매니저 노드는 토큰을 순차적으로 생성하므로, 슈퍼바이저 위임은 홉마다 상당한 오케스트레이션 오버헤드를 초래할 수 있습니다. 생성형 슈퍼바이저를 비자기회귀 의사결정 모델로 교체하면 위임이 한 번의 포워드 패스로 이루어져 에이전트 간에 결정론적 라우팅을 제공합니다.
Laya는 널리 쓰이는 오케스트레이션 프레임워크를 위한 자사 통합을 제공합니다:
- CrewAI: 계층형 멀티 에이전트 작업 라우팅과 가드레일에는
LayaCrewRouter를 사용하십시오. - LlamaIndex: 단일 포워드 패스 라우터 쿼리 엔진에는
LayaSingleSelector를 사용하십시오. - LangChain / LangGraph: 조건부 엣지 디스패치에는
laya.integrations.langchain을 사용하십시오.
아키텍처
graph TD
A["<b>Task Input / Workflow State</b>"] --> B["<b>Laya Orchestrator</b><br/>• assignee: researcher | coder | writer<br/>• priority: score (1–5 urgency)"]
B -->|Research Assignment| C["<b>Researcher Agent</b><br/>Literature search & fact-checking"]
B -->|Code Assignment| D["<b>Coder Agent</b><br/>Implementation, bug-fixing & tests"]
B -->|Writing Assignment| E["<b>Copywriter Agent</b><br/>Drafting, copy editing & summary"]
구현(CrewAI / LangGraph 예시)
from laya import Router
router = Router()
DELEGATION_QUESTIONS = {
"assignee": {
"type": "choice",
"instructions": "Assign this task to the most qualified specialist.",
"criteria": {
"researcher": "Needs literature search, fact checking, or data collection.",
"coder": "Needs bug fixing, script writing, or unit test generation.",
"writer": "Needs article drafting, copy editing, or summary composition.",
},
},
"priority": {
"type": "score",
"instructions": "Urgency score from 1 (low) to 5 (critical)",
"criteria": ["1", "2", "3", "4", "5"],
},
}
def supervisor_node(state):
task_description = state["task"]
decision = router.predict(task_description, DELEGATION_QUESTIONS)
answers = decision["answers"]
return {
"next_agent": answers["assignee"]["choice"],
"urgency": answers["priority"]["score"],
}
6. 고처리량 티켓과 지원 트리아지
고객 지원 조직과 운영 센터는 매일 대량의 티켓, 이메일, 알림을 처리합니다. 범주형 트리아지에 호스팅 생성 LLM API를 사용하면 다음이 생길 수 있습니다:
- 네트워크 속도 제한: 갑작스러운 트래픽 급증 시 스로틀링.
- 비용 증폭: 이산적 분류만을 위해 가변적인 토큰 비용이 발생.
- 스키마 드리프트: 생성 모델이 잘못된 JSON이나 markdown 코드 블록을 반환.
배치 처리 파이프라인은 predict_batch나 decide_batch()를 사용해 공유 포워드 패스에서 티켓
스트림을 평가하고, 애플리케이션 스키마에 직접 부합하는 엄격한 타입의 데이터를 출력할 수 있습니다.
아키텍처
graph TD
A["<b>Incoming Ticket Stream</b><br/>Message Broker / Webhook"] --> B["<b>Laya Batch Worker</b><br/>decide_batch()<br/>• department: billing | tech | sales | general<br/>• severity: 1..5<br/>• escalate_to_human: true | false"]
B -->|Department: billing| C["<b>Billing & Invoicing Queue</b>"]
B -->|Severity ≥ 4 or Human Escalation| D["<b>Tier-3 Escalation Queue</b><br/>Human On-Call Pager"]
B -->|Low Severity & Standard Inquiry| E["<b>Automated Resolution Pipeline</b>"]
구현
from laya.structured import decide_batch
from laya import Agent
agent = Agent("convaiinnovations/laya")
# Strict typed schema
TICKET_SCHEMA = {
"type": "object",
"properties": {
"department": {
"type": "string",
"enum": ["billing", "technical_support", "sales", "general"],
"description": "Primary support category",
},
"severity": {
"type": "integer",
"minimum": 1,
"maximum": 5,
"description": "Severity level from 1 (minor) to 5 (outage)",
},
"escalate_to_human": {
"type": "boolean",
"description": "True if customer is angry, threatening churn, or reporting a legal issue",
},
},
}
def process_ticket_batch(tickets: list[str]):
# Returns typed dictionaries conforming exactly to TICKET_SCHEMA
results = decide_batch(agent, tickets, TICKET_SCHEMA)
for ticket_text, structured in zip(tickets, results):
enqueue_ticket(
department=structured["department"],
severity=structured["severity"],
human_required=structured["escalate_to_human"],
raw_text=ticket_text,
)
프로덕션 배포 체크리스트
위 패턴 중 하나를 프로덕션에 출시하기 전에 다음을 확인하십시오:
- 하드웨어 사이징: 상주 모델 가중치를 위한 호스트 메모리가 충분한지 확인하십시오. CPU에서는 스레드 풀을 적절히 구성하십시오(
torch.set_num_threads). - 신뢰도 임계값: 중요한 게이트에
min_confidence(예:0.80–0.90)를 설정해 질의가 모호할 때 시스템이 안전하게 폴백하도록 하십시오. - 다국어 라우팅: 사용자 트래픽에 혼합 입력이나 비영어 입력이 포함되면 정적
Agent()대신Router()를 사용하십시오. - 단계적 롤아웃: 단계적 도입 가이드를 따라 결정을 공식화하기 전에 프로덕션 트래픽으로 섀도 운영하십시오.