스키마 기반 의사결정
JSON 스키마나 pydantic 모델을 Laya 질문으로 바꾸고, 캘리브레이션된 신뢰도와 함께 타입이 지정된 값을 돌려받습니다. 이것이 Laya를 구조화 출력 엔진으로 만드는 다리입니다. 원하는 형태를 서술하면 Laya가 한 번의 포워드 패스로 그것을 답합니다.
import laya
schema = {
"type": "object",
"properties": {
"department": {"type": "string", "enum": ["billing", "support", "sales"],
"description": "Which team should handle this?"},
"urgency": {"type": "integer", "minimum": 0, "maximum": 2},
"needs_human": {"type": "boolean"},
},
}
agent = laya.load("convaiinnovations/laya")
values = agent.decide("I was charged twice, refund me.", schema=schema)
# {"department": "billing", "urgency": 2, "needs_human": True}
pydantic을 쓰려면(laya[structured] 설치):
from typing import Literal
from pydantic import BaseModel
class Ticket(BaseModel):
department: Literal["billing", "support", "sales"]
urgency: Literal[0, 1, 2]
needs_human: bool
ticket = agent.decide("I was charged twice, refund me.", schema=Ticket)
지원되는 부분집합
최상위는 properties를 가진 객체여야 합니다. 각 속성은 질문 하나가 됩니다.
아래 각 행은 실제 스키마입니다. tests/test_structured_docs.py가 첫 번째 열을 컴파일하고
컴파일러가 실제로 만들어 내는 질문을 검증하므로, 이 표는 코드에서 어긋날 수 없습니다. 각 칸은
속성 스키마 자체이거나 진입점 호출입니다.
| JSON 스키마 | Laya 질문 | 반환되는 값 |
|---|---|---|
{"enum": ["billing", "support"]} |
choice |
원래 타입 그대로의, 선택된 값 |
{"const": "billing"} |
choice |
그 하나의 값 |
{"type": "boolean"} |
noul |
true / false |
{"type": "integer", "minimum": 0, "maximum": 5} |
score |
확률이 가장 높은 단계, 정수로 |
{"type": "number", "minimum": 0, "maximum": 5} |
score |
단계, 정수로 |
{"type": "string", "enum": ["low", "high"], "description": "How urgent?"} |
choice |
How urgent?가 질문 문구입니다 |
{"anyOf": [{"enum": ["x", "y"]}, {"type": "null"}]} |
choice |
일반 enum 행과 같으며, 답이 없으면 키가 빠집니다 |
{"oneOf": [{"type": "boolean"}, {"type": "null"}]} |
noul |
일반 boolean 행과 같습니다 |
{"type": ["integer", "null"], "minimum": 1, "maximum": 3} |
score |
일반 범위 정수 행과 같습니다 |
Literal[...]과 Optional[...]는 enum과 anyOf 행의 pydantic 표기입니다.
questions_from_pydantic이 이들을 그 형태로 렌더링하며 같은 행들이 적용됩니다.
title은 읽지 않습니다. pydantic은 요청 여부와 관계없이 model_json_schema()의 모든 필드에
title을 붙이는데, 속성 단위의 이름은 질문을 구성하는 선택지별 choice에 레이블을 달 수 없습니다.
따라서 문구를 조정하는 지렛대는 description입니다 —— 아래 내부 매핑 방식을 참고하십시오.
투영은 정확합니다. enum: [1, 2, 3]은 2를 돌려주지 "2"가 아니며, 범위가 있는 정수는
minimum과 maximum 사이의 단계를 돌려주고, 불리언은 noul >= 0.5입니다.
거부
고정된 선택지 집합으로 답할 수 없는 스키마는 정확한 경로를 밝히며
laya.structured.SchemaError(ValueError)를 발생시킵니다. 각 행도 필드 이름을 name으로 두고
실행됩니다.
| 속성 스키마 | 메시지 |
|---|---|
{"type": "string"} |
properties.name: a free string cannot be a fixed option set; use 'enum' or a boolean |
{"type": "array", "items": {"type": "string"}} |
properties.name: arrays are not supported; ask one field per element |
{"type": "object", "properties": {"inner": {"type": "boolean"}}} |
properties.name: nested objects are not supported; flatten the schema |
{"$ref": "#/definitions/node"} |
properties.name: $ref/recursion is not supported; flatten the schema |
{"enum": [1, "1"]} |
properties.name: enum values produce duplicate choice labels |
{"enum": []} |
properties.name: 'enum' must not be empty |
{"type": "number"} |
properties.name: a numeric field needs integer 'minimum' and 'maximum' to become a score |
{"type": "integer", "minimum": 5, "maximum": 2} |
properties.name: 'maximum' 2 is below 'minimum' 5 |
{"type": "integer", "minimum": 0, "maximum": 10} |
properties.name: 11 levels exceeds MAX_SCORE_LEVELS=10; narrow the range or use an enum |
{"anyOf": [{"type": "string"}, {"type": "integer"}]} |
properties.name: only 'Optional[...]' unions (one non-null branch) are supported, got 2 |
{"type": ["string", "integer"]} |
properties.name: 'type' has multiple non-null types; unions are not supported |
{"format": "date"} |
properties.name: unsupported schema {'format': 'date'} |
"boolean" |
properties.name: property must be an object, got str |
진입점 자체는 다음을 거부합니다.
| 호출 | 메시지 |
|---|---|
plan_from_json_schema("not a schema") |
expected a JSON schema object, got str |
plan_from_json_schema({"type": "object"}) |
the top level must be an object with 'properties' |
plan_from_json_schema({"type": "object", "properties": {}}) |
'properties' must be a non-empty object |
plan_from_json_schema({"type": "object", "properties": {"p%d" % i: {"type": "boolean"} for i in range(33)}}) |
33 properties exceeds MAX_PROPERTIES=32 |
plan_from_json_schema({"type": "object", "properties": {"name": {"enum": ["v%d" % i for i in range(33)]}}}) |
properties.name: 33 options exceeds MAX_OPTIONS=32 |
decide(None, "I was charged twice.", schema=42) |
expected a JSON schema dict or a pydantic model, got int |
한도: MAX_PROPERTIES = 32, MAX_OPTIONS = 32, MAX_SCORE_LEVELS = 10.
API
| 함수 | 용도 |
|---|---|
laya.decide(runner, state, schema=..., *, questions=..., return_details=..., min_confidence=..., **predict_kwargs) |
자유 함수로, Agent와 Router 모두에서 동작합니다 |
agent.decide(state, schema=..., ...) / router.decide(state, schema=..., ...) |
편의 메서드 |
laya.decide_batch(runner, states, schema=..., ...) / agent.decide_batch(...) / router.decide_batch(...) |
여러 상태에 같은 작업을, 한 번의 배치 호출로 |
questions_from_json_schema(schema) |
스키마를 Laya 질문으로 |
questions_from_pydantic(model) |
pydantic 모델을 질문으로 (pydantic 필요) |
answers_to_json(answers, schema) |
원시 답을 스키마 값으로 투영 |
answer_to_pydantic(model, answers) |
원시 답을 pydantic 인스턴스로 투영 |
plan_from_json_schema(schema) |
검증된 필드 계획 (고급) |
schema와 questions 중 정확히 하나만 전달하십시오. questions를 쓰면 decide는 투영하지 않고
원시 답을 돌려줍니다. 추가 키워드 인수는 predict로 전달되므로 훅, model=, task=, 토큰 예산이
모두 동작합니다:
router.decide(state, schema=Ticket, model="multilingual", hooks=[Metrics()])
여러 상태 채점
decide_batch는 처리량 지향 형태입니다. 스키마는 한 번만 계획되고 그 질문들이 predict_batch를
통해 모든 상태에 대해 실행되므로, 상태들이 호출마다 하나씩이 아니라 포워드 패스를 공유합니다. 결과는
입력 순서대로 돌아오고 decide가 투영하는 방식과 똑같이 투영되며, return_details=True는 상태마다
DecisionResult 하나를 줍니다:
values = agent.decide_batch(ticket_texts, schema=Ticket) # values[i] matches ticket_texts[i]
results = router.decide_batch(states, schema=Ticket, return_details=True, batch_size=64)
Router에서는 각 상태가 여전히 개별적으로 라우팅되므로, 한 번의 호출이 여러 체크포인트에 걸칠 수
있습니다. 키워드 인수는 predict_batch에 도달하므로 batch_size=, model=, 훅이 decide에서와
같이 동작합니다. Agent, ONNXAgent, Router 모두 이 메서드를 가집니다. predict_batch가 없는
러너는 조용히 루프로 폴백하지 않고 TypeError를 발생시키므로, 그런 경우 상태마다 decide를
호출하십시오. 배치 처리는 predict_batch가 그렇듯 경계선상의 argmax를 옮길 수 있습니다. README에 두
장치에서 측정한 속도 향상이 기록되어 있습니다.
신뢰도와 확률
기본적으로 decide는 값만 돌려줍니다. return_details=True를 전달하면 필드별 신뢰도, 확률, 원시
답, 호출의 사용량과 라우팅을 담은 DecisionResult를 받습니다:
result = agent.decide(state, schema=Ticket, return_details=True)
result.values["department"] # "billing"
result.answer_confidence["department"] # 0.94 max(p): the quantity min_confidence gates on
result.confidence["department"] # 0.71 normalized entropy, which depends on label count
result.probabilities["department"] # {"billing": 0.94, "support": 0.06, "sales": 0.0}
result.usage # {"input_tokens": ..., "output_tokens": 0,
# "state_tokens": ..., "state_tokens_dropped": ...,
# "truncated": ..., "truncated_questions": [...]}
result.routing # the Router decision, when a Router answered
usage는 predict()가 구성한 블록으로, 그대로 전달됩니다. input_tokens는 질문마다 한 행씩 state를 합산하고, output_tokens는 아무것도 생성되지 않으므로 항상 0이며, 나머지 네 키는 잘림 보고서(#174)입니다. state_tokens는 직렬화된 state 전체에 필요한 양, state_tokens_dropped는 개별 질문의 head가 포기한 최대량, truncated / truncated_questions는 어느 것인지를 나타냅니다. 일곱 번째 키 options는 head 예산으로 인해 어떤 질문의 선택지가 하나의 토큰 구간을 공유하게 된 경우에만 존재합니다(#538). docs/http-api.md는 이 블록과 그 필드를 모두 응답 키로 문서화하며, tests/test_structured_docs.py는 이 페이지의 목록을 그것을 구성하는 코드에 고정합니다.
confidence와 answer_confidence는 서로 다른 양이며, 이름은 정의를 따릅니다. answer_confidence는
max(p), 즉 보고되는 답에 실린 확률 질량입니다. 온도 스케일링이 적합하는 대상이고, 이 저장소의
모든 캘리브레이션 수치가 이것을 기준으로 계산되며, min_confidence가 비교되는 대상입니다 ——
그래서 confidence가 아니라 이것으로 게이팅해야 합니다. confidence는 정규화 엔트로피로, 질문에
선택지가 몇 개였는지에 따라 달라집니다. tests/test_confidence.py는 선택지가 두 개인 분포가
noul에서는 0.90, 동등한 choice에서는 0.53으로 돌아온다고 고정하므로, 이것은 임계값과 비교할
대상이 아닙니다. 사용할 수 있는 answer_confidence를 보고하지 않은 필드는 None에 매핑되며, 이는
보고된 0.0과 다릅니다.
게이트가 사용하는 것과 같은 양으로 게이팅하십시오:
if result.answer_confidence["department"] < 0.6:
result.values["department"] = "human-review"
answer_confidence가 필터링에 알맞은 숫자라는 것과 신뢰할 수 있는 확률이라는 것은 다릅니다.
“c로 반환된 답 가운데 약 c가 맞다”라고 읽는 것은, 해당 체크포인트와 질문 형태에 대해 홀드아웃
데이터로 온도를 적합하고 검증한 뒤에만 성립합니다. 배포되는 체크포인트는 배포된 상태 그대로
과신하며, laya-multilingual은 적합된 온도를 전혀 갖고 있지 않습니다 —— README의
캘리브레이션과
정직한 한계 절, 그리고 적합 루프는
파인튜닝 노트북을
참고하십시오. 그 수준에 의존하기 전에 적합하십시오. 게이트와 평가 하네스가 모두 사용하는 양이므로
보고하십시오.
내부 매핑 방식
- Enum과
Literal은 값을 문자열 레이블로 삼아choice질문이 됩니다. 나올 때 레이블은 원래 값으로 되매핑되므로 정수는 정수로 남습니다. - 범위가 있는 정수는 값마다 단계 하나를 가진
score질문이 됩니다. 반환되는 값은minimum + argmax입니다. - 불리언은
noul질문이 됩니다. 값은noul >= 0.5입니다. description은 질문 instructions가 되므로, 좋은 description이 의사결정을 정확하게 만듭니다. 이는 훅 가이드와 같은 규칙을 따릅니다. 각 선택지가 무엇을 뜻하는지 명시하십시오.null분기는 필드가 계획되기 전에 제거되므로,Optional[X]는X가 던지는 질문을 그대로 던집니다. 그 필드에 대한 답이 없으면 값에서 키가 그냥 빠지며, 그래서 모델이 보는 것을 바꾸지 않고도 필드를 선택 사항으로 선언해도 안전합니다.
함께 보기
- 예측 훅: 여기서 나오는 의사결정을 관찰·변형·캐시·게이팅합니다.
- 의사결정 프리미티브:
choice,score,noul을 깊이 다룹니다. - LangChain과 LangGraph:
LayaDecision은 이 호출을 체인의 러너블로 만든 것입니다.