문서

스키마 기반 의사결정

JSON 스키마나 pydantic 모델을 Laya 질문으로 바꾸고, 캘리브레이션된 신뢰도와 함께 타입이 지정된 값을 돌려받습니다. 이것이 Laya를 구조화 출력 엔진으로 만드는 다리입니다. 원하는 형태를 서술하면 Laya가 한 번의 포워드 패스로 그것을 답합니다.

import laya

schema = {
    "type": "object",
    "properties": {
        "department": {"type": "string", "enum": ["billing", "support", "sales"],
                       "description": "Which team should handle this?"},
        "urgency": {"type": "integer", "minimum": 0, "maximum": 2},
        "needs_human": {"type": "boolean"},
    },
}

agent = laya.load("convaiinnovations/laya")
values = agent.decide("I was charged twice, refund me.", schema=schema)
# {"department": "billing", "urgency": 2, "needs_human": True}

pydantic을 쓰려면(laya[structured] 설치):

from typing import Literal
from pydantic import BaseModel

class Ticket(BaseModel):
    department: Literal["billing", "support", "sales"]
    urgency: Literal[0, 1, 2]
    needs_human: bool

ticket = agent.decide("I was charged twice, refund me.", schema=Ticket)

지원되는 부분집합

최상위는 properties를 가진 객체여야 합니다. 각 속성은 질문 하나가 됩니다.

아래 각 행은 실제 스키마입니다. tests/test_structured_docs.py가 첫 번째 열을 컴파일하고 컴파일러가 실제로 만들어 내는 질문을 검증하므로, 이 표는 코드에서 어긋날 수 없습니다. 각 칸은 속성 스키마 자체이거나 진입점 호출입니다.

JSON 스키마 Laya 질문 반환되는 값
{"enum": ["billing", "support"]} choice 원래 타입 그대로의, 선택된 값
{"const": "billing"} choice 그 하나의 값
{"type": "boolean"} noul true / false
{"type": "integer", "minimum": 0, "maximum": 5} score 확률이 가장 높은 단계, 정수로
{"type": "number", "minimum": 0, "maximum": 5} score 단계, 정수로
{"type": "string", "enum": ["low", "high"], "description": "How urgent?"} choice How urgent?가 질문 문구입니다
{"anyOf": [{"enum": ["x", "y"]}, {"type": "null"}]} choice 일반 enum 행과 같으며, 답이 없으면 키가 빠집니다
{"oneOf": [{"type": "boolean"}, {"type": "null"}]} noul 일반 boolean 행과 같습니다
{"type": ["integer", "null"], "minimum": 1, "maximum": 3} score 일반 범위 정수 행과 같습니다

Literal[...]과 Optional[...]는 enum과 anyOf 행의 pydantic 표기입니다. questions_from_pydantic이 이들을 그 형태로 렌더링하며 같은 행들이 적용됩니다.

title은 읽지 않습니다. pydantic은 요청 여부와 관계없이 model_json_schema()의 모든 필드에 title을 붙이는데, 속성 단위의 이름은 질문을 구성하는 선택지별 choice에 레이블을 달 수 없습니다. 따라서 문구를 조정하는 지렛대는 description입니다 —— 아래 내부 매핑 방식을 참고하십시오.

투영은 정확합니다. enum: [1, 2, 3]은 2를 돌려주지 "2"가 아니며, 범위가 있는 정수는 minimum과 maximum 사이의 단계를 돌려주고, 불리언은 noul >= 0.5입니다.

거부

고정된 선택지 집합으로 답할 수 없는 스키마는 정확한 경로를 밝히며 laya.structured.SchemaError(ValueError)를 발생시킵니다. 각 행도 필드 이름을 name으로 두고 실행됩니다.

속성 스키마 메시지
{"type": "string"} properties.name: a free string cannot be a fixed option set; use 'enum' or a boolean
{"type": "array", "items": {"type": "string"}} properties.name: arrays are not supported; ask one field per element
{"type": "object", "properties": {"inner": {"type": "boolean"}}} properties.name: nested objects are not supported; flatten the schema
{"$ref": "#/definitions/node"} properties.name: $ref/recursion is not supported; flatten the schema
{"enum": [1, "1"]} properties.name: enum values produce duplicate choice labels
{"enum": []} properties.name: 'enum' must not be empty
{"type": "number"} properties.name: a numeric field needs integer 'minimum' and 'maximum' to become a score
{"type": "integer", "minimum": 5, "maximum": 2} properties.name: 'maximum' 2 is below 'minimum' 5
{"type": "integer", "minimum": 0, "maximum": 10} properties.name: 11 levels exceeds MAX_SCORE_LEVELS=10; narrow the range or use an enum
{"anyOf": [{"type": "string"}, {"type": "integer"}]} properties.name: only 'Optional[...]' unions (one non-null branch) are supported, got 2
{"type": ["string", "integer"]} properties.name: 'type' has multiple non-null types; unions are not supported
{"format": "date"} properties.name: unsupported schema {'format': 'date'}
"boolean" properties.name: property must be an object, got str

진입점 자체는 다음을 거부합니다.

호출 메시지
plan_from_json_schema("not a schema") expected a JSON schema object, got str
plan_from_json_schema({"type": "object"}) the top level must be an object with 'properties'
plan_from_json_schema({"type": "object", "properties": {}}) 'properties' must be a non-empty object
plan_from_json_schema({"type": "object", "properties": {"p%d" % i: {"type": "boolean"} for i in range(33)}}) 33 properties exceeds MAX_PROPERTIES=32
plan_from_json_schema({"type": "object", "properties": {"name": {"enum": ["v%d" % i for i in range(33)]}}}) properties.name: 33 options exceeds MAX_OPTIONS=32
decide(None, "I was charged twice.", schema=42) expected a JSON schema dict or a pydantic model, got int

한도: MAX_PROPERTIES = 32, MAX_OPTIONS = 32, MAX_SCORE_LEVELS = 10.

API

함수 용도
laya.decide(runner, state, schema=..., *, questions=..., return_details=..., min_confidence=..., **predict_kwargs) 자유 함수로, Agent와 Router 모두에서 동작합니다
agent.decide(state, schema=..., ...) / router.decide(state, schema=..., ...) 편의 메서드
laya.decide_batch(runner, states, schema=..., ...) / agent.decide_batch(...) / router.decide_batch(...) 여러 상태에 같은 작업을, 한 번의 배치 호출로
questions_from_json_schema(schema) 스키마를 Laya 질문으로
questions_from_pydantic(model) pydantic 모델을 질문으로 (pydantic 필요)
answers_to_json(answers, schema) 원시 답을 스키마 값으로 투영
answer_to_pydantic(model, answers) 원시 답을 pydantic 인스턴스로 투영
plan_from_json_schema(schema) 검증된 필드 계획 (고급)

schema와 questions 중 정확히 하나만 전달하십시오. questions를 쓰면 decide는 투영하지 않고 원시 답을 돌려줍니다. 추가 키워드 인수는 predict로 전달되므로 훅, model=, task=, 토큰 예산이 모두 동작합니다:

router.decide(state, schema=Ticket, model="multilingual", hooks=[Metrics()])

여러 상태 채점

decide_batch는 처리량 지향 형태입니다. 스키마는 한 번만 계획되고 그 질문들이 predict_batch를 통해 모든 상태에 대해 실행되므로, 상태들이 호출마다 하나씩이 아니라 포워드 패스를 공유합니다. 결과는 입력 순서대로 돌아오고 decide가 투영하는 방식과 똑같이 투영되며, return_details=True는 상태마다 DecisionResult 하나를 줍니다:

values = agent.decide_batch(ticket_texts, schema=Ticket)          # values[i] matches ticket_texts[i]
results = router.decide_batch(states, schema=Ticket, return_details=True, batch_size=64)

Router에서는 각 상태가 여전히 개별적으로 라우팅되므로, 한 번의 호출이 여러 체크포인트에 걸칠 수 있습니다. 키워드 인수는 predict_batch에 도달하므로 batch_size=, model=, 훅이 decide에서와 같이 동작합니다. Agent, ONNXAgent, Router 모두 이 메서드를 가집니다. predict_batch가 없는 러너는 조용히 루프로 폴백하지 않고 TypeError를 발생시키므로, 그런 경우 상태마다 decide를 호출하십시오. 배치 처리는 predict_batch가 그렇듯 경계선상의 argmax를 옮길 수 있습니다. README에 두 장치에서 측정한 속도 향상이 기록되어 있습니다.

신뢰도와 확률

기본적으로 decide는 값만 돌려줍니다. return_details=True를 전달하면 필드별 신뢰도, 확률, 원시 답, 호출의 사용량과 라우팅을 담은 DecisionResult를 받습니다:

result = agent.decide(state, schema=Ticket, return_details=True)
result.values["department"]            # "billing"
result.answer_confidence["department"] # 0.94  max(p): the quantity min_confidence gates on
result.confidence["department"]        # 0.71  normalized entropy, which depends on label count
result.probabilities["department"]     # {"billing": 0.94, "support": 0.06, "sales": 0.0}
result.usage                           # {"input_tokens": ..., "output_tokens": 0,
                                       #  "state_tokens": ..., "state_tokens_dropped": ...,
                                       #  "truncated": ..., "truncated_questions": [...]}
result.routing                         # the Router decision, when a Router answered

usage는 predict()가 구성한 블록으로, 그대로 전달됩니다. input_tokens는 질문마다 한 행씩 state를 합산하고, output_tokens는 아무것도 생성되지 않으므로 항상 0이며, 나머지 네 키는 잘림 보고서(#174)입니다. state_tokens는 직렬화된 state 전체에 필요한 양, state_tokens_dropped는 개별 질문의 head가 포기한 최대량, truncated / truncated_questions는 어느 것인지를 나타냅니다. 일곱 번째 키 options는 head 예산으로 인해 어떤 질문의 선택지가 하나의 토큰 구간을 공유하게 된 경우에만 존재합니다(#538). docs/http-api.md는 이 블록과 그 필드를 모두 응답 키로 문서화하며, tests/test_structured_docs.py는 이 페이지의 목록을 그것을 구성하는 코드에 고정합니다.

confidence와 answer_confidence는 서로 다른 양이며, 이름은 정의를 따릅니다. answer_confidence는 max(p), 즉 보고되는 답에 실린 확률 질량입니다. 온도 스케일링이 적합하는 대상이고, 이 저장소의 모든 캘리브레이션 수치가 이것을 기준으로 계산되며, min_confidence가 비교되는 대상입니다 —— 그래서 confidence가 아니라 이것으로 게이팅해야 합니다. confidence는 정규화 엔트로피로, 질문에 선택지가 몇 개였는지에 따라 달라집니다. tests/test_confidence.py는 선택지가 두 개인 분포가 noul에서는 0.90, 동등한 choice에서는 0.53으로 돌아온다고 고정하므로, 이것은 임계값과 비교할 대상이 아닙니다. 사용할 수 있는 answer_confidence를 보고하지 않은 필드는 None에 매핑되며, 이는 보고된 0.0과 다릅니다.

게이트가 사용하는 것과 같은 양으로 게이팅하십시오:

if result.answer_confidence["department"] < 0.6:
    result.values["department"] = "human-review"

answer_confidence가 필터링에 알맞은 숫자라는 것과 신뢰할 수 있는 확률이라는 것은 다릅니다. “c로 반환된 답 가운데 약 c가 맞다”라고 읽는 것은, 해당 체크포인트와 질문 형태에 대해 홀드아웃 데이터로 온도를 적합하고 검증한 뒤에만 성립합니다. 배포되는 체크포인트는 배포된 상태 그대로 과신하며, laya-multilingual은 적합된 온도를 전혀 갖고 있지 않습니다 —— README의 캘리브레이션과 정직한 한계 절, 그리고 적합 루프는 파인튜닝 노트북을 참고하십시오. 그 수준에 의존하기 전에 적합하십시오. 게이트와 평가 하네스가 모두 사용하는 양이므로 보고하십시오.

내부 매핑 방식

  • Enum과 Literal은 값을 문자열 레이블로 삼아 choice 질문이 됩니다. 나올 때 레이블은 원래 값으로 되매핑되므로 정수는 정수로 남습니다.
  • 범위가 있는 정수는 값마다 단계 하나를 가진 score 질문이 됩니다. 반환되는 값은 minimum + argmax입니다.
  • 불리언은 noul 질문이 됩니다. 값은 noul >= 0.5입니다.
  • description은 질문 instructions가 되므로, 좋은 description이 의사결정을 정확하게 만듭니다. 이는 훅 가이드와 같은 규칙을 따릅니다. 각 선택지가 무엇을 뜻하는지 명시하십시오.
  • null 분기는 필드가 계획되기 전에 제거되므로, Optional[X]는 X가 던지는 질문을 그대로 던집니다. 그 필드에 대한 답이 없으면 값에서 키가 그냥 빠지며, 그래서 모델이 보는 것을 바꾸지 않고도 필드를 선택 사항으로 선언해도 안전합니다.

함께 보기