TypeSafe로 구축하는 방법
제어권을 코드에 두고 System One에 좁고 구조화된 결정을 맡겨, AI 기반 소프트웨어를 설계합니다.
System One은 에이전트가 아니라 AI 기반 소프트웨어를 구축하기 위한 TypeSafe의 모델입니다. 코드를 생성하거나 스스로 다음 행동을 고르지 않습니다. 소프트웨어에 내장되는 AI 프리미티브를 제공하므로, 코드가 제어권을 유지한 채 모델이 비정형 데이터에 대한 상식적 판단을 맡습니다.
세 가지 소프트웨어 아키텍처
TypeSafe는 코드가 워크플로를 소유하고 AI가 좁고 구조화된 결정을 맡는 AI 기반 소프트웨어를 구축하도록 설계되었습니다.
전통적 코드는 단순한 소프트웨어 프리미티브로 만든 복잡한 결정 트리입니다. 각 프리미티브가 신뢰할 수 있으므로, 개발자는 그것들을 더 높은 수준의 추상화로 조합할 수 있습니다.
에이전트는 지시를 처리하고 스스로 다음 단계를 고릅니다. 사람이 과정을 지켜볼 때는 잘 작동하지만, 루프를 한 번 돌 때마다 탈선할 기회가 하나씩 늘어납니다.
코드가 결정적 작업을 처리하고 제어 흐름을 소유합니다. 모델은 시스템에 프로그래밍 가능한 상식이 필요하거나 비정형 데이터를 해석해야 할 때만 등장합니다. 각 AI 작업은 원자적이고 제약된 상태로 유지됩니다.


System One을 조합 가능하게 만드는 것
구조화
System One은 구성상 타입 안전합니다. 결정과 확률은 코드가 기대하는 구조화된 소프트웨어 타입과 JSON schema를 따르므로, 생성된 산문에서 값을 되살릴 필요가 없습니다.
병렬
질문은 독립적이고 병렬로 평가됩니다. 한 프리미티브의 결과가 숨은 컨텍스트가 되어 다른 프리미티브의 결과를 바꾸지 않습니다.
비교 가능
출력은 정렬할 수 있고, 영리한 if 문과 임계값, 비교를 구동할 수 있습니다.
빠름
대부분의 쿼리는 약 100 ms 안에 완료됩니다. System One은 실시간 요청 경로와 사용자 인터페이스에 쓸 만큼 빠릅니다.
캘리브레이션된 신뢰도
RLCD는 과신으로 기울기보다 캘리브레이션된 확률로 불확실성을 전달합니다.
자기 일관성
System One은 반복 평가에서도 안정적인 답을 반환하도록 설계되었습니다. 자기 일관성 쿡북을 참조하십시오.
모든 출력이 주어진 옵션으로 제약되므로, 모델은 schema 밖의 값을 지어내는 대신 그 옵션들에 대한 완전한 확률 분포를 반환합니다. TypeSafe의 목표는 지능 대비 속도·비용 비율 100배 이상입니다. 그 밑에 깔린 베팅은 더 저렴한 지능이 훨씬 더 많은 수요를 만든다는 것입니다.
System One 워크플로 설계하기
가능하면 코드를 사용하십시오
결정적 작업은 코드에 두십시오. 그것이 신뢰할 수 있고 저렴합니다. 소프트웨어 워크플로가 같은 동작을 표현할 수 있다면 에이전트
while루프는 피하십시오.예시: 결정적 규칙을 코드에 두기
days_overdue = (today - invoice.due_date).days if days_overdue > 30: route_to_collections(invoice)모델 결정을 코드와 조합하는 경계 지어진 방법들은 System One 패턴에서 살펴보십시오.
입력 상태를 분해하십시오
현재 질문에 관련된 컨텍스트만 포함하십시오. 이것은 모델이 방해 요소와 컨텍스트 부패를 피하도록 돕습니다. 최신 정보를 자체 지식 베이스에서 가져올 수 있을 때는 모델 가중치에 저장된 지식에 의존하지 마십시오.
예시: 관련 컨텍스트만 보내기
request { "state": { "ticket_message": "My flight was cancelled. Can I get a refund?", "refund_policy": "Cancelled flights are eligible for a full refund." }, "questions": { "policy_supports_refund": { "type": "noul", "instructions": "Does the refund policy support the refund requested in the ticket?" } } }입력 상태에 구조를 사용하십시오
state와questions필드에 중첩 JSON을 사용하십시오. 모호함이 사라진다면 질문이 특정 값을 가리키게 하고, 질문 안의 각 경로를 백틱 문자로 감싸십시오.예시: 중첩 값 참조하기
백틱으로 감싼 점과 인덱스 경로를 사용해, 질문이
support.tickets[0].message같은 특정 중첩 값을 가리키게 하십시오.request { "state": { "support": { "tickets": [ { "message": "I was charged twice for order A-104." }, { "message": "How do I reset my password?" } ] }, "commerce": { "orders": [ { "id": "A-104", "charges": [ { "amount_usd": 49, "status": "captured" }, { "amount_usd": 49, "status": "captured" } ] } ] }, "account": { "security": { "password_reset": "Email a reset link to the address on file." } } }, "questions": { "duplicate_charge": { "type": "noul", "instructions": "Do `support.tickets[0].message` and `commerce.orders[0].charges` indicate a duplicate charge?" }, "password_reset_supported": { "type": "noul", "instructions": "Can `account.security.password_reset` resolve the request in `support.tickets[1].message`?" } } }질문을 분해하십시오
가능한 한 명시적이고 좁고 구체적인 원자적 질문을 던지십시오. 복잡하거나 잘 정의되지 않은 질문을, 각각 하나의 속성만 평가하는 별개의 질문으로 나누십시오.
예시: 스팸 탐지 분해하기
하나의 넓은 질문 (나쁨) { "is_spam": { "type": "noul", "instructions": "Is `message` spam?" } }분해한 질문 (좋음) { "requests_credentials": { "type": "noul", "instructions": "Does `message.body` ask the recipient to provide a password or other login credential?" }, "offers_unexpected_reward": { "type": "noul", "instructions": "Does `message.body` claim the recipient received an unexpected prize, payment, or reward?" }, "creates_time_pressure": { "type": "noul", "instructions": "Does `message.subject` or `message.body` pressure the recipient to act quickly?" }, "sender_identity_mismatch": { "type": "noul", "instructions": "Does the organization named in `message.sender.display_name` conflict with the domain in `message.sender.email`?" }, "link_domain_mismatch": { "type": "noul", "instructions": "Does the domain in `message.links[0].url` conflict with the organization named in `message.sender.display_name`?" }, "disguises_link_destination": { "type": "noul", "instructions": "Does `message.links[0].text` conceal or misrepresent the destination in `message.links[0].url`?" } }예시: 도구 호출 트레이스 검증하기
하나의 넓은 질문 (나쁨) { "tool_calls_are_correct": { "type": "noul", "instructions": "Is `trace.tool_calls` correct for `request` and `available_tools`?" } }분해한 질문 (좋음) { "geocode_tool_is_relevant": { "type": "noul", "instructions": "Is `trace.tool_calls[0].name` an appropriate tool for resolving `request.location`?" }, "geocode_location_matches": { "type": "noul", "instructions": "Does `trace.tool_calls[0].arguments.city` match `request.location`?" }, "geocode_arguments_match_schema": { "type": "noul", "instructions": "Does `trace.tool_calls[0].arguments` conform to `available_tools.geocode_city.parameters`?" }, "geocode_result_matches_call": { "type": "noul", "instructions": "Does `trace.tool_results[0].tool_call_id` match `trace.tool_calls[0].id`?" }, "weather_tool_is_relevant": { "type": "noul", "instructions": "Is `trace.tool_calls[1].name` an appropriate tool for answering `request.text`?" }, "weather_arguments_match_schema": { "type": "noul", "instructions": "Does `trace.tool_calls[1].arguments` conform to `available_tools.get_weather.parameters`?" }, "weather_uses_geocoded_coordinates": { "type": "noul", "instructions": "Do the coordinates in `trace.tool_calls[1].arguments` match those in `trace.tool_results[0].output`?" }, "weather_date_matches": { "type": "noul", "instructions": "Does `trace.tool_calls[1].arguments.date` match `request.date`?" }, "weather_unit_matches": { "type": "noul", "instructions": "Does `trace.tool_calls[1].arguments.unit` match `request.unit`?" } }질문에 구조를 사용하십시오
질문은 짧게 유지하십시오.
instructions와criteria는 보통 문자열이며, 짧고 모호하지 않은 질문에는 문자열이면 충분합니다. 객체나 배열일 수도 있습니다. 질문은 한 필드에, 질문을 이끄는 데이터는 다른 필드에 두십시오.구조는 다음 상황에서 도움이 됩니다.
- 질문에 컨텍스트나 예시가 필요할 때. 길게 늘어선 배경 정보 문장이나 예시 입력 목록은 질문 옆의 이름 있는 필드에 두면, 질문을 다시 쓰지 않고도 코드에서 거기에 더하거나 교체할 수 있습니다.
- 질문의 일부가 코드에서 올 때. 값이 데이터베이스에서 오면, 문자열 템플릿에 끼워 넣는 대신 별도 필드에 두십시오.
- 여러 질문이 비슷한 instructions를 가질 때. 한 요청은 하나의 상태를 받고 여러 질문을 포함할 수 있습니다. 보충 데이터를 더하면 질문을 서로 구별하는 데 도움이 됩니다.
예시: 코드의 레코드 참조하기
이 Noul은 상태의 이력서를 후보 데이터베이스의 레코드와 비교합니다. 레코드는 그대로
potential_duplicate에 들어가고, 질문은 이름으로 그것을 가리킵니다.questions { "same_as_record_18": { "type": "noul", "instructions": { "potential_duplicate": { "name": "John Smith", "location": "Oakland, California", "last_employer": "Google" }, "question": "Is the resume for the same person as `potential_duplicate`?" } } }코드에서 온 “potential_duplicate” 데이터는 시간이 지나면 바뀔 수 있습니다. “question”은 백틱을 사용해 그것을 가리킵니다.
criteria안의 설명도 객체일 수 있습니다. Choice의 경우 각 옵션의 설명은 그 옵션이 무엇을 다루는지, 무엇이 다른 옵션에 속하는지, 그리고 몇 가지 예시를 담은 객체일 수 있습니다. 모델이 직접 비교할 수 있도록 옵션 전반에서 같은 필드 이름을 사용하십시오.예시: 대조적인 Choice 기준 정의하기
questions { "card_help_topic": { "type": "choice", "instructions": { "question": "Which disposable virtual card topic is the user asking about?", "focus": "Classify the information the user wants." }, "criteria": { "get_disposable_virtual_card": { "what": "Purpose, eligibility, or setup", "not_for": "Quantity, transaction, or merchant restrictions", "examples": [ "How can I get a disposable virtual card?", "What are disposable cards for?" ] }, "disposable_card_limits": { "what": "Quantity, transaction, or merchant restrictions", "not_for": "Purpose, eligibility, or setup", "examples": [ "How many disposable cards can I make per day?", "Where can I use a disposable card?" ] } } } }각 질문 유형 페이지에는 구체적인 예제가 있습니다.
- Noul은 이력서 하나를 여러 후보 레코드와 비교하며, 레코드마다 질문 하나를 두고, 질문은 코드에서 만듭니다.
- Choice는 혼동하기 쉬운 두 옵션을, 각각이 무엇을 다루는지, 무엇을 위한 것이 아닌지, 예시와 함께 설명합니다.
- Score는 각 레벨에 설명과 예시 상황을 부여합니다.
구조화 데이터 추출 캐스케이드 쿡북은 추출된 레코드의 모든 필드에 같은 질문 묶음을 던지는, 공통 문구를 쓰는 경우를 보여 줍니다.
짧고 모호하지 않은 질문이나 기준은 문자열로 남겨 둘 수 있습니다. 그렇지 않으면 뒤섞일 안내를 구분해 줄 때 구조를 추가하십시오. 구조가 허용되는 모든 위치는 고급: 구조를 참조하십시오.
질문을 많이 던지십시오
같은 상태에 대해 좁고 독립적인 질문을 한 요청에 많이 던지십시오. 이것이 API로 비용 대비 효과와 지능을 최대화하는 방법입니다. 질문은 병렬로 실행되고, 코드는 직렬 모델 왕복을 늘리지 않고도 그 신호를 결합할 수 있습니다.
투기적 팬아웃 패턴과 병렬 질문 쿡북을 참조하십시오.
질문 출력을 코드에서 결합하십시오 (또는 고전 ML 모델에 넣으십시오)
독립적인 답을 결정적 규칙이나 가중합으로 결합하십시오. 학습된 조합에는 확률을 다운스트림 고전 머신러닝 모델의 특징으로 사용하십시오.
예시: 가중 점수로 신호 결합하기
answers = response.answers # Combine independent signals into one application-specific score. quality = ( 0.4 * answers["answers_request"].noul + 0.4 * answers["citations_are_supported"].noul + 0.2 * (1 - answers["contradicts_context"].noul) )복합 스코어링은 개별 판단을 보존하면서 결합하는 방법을 보여 줍니다. 다운스트림 모델의 레이블이 없다면 비싼 추론 모델의 앙상블로 레이블을 생성하십시오. AutoResearch 쿡북은 System One 출력으로 고전 모델을 훈련하는 방법을 보여 줍니다.
불확실성에 따라 라우팅하십시오
확신 있는 답과 그렇지 않은 답에 대해 코드가 다른 행동을 하도록 만드십시오. 불확실한 사례는 사람이나 더 비싼 추론 모델로 에스컬레이션하십시오. 자체 데이터에서 신뢰도 대 정확도를 그려 임계값을 테스트하십시오.
예시: 신뢰도로 라우팅하기
answer = response.answers["card_help_topic"] if answer.confidence < 0.8: route_to_human_review(ticket) else: route_to_handler(answer.choice, ticket)임계값을 고르고 각 행동의 위험에 맞추는 방법은 신뢰도와 신뢰도 게이팅 라우팅을 참조하십시오.
모두 합치기
이 지원 티켓 워크플로는 결정적 작업을 코드에 두고, 관련된 구조화 컨텍스트만 보내고, 많은 원자적 질문을 한 요청에서 평가하고, 명시적 신뢰도 게이트로 답을 조합합니다.
triage_ticket.py
from typesafe_sdk import Choice, Noul, NoulCriteria, Score, TypeSafeClient
def triage_ticket(ticket, customer):
# Handle deterministic states without calling a model.
if ticket["status"] == "closed":
return "no_action"
open_orders = [
order for order in customer["orders"] if order["status"] != "delivered"
]
# Include only the structured context needed by the questions below.
state = {
"ticket": {
"message": ticket["message"],
"sender": ticket["sender"],
"links": ticket["links"],
},
"customer": {
"plan": customer["plan"],
"open_orders": open_orders,
},
"policy": {
"sensitive_credentials": ["password", "security code", "API key"],
},
}
# Ask structured, atomic questions together so they run in parallel.
questions = {
"topic": Choice(
instructions={
"question": "Which team should handle `ticket.message`?",
"focus": "Classify the customer's primary request.",
},
criteria={
"billing": {
"what": "Charges, invoices, refunds, or subscriptions",
"not_for": "Order tracking or account access",
"examples": ["I was charged twice", "Where is my refund?"],
},
"orders": {
"what": "Order status, delivery, cancellation, or returns",
"not_for": "Charges or account access",
"examples": ["Where is my order?", "Cancel my shipment"],
},
"account": {
"what": "Login, profile, permissions, or security",
"not_for": "Charges or order tracking",
"examples": ["Reset my password", "I cannot sign in"],
},
},
),
"requests_credentials": Noul(
instructions={
"question": "Does the message request a sensitive credential?",
"compare": [
"`ticket.message`",
"`policy.sensitive_credentials`",
],
"focus": "Look for a request to disclose the credential itself.",
},
criteria=NoulCriteria(
true={
"what": "Asks the recipient to disclose a listed credential",
"examples": [
"Reply with your password",
"Send us your API key",
],
},
false={
"what": "Does not ask the recipient to disclose a credential",
"not_for": "A legitimate instruction to reset a credential",
"examples": ["Use this link to reset your password"],
},
),
),
"sender_identity_mismatch": Noul(
instructions={
"question": "Does the claimed sender identity conflict with its domain?",
"compare": [
"`ticket.sender.display_name`",
"`ticket.sender.email`",
],
"focus": "Compare the named organization with the email domain.",
},
criteria=NoulCriteria(
true={
"what": "Claims an organization unrelated to the email domain",
"examples": ["Acme Payroll sent from claim-bonus.example"],
},
false={
"what": "The identity and domain agree or make no conflicting claim",
"examples": ["Acme Payroll sent from acme.example"],
},
),
),
"unexpected_reward": Noul(
instructions={
"question": "Does the message announce an unexpected reward?",
"inspect": "`ticket.message`",
"focus": "Look for an unsolicited prize, payment, or reward claim.",
},
criteria=NoulCriteria(
true={
"what": "Announces an unrequested prize, payment, or reward",
"examples": ["You were selected for a $1,000 bonus"],
},
false={
"what": "Contains no reward claim or discusses an expected payment",
"not_for": "A customer asking about a known refund or payroll deposit",
"examples": ["When will my approved refund arrive?"],
},
),
),
"refund_requested": Noul(
instructions={
"question": "Does the customer explicitly request a refund or credit?",
"inspect": "`ticket.message`",
"focus": "Require a requested remedy, not a billing complaint alone.",
},
criteria=NoulCriteria(
true={
"what": "Directly asks for money back or an account credit",
"examples": ["Please refund the duplicate charge"],
},
false={
"what": "Does not ask for a refund or credit",
"not_for": "A complaint or billing question without a requested remedy",
"examples": ["Why was I charged twice?"],
},
),
),
"mentions_open_order": Noul(
instructions={
"question": "Does the message refer to a supplied open order?",
"compare": [
"`ticket.message`",
"`customer.open_orders`",
],
"focus": "Match an order id or other identifying details.",
},
criteria=NoulCriteria(
true={
"what": "Refers to an open order by id or identifying details",
"examples": ["Where is order A-104?"],
},
false={
"what": "Does not identify any supplied open order",
"not_for": "A generic order question with no matching details",
"examples": ["How long does shipping usually take?"],
},
),
),
"frustration": Score(
instructions={
"question": "How frustrated does the customer appear?",
"inspect": "`ticket.message`",
"focus": "Judge expressed frustration, not issue severity.",
},
criteria=[
{
"what": "Calm and matter-of-fact",
"signals": ["Neutral wording", "No complaint about the experience"],
},
{
"what": "Frustrated but civil",
"signals": ["Expresses annoyance", "Remains constructive"],
},
{
"what": "Very angry or threatening to leave",
"signals": ["Hostile language", "Threatens cancellation or churn"],
},
],
),
}
with TypeSafeClient() as client:
response = client.system_one(
state=state,
questions=questions,
)
# Compose independent spam signals with weights controlled by code.
answers = response.answers
spam_risk = (
0.45 * answers["requests_credentials"].noul
+ 0.30 * answers["sender_identity_mismatch"].noul
+ 0.25 * answers["unexpected_reward"].noul
)
# Escalate uncertain judgments instead of guessing.
spam_is_uncertain = 0.4 < spam_risk < 0.6
if spam_is_uncertain or answers["topic"].confidence < 0.75:
return route_to_human_review(ticket)
if spam_risk >= 0.6:
return quarantine_as_spam(ticket)
# Let code decide which speculative answers matter on this path.
if answers["topic"].choice == "billing":
return route_to_billing(
ticket,
refund_requested=answers["refund_requested"].noul >= 0.7,
)
if answers["topic"].choice == "orders":
return route_to_orders(
ticket,
mentions_open_order=answers["mentions_open_order"].noul >= 0.7,
)
priority = (
"high"
if answers["frustration"].confidence >= 0.7
and answers["frustration"].score >= 1.5
else "normal"
)
return route_to_account_support(ticket, priority=priority)