如何用 TypeSafe 构建
如何用 TypeSafe 构建
把控制权留在代码里,把狭窄的结构化决策交给 System One,从而设计 AI 驱动的软件。
System One 是 TypeSafe 用于构建 AI 驱动软件(而不是 agent)的模型。它不生成代码,也不自己选择下一步动作。它提供可嵌入软件中的 AI 原语,让代码保持控制权,由模型负责对非结构化数据做常识判断。
三种软件架构
TypeSafe 是为构建 AI 驱动软件 而设计的:代码掌握工作流,AI 负责狭窄的结构化决策。
传统代码是用简单的软件原语搭成的复杂决策树。因为每个原语都可靠,开发者可以把它们组合成更高层的抽象。
agent 处理指令并选择自己的下一步。有人在旁盯着时这没问题,但每多一次循环,就多一次跑偏的机会。
代码负责确定性工作,掌握控制流。只有当系统需要可编程的常识、或需要解读非结构化数据时,模型才出现。每个 AI 任务都被保持为原子且受约束。


System One 为什么可组合
结构化
System One 在构造上就是类型安全的。决策和概率符合你代码所期望的结构化软件类型和 JSON schema,因此它从不需要从生成的文本里再去还原一个值。
并行
问题彼此独立、并行求值。一个原语的结果不会变成隐藏的上下文,去改变另一个原语的结果。
可比较
输出可排序,能驱动智能的 if 语句、阈值和比较。
快速
大多数查询在约 100 ms 内完成。System One 快到足以用在实时请求路径和用户界面里。
校准的置信度
RLCD 通过校准过的概率来传达不确定性,而不是倾向于过度自信。
自洽
System One 被设计为在反复求值下返回稳定的答案。见自洽性 cookbook。
因为每个输出都受限于给定的选项,模型会返回这些选项上的完整概率分布,而不会造出一个 schema 之外的值。TypeSafe 的目标是超过 100 倍的智能与速度成本比;背后的赌注是:更便宜的智能会带来大得多的需求。
设计一个 System One 工作流
能用代码就用代码
把确定性工作留在代码里。它可靠又便宜。当软件工作流能表达同样的行为时,就避免 agent 的
while循环。示例:把确定性规则留在代码里
days_overdue = (today - invoice.due_date).days if days_overdue > 30: route_to_collections(invoice)浏览 System One 模式,了解把模型决策与代码组合起来的有界方式。
拆解输入状态
只纳入与当前问题相关的上下文。这能帮模型避开干扰和上下文腐化(context rot)。当最新信息可以来自你自己的知识库时,不要依赖模型权重里存的知识。
示例:只发送相关上下文
request { "state": { "ticket_message": "My flight was cancelled. Can I get a refund?", "refund_policy": "Cancelled flights are eligible for a full refund." }, "questions": { "policy_supports_refund": { "type": "noul", "instructions": "Does the refund policy support the refund requested in the ticket?" } } }在输入状态里使用结构
state和questions字段使用嵌套 JSON。当指向具体值能消除歧义时,就让问题指向那些值;问题内部的每个路径都要写上反引号。示例:引用嵌套值
用带反引号的「点号加下标」路径,把问题指向某个具体的嵌套值,例如
support.tickets[0].message。request { "state": { "support": { "tickets": [ { "message": "I was charged twice for order A-104." }, { "message": "How do I reset my password?" } ] }, "commerce": { "orders": [ { "id": "A-104", "charges": [ { "amount_usd": 49, "status": "captured" }, { "amount_usd": 49, "status": "captured" } ] } ] }, "account": { "security": { "password_reset": "Email a reset link to the address on file." } } }, "questions": { "duplicate_charge": { "type": "noul", "instructions": "Do `support.tickets[0].message` and `commerce.orders[0].charges` indicate a duplicate charge?" }, "password_reset_supported": { "type": "noul", "instructions": "Can `account.security.password_reset` resolve the request in `support.tickets[1].message`?" } } }拆解问题
尽量提出最明确、最狭窄、最具体、最原子的问题。把复杂或定义不清的问题拆成一个个独立问题,每个只评估一个属性。
示例:拆解垃圾邮件检测
One broad question (bad) { "is_spam": { "type": "noul", "instructions": "Is `message` spam?" } }Decomposed questions (good) { "requests_credentials": { "type": "noul", "instructions": "Does `message.body` ask the recipient to provide a password or other login credential?" }, "offers_unexpected_reward": { "type": "noul", "instructions": "Does `message.body` claim the recipient received an unexpected prize, payment, or reward?" }, "creates_time_pressure": { "type": "noul", "instructions": "Does `message.subject` or `message.body` pressure the recipient to act quickly?" }, "sender_identity_mismatch": { "type": "noul", "instructions": "Does the organization named in `message.sender.display_name` conflict with the domain in `message.sender.email`?" }, "link_domain_mismatch": { "type": "noul", "instructions": "Does the domain in `message.links[0].url` conflict with the organization named in `message.sender.display_name`?" }, "disguises_link_destination": { "type": "noul", "instructions": "Does `message.links[0].text` conceal or misrepresent the destination in `message.links[0].url`?" } }示例:验证工具调用轨迹
One broad question (bad) { "tool_calls_are_correct": { "type": "noul", "instructions": "Is `trace.tool_calls` correct for `request` and `available_tools`?" } }Decomposed questions (good) { "geocode_tool_is_relevant": { "type": "noul", "instructions": "Is `trace.tool_calls[0].name` an appropriate tool for resolving `request.location`?" }, "geocode_location_matches": { "type": "noul", "instructions": "Does `trace.tool_calls[0].arguments.city` match `request.location`?" }, "geocode_arguments_match_schema": { "type": "noul", "instructions": "Does `trace.tool_calls[0].arguments` conform to `available_tools.geocode_city.parameters`?" }, "geocode_result_matches_call": { "type": "noul", "instructions": "Does `trace.tool_results[0].tool_call_id` match `trace.tool_calls[0].id`?" }, "weather_tool_is_relevant": { "type": "noul", "instructions": "Is `trace.tool_calls[1].name` an appropriate tool for answering `request.text`?" }, "weather_arguments_match_schema": { "type": "noul", "instructions": "Does `trace.tool_calls[1].arguments` conform to `available_tools.get_weather.parameters`?" }, "weather_uses_geocoded_coordinates": { "type": "noul", "instructions": "Do the coordinates in `trace.tool_calls[1].arguments` match those in `trace.tool_results[0].output`?" }, "weather_date_matches": { "type": "noul", "instructions": "Does `trace.tool_calls[1].arguments.date` match `request.date`?" }, "weather_unit_matches": { "type": "noul", "instructions": "Does `trace.tool_calls[1].arguments.unit` match `request.unit`?" } }在问题里使用结构
保持问题简短。
instructions和criteria通常是字符串;对一个简短、无歧义的问题,一个字符串就够了。它们也可以是对象或数组。把问题放在一个字段里,把引导这个问题的数据放在其它字段里。在下列情形里,结构化会更有帮助:
- 问题需要上下文或示例。一长句背景信息,或一列示例输入,应该放在问题旁边的具名字段里,这样你的代码可以增删或替换它们,而不用重写问题。
- 问题的一部分来自你的代码。当某个值来自数据库时,把它放进单独的字段,而不是拼接到字符串模板里。
- 多个问题的 instructions 相近。一个请求接受一个 state,可以包含多个问题。加入补充数据有助于让问题彼此区分。
示例:引用来自你代码的一条记录
这个 Noul 把 state 里的一份简历和候选人数据库里的一条记录做比较。记录原样放进
potential_duplicate,问题通过名字引用它。questions { "same_as_record_18": { "type": "noul", "instructions": { "potential_duplicate": { "name": "John Smith", "location": "Oakland, California", "last_employer": "Google" }, "question": "Is the resume for the same person as `potential_duplicate`?" } } }来自代码的 “potential_duplicate” 数据会随时间变化。“question” 用反引号引用它。
criteria里的描述也可以是对象。对 Choice 来说,每个选项的描述可以是一个对象,说明这个选项涵盖什么、什么属于另一个选项,再加几个示例。各选项使用相同的字段名,这样模型能直接比较。示例:定义对比式的 Choice 判定标准
questions { "card_help_topic": { "type": "choice", "instructions": { "question": "Which disposable virtual card topic is the user asking about?", "focus": "Classify the information the user wants." }, "criteria": { "get_disposable_virtual_card": { "what": "Purpose, eligibility, or setup", "not_for": "Quantity, transaction, or merchant restrictions", "examples": [ "How can I get a disposable virtual card?", "What are disposable cards for?" ] }, "disposable_card_limits": { "what": "Quantity, transaction, or merchant restrictions", "not_for": "Purpose, eligibility, or setup", "examples": [ "How many disposable cards can I make per day?", "Where can I use a disposable card?" ] } } } }每种问题类型的页面都有一个完整示例:
- Noul 把一份简历和若干候选人记录逐一比较,每条记录一个问题,问题在代码里构建。
- Choice 用「各自涵盖什么、不适用于什么、示例」来描述两个容易混淆的选项。
- Score 给每一档一个描述和示例场景。
结构化数据抽取级联 cookbook 展示了共用措辞的情形:对抽取出的记录的每个字段,都问同一组问题。
简短、无歧义的问题或判定标准可以继续用字符串。当结构化能把本会混在一起的引导区分开时,就加上它。关于哪些地方可以接受结构化,见进阶:结构化。
提出大量问题
在一个请求里,针对同一个 state 提出许多狭窄、独立的问题。这就是用这个 API 把效果和「每美元智能」最大化的方式:问题并行运行,代码可以组合它们的信号,而不用增加串行的模型往返。
在代码里组合问题的输出(或喂给经典 ML 模型)
用确定性规则或加权和来组合独立的答案。想要学习式组合时,把概率分布当作下游经典机器学习模型的特征。
示例:用加权分数组合信号
answers = response.answers # Combine independent signals into one application-specific score. quality = ( 0.4 * answers["answers_request"].noul + 0.4 * answers["citations_are_supported"].noul + 0.2 * (1 - answers["contradicts_context"].noul) )组合评分 展示了如何在组合各个判断的同时保留它们。如果下游模型没有标签,就用一组昂贵的推理模型来生成标签;AutoResearch cookbook 展示了如何用 System One 的输出训练一个经典模型。
按不确定性路由
让代码对高置信度和低置信度的答案采取不同的行动。把不确定的案例升级给人工或更昂贵的推理模型。通过在数据上画「置信度—准确率」曲线来测试阈值。
示例:按置信度路由
answer = response.answers["card_help_topic"] if answer.confidence < 0.8: route_to_human_review(ticket) else: route_to_handler(answer.choice, ticket)
把它全部串起来
这个工单处理工作流把确定性工作留在代码里,只发送相关的结构化上下文,在一个请求里评估许多原子问题,并用明确的置信度门控来组合答案。
triage_ticket.py
from typesafe_sdk import Choice, Noul, NoulCriteria, Score, TypeSafeClient
def triage_ticket(ticket, customer):
# Handle deterministic states without calling a model.
if ticket["status"] == "closed":
return "no_action"
open_orders = [
order for order in customer["orders"] if order["status"] != "delivered"
]
# Include only the structured context needed by the questions below.
state = {
"ticket": {
"message": ticket["message"],
"sender": ticket["sender"],
"links": ticket["links"],
},
"customer": {
"plan": customer["plan"],
"open_orders": open_orders,
},
"policy": {
"sensitive_credentials": ["password", "security code", "API key"],
},
}
# Ask structured, atomic questions together so they run in parallel.
questions = {
"topic": Choice(
instructions={
"question": "Which team should handle `ticket.message`?",
"focus": "Classify the customer's primary request.",
},
criteria={
"billing": {
"what": "Charges, invoices, refunds, or subscriptions",
"not_for": "Order tracking or account access",
"examples": ["I was charged twice", "Where is my refund?"],
},
"orders": {
"what": "Order status, delivery, cancellation, or returns",
"not_for": "Charges or account access",
"examples": ["Where is my order?", "Cancel my shipment"],
},
"account": {
"what": "Login, profile, permissions, or security",
"not_for": "Charges or order tracking",
"examples": ["Reset my password", "I cannot sign in"],
},
},
),
"requests_credentials": Noul(
instructions={
"question": "Does the message request a sensitive credential?",
"compare": [
"`ticket.message`",
"`policy.sensitive_credentials`",
],
"focus": "Look for a request to disclose the credential itself.",
},
criteria=NoulCriteria(
true={
"what": "Asks the recipient to disclose a listed credential",
"examples": [
"Reply with your password",
"Send us your API key",
],
},
false={
"what": "Does not ask the recipient to disclose a credential",
"not_for": "A legitimate instruction to reset a credential",
"examples": ["Use this link to reset your password"],
},
),
),
"sender_identity_mismatch": Noul(
instructions={
"question": "Does the claimed sender identity conflict with its domain?",
"compare": [
"`ticket.sender.display_name`",
"`ticket.sender.email`",
],
"focus": "Compare the named organization with the email domain.",
},
criteria=NoulCriteria(
true={
"what": "Claims an organization unrelated to the email domain",
"examples": ["Acme Payroll sent from claim-bonus.example"],
},
false={
"what": "The identity and domain agree or make no conflicting claim",
"examples": ["Acme Payroll sent from acme.example"],
},
),
),
"unexpected_reward": Noul(
instructions={
"question": "Does the message announce an unexpected reward?",
"inspect": "`ticket.message`",
"focus": "Look for an unsolicited prize, payment, or reward claim.",
},
criteria=NoulCriteria(
true={
"what": "Announces an unrequested prize, payment, or reward",
"examples": ["You were selected for a $1,000 bonus"],
},
false={
"what": "Contains no reward claim or discusses an expected payment",
"not_for": "A customer asking about a known refund or payroll deposit",
"examples": ["When will my approved refund arrive?"],
},
),
),
"refund_requested": Noul(
instructions={
"question": "Does the customer explicitly request a refund or credit?",
"inspect": "`ticket.message`",
"focus": "Require a requested remedy, not a billing complaint alone.",
},
criteria=NoulCriteria(
true={
"what": "Directly asks for money back or an account credit",
"examples": ["Please refund the duplicate charge"],
},
false={
"what": "Does not ask for a refund or credit",
"not_for": "A complaint or billing question without a requested remedy",
"examples": ["Why was I charged twice?"],
},
),
),
"mentions_open_order": Noul(
instructions={
"question": "Does the message refer to a supplied open order?",
"compare": [
"`ticket.message`",
"`customer.open_orders`",
],
"focus": "Match an order id or other identifying details.",
},
criteria=NoulCriteria(
true={
"what": "Refers to an open order by id or identifying details",
"examples": ["Where is order A-104?"],
},
false={
"what": "Does not identify any supplied open order",
"not_for": "A generic order question with no matching details",
"examples": ["How long does shipping usually take?"],
},
),
),
"frustration": Score(
instructions={
"question": "How frustrated does the customer appear?",
"inspect": "`ticket.message`",
"focus": "Judge expressed frustration, not issue severity.",
},
criteria=[
{
"what": "Calm and matter-of-fact",
"signals": ["Neutral wording", "No complaint about the experience"],
},
{
"what": "Frustrated but civil",
"signals": ["Expresses annoyance", "Remains constructive"],
},
{
"what": "Very angry or threatening to leave",
"signals": ["Hostile language", "Threatens cancellation or churn"],
},
],
),
}
with TypeSafeClient() as client:
response = client.system_one(
state=state,
questions=questions,
)
# Compose independent spam signals with weights controlled by code.
answers = response.answers
spam_risk = (
0.45 * answers["requests_credentials"].noul
+ 0.30 * answers["sender_identity_mismatch"].noul
+ 0.25 * answers["unexpected_reward"].noul
)
# Escalate uncertain judgments instead of guessing.
spam_is_uncertain = 0.4 < spam_risk < 0.6
if spam_is_uncertain or answers["topic"].confidence < 0.75:
return route_to_human_review(ticket)
if spam_risk >= 0.6:
return quarantine_as_spam(ticket)
# Let code decide which speculative answers matter on this path.
if answers["topic"].choice == "billing":
return route_to_billing(
ticket,
refund_requested=answers["refund_requested"].noul >= 0.7,
)
if answers["topic"].choice == "orders":
return route_to_orders(
ticket,
mentions_open_order=answers["mentions_open_order"].noul >= 0.7,
)
priority = (
"high"
if answers["frustration"].confidence >= 0.7
and answers["frustration"].score >= 1.5
else "normal"
)
return route_to_account_support(ticket, priority=priority)