文件導航

如何用 TypeSafe 構建

把控制權留在程式碼裡,把狹窄的結構化決策交給 System One,從而設計 AI 驅動的軟體。

System One 是 TypeSafe 用於構建 AI 驅動軟體(而不是 agent)的模型。它不生成程式碼,也不自己選擇下一步動作。它提供可嵌入軟體中的 AI 原語,讓程式碼保持控制權,由模型負責對非結構化資料做常識判斷。

三種軟體架構

TypeSafe 是為構建 AI 驅動軟體 而設計的:程式碼掌握工作流,AI 負責狹窄的結構化決策。

傳統程式碼是用簡單的軟體原語搭成的複雜決策樹。因為每個原語都可靠,開發者可以把它們組合成更高層的抽象。

agent 處理指令並選擇自己的下一步。有人在旁盯著時這沒問題,但每多一次迴圈,就多一次跑偏的機會。

程式碼負責確定性工作,掌握控制流。只有當系統需要可程式設計的常識、或需要解讀非結構化資料時,模型才出現。每個 AI 任務都被保持為原子且受約束。

傳統軟體、agent 與 AI 驅動軟體,展示為三種不同的系統架構。

System One 為什麼可組合

結構化

System One 在構造上就是型別安全的。決策和機率符合你程式碼所期望的結構化軟體型別和 JSON schema,因此它從不需要從生成的文本里再去還原一個值。

並行

問題彼此獨立、並行求值。一個原語的結果不會變成隱藏的上下文,去改變另一個原語的結果。

可比較

輸出可排序,能驅動智慧的 if 語句、閾值和比較。

快速

大多數查詢在約 100 ms 內完成。System One 快到足以用在即時請求路徑和使用者介面裡。

校準的置信度

RLCD 通過校準過的機率來傳達不確定性,而不是傾向於過度自信。

自洽

System One 被設計為在反覆求值下返回穩定的答案。見自洽性 cookbook。

因為每個輸出都受限於給定的選項,模型會返回這些選項上的完整機率分佈,而不會造出一個 schema 之外的值。TypeSafe 的目標是超過 100 倍的智慧與速度成本比;背後的賭注是:更便宜的智慧會帶來大得多的需求。

設計一個 System One 工作流

  1. 能用程式碼就用程式碼

    把確定性工作留在程式碼裡。它可靠又便宜。當軟體工作流能表達同樣的行為時,就避免 agent 的 while 迴圈。

    示例:把確定性規則留在程式碼裡
    days_overdue = (today - invoice.due_date).days
    
    if days_overdue > 30:
        route_to_collections(invoice)

    瀏覽 System One 模式,瞭解把模型決策與程式碼組合起來的有界方式。

  2. 拆解輸入狀態

    只納入與當前問題相關的上下文。這能幫模型避開干擾和上下文腐化(context rot)。當最新資訊可以來自你自己的知識庫時,不要依賴模型權重裡存的知識。

    示例:只發送相關上下文
    request
    {
      "state": {
        "ticket_message": "My flight was cancelled. Can I get a refund?",
        "refund_policy": "Cancelled flights are eligible for a full refund."
      },
      "questions": {
        "policy_supports_refund": {
          "type": "noul",
          "instructions": "Does the refund policy support the refund requested in the ticket?"
        }
      }
    }
  3. 在輸入狀態裡使用結構

    state 和 questions 欄位使用巢狀 JSON。當指向具體值能消除歧義時,就讓問題指向那些值;問題內部的每個路徑都要寫上反引號。

    示例:引用巢狀值

    用帶反引號的「點號加下標」路徑,把問題指向某個具體的巢狀值,例如 support.tickets[0].message。

    request
    {
      "state": {
        "support": {
          "tickets": [
            {
              "message": "I was charged twice for order A-104."
            },
            {
              "message": "How do I reset my password?"
            }
          ]
        },
        "commerce": {
          "orders": [
            {
              "id": "A-104",
              "charges": [
                {
                  "amount_usd": 49,
                  "status": "captured"
                },
                {
                  "amount_usd": 49,
                  "status": "captured"
                }
              ]
            }
          ]
        },
        "account": {
          "security": {
            "password_reset": "Email a reset link to the address on file."
          }
        }
      },
      "questions": {
        "duplicate_charge": {
          "type": "noul",
          "instructions": "Do `support.tickets[0].message` and `commerce.orders[0].charges` indicate a duplicate charge?"
        },
        "password_reset_supported": {
          "type": "noul",
          "instructions": "Can `account.security.password_reset` resolve the request in `support.tickets[1].message`?"
        }
      }
    }
  4. 拆解問題

    儘量提出最明確、最狹窄、最具體、最原子的問題。把複雜或定義不清的問題拆成一個個獨立問題,每個只評估一個屬性。

    示例:拆解垃圾郵件檢測
    One broad question (bad)
    {
      "is_spam": {
        "type": "noul",
        "instructions": "Is `message` spam?"
      }
    }
    Decomposed questions (good)
    {
      "requests_credentials": {
        "type": "noul",
        "instructions": "Does `message.body` ask the recipient to provide a password or other login credential?"
      },
      "offers_unexpected_reward": {
        "type": "noul",
        "instructions": "Does `message.body` claim the recipient received an unexpected prize, payment, or reward?"
      },
      "creates_time_pressure": {
        "type": "noul",
        "instructions": "Does `message.subject` or `message.body` pressure the recipient to act quickly?"
      },
      "sender_identity_mismatch": {
        "type": "noul",
        "instructions": "Does the organization named in `message.sender.display_name` conflict with the domain in `message.sender.email`?"
      },
      "link_domain_mismatch": {
        "type": "noul",
        "instructions": "Does the domain in `message.links[0].url` conflict with the organization named in `message.sender.display_name`?"
      },
      "disguises_link_destination": {
        "type": "noul",
        "instructions": "Does `message.links[0].text` conceal or misrepresent the destination in `message.links[0].url`?"
      }
    }
    示例:驗證工具呼叫軌跡
    One broad question (bad)
    {
      "tool_calls_are_correct": {
        "type": "noul",
        "instructions": "Is `trace.tool_calls` correct for `request` and `available_tools`?"
      }
    }
    Decomposed questions (good)
    {
      "geocode_tool_is_relevant": {
        "type": "noul",
        "instructions": "Is `trace.tool_calls[0].name` an appropriate tool for resolving `request.location`?"
      },
      "geocode_location_matches": {
        "type": "noul",
        "instructions": "Does `trace.tool_calls[0].arguments.city` match `request.location`?"
      },
      "geocode_arguments_match_schema": {
        "type": "noul",
        "instructions": "Does `trace.tool_calls[0].arguments` conform to `available_tools.geocode_city.parameters`?"
      },
      "geocode_result_matches_call": {
        "type": "noul",
        "instructions": "Does `trace.tool_results[0].tool_call_id` match `trace.tool_calls[0].id`?"
      },
      "weather_tool_is_relevant": {
        "type": "noul",
        "instructions": "Is `trace.tool_calls[1].name` an appropriate tool for answering `request.text`?"
      },
      "weather_arguments_match_schema": {
        "type": "noul",
        "instructions": "Does `trace.tool_calls[1].arguments` conform to `available_tools.get_weather.parameters`?"
      },
      "weather_uses_geocoded_coordinates": {
        "type": "noul",
        "instructions": "Do the coordinates in `trace.tool_calls[1].arguments` match those in `trace.tool_results[0].output`?"
      },
      "weather_date_matches": {
        "type": "noul",
        "instructions": "Does `trace.tool_calls[1].arguments.date` match `request.date`?"
      },
      "weather_unit_matches": {
        "type": "noul",
        "instructions": "Does `trace.tool_calls[1].arguments.unit` match `request.unit`?"
      }
    }
  5. 在問題裡使用結構

    保持問題簡短。instructions 和 criteria 通常是字串;對一個簡短、無歧義的問題,一個字串就夠了。它們也可以是物件或陣列。把問題放在一個欄位裡,把引導這個問題的資料放在其它欄位裡。

    在下列情形裡,結構化會更有幫助:

    • 問題需要上下文或示例。一長句背景資訊,或一列示例輸入,應該放在問題旁邊的具名欄位裡,這樣你的程式碼可以增刪或替換它們,而不用重寫問題。
    • 問題的一部分來自你的程式碼。當某個值來自資料庫時,把它放進單獨的欄位,而不是拼接到字串模板裡。
    • 多個問題的 instructions 相近。一個請求接受一個 state,可以包含多個問題。加入補充資料有助於讓問題彼此區分。
    示例:引用來自你程式碼的一條記錄

    這個 Noul 把 state 裡的一份簡歷和候選人資料庫裡的一條記錄做比較。記錄原樣放進 potential_duplicate,問題通過名字引用它。

    questions
    {
      "same_as_record_18": {
        "type": "noul",
        "instructions": {
          "potential_duplicate": {
            "name": "John Smith",
            "location": "Oakland, California",
            "last_employer": "Google"
          },
          "question": "Is the resume for the same person as `potential_duplicate`?"
        }
      }
    }

    來自程式碼的 “potential_duplicate” 資料會隨時間變化。“question” 用反引號引用它。

    criteria 裡的描述也可以是物件。對 Choice 來說,每個選項的描述可以是一個物件,說明這個選項涵蓋什麼、什麼屬於另一個選項,再加幾個示例。各選項使用相同的欄位名,這樣模型能直接比較。

    示例:定義對比式的 Choice 判定標準
    questions
    {
      "card_help_topic": {
        "type": "choice",
        "instructions": {
          "question": "Which disposable virtual card topic is the user asking about?",
          "focus": "Classify the information the user wants."
        },
        "criteria": {
          "get_disposable_virtual_card": {
            "what": "Purpose, eligibility, or setup",
            "not_for": "Quantity, transaction, or merchant restrictions",
            "examples": [
              "How can I get a disposable virtual card?",
              "What are disposable cards for?"
            ]
          },
          "disposable_card_limits": {
            "what": "Quantity, transaction, or merchant restrictions",
            "not_for": "Purpose, eligibility, or setup",
            "examples": [
              "How many disposable cards can I make per day?",
              "Where can I use a disposable card?"
            ]
          }
        }
      }
    }

    每種問題型別的頁面都有一個完整示例:

    • Noul 把一份簡歷和若干候選人記錄逐一比較,每條記錄一個問題,問題在程式碼裡構建。
    • Choice 用「各自涵蓋什麼、不適用於什麼、示例」來描述兩個容易混淆的選項。
    • Score 給每一檔一個描述和示例場景。

    結構化資料抽取級聯 cookbook 展示了共用措辭的情形:對抽取出的記錄的每個欄位,都問同一組問題。

    簡短、無歧義的問題或判定標準可以繼續用字串。當結構化能把本會混在一起的引導區分開時,就加上它。關於哪些地方可以接受結構化,見進階:結構化。

  6. 提出大量問題

    在一個請求裡,針對同一個 state 提出許多狹窄、獨立的問題。這就是用這個 API 把效果和「每美元智慧」最大化的方式:問題並行執行,程式碼可以組合它們的訊號,而不用增加序列的模型往返。

    見推測式扇出模式和並行提問 cookbook。

  7. 在程式碼裡組合問題的輸出(或餵給經典 ML 模型)

    用確定性規則或加權和來組合獨立的答案。想要學習式組合時,把機率分佈當作下游經典機器學習模型的特徵。

    示例:用加權分數組合訊號
    answers = response.answers
    
    # Combine independent signals into one application-specific score.
    quality = (
        0.4 * answers["answers_request"].noul
        + 0.4 * answers["citations_are_supported"].noul
        + 0.2 * (1 - answers["contradicts_context"].noul)
    )

    組合評分 展示瞭如何在組合各個判斷的同時保留它們。如果下游模型沒有標籤,就用一組昂貴的推理模型來生成標籤;AutoResearch cookbook 展示瞭如何用 System One 的輸出訓練一個經典模型。

  8. 按不確定性路由

    讓程式碼對高置信度和低置信度的答案採取不同的行動。把不確定的案例升級給人工或更昂貴的推理模型。通過在資料上畫「置信度—準確率」曲線來測試閾值。

    示例:按置信度路由
    answer = response.answers["card_help_topic"]
    
    if answer.confidence < 0.8:
        route_to_human_review(ticket)
    else:
        route_to_handler(answer.choice, ticket)

    如何選擇閾值並把它們對應到每個行動的風險,見置信度和按置信度門控的路由。

把它全部串起來

這個工單處理工作流把確定性工作留在程式碼裡,只發送相關的結構化上下文,在一個請求裡評估許多原子問題,並用明確的置信度門控來組合答案。

triage_ticket.py

from typesafe_sdk import Choice, Noul, NoulCriteria, Score, TypeSafeClient

def triage_ticket(ticket, customer):
    # Handle deterministic states without calling a model.
    if ticket["status"] == "closed":
        return "no_action"

    open_orders = [
        order for order in customer["orders"] if order["status"] != "delivered"
    ]

    # Include only the structured context needed by the questions below.
    state = {
        "ticket": {
            "message": ticket["message"],
            "sender": ticket["sender"],
            "links": ticket["links"],
        },
        "customer": {
            "plan": customer["plan"],
            "open_orders": open_orders,
        },
        "policy": {
            "sensitive_credentials": ["password", "security code", "API key"],
        },
    }

    # Ask structured, atomic questions together so they run in parallel.
    questions = {
        "topic": Choice(
            instructions={
                "question": "Which team should handle `ticket.message`?",
                "focus": "Classify the customer's primary request.",
            },
            criteria={
                "billing": {
                    "what": "Charges, invoices, refunds, or subscriptions",
                    "not_for": "Order tracking or account access",
                    "examples": ["I was charged twice", "Where is my refund?"],
                },
                "orders": {
                    "what": "Order status, delivery, cancellation, or returns",
                    "not_for": "Charges or account access",
                    "examples": ["Where is my order?", "Cancel my shipment"],
                },
                "account": {
                    "what": "Login, profile, permissions, or security",
                    "not_for": "Charges or order tracking",
                    "examples": ["Reset my password", "I cannot sign in"],
                },
            },
        ),
        "requests_credentials": Noul(
            instructions={
                "question": "Does the message request a sensitive credential?",
                "compare": [
                    "`ticket.message`",
                    "`policy.sensitive_credentials`",
                ],
                "focus": "Look for a request to disclose the credential itself.",
            },
            criteria=NoulCriteria(
                true={
                    "what": "Asks the recipient to disclose a listed credential",
                    "examples": [
                        "Reply with your password",
                        "Send us your API key",
                    ],
                },
                false={
                    "what": "Does not ask the recipient to disclose a credential",
                    "not_for": "A legitimate instruction to reset a credential",
                    "examples": ["Use this link to reset your password"],
                },
            ),
        ),
        "sender_identity_mismatch": Noul(
            instructions={
                "question": "Does the claimed sender identity conflict with its domain?",
                "compare": [
                    "`ticket.sender.display_name`",
                    "`ticket.sender.email`",
                ],
                "focus": "Compare the named organization with the email domain.",
            },
            criteria=NoulCriteria(
                true={
                    "what": "Claims an organization unrelated to the email domain",
                    "examples": ["Acme Payroll sent from claim-bonus.example"],
                },
                false={
                    "what": "The identity and domain agree or make no conflicting claim",
                    "examples": ["Acme Payroll sent from acme.example"],
                },
            ),
        ),
        "unexpected_reward": Noul(
            instructions={
                "question": "Does the message announce an unexpected reward?",
                "inspect": "`ticket.message`",
                "focus": "Look for an unsolicited prize, payment, or reward claim.",
            },
            criteria=NoulCriteria(
                true={
                    "what": "Announces an unrequested prize, payment, or reward",
                    "examples": ["You were selected for a $1,000 bonus"],
                },
                false={
                    "what": "Contains no reward claim or discusses an expected payment",
                    "not_for": "A customer asking about a known refund or payroll deposit",
                    "examples": ["When will my approved refund arrive?"],
                },
            ),
        ),
        "refund_requested": Noul(
            instructions={
                "question": "Does the customer explicitly request a refund or credit?",
                "inspect": "`ticket.message`",
                "focus": "Require a requested remedy, not a billing complaint alone.",
            },
            criteria=NoulCriteria(
                true={
                    "what": "Directly asks for money back or an account credit",
                    "examples": ["Please refund the duplicate charge"],
                },
                false={
                    "what": "Does not ask for a refund or credit",
                    "not_for": "A complaint or billing question without a requested remedy",
                    "examples": ["Why was I charged twice?"],
                },
            ),
        ),
        "mentions_open_order": Noul(
            instructions={
                "question": "Does the message refer to a supplied open order?",
                "compare": [
                    "`ticket.message`",
                    "`customer.open_orders`",
                ],
                "focus": "Match an order id or other identifying details.",
            },
            criteria=NoulCriteria(
                true={
                    "what": "Refers to an open order by id or identifying details",
                    "examples": ["Where is order A-104?"],
                },
                false={
                    "what": "Does not identify any supplied open order",
                    "not_for": "A generic order question with no matching details",
                    "examples": ["How long does shipping usually take?"],
                },
            ),
        ),
        "frustration": Score(
            instructions={
                "question": "How frustrated does the customer appear?",
                "inspect": "`ticket.message`",
                "focus": "Judge expressed frustration, not issue severity.",
            },
            criteria=[
                {
                    "what": "Calm and matter-of-fact",
                    "signals": ["Neutral wording", "No complaint about the experience"],
                },
                {
                    "what": "Frustrated but civil",
                    "signals": ["Expresses annoyance", "Remains constructive"],
                },
                {
                    "what": "Very angry or threatening to leave",
                    "signals": ["Hostile language", "Threatens cancellation or churn"],
                },
            ],
        ),
    }

    with TypeSafeClient() as client:
        response = client.system_one(
            state=state,
            questions=questions,
        )

    # Compose independent spam signals with weights controlled by code.
    answers = response.answers
    spam_risk = (
        0.45 * answers["requests_credentials"].noul
        + 0.30 * answers["sender_identity_mismatch"].noul
        + 0.25 * answers["unexpected_reward"].noul
    )

    # Escalate uncertain judgments instead of guessing.
    spam_is_uncertain = 0.4 < spam_risk < 0.6
    if spam_is_uncertain or answers["topic"].confidence < 0.75:
        return route_to_human_review(ticket)
    if spam_risk >= 0.6:
        return quarantine_as_spam(ticket)

    # Let code decide which speculative answers matter on this path.
    if answers["topic"].choice == "billing":
        return route_to_billing(
            ticket,
            refund_requested=answers["refund_requested"].noul >= 0.7,
        )
    if answers["topic"].choice == "orders":
        return route_to_orders(
            ticket,
            mentions_open_order=answers["mentions_open_order"].noul >= 0.7,
        )

    priority = (
        "high"
        if answers["frustration"].confidence >= 0.7
        and answers["frustration"].score >= 1.5
        else "normal"
    )
    return route_to_account_support(ticket, priority=priority)