ドキュメント

TypeSafe での構築方法

TypeSafe での構築方法

コードに制御を残し、System One に狭く構造化された意思決定を与えることで、AI 駆動のソフトウェアを設計する。

System One は、agent ではなく AI 駆動のソフトウェアを構築するための TypeSafe のモデルです。コードを生成したり、自分で次の行動を選んだりしません。ソフトウェアに埋め込む AI プリミティブを提供し、コードが制御を保ったまま、非構造化データに対する常識的な判断をモデルが担います。

三つのソフトウェアアーキテクチャ

TypeSafe は AI 駆動のソフトウェアの構築のために設計されています。コードがワークフローを握り、AI が狭く構造化された意思決定を担います。

従来のコードは、単純なソフトウェアプリミティブから作られた複雑な決定木です。各プリミティブが信頼できるので、開発者はそれらをより高次の抽象に組み合わせられます。

agent は指示を処理し、次のステップを自分で選びます。人がプロセスを監視しているときはうまく機能しますが、ループのたびに脱線する機会が増えます。

コードが決定的な作業を担い、制御フローを握ります。モデルは、システムがプログラム可能な常識を必要とするとき、または非構造化データを解釈する必要があるときにだけ現れます。各 AI タスクは原子的で制約されたものに保たれます。

従来のソフトウェア、agent、AI 駆動ソフトウェアを三つの異なるシステムアーキテクチャとして示したもの。

System One を組み合わせ可能にしているもの

構造化

System One は構造上型安全です。意思決定と確率は、コードが期待する構造化されたソフトウェア型と JSON スキーマに適合するので、生成された文章から値を復元する必要が一切ありません。

並列

質問は独立に、並列に評価されます。あるプリミティブの結果が、別のプリミティブの結果を変える隠れたコンテキストになることはありません。

比較可能

出力は並べ替え可能で、賢い if 文、しきい値、比較を駆動できます。

高速

ほとんどのクエリは約 100 ms で完了します。System One はリアルタイムのリクエスト経路やユーザーインターフェースに使えるほど高速です。

キャリブレーションされた信頼度

RLCD は、過信に傾くのではなく、キャリブレーションされた確率を通じて不確実性を伝えます。

自己一貫性

System One は、繰り返しの評価に対して安定した答えを返すように設計されています。自己一貫性クックブックを参照してください。

すべての出力が与えられた選択肢に制約されているので、モデルはスキーマの外の値をでっち上げるのではなく、それらの選択肢にわたる完全な確率分布を返します。TypeSafe の目標は 100 倍を超える知能対速度・コスト比です。その根底にある賭けは、より安価な知能がはるかに大きな需要を生むというものです。

System One のワークフローを設計する

  1. できる限りコードを使う

    決定的な作業はコードに残します。それは信頼でき、安価です。ソフトウェアワークフローが同じ振る舞いを表現できるときは、agent の while ループを避けます。

    例:決定的な規則をコードに残す
    days_overdue = (today - invoice.due_date).days
    
    if days_overdue > 30:
        route_to_collections(invoice)

    モデルの意思決定をコードと組み合わせる、範囲の定まった方法については System One のパターン を見てください。

  2. 入力 state を分解する

    現在の質問に関連するコンテキストだけを含めます。これによりモデルが気を散らしたりコンテキストが腐敗したりするのを防ぎます。最新の情報を自分の知識ベースから得られるのに、モデルの重みに保存された知識に頼らないでください。

    例:関連するコンテキストだけを送る
    request
    {
      "state": {
        "ticket_message": "My flight was cancelled. Can I get a refund?",
        "refund_policy": "Cancelled flights are eligible for a full refund."
      },
      "questions": {
        "policy_supports_refund": {
          "type": "noul",
          "instructions": "Does the refund policy support the refund requested in the ticket?"
        }
      }
    }
  3. 入力 state に構造を使う

    state と questions フィールドにはネストした JSON を使います。曖昧さを取り除けるときは、質問を特定の値に向け、質問内の各パスの周りにバッククォート文字を付けます。

    例:ネストした値を参照する

    バッククォート付きのドットと添字のパスを使って、質問を特定のネストした値(たとえば support.tickets[0].message)に向けます。

    request
    {
      "state": {
        "support": {
          "tickets": [
            {
              "message": "I was charged twice for order A-104."
            },
            {
              "message": "How do I reset my password?"
            }
          ]
        },
        "commerce": {
          "orders": [
            {
              "id": "A-104",
              "charges": [
                {
                  "amount_usd": 49,
                  "status": "captured"
                },
                {
                  "amount_usd": 49,
                  "status": "captured"
                }
              ]
            }
          ]
        },
        "account": {
          "security": {
            "password_reset": "Email a reset link to the address on file."
          }
        }
      },
      "questions": {
        "duplicate_charge": {
          "type": "noul",
          "instructions": "Do `support.tickets[0].message` and `commerce.orders[0].charges` indicate a duplicate charge?"
        },
        "password_reset_supported": {
          "type": "noul",
          "instructions": "Can `account.security.password_reset` resolve the request in `support.tickets[1].message`?"
        }
      }
    }
  4. 質問を分解する

    可能な限り最も明示的で、狭く、具体的で、原子的な質問を尋ねます。複雑または不明瞭な質問を、それぞれが一つの性質を評価する別々の質問に分解します。

    例:スパム検出を分解する
    One broad question (bad)
    {
      "is_spam": {
        "type": "noul",
        "instructions": "Is `message` spam?"
      }
    }
    Decomposed questions (good)
    {
      "requests_credentials": {
        "type": "noul",
        "instructions": "Does `message.body` ask the recipient to provide a password or other login credential?"
      },
      "offers_unexpected_reward": {
        "type": "noul",
        "instructions": "Does `message.body` claim the recipient received an unexpected prize, payment, or reward?"
      },
      "creates_time_pressure": {
        "type": "noul",
        "instructions": "Does `message.subject` or `message.body` pressure the recipient to act quickly?"
      },
      "sender_identity_mismatch": {
        "type": "noul",
        "instructions": "Does the organization named in `message.sender.display_name` conflict with the domain in `message.sender.email`?"
      },
      "link_domain_mismatch": {
        "type": "noul",
        "instructions": "Does the domain in `message.links[0].url` conflict with the organization named in `message.sender.display_name`?"
      },
      "disguises_link_destination": {
        "type": "noul",
        "instructions": "Does `message.links[0].text` conceal or misrepresent the destination in `message.links[0].url`?"
      }
    }
    例:ツール呼び出しのトレースを検証する
    One broad question (bad)
    {
      "tool_calls_are_correct": {
        "type": "noul",
        "instructions": "Is `trace.tool_calls` correct for `request` and `available_tools`?"
      }
    }
    Decomposed questions (good)
    {
      "geocode_tool_is_relevant": {
        "type": "noul",
        "instructions": "Is `trace.tool_calls[0].name` an appropriate tool for resolving `request.location`?"
      },
      "geocode_location_matches": {
        "type": "noul",
        "instructions": "Does `trace.tool_calls[0].arguments.city` match `request.location`?"
      },
      "geocode_arguments_match_schema": {
        "type": "noul",
        "instructions": "Does `trace.tool_calls[0].arguments` conform to `available_tools.geocode_city.parameters`?"
      },
      "geocode_result_matches_call": {
        "type": "noul",
        "instructions": "Does `trace.tool_results[0].tool_call_id` match `trace.tool_calls[0].id`?"
      },
      "weather_tool_is_relevant": {
        "type": "noul",
        "instructions": "Is `trace.tool_calls[1].name` an appropriate tool for answering `request.text`?"
      },
      "weather_arguments_match_schema": {
        "type": "noul",
        "instructions": "Does `trace.tool_calls[1].arguments` conform to `available_tools.get_weather.parameters`?"
      },
      "weather_uses_geocoded_coordinates": {
        "type": "noul",
        "instructions": "Do the coordinates in `trace.tool_calls[1].arguments` match those in `trace.tool_results[0].output`?"
      },
      "weather_date_matches": {
        "type": "noul",
        "instructions": "Does `trace.tool_calls[1].arguments.date` match `request.date`?"
      },
      "weather_unit_matches": {
        "type": "noul",
        "instructions": "Does `trace.tool_calls[1].arguments.unit` match `request.unit`?"
      }
    }
  5. 質問に構造を使う

    質問は短く保ちます。instructions と criteria は通常文字列で、短く曖昧でない質問には文字列で十分です。オブジェクトや配列にもできます。質問を一つのフィールドに、質問を導くデータを他のフィールドに置きます。

    構造が役立つのは次の状況です:

    • 質問がコンテキストや例を必要とする。長い背景説明の文や入力例のリストは、質問の隣の名前付きフィールドに置きます。そうすればコードが質問を書き直さずにそれらを追加・入れ替えできます。
    • 質問の一部がコードから来る。値がデータベースから来るときは、文字列テンプレートに継ぎ込むのではなく、独自のフィールドに入れます。
    • 複数の質問が似た instructions を持つ。リクエストは一つの状態を取り、複数の質問を含められます。補足データを加えると、質問を区別しやすくなります。
    例:コードから来るレコードを参照する

    この Noul は、状態内の履歴書を候補者データベースのレコードと比較します。レコードはそのまま potential_duplicate に入り、質問はそれを名前で参照します。

    questions
    {
      "same_as_record_18": {
        "type": "noul",
        "instructions": {
          "potential_duplicate": {
            "name": "John Smith",
            "location": "Oakland, California",
            "last_employer": "Google"
          },
          "question": "Is the resume for the same person as `potential_duplicate`?"
        }
      }
    }

    コードから供給される “potential_duplicate” データは時間とともに変わります。“question” はバッククォートを使ってそれを参照します。

    criteria 内の説明もオブジェクトにできます。Choice では、各選択肢の説明を、その選択肢がカバーするもの、別の選択肢に属するもの、いくつかの例を述べるオブジェクトにできます。モデルが直接比較できるよう、選択肢間で同じフィールド名を使ってください。

    例:対比的な Choice の判定基準を定義する
    questions
    {
      "card_help_topic": {
        "type": "choice",
        "instructions": {
          "question": "Which disposable virtual card topic is the user asking about?",
          "focus": "Classify the information the user wants."
        },
        "criteria": {
          "get_disposable_virtual_card": {
            "what": "Purpose, eligibility, or setup",
            "not_for": "Quantity, transaction, or merchant restrictions",
            "examples": [
              "How can I get a disposable virtual card?",
              "What are disposable cards for?"
            ]
          },
          "disposable_card_limits": {
            "what": "Quantity, transaction, or merchant restrictions",
            "not_for": "Purpose, eligibility, or setup",
            "examples": [
              "How many disposable cards can I make per day?",
              "Where can I use a disposable card?"
            ]
          }
        }
      }
    }

    各質問タイプのページに動作する例があります:

    • Noul は、一つの履歴書を複数の候補者レコードと、レコードごとに一つの質問で比較し、質問はコードで組み立てます。
    • Choice は、混同しやすい二つの選択肢を、それぞれがカバーするもの、何のためでないか、例で説明します。
    • Score は、各レベルに説明と例の状況を与えます。

    構造化データ抽出カスケードクックブックは、共用の言い回しのケースを示します。抽出されたレコードのすべてのフィールドについて、同じ質問の組を尋ねます。

    短く曖昧でない質問や判定基準は文字列のままで構いません。そうしなければ混ざり合ってしまう指針を分離できるときに、構造を加えます。構造が受け付けられる場所の全体は、発展:構造を参照してください。

  6. 多くの質問をする

    同じ状態について、多くの狭く独立した質問を一つのリクエストで尋ねます。これが、この API で効果と「1 ドルあたりの知能」を最大化する方法です。質問は並列に走り、コードは直列のモデル往復を増やさずにそれらの信号を組み合わせられます。

    投機的ファンアウトパターンと並列質問クックブックを参照してください。

  7. 質問の出力をコードで組み合わせる(または古典的 ML モデルに渡す)

    独立した答えを、決定的な規則や加重和で組み合わせます。学習による組み合わせには、確率分布を下流の古典的機械学習モデルの特徴として使います。

    例:加重スコアで信号を組み合わせる
    answers = response.answers
    
    # Combine independent signals into one application-specific score.
    quality = (
        0.4 * answers["answers_request"].noul
        + 0.4 * answers["citations_are_supported"].noul
        + 0.2 * (1 - answers["contradicts_context"].noul)
    )

    複合スコアリングは、個々の判断を保ちながらそれらを組み合わせる方法を示しています。下流モデルのラベルがないなら、高価な推論モデルのアンサンブルでラベルを生成します。AutoResearch クックブックは、System One の出力で古典的モデルを訓練する方法を示しています。

  8. 不確実性でルーティングする

    高信頼度の答えと低信頼度の答えでコードに異なるアクションを取らせます。不確かなケースは人か、より高価な推論モデルにエスカレーションします。自分のデータで信頼度と正解率をプロットしてしきい値をテストします。

    例:信頼度でルーティングする
    answer = response.answers["card_help_topic"]
    
    if answer.confidence < 0.8:
        route_to_human_review(ticket)
    else:
        route_to_handler(answer.choice, ticket)

    しきい値の選び方と、各アクションのリスクへの対応付けについては、信頼度と信頼度で門控するルーティングを参照してください。

すべてをまとめる

このサポートチケットのワークフローは、決定的な作業をコードに残し、関連する構造化コンテキストだけを送り、多くの原子的な質問を一つのリクエストで評価し、明示的な信頼度の門で答えを組み合わせます。

triage_ticket.py

from typesafe_sdk import Choice, Noul, NoulCriteria, Score, TypeSafeClient

def triage_ticket(ticket, customer):
    # Handle deterministic states without calling a model.
    if ticket["status"] == "closed":
        return "no_action"

    open_orders = [
        order for order in customer["orders"] if order["status"] != "delivered"
    ]

    # Include only the structured context needed by the questions below.
    state = {
        "ticket": {
            "message": ticket["message"],
            "sender": ticket["sender"],
            "links": ticket["links"],
        },
        "customer": {
            "plan": customer["plan"],
            "open_orders": open_orders,
        },
        "policy": {
            "sensitive_credentials": ["password", "security code", "API key"],
        },
    }

    # Ask structured, atomic questions together so they run in parallel.
    questions = {
        "topic": Choice(
            instructions={
                "question": "Which team should handle `ticket.message`?",
                "focus": "Classify the customer's primary request.",
            },
            criteria={
                "billing": {
                    "what": "Charges, invoices, refunds, or subscriptions",
                    "not_for": "Order tracking or account access",
                    "examples": ["I was charged twice", "Where is my refund?"],
                },
                "orders": {
                    "what": "Order status, delivery, cancellation, or returns",
                    "not_for": "Charges or account access",
                    "examples": ["Where is my order?", "Cancel my shipment"],
                },
                "account": {
                    "what": "Login, profile, permissions, or security",
                    "not_for": "Charges or order tracking",
                    "examples": ["Reset my password", "I cannot sign in"],
                },
            },
        ),
        "requests_credentials": Noul(
            instructions={
                "question": "Does the message request a sensitive credential?",
                "compare": [
                    "`ticket.message`",
                    "`policy.sensitive_credentials`",
                ],
                "focus": "Look for a request to disclose the credential itself.",
            },
            criteria=NoulCriteria(
                true={
                    "what": "Asks the recipient to disclose a listed credential",
                    "examples": [
                        "Reply with your password",
                        "Send us your API key",
                    ],
                },
                false={
                    "what": "Does not ask the recipient to disclose a credential",
                    "not_for": "A legitimate instruction to reset a credential",
                    "examples": ["Use this link to reset your password"],
                },
            ),
        ),
        "sender_identity_mismatch": Noul(
            instructions={
                "question": "Does the claimed sender identity conflict with its domain?",
                "compare": [
                    "`ticket.sender.display_name`",
                    "`ticket.sender.email`",
                ],
                "focus": "Compare the named organization with the email domain.",
            },
            criteria=NoulCriteria(
                true={
                    "what": "Claims an organization unrelated to the email domain",
                    "examples": ["Acme Payroll sent from claim-bonus.example"],
                },
                false={
                    "what": "The identity and domain agree or make no conflicting claim",
                    "examples": ["Acme Payroll sent from acme.example"],
                },
            ),
        ),
        "unexpected_reward": Noul(
            instructions={
                "question": "Does the message announce an unexpected reward?",
                "inspect": "`ticket.message`",
                "focus": "Look for an unsolicited prize, payment, or reward claim.",
            },
            criteria=NoulCriteria(
                true={
                    "what": "Announces an unrequested prize, payment, or reward",
                    "examples": ["You were selected for a $1,000 bonus"],
                },
                false={
                    "what": "Contains no reward claim or discusses an expected payment",
                    "not_for": "A customer asking about a known refund or payroll deposit",
                    "examples": ["When will my approved refund arrive?"],
                },
            ),
        ),
        "refund_requested": Noul(
            instructions={
                "question": "Does the customer explicitly request a refund or credit?",
                "inspect": "`ticket.message`",
                "focus": "Require a requested remedy, not a billing complaint alone.",
            },
            criteria=NoulCriteria(
                true={
                    "what": "Directly asks for money back or an account credit",
                    "examples": ["Please refund the duplicate charge"],
                },
                false={
                    "what": "Does not ask for a refund or credit",
                    "not_for": "A complaint or billing question without a requested remedy",
                    "examples": ["Why was I charged twice?"],
                },
            ),
        ),
        "mentions_open_order": Noul(
            instructions={
                "question": "Does the message refer to a supplied open order?",
                "compare": [
                    "`ticket.message`",
                    "`customer.open_orders`",
                ],
                "focus": "Match an order id or other identifying details.",
            },
            criteria=NoulCriteria(
                true={
                    "what": "Refers to an open order by id or identifying details",
                    "examples": ["Where is order A-104?"],
                },
                false={
                    "what": "Does not identify any supplied open order",
                    "not_for": "A generic order question with no matching details",
                    "examples": ["How long does shipping usually take?"],
                },
            ),
        ),
        "frustration": Score(
            instructions={
                "question": "How frustrated does the customer appear?",
                "inspect": "`ticket.message`",
                "focus": "Judge expressed frustration, not issue severity.",
            },
            criteria=[
                {
                    "what": "Calm and matter-of-fact",
                    "signals": ["Neutral wording", "No complaint about the experience"],
                },
                {
                    "what": "Frustrated but civil",
                    "signals": ["Expresses annoyance", "Remains constructive"],
                },
                {
                    "what": "Very angry or threatening to leave",
                    "signals": ["Hostile language", "Threatens cancellation or churn"],
                },
            ],
        ),
    }

    with TypeSafeClient() as client:
        response = client.system_one(
            state=state,
            questions=questions,
        )

    # Compose independent spam signals with weights controlled by code.
    answers = response.answers
    spam_risk = (
        0.45 * answers["requests_credentials"].noul
        + 0.30 * answers["sender_identity_mismatch"].noul
        + 0.25 * answers["unexpected_reward"].noul
    )

    # Escalate uncertain judgments instead of guessing.
    spam_is_uncertain = 0.4 < spam_risk < 0.6
    if spam_is_uncertain or answers["topic"].confidence < 0.75:
        return route_to_human_review(ticket)
    if spam_risk >= 0.6:
        return quarantine_as_spam(ticket)

    # Let code decide which speculative answers matter on this path.
    if answers["topic"].choice == "billing":
        return route_to_billing(
            ticket,
            refund_requested=answers["refund_requested"].noul >= 0.7,
        )
    if answers["topic"].choice == "orders":
        return route_to_orders(
            ticket,
            mentions_open_order=answers["mentions_open_order"].noul >= 0.7,
        )

    priority = (
        "high"
        if answers["frustration"].confidence >= 0.7
        and answers["frustration"].score >= 1.5
        else "normal"
    )
    return route_to_account_support(ticket, priority=priority)