文件導航

Choice

Choice 是 System One 的一種問題型別,用於從給定的一組選項中選出一個。答案包含選中的選項、每個選項的機率,以及置信度。

當答案是固定的一組選項之一時,就用 Choice。比如哪支團隊處理這張工單、某件商品屬於哪個品類,或某段程式碼是用什麼語言寫的。如果答案落在一個譜系上的某個位置,用 Score。如果是是非題,用 Noul。選擇問題型別對三者做了比較。

Choice 的答案就是 choice 裡選中的那個選項。模型還會在 probabilities 裡返回每個選項的機率,併為選中的選項給出一個 confidence 值。

示例問題:

"What programming language is this code written in"
  → options: python, javascript, typescript, go, rust, other

"What type of meeting is this based on the title and description"
  → options: standup, planning, retrospective, one on one, brainstorm, none of the above

"Which product category does this item belong to"
  → options: electronics, clothing, home garden, food and beverage

請求結構

發往 TypeSafe API 的 POST 請求體有固定的結構。最外層有三個欄位:state,要評估的內容;model;以及 questions,一個從你自選的問題 ID 到問題物件的對映。每個 Choice 問題都有以下欄位:

  • type:始終為 "choice"。
  • instructions:模型要回答的問題。
  • criteria:作為對映給出的備選答案。每個鍵是一個選項名,每個值是該選項的描述。

下面這次請求的狀態,是某家線上鞋店的一張客服工單,問的是該由哪支團隊處理:

request
{
  "state": "My running shoes arrived in the wrong size. Can I swap them for a size 10?",
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this?",
      "criteria": {
        "returns": "Exchanges, wrong or damaged items",
        "shipping": "Delivery status, delays, lost packages",
        "billing": "Charges, invoices, payment problems"
      }
    }
  }
}

問題 ID 由你自己選,這裡是 department。答案也會以同一個 ID 返回。模型看不到問題 ID。選項名和它們的描述都會發給模型,所以描述要寫得能把各個選項區分開。

我們的客戶端 SDK提供型別化的問題。在 Python 裡,同一個問題寫成 Choice:

from typesafe_sdk import Choice, TypeSafeClient

with TypeSafeClient() as client:
    response = client.system_one(
        state="My running shoes arrived in the wrong size. Can I swap them for a size 10?",
        questions={
            "department": Choice(
                instructions="Which team should handle this?",
                criteria={
                    "returns": "Exchanges, wrong or damaged items",
                    "shipping": "Delivery status, delays, lost packages",
                    "billing": "Charges, invoices, payment problems",
                },
            ),
        },
    )

    print(response.answers["department"].choice)

用 system_one 方法或 https://api.typesafe.ai/v1/systemone 端點來呼叫 System One 模型。model 欄位決定由哪個模型處理請求。如何用 TypeSafe 構建講了該在程式碼的哪個位置呼叫它。

用我們的某個客戶端 SDK,或者直接呼叫 HTTP API。如果是編碼智慧體替你寫整合,先裝上 TypeSafe Agent 技能,它才知道請求和響應的形狀。

響應結構

響應裡每個問題在 answers 下各有一項,鍵就是請求裡的 ID。上面那次示例請求的響應是:

{
  "model": "jev-1.13.0",
  "answers": {
    "department": {
      "type": "choice",
      "choice": "returns",
      "confidence": 1.0,
      "probabilities": {
        "shipping": 0.0,
        "returns": 1.0,
        "billing": 0.0
      }
    }
  },
  "usage": {
    "input_tokens": 328,
    "output_tokens": 34
  }
}

除了 type,每個 Choice 答案還有三個值:

  • choice:機率最高的那個選項。
  • probabilities:在所有選項上的完整機率分佈。所有值之和為 1。
  • confidence:一個 0 到 1 之間的數,由 probabilities 的分佈形狀算出。分佈很平、機率攤在好幾個選項上,意味著低置信度;只在某一個選項上有一個尖峰,意味著高置信度。

這張工單很好判斷,所以全部機率都落在 returns 上,置信度為 1.0。如果一張工單同時提到尺寸不對和退款沒到賬,機率就會分散在 returns 和 billing 之間,置信度隨之下降。

最佳實踐:一次呼叫問多個問題

把程式碼可能用到的每個 Choice 問題都放進一次請求,而不是每個問題發一次。問題會並行求值。增加問題幾乎不改變響應時間,程式碼也可以忽略用不到的答案。多出來的問題仍然要花 token。一次提出多個問題完整講了這件事;下一節展示了在一次呼叫裡問五個 Choice 問題。

同一個道理也適用於單個 Choice 問題內部的選項。一個 Choice 問題最多接受 255 個選項,每加一個選項只多花幾個 token,所以把團隊、品類或產品的完整列表交給模型,而不是隻給一個候選短名單。當列表未必能覆蓋所有輸入時,加上一個 other 或 none of the above 選項,好讓模型能說其它選項都不合適。

要沿著深層層級或龐大的分類體系給文件分類,就把 Choice 問題逐層串起來。層級分類 cookbook展示瞭如何在 Choice 機率上跑束搜尋,在每一層保留最好的 K 條候選路徑,而不是隻認定一條貪心路徑。

一個更復雜的例子

上面的基礎例子把一張工單路由到某支團隊。更大的客服系統可能還需要知道退貨原因、配送問題、客戶想要什麼,以及客戶的語氣。

下面這次請求針對一張比剛才更含糊的工單提了五個 Choice 問題:它牽扯到三支團隊,而且沒說客戶想要什麼。

request
{
  "state": "Shoes arrived two weeks late and in the wrong size. Also I see two charges of $120 on my card. What are you going to do about this?",
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this?",
      "criteria": {
        "returns": "Exchanges, wrong or damaged items",
        "shipping": "Delivery status, delays, lost packages",
        "billing": "Charges, invoices, payment problems"
      }
    },
    "return_reason": {
      "type": "choice",
      "instructions": "If the customer wants to return something, why?",
      "criteria": {
        "wrong_size": "The item doesn't fit",
        "wrong_item": "A different product was delivered",
        "damaged": "The item arrived broken or faulty",
        "changed_mind": "The item is fine, the customer no longer wants it",
        "other": "A return reason that fits none of the above"
      }
    },
    "shipping_issue": {
      "type": "choice",
      "instructions": "If this is a shipping problem, which kind is it?",
      "criteria": {
        "not_delivered": "The package never arrived",
        "delayed": "The package is late but still on its way",
        "wrong_address": "The package went to the wrong place",
        "damaged_in_transit": "The package arrived damaged",
        "other": "A shipping problem that fits none of the above"
      }
    },
    "requested_resolution": {
      "type": "choice",
      "instructions": "What does the customer want to happen?",
      "criteria": {
        "exchange": "Swap the item for a different one",
        "refund": "Money back",
        "replacement": "The same item sent again",
        "information": "Just an answer, no action needed"
      }
    },
    "tone": {
      "type": "choice",
      "instructions": "What is the customer's tone?",
      "criteria": {
        "calm": null,
        "frustrated": null,
        "angry": null
      }
    }
  }
}

這些 Choice 問題裡有兩個是推測性的:return_reason 只有在 department 是 returns 時才有意義,shipping_issue 只有在它是 shipping 時才有意義。tone 問題用的是 null 描述,因為這些選項名本身就夠清楚。

TypeSafe 的響應:

{
  "model": "jev-1.13.0",
  "answers": {
    "department": {
      "type": "choice",
      "choice": "returns",
      "confidence": 0.42,
      "probabilities": {
        "shipping": 0.04,
        "billing": 0.35,
        "returns": 0.61
      }
    },
    "return_reason": {
      "type": "choice",
      "choice": "wrong_size",
      "confidence": 1.0,
      "probabilities": {
        "other": 0.0,
        "wrong_size": 1.0,
        "changed_mind": 0.0,
        "damaged": 0.0,
        "wrong_item": 0.0
      }
    },
    "shipping_issue": {
      "type": "choice",
      "choice": "delayed",
      "confidence": 0.67,
      "probabilities": {
        "wrong_address": 0.0,
        "other": 0.26,
        "not_delivered": 0.0,
        "damaged_in_transit": 0.0,
        "delayed": 0.74
      }
    },
    "requested_resolution": {
      "type": "choice",
      "choice": "refund",
      "confidence": 0.2,
      "probabilities": {
        "replacement": 0.34,
        "refund": 0.4,
        "information": 0.02,
        "exchange": 0.24
      }
    },
    "tone": {
      "type": "choice",
      "choice": "frustrated",
      "confidence": 0.76,
      "probabilities": {
        "frustrated": 0.84,
        "angry": 0.16,
        "calm": 0.0
      }
    }
  },
  "usage": {
    "input_tokens": 589,
    "output_tokens": 212
  }
}

每個問題都獨立地針對這張工單作答:

  • department 的答案是 returns,機率 0.61;但由於重複扣款,billing 也有 0.35。這張工單同時屬於兩支團隊,置信度只有 0.42,正反映了這種分叉。
  • return_reason 是 wrong_size,置信度 1.0。這在意料之中,因為工單裡把這點說得很清楚。
  • shipping_issue 的答案在 delayed 和 other 之間分叉。它是個推測性問題,而 department 也不返回 shipping,所以程式碼可以忽略它,如下面的示例程式碼所示。
  • requested_resolution 偏向 refund,機率 0.40,其餘大部分由 replacement 和 exchange 分走,置信度為 0.20。重複扣款暗示要退款,尺寸不對暗示要換貨,而客戶從沒說自己想要哪個。
  • tone 的答案是 frustrated,機率 0.84,置信度 0.76。

下面的示例程式碼只讀取它需要的答案,忽略其餘的,並把低置信度的答案當作「先問、別行動」的理由:

from typesafe_sdk import Choice, TypeSafeClient

TRIAGE_QUESTIONS = {
    "department": Choice(
        instructions="Which team should handle this?",
        criteria={
            "returns": "Exchanges, wrong or damaged items",
            "shipping": "Delivery status, delays, lost packages",
            "billing": "Charges, invoices, payment problems",
        },
    ),
    "return_reason": Choice(
        instructions="If the customer wants to return something, why?",
        criteria={
            "wrong_size": "The item doesn't fit",
            "wrong_item": "A different product was delivered",
            "damaged": "The item arrived broken or faulty",
            "changed_mind": "The item is fine, the customer no longer wants it",
            "other": "A return reason that fits none of the above",
        },
    ),
    "shipping_issue": Choice(
        instructions="If this is a shipping problem, which kind is it?",
        criteria={
            "not_delivered": "The package never arrived",
            "delayed": "The package is late but still on its way",
            "wrong_address": "The package went to the wrong place",
            "damaged_in_transit": "The package arrived damaged",
            "other": "A shipping problem that fits none of the above",
        },
    ),
    "requested_resolution": Choice(
        instructions="What does the customer want to happen?",
        criteria={
            "exchange": "Swap the item for a different one",
            "refund": "Money back",
            "replacement": "The same item sent again",
            "information": "Just an answer, no action needed",
        },
    ),
    "tone": Choice(
        instructions="What is the customer's tone?",
        criteria={"calm": None, "frustrated": None, "angry": None},
    ),
}

def triage(ticket: str) -> None:
    with TypeSafeClient() as client:
        response = client.system_one(
            state=ticket,
            questions=TRIAGE_QUESTIONS,
        )
    answers = response.answers

    department = answers["department"]
    if department.confidence < 0.3:
        # Not clear which team to send to. Let a person decide.
        send_to_manual_triage(ticket)
        return

    if department.choice == "returns":
        # return_reason answer is only used here
        assign(ticket, team="returns", issue=answers["return_reason"].choice)
    elif department.choice == "shipping":
        # shipping_issue answer is only used here
        assign(ticket, team="shipping", issue=answers["shipping_issue"].choice)
    else:
        assign(ticket, team="billing")

    # A second team with a real share of the probability gets a copy
    for team, probability in department.probabilities.items():
        if team != department.choice and probability > 0.25:
            notify(ticket, team=team)

    resolution = answers["requested_resolution"]
    if resolution.confidence < 0.5:
        # The customer hasn't said what they want. Ask, don't guess.
        ask_customer_what_they_want(ticket)
    elif resolution.choice == "refund":
        flag_for_refund_approval(ticket)

    if answers["tone"].choice == "angry":
        flag_for_senior_agent(ticket)

對上面這張工單,這段程式碼會把它派給退貨團隊,原因標為 wrong_size;給賬單團隊也抄送一份,因為它 0.35 的份額超過了 0.25 的閾值;並去問客戶想要什麼,因為解決方案的置信度 0.20 低於 0.5。程式碼沒有用到 shipping_issue 的答案。

一次請求,五個答案,而路由邏輯就是普通的 if 語句。以後如果需要知道客戶的語言,或這張工單涉及哪個產品,就往 TRIAGE_QUESTIONS 裡再加一個 Choice 問題;請求次數仍然是一次。

智慧家居助手 demo在一次呼叫裡用一長串 Choice 問題評估每個使用者請求:請求類別、房間、裝置,以及動作。這些問題大多和任何一個具體請求都無關,程式碼會忽略它們。

結構化的 instructions 和 criteria

先給每個選項寫一行描述。當兩個選項很像、模型老是搞混時,就改用物件而不是字串來描述它們。給它幾個欄位:這個選項覆蓋什麼、什麼其實屬於相鄰的另一個選項,以及幾個示例輸入。

下面這兩個備選答案 return_policy 和 return_status 很容易混淆。關於其中任何一個的工單都可能提到退貨和退款,所以每個選項都寫明它不適用於什麼。

request
{
  "state": "I sent the shoes back a week ago. When do I get my money?",
  "questions": {
    "return_topic": {
      "type": "choice",
      "instructions": {
        "question": "Which returns topic is the customer asking about?",
        "focus": "Classify the information the customer wants."
      },
      "criteria": {
        "return_policy": {
          "what": "Whether and how an item can be returned",
          "not_for": "Progress of a return already sent",
          "examples": [
            "Can I return shoes I've worn once?",
            "How long do I have to return an order?"
          ]
        },
        "return_status": {
          "what": "Progress of a return already sent",
          "not_for": "Whether and how an item can be returned",
          "examples": [
            "Has my return arrived yet?",
            "When will my refund be paid?"
          ]
        }
      }
    }
  }
}

響應是 return_status,置信度 1.0:

{
  "model": "jev-1.13.0",
  "answers": {
    "return_topic": {
      "type": "choice",
      "choice": "return_status",
      "confidence": 1.0,
      "probabilities": {
        "return_policy": 0.0,
        "return_status": 1.0
      }
    }
  },
  "usage": {
    "input_tokens": 407,
    "output_tokens": 32
  }
}

欄位名 question、focus、what、not_for 和 examples 都不屬於 API 的一部分,也都不保留。它們由你自選,就像選項名一樣。模型會連同值一起看到這些名字,所以用簡短、能說明後面內容的欄位名。