ドキュメント

Noul

Noul

Noul の質問は、TypeSafe モデルに はい/いいえ の質問を評価させ、答えが はい である確率を返させます。

答えが はい か いいえ のときに Noul を使います。たとえば、このメッセージは返金を求めているか、この履歴書は分散システムに言及しているか、このコメントは個人データを含むか。答えが複数の選択肢のいずれかなら Choice を使います。スペクトル上の位置なら Score を使います。質問タイプを選ぶで三つを比較しています。

Noul の答えは、答えが はい である確率を表す単一の数値で、0 は いいえ、1 は はい を意味します。

リクエストの構造

TypeSafe API への POST リクエストボディは、他のどの質問タイプとも同じ三つのトップレベルフィールドを持ちます。評価する内容である state、model、questions です。各 Noul 質問は次のフィールドを持ちます:

  • type:常に "noul"。
  • instructions:モデルが答える はい/いいえ の質問、またはそれが判断する文。
  • criteria:任意。はい と いいえ が何を意味するかの true と false の説明を持つオブジェクト。

以下は、状態がサポートメッセージで、二つの質問が顧客が人間を求めているか、以前にサポートに連絡したことがあるかであるリクエストです:

request
{
  "state": "I have asked three times now. Can I please just talk to a real person?",
  "questions": {
    "is_human_escalation": {
      "type": "noul",
      "instructions": "Is the customer asking for a human agent?"
    },
    "is_repeat_contact": {
      "type": "noul",
      "instructions": "Has the customer contacted support about this before?",
      "criteria": {
        "true": "Mentions a prior attempt, ticket, or that they have asked before",
        "false": "No sign of any previous contact"
      }
    }
  }
}

質問 ID は自分で選び、ここでは is_human_escalation と is_repeat_contact です。ID はモデルには送られません。各答えは同じ ID の下に返ります。最初の質問は instructions だけに頼っています。二番目は criteria を加えて、何が はい で何が いいえ かを述べています。

Python SDK では、同じ質問は Noul オブジェクトです:

from typesafe_sdk import Noul, NoulCriteria, TypeSafeClient

with TypeSafeClient() as client:
    response = client.system_one(
        model="jev-latest",
        state="I have asked three times now. Can I please just talk to a real person?",
        questions={
            "is_human_escalation": Noul(
                instructions="Is the customer asking for a human agent?",
            ),
            "is_repeat_contact": Noul(
                instructions="Has the customer contacted support about this before?",
                criteria=NoulCriteria(
                    true="Mentions a prior attempt, ticket, or that they have asked before",
                    false="No sign of any previous contact",
                ),
            ),
        },
    )

    print(response.answers["is_human_escalation"].noul)
    print(response.answers["is_repeat_contact"].noul)

system_one メソッドと https://api.typesafe.ai/v1/systemone エンドポイントはどちらも、TypeSafe の AI モデルである System One にちなんで名付けられています。TypeSafe での構築方法で、コードのどこで使うかを扱っています。

コーディングエージェントを使っているなら、先に TypeSafe agent skill をインストールして、リクエストと応答の形を把握させてください。

応答の構造

応答は、リクエストの ID の下に、質問ごとに answers のエントリを一つ持ちます:

{
  "model": "jev-1.13.0",
  "answers": {
    "is_human_escalation": {
      "type": "noul",
      "noul": 0.99
    },
    "is_repeat_contact": {
      "type": "noul",
      "noul": 0.93
    }
  },
  "usage": {
    "input_tokens": 360,
    "output_tokens": 39
  }
}

ここではどちらの答えも 1 に近いです。顧客は「talk to a real person」と言っているので is_human_escalation は 0.99 です。「I have asked three times now」は is_repeat_contact の true の説明に合致するので 0.93 です。

Noul を読む

その数値が答えと確実性を一つにしたものです。1 に近い値は強い はい です。0 に近い値は強い いいえ です。0.5 に近い値は、モデルが はい と いいえ に同程度の確率を与えていることを意味します。

下の表は、異なる顧客メッセージに対する is_human_escalation 質問への、記録された jev-1.13.0 の答えを示します:

状態 noul
Thanks, that fixed it! 0.02
How do I reset my password? 0.07
I need this sorted today, whatever it takes. 0.26
Are you a bot? 0.40
Is there any way to speak to someone about my invoice? 0.84
I have asked three times now. Can I please just talk to a real person? 0.99

最初の二つと最後の二つは明確です。「I need this sorted today」は緊急性はありますが人を求めてはおらず、0.26 です。「Are you a bot?」は人を明示的に求めてはいないものの人間を望むことをほのめかし、モデルはほぼ均等に割って 0.40 です。どちらも、コード内のしきい値に基づいて判断を下す必要がある種類のメッセージです。

Choice や Score と違い、Noul には別個の confidence 値はありません。Noul の確率分布は はい と いいえ の二つの結果しかないので、単一の noul 値がそれを完全に表します。Choice や Score は確率を複数の選択肢やレベルに広げ、confidence がその広がりを要約します。

多くの場合、コードは noul をブール値にしきい値処理します:

wants_human = response.answers["is_human_escalation"].noul > 0.9

if wants_human:
    route_to_agent(ticket)
else:
    route_to_bot(ticket)

しきい値をどこに置くかは、誤ることのコストによります。はい と いいえ が同じくらい行動しやすいときは 0.5 を使います。誤った はい に基づいて行動するのが高くつくとき(誰かを呼び出す、返金を発行するなど)は上げます。真の はい を見逃すのが高くつくとき(安全上の問題にフラグを立て損ねるなど)は下げます。中間の値は、どちらのコード経路でもなく人に回せます。これは、信頼度ページが Choice と Score の答えについて説明しているのと同じ三方向の分岐です。

Noul の値は 0 から 1 の範囲ですが、尋ねたものの尺度ではありません。答えが はい である確率です。質問が実質的に程度についてなら、その値は程度を測っていません。以下では「Is the candidate strong in Python?」を四人の候補者について尋ね、四つのレベル(経験なし、多少の知識、仕事での日常的な使用、深い専門性)を持つ Score と並べています。

候補者 Noul:「Is the candidate strong in Python?」 Score:「How much Python experience does the candidate have?」
My experience is in Java and Go. I have not used Python. 0.03 0.0 (No experience)
I have used Python occasionally for small scripts alongside my main Java work. 0.14 1.0 (Some familiarity)
I used Python every day for two years in my last job, mostly data pipelines. 0.81 2.05 (Regular use in a job)
I have written Python daily for eight years, including maintaining a large Django codebase. 0.92 2.89 (Deep expertise)

Noul は「strong」という一つの命題を判断し、値はそれがどのくらいありそうかを示します。コード内で 0 から 1 の範囲にレベルを作ることもできます(たとえば 0.3 から 0.7 を「多少の経験」とする)が、モデルはそれを見ないので、答えの中の何もそれに照らして判断されたわけではありません。中間の値は、中程度の経験か不明確なケースかを意味しえ、候補者間の間隔はあなたが選んだものではありません。Score は各レベルの説明を単独で判断するので、どの候補者もあなたが書いたレベルの上かその近くに着地し、返された確率はモデルが判断をレベル間でどう分けたかを示します。納得できないなら、レベルの言葉を変えて再実行してください。質問タイプを選ぶがこの違いを説明しています。

Noul の質問を書く

Noul 一つにつき はい/いいえ の質問を一つ尋ねます。質問が二つの条件を持つ場合(たとえば「Is the customer angry and asking for a refund?」)、モデルは両方を同時に判断しなければならず、値の意味が薄れます。Noul を二つ尋ね、コードで組み合わせます。

高い値が はい を意味するように質問を表現してください。「Does the message contain personal data?」は明確です。「Is the message free of personal data?」は意味を反転させ、後でそれを読むコードは逆の結果を得ます。

文も質問と同様に機能します。「The customer is requesting a refund」では、1 に近い値はその文が真であることを意味します。自分のデータで両方の言い回しを試し、どちらがより良いかを見てください。

はい と いいえ の境界を曖昧でなくしてください。「Does this candidate have any Python experience?」は「any」が中間を残さないのでうまく機能します。境界が微妙なときは、上の is_repeat_contact 質問がするように、true と false の説明を持つ criteria を加えます。ほとんどの Noul には instructions で十分なので、質問を criteria ありとなしで試し、あなたの文書でより良い答えを出す方を採用してください。

良い実践:一回の呼び出しで複数の質問をする

条件のチェックリストには、一つのリクエストで多くの Noul 質問を尋ねます。条件ごとに一つの質問で、その組み合わせが何を意味するかはコードが決めます。質問は並列に評価されるので、Noul を増やしても応答時間はほとんど変わりません。複数の質問をまとめて尋ねるでより詳しく説明しています。

複数の Noul の答えをコードで扱う

上の二つの質問のリクエストは、メッセージをルーティングするのに十分なものをコードに与えます。下の例は、顧客が人を求めているときに人へエスカレーションし、以前に連絡したことがあるときに優先度を上げます。どちらかの質問で中間の値は、コード経路ではなくレビュアーに回ります:

from typesafe_sdk import Noul, NoulCriteria, TypeSafeClient

SUPPORT_QUESTIONS = {
    "is_human_escalation": Noul(
        instructions="Is the customer asking for a human agent?",
    ),
    "is_repeat_contact": Noul(
        instructions="Has the customer contacted support about this before?",
        criteria=NoulCriteria(
            true="Mentions a prior attempt, ticket, or that they have asked before",
            false="No sign of any previous contact",
        ),
    ),
}

YES = 0.8
NO = 0.2

def route(message: str) -> None:
    with TypeSafeClient() as client:
        response = client.system_one(
            model="jev-latest",
            state=message,
            questions=SUPPORT_QUESTIONS,
        )
    answers = response.answers

    wants_human = answers["is_human_escalation"].noul
    repeat = answers["is_repeat_contact"].noul

    if NO < wants_human < YES or NO < repeat < YES:
        # The model isn't sure either way. Let a person decide.
        send_to_review(message)
        return

    priority = "high" if repeat > YES else "normal"
    if wants_human > YES:
        route_to_agent(message, priority=priority)
    else:
        route_to_bot(message, priority=priority)

上のメッセージでは、is_human_escalation の noul の答えの値は 0.99、is_repeat_contact は 0.93 なので、コードはそれを優先度 high でエージェントにルーティングします。「How do I reset my password?」というメッセージは両方の質問で 0.07 で、ボットにルーティングされます。

しきい値はあなたのコードの中にあります。レビュアーが目にするメッセージが多すぎるなら、NO と YES の間の幅を狭めます。誤ったルーティングが通り過ぎるのが多すぎるなら、広げます。後でメッセージが支払いに言及しているか、個人データを含むかを知る必要が出たら、SUPPORT_QUESTIONS に別の Noul を加えます。リクエスト数は 1 のままです。

構造化された instructions

instructions は文字列ではなくオブジェクトにもでき、質問を一つのフィールドに、補足データを他のフィールドに置きます。質問で構造を使うで、それが役立つ場合を扱っています。ここではコードで組み立てた質問に使います。ちょうど届いた履歴書を、同じ人物かもしれない候補者データベース内のレコードと比較します。各レコードはそのまま potential_duplicate フィールドに入り、question はすべてのレコードで同じで、すべてのレコードが一つのリクエストで確認されます。コードが生成する質問キーには、各レコードのデータベース ID が含まれます:

request
{
  "state": {
    "resume": {
      "name": "John Smith",
      "location": "Oakland, CA",
      "summary": "Backend engineer with eight years of Python and Go experience.",
      "experience": [
        {
          "employer": "Google",
          "title": "Senior Backend Engineer",
          "years": "2021-2025"
        },
        {
          "employer": "Microsoft",
          "title": "Software Engineer",
          "years": "2017-2021"
        }
      ]
    }
  },
  "questions": {
    "same_as_record_18": {
      "type": "noul",
      "instructions": {
        "potential_duplicate": {
          "name": "Jon Smith",
          "location": "Oakland, CA",
          "last_employer": "Google"
        },
        "question": "Is the resume for the same person as `potential_duplicate`?"
      }
    },
    "same_as_record_42": {
      "type": "noul",
      "instructions": {
        "potential_duplicate": {
          "name": "John Smith",
          "location": "Austin, TX",
          "last_employer": "Lone Star Freight"
        },
        "question": "Is the resume for the same person as `potential_duplicate`?"
      }
    },
    "same_as_record_77": {
      "type": "noul",
      "instructions": {
        "potential_duplicate": {
          "name": "John Smithers",
          "location": "Oakland, CA",
          "last_employer": "Bay Health Clinic"
        },
        "question": "Is the resume for the same person as `potential_duplicate`?"
      }
    }
  }
}

応答:

{
  "model": "jev-1.13.0",
  "answers": {
    "same_as_record_18": {
      "type": "noul",
      "noul": 0.74
    },
    "same_as_record_42": {
      "type": "noul",
      "noul": 0.09
    },
    "same_as_record_77": {
      "type": "noul",
      "noul": 0.08
    }
  },
  "usage": {
    "input_tokens": 535,
    "output_tokens": 58
  }
}

各答えは、履歴書がそのレコードの人物のものである確率です。レコード 18 は名前の綴りが違いますが、所在地と雇用主が一致し、0.74 です。レコード 42 は同じ名前で別の都市、別の雇用主で、0.09 です。レコード 77 は似た名前で同じ所在地、別の雇用主で、0.08 です。各値を複数の Noul の答えをコードで扱うのようにコードでしきい値処理し、中間の値は人に回します。

Python SDK では、質問は候補者レコードから組み立てられます。質問のテキストは固定で、レコードが変わります:

from typesafe_sdk import Noul, TypeSafeClient

SAME_PERSON = "Is the resume for the same person as `potential_duplicate`?"

def duplicate_questions(candidates: list[dict]) -> dict[str, Noul]:
    """One Noul per candidate record, all asking the same question."""
    return {
        f"same_as_record_{candidate['id']}": Noul(
            instructions={
                "potential_duplicate": {
                    "name": candidate["name"],
                    "location": candidate["location"],
                    "last_employer": candidate["last_employer"],
                },
                "question": SAME_PERSON,
            },
        )
        for candidate in candidates
    }

def find_duplicates(resume: dict, candidates: list[dict]) -> list[str]:
    with TypeSafeClient() as client:
        response = client.system_one(
            model="jev-latest",
            state={"resume": resume},
            questions=duplicate_questions(candidates),
        )
    return [
        question_id
        for question_id, answer in response.answers.items()
        if answer.noul > 0.7
    ]

構造化データ抽出カスケードクックブックは、構造化された instructions を使って抽出されたレコードを検証します。すべてのフィールドが同じ質問の組を受け取ります。各質問の instructions オブジェクトは、質問のテキストを main_question プロパティに持ちます。さらに field_spec と extracted_field プロパティがあり、これらはフィールドごとに変わります。

クックブックの Noul

クックブックを見て、Noul 質問を使うアプリを確認してください:

  • 並列質問は、13 問の規制チェックリストを一つの記事に対して一つのリクエストで実行します。
  • 自己一貫性:noulは、保険金請求を 15 問のルーブリックに照らしてスコアリングし、値が実行間でどれほど安定しているかを測ります。
  • 再ランキングは、しきい値ではなく確率そのものを使います。クエリと候補のペアごとに一つの Noul を置き、候補を値で並べ替えます。
  • 行単位の検索は、一致する行を見つける Choice と、文書に答えがそもそも含まれるかを確認する Noul を組み合わせます。
  • 構造復元は、行のペアごとに、改行が文を分割したかどうかを問う一つの Noul を尋ね、プレーンテキストから段落を再構築します。