文档导航

Noul

Noul

Noul 问题让 TypeSafe 模型评估一个是/否问题,并返回答案为「是」的概率。

当答案只有「是」或「否」时,就用 Noul。比如:这条消息是在要求退款吗、这份简历提到分布式系统了吗、这条评论含有个人数据吗。如果答案是一组选项之一,就用 Choice。如果答案是一个连续谱系上的位置,就用 Score。选择题类型对三者做了比较。

Noul 的答案是一个数,表示答案为「是」的概率,其中 0 表示否,1 表示是。

请求结构

发往 TypeSafe API 的 POST 请求体,和任何其它问题类型一样有同样的三个顶层字段:state,即要评估的内容;model;以及 questions。每个 Noul 问题都有以下字段:

  • type:固定为 "noul"。
  • instructions:模型要回答的是/否问题,或者要它判断的一个陈述。
  • criteria:可选。一个对象,用 true 和 false 两个描述说明「是」和「否」分别意味着什么。

下面这个请求里,状态是一条客服消息,两个问题分别是:客户是否想找真人,以及客户之前是否联系过客服:

request
{
  "state": "I have asked three times now. Can I please just talk to a real person?",
  "questions": {
    "is_human_escalation": {
      "type": "noul",
      "instructions": "Is the customer asking for a human agent?"
    },
    "is_repeat_contact": {
      "type": "noul",
      "instructions": "Has the customer contacted support about this before?",
      "criteria": {
        "true": "Mentions a prior attempt, ticket, or that they have asked before",
        "false": "No sign of any previous contact"
      }
    }
  }
}

问题 id,也就是这里的 is_human_escalation 和 is_repeat_contact,由你选择。这些 id 不会发送给模型。每个答案都以相同的 id 返回。第一个问题只靠 instructions。第二个额外用 criteria 说明什么算「是」、什么算「否」。

用 Python SDK 时,同样的问法是 Noul 对象:

from typesafe_sdk import Noul, NoulCriteria, TypeSafeClient

with TypeSafeClient() as client:
    response = client.system_one(
        model="jev-latest",
        state="I have asked three times now. Can I please just talk to a real person?",
        questions={
            "is_human_escalation": Noul(
                instructions="Is the customer asking for a human agent?",
            ),
            "is_repeat_contact": Noul(
                instructions="Has the customer contacted support about this before?",
                criteria=NoulCriteria(
                    true="Mentions a prior attempt, ticket, or that they have asked before",
                    false="No sign of any previous contact",
                ),
            ),
        },
    )

    print(response.answers["is_human_escalation"].noul)
    print(response.answers["is_repeat_contact"].noul)

system_one 方法和 https://api.typesafe.ai/v1/systemone 端点,都是以 TypeSafe 的 AI 模型 System One 命名的。如何用 TypeSafe 构建介绍了该在代码里的哪些地方用它。

如果你在用 coding agent,先安装 TypeSafe agent skill,这样它就知道请求和响应长什么样。

响应结构

响应里的 answers 为每个问题各有一个条目,键是请求里用的 id:

{
  "model": "jev-1.13.0",
  "answers": {
    "is_human_escalation": {
      "type": "noul",
      "noul": 0.99
    },
    "is_repeat_contact": {
      "type": "noul",
      "noul": 0.93
    }
  },
  "usage": {
    "input_tokens": 360,
    "output_tokens": 39
  }
}

这里两个答案都接近 1。客户说的是 “talk to a real person”,所以 is_human_escalation 是 0.99。“I have asked three times now” 与 is_repeat_contact 的 true 描述相符,所以是 0.93。

如何解读 Noul

这个数同时给出了答案和确定性。接近 1 是强「是」。接近 0 是强「否」。接近 0.5 表示模型给「是」和「否」的概率差不多。

下表是 jev-1.13.0 对不同客户消息在 is_human_escalation 问题上记录下来的答案:

状态 noul
Thanks, that fixed it! 0.02
How do I reset my password? 0.07
I need this sorted today, whatever it takes. 0.26
Are you a bot? 0.40
Is there any way to speak to someone about my invoice? 0.84
I have asked three times now. Can I please just talk to a real person? 0.99

前两条和后两条都很明确。“I need this sorted today” 很急,但从未要求找真人,得到 0.26。“Are you a bot?” 暗示想要真人,却没有明说,模型几乎五五开,给出 0.40。这两类消息,都需要靠代码里的阈值来做决定。

与 Choice 或 Score 不同,Noul 没有单独的 confidence 值。Noul 的概率分布只有两个结果,是和非,所以单个 noul 值就完整描述了它。Choice 或 Score 把概率铺在多个选项或档位上,confidence 则是对这种铺开的概括。

最常见的是,代码把 noul 取个阈值变成布尔值:

wants_human = response.answers["is_human_escalation"].noul > 0.9

if wants_human:
    route_to_agent(ticket)
else:
    route_to_bot(ticket)

阈值设在哪,取决于判断错了要付出多大代价。当「是」和「否」都同样容易处理时,用 0.5。把假「是」当真代价很大时(比如呼叫值班人员或退款),就调高。漏掉真「是」代价很大时(比如没能标出一个安全问题),就调低。中间的值可以交给人工,而不走任何一条代码分支。这和在置信度页里为 Choice 和 Score 答案描述的三路分流是一样的。

Noul 的取值从 0 到 1,但它不是你问的那件事本身的刻度。它是答案为「是」的概率。如果问题实际上问的是程度,这个值并不衡量程度。下面,对四位候选人问了 “Is the candidate strong in Python?”,并排给出一个含四个档位的 Score:no experience、some familiarity、regular use in a job、deep expertise。

候选人 Noul: “Is the candidate strong in Python?” Score: “How much Python experience does the candidate have?”
My experience is in Java and Go. I have not used Python. 0.03 0.0 (No experience)
I have used Python occasionally for small scripts alongside my main Java work. 0.14 1.0 (Some familiarity)
I used Python every day for two years in my last job, mostly data pipelines. 0.81 2.05 (Regular use in a job)
I have written Python daily for eight years, including maintaining a large Django codebase. 0.92 2.89 (Deep expertise)

Noul 只判断一个命题——“strong”,值就是它成立的可能性有多大。你可以在代码里给 0 到 1 划分区间,比如把 0.3 到 0.7 当作 “some experience”,但模型看不到这些区间,所以答案里没有任何东西是照着它们判断的。中间值可能意味着中等经验,也可能意味着情况不明;而且候选人之间的间隔也不是你选的。Score 则逐条判断每个档位描述,所以每位候选人都落到了你写的某个档位之上或附近,返回的概率显示了模型如何在各档位之间分配它的判断。如果你不认同,就改写某个档位再跑一次。选择题类型解释了其中的区别。

如何写 Noul 问题

一个 Noul 只问一个是/否问题。如果一个问题里有两个条件,比如 “Is the customer angry and asking for a refund?”,模型就得同时判断两者,值的信息量就变小了。应该问两个 Noul,再在代码里把它们组合起来。

把问题措辞成「值高就代表是」。“Does the message contain personal data?” 很清晰。“Is the message free of personal data?” 则把含义反了过来,之后读它的代码会弄反。

用陈述句和用疑问句一样有效。对于 “The customer is requesting a refund”,接近 1 的值表示该陈述为真。拿你自己的数据把两种措辞都试一遍,看看哪种更好。

把「是」和「否」之间的边界划得毫不含糊。“Does this candidate have any Python experience?” 效果不错,因为 “any” 不留中间地带。当边界很微妙时,就加上带 true 和 false 描述的 criteria,就像上面 is_repeat_contact 那样。对大多数 Noul 来说 instructions 就够了,所以把你的问题分别在有、没有 criteria 的情况下各试一遍,留下在你的文档上答案更好的那种。

好做法:一次调用问多个问题

对于一张条件清单,就在一次请求里问许多 Noul 问题:每个条件一个问题,由代码来决定这些结果的组合意味着什么。问题是并行评估的,所以增加 Noul 几乎不改变响应时间。一起问多个问题有更详细的说明。

在代码里处理多个 Noul 答案

上面那个两问题的请求,已经足够让代码把消息分流了。下面的示例在客户要求找真人时升级给人工,在客户之前联系过时提高优先级。任一问题取到中间值时,就交给复核人,而不走代码分支:

from typesafe_sdk import Noul, NoulCriteria, TypeSafeClient

SUPPORT_QUESTIONS = {
    "is_human_escalation": Noul(
        instructions="Is the customer asking for a human agent?",
    ),
    "is_repeat_contact": Noul(
        instructions="Has the customer contacted support about this before?",
        criteria=NoulCriteria(
            true="Mentions a prior attempt, ticket, or that they have asked before",
            false="No sign of any previous contact",
        ),
    ),
}

YES = 0.8
NO = 0.2

def route(message: str) -> None:
    with TypeSafeClient() as client:
        response = client.system_one(
            model="jev-latest",
            state=message,
            questions=SUPPORT_QUESTIONS,
        )
    answers = response.answers

    wants_human = answers["is_human_escalation"].noul
    repeat = answers["is_repeat_contact"].noul

    if NO < wants_human < YES or NO < repeat < YES:
        # The model isn't sure either way. Let a person decide.
        send_to_review(message)
        return

    priority = "high" if repeat > YES else "normal"
    if wants_human > YES:
        route_to_agent(message, priority=priority)
    else:
        route_to_bot(message, priority=priority)

对上面那条消息,is_human_escalation 的 noul 答案是 0.99,is_repeat_contact 是 0.93,所以代码以高优先级把它路由给人工客服。“How do I reset my password?” 这条消息在两个问题上都是 0.07,被路由给机器人。

阈值在你自己的代码里。如果复核人收到的消息太多,就缩小 NO 和 YES 之间的间距。如果放过去的错路由太多,就把它拉大。如果以后还需要知道消息是否提到付款、是否含有个人数据,就往 SUPPORT_QUESTIONS 里再加一个 Noul。请求次数仍然是一次。

结构化 instructions

Instructions 可以是一个对象而不是字符串,问题放在一个字段里,补充数据放在其它字段里。在问题中使用结构化数据讲了它什么时候有用。这里它用于一个用代码拼出来的问题:把一份刚到的简历,与候选人库里可能是同一个人的记录逐一比对。每条记录都原样放进一个 potential_duplicate 字段,question 对每条记录都一样,所有记录都在一次请求里检查。代码生成的问题键里含有每条记录的数据库 ID:

request
{
  "state": {
    "resume": {
      "name": "John Smith",
      "location": "Oakland, CA",
      "summary": "Backend engineer with eight years of Python and Go experience.",
      "experience": [
        {
          "employer": "Google",
          "title": "Senior Backend Engineer",
          "years": "2021-2025"
        },
        {
          "employer": "Microsoft",
          "title": "Software Engineer",
          "years": "2017-2021"
        }
      ]
    }
  },
  "questions": {
    "same_as_record_18": {
      "type": "noul",
      "instructions": {
        "potential_duplicate": {
          "name": "Jon Smith",
          "location": "Oakland, CA",
          "last_employer": "Google"
        },
        "question": "Is the resume for the same person as `potential_duplicate`?"
      }
    },
    "same_as_record_42": {
      "type": "noul",
      "instructions": {
        "potential_duplicate": {
          "name": "John Smith",
          "location": "Austin, TX",
          "last_employer": "Lone Star Freight"
        },
        "question": "Is the resume for the same person as `potential_duplicate`?"
      }
    },
    "same_as_record_77": {
      "type": "noul",
      "instructions": {
        "potential_duplicate": {
          "name": "John Smithers",
          "location": "Oakland, CA",
          "last_employer": "Bay Health Clinic"
        },
        "question": "Is the resume for the same person as `potential_duplicate`?"
      }
    }
  }
}

响应:

{
  "model": "jev-1.13.0",
  "answers": {
    "same_as_record_18": {
      "type": "noul",
      "noul": 0.74
    },
    "same_as_record_42": {
      "type": "noul",
      "noul": 0.09
    },
    "same_as_record_77": {
      "type": "noul",
      "noul": 0.08
    }
  },
  "usage": {
    "input_tokens": 535,
    "output_tokens": 58
  }
}

每个答案都是「这份简历就是该记录里那个人」的概率。记录 18 名字拼写不同,但地点和雇主对得上,得到 0.74。记录 42 同名,但城市不同、雇主也不同,得到 0.09。记录 77 名字相近、地点相同,但雇主不同,得到 0.08。在代码里给每个值取阈值,就像在代码里处理多个 Noul 答案那样,把中间值交给人工。

用 Python SDK 时,问题是根据候选人记录拼出来的。问题文本固定,记录会变:

from typesafe_sdk import Noul, TypeSafeClient

SAME_PERSON = "Is the resume for the same person as `potential_duplicate`?"

def duplicate_questions(candidates: list[dict]) -> dict[str, Noul]:
    """One Noul per candidate record, all asking the same question."""
    return {
        f"same_as_record_{candidate['id']}": Noul(
            instructions={
                "potential_duplicate": {
                    "name": candidate["name"],
                    "location": candidate["location"],
                    "last_employer": candidate["last_employer"],
                },
                "question": SAME_PERSON,
            },
        )
        for candidate in candidates
    }

def find_duplicates(resume: dict, candidates: list[dict]) -> list[str]:
    with TypeSafeClient() as client:
        response = client.system_one(
            model="jev-latest",
            state={"resume": resume},
            questions=duplicate_questions(candidates),
        )
    return [
        question_id
        for question_id, answer in response.answers.items()
        if answer.noul > 0.7
    ]

结构化数据抽取级联 cookbook用结构化的 instructions 来校验一条抽取出来的记录。每个字段都得到同一组问题。每个问题的 instructions 对象把问题文本放在 main_question 属性里。还有 field_spec 和 extracted_field 两个属性,会随字段变化。

cookbook 里的 Noul

看看我们的 cookbooks,里面有使用 Noul 问题的应用:

  • 并行问题 在一次请求里对一篇文章跑一遍 13 个问题的合规清单。
  • 自洽性:noul 用一份 15 个问题的量规给一笔保险索赔打分,并测量多次运行之间取值的稳定程度。
  • 重排序 直接用概率本身,而不是阈值:每个“查询-候选”对一个 Noul,然后按值给候选排序。
  • 逐行搜索 把一个找出匹配行的 Choice 与一个检查文档中到底有没有答案的 Noul 配对使用。
  • 结构还原 对每一对相邻行问一个 Noul——判断换行是否把一句话拆开了——从而从纯文本重建段落。