Docs

Code owns the threshold, the model only gives a probability

Put a few hundred real projects side by side and the most stable common trait is this:

The model answers “how likely”. Your code answers “is that enough”.

This is not a style preference. It is the only safe conclusion to draw from the fact that self-reported confidence cannot be trusted — and there is independent measurement behind that in Can the probabilities be trusted.

What it looks like concretely

Choice and Score return an answer with a confidence attached; Noul returns a 0-to-1 probability directly. The interface itself makes no judgement for you — the threshold is yours. In real projects that turns into something very specific:

  • A profanity filter running on a Cloudflare Worker keeps its threshold, its max() aggregation policy and its public OpenAPI schema all in the Worker’s own code. The model is just one function it calls.
  • A browser extension that hides comments keeps 0.85 / 0.7 / 0.5 pinned in the extension, and stays hidden when the check fails — that preference is part of the threshold policy too.
  • An incident triage tool defines “page someone” as P(SEV1) + P(SEV2) ≥ 0.80, and that 0.80 lives in its own config, not in the model.

All three share one property: swap the model and the threshold does not move; retune the threshold and the model does not move. Hand the threshold to the model and those two things become one — at which point “tuning my threshold” means editing a prompt.

Why the principle matters

Because it keeps two kinds of failure apart:

  • The model answered badly — change the question, change the state, change the model.
  • The threshold is set badly — change the number in the code.

Mix them and “accuracy isn’t good enough” could be either one, with no way to tell which. The triage tool can put 0.80 straight into config precisely because that number can be reviewed on its own.

The counterexample you must know about

The principle rests on one assumption: that the number you receive is a calibrated probability.

An independent recomputation found that the confidence returned for Score questions is frequently just the fractional part of the score, with those outputs tightly clustered (121 of them). That is not “the model is confident” — it looks like a value derived from the score itself.

The difference matters:

  • If it is a calibrated probability, 0.85 supports reasoning like “I will be right in about 85% of cases like this”.
  • If it is a function of the score, 0.85 only means “this one scored high” and says nothing about how often it is right — and setting a threshold on it means setting a bar for a number with no probabilistic meaning.

Separately, an independent measurement report checked the model’s self-reported fields against ground truth and found the values far less granular than the description suggests (see Public criticism).

What to do instead

Three steps, and the order matters:

  1. Use the threshold operationally, not semantically. Whether 0.8 means “80% correct” is unimportant. What matters is whether, on your data, the cases above 0.8 are measurably more accurate. That you can check against a holdout.
  2. Fit the threshold on your own labels. Tooling exists for exactly this: take your labelled data, fit the per-question threshold that reaches your target accuracy, validate on a holdout — and fail CI when a model update breaks a locked threshold.
  3. Put the threshold in code and give it a review date. Thresholds expire with model versions, just like prices. It is something a person has to look at again, not something you set once.

In one sentence

The model’s trustworthiness should show up in your threshold, not in its own reported number. You can measure the first on your own data; you cannot measure the second.