Docs

Can the probabilities be trusted?

This page answers one question: can those confidence figures be trusted?

The conclusion up front: “the shape of the output is guaranteed” and “the probability is accurate” are two entirely different claims. The first has been verified from several directions. The second is being publicly contested. Anyone putting a decision model behind a production gate should treat vendor or project-reported confidence as a prior to be calibrated, not a guarantee.

What calibration is

A model says “I am 80% sure”. If you collect everything it said at 80% and exactly 80% of those are right, the number is calibrated. If fewer are right, it is overconfident.

The usual metric is ECE (Expected Calibration Error): bin the predictions by confidence, compare each bin’s average confidence against its actual accuracy, and take the bin-size-weighted average. Lower is better.

Calibration and accuracy are independent axes. A highly accurate model can be badly overconfident (especially on the hardest slice), and a mediocre model can be well calibrated. You will see both below.

Fix 1: temperature scaling

The cheapest and most general fix. The idea: the model’s ordering of options is usually right, but the scale is wrong — overconfidence means everything sits too close to 0 and 1. One temperature parameter, fitted on a holdout, flattens or sharpens the distribution.

A replica that ships a runnable offline evaluation reports:

Temperature scaling plus conformal abstention takes ECE from 0.170 down to 0.071 (cross-validated).

Note “cross-validated”: it was not fitted and evaluated on the same data. That is the difference between a number worth trusting and one that is not.

Fix 2: conformal abstention

Temperature scaling fixes scale. Conformal methods decide when not to answer.

Rather than making every probability correct, they produce an abstention set and guarantee that the true answer falls inside it at least some stated fraction of the time (a coverage guarantee). Cases that do not clear the bar are escalated rather than auto-processed.

In real projects:

  • A 118M open alternative uses temperature scaling plus a split-conformal abstain set and reports ECE 0.01–0.03 on public suites — while stating plainly that on typed-decision benchmarks it loses to Laya (0.71 vs 0.77). A project that publishes its losses is more informative than one that only publishes wins.
  • A medical implementation uses two frozen local readers plus a fit-free router and a split-conformal candidate set to bound the error, landing within 2 points of the hosted service on three 600-item national licensing exams, with no fine-tuning and no distillation.

What both have in common: neither claims the probabilities became accurate. They claim to know when they are not. In a real system, the second is what you can gate on.

One counterintuitive independent result: the bias has opposite signs

An independent calibration test used 900 rule-generated tickets (which the model cannot have seen) plus three public benchmarks, published every raw response and an ECE against a simulated noise floor, and arrived at a conclusion about sign:

Primitive Direction of bias
Choice systematically overconfident
Score systematically overconfident
Boolean (Noul) systematically underconfident

The value is in the sign. If the bias were random, one global temperature would fix it. Opposite directions mean a single global correction cannot — you need to calibrate per primitive at minimum.

It also explains a specific design failure: if a system compares Noul and Choice confidences against the same threshold, it is measuring the same thing with one ruler that is too short and another that is too long.

Input language affects calibration too

Another measurement, on a Spanish corpus of 3,200 human-labelled items:

Writing the state in Spanish cost 3.0–6.4 accuracy points and roughly doubled the ECE on XNLI / PAWS-X, while writing the instructions in Spanish had no effect.

The state/instruction distinction is the point: the state is content, the instructions are metadata. This matters especially for Chinese and Japanese readers — their states will very likely take a similar path, and nobody has published a measurement of it. There is no reason to assume it is not happening in your language.

So how is the threshold chosen?

Everything above lands on the same question: where did that 0.8 come from?

The reproducible answer is measure, do not guess: take your own labelled data, fit the per-question threshold that reaches your target accuracy, and validate it on a holdout.

One tool does exactly that and adds the step that matters most: it fails CI when a model update breaks a locked threshold. Thresholds expire like prices do — except a wrong price gets noticed and a wrong threshold does not.

In one sentence

Treat confidence as a prior, not a guarantee. Measure the trustworthy band on your own data first, then write that band into code.