How to evaluate a decision model
This page answers a practical question: when two sets of numbers disagree, which one do you believe?
What a public evaluation looks like
Start with the ones a third party can re-run. Their design is more useful than their numbers.
Shuffled options, repeated runs. An evaluation with published answers and a reproduction runner uses 108 four-option questions, no documents and no tools, with 3 runs per question and shuffled option order; random guessing is 25%. It reports the current public model at 84.6%, a 199 ms median and $0.0088, against the strongest comparable model at 98.0% but 2.12 s and $1.59.
⚠️ The unit on those two figures is not stated in the source (it says per full run) — do not turn them into a per-call price.
The point is not who wins, it is that all three axes are on the table — 84.6% at 199 ms and 98.0% at 2.12 s are two different products, not one that is worse than the other.
Shuffling exists to test one specific thing: if the model is sensitive to option order, it is an unstable component in production. That evaluation reports that axis.
The base rate is a ceiling. Another independent leaderboard runs PubMedQA, Banking77 and HelpSteer2 with 300 items × 5 runs per question and publishes every per-decision probability under CC BY 4.0. Two of its findings matter:
- On PubMedQA the model ties for first while costing 1/28 as much.
- On HelpSteer2, no model beats the label base rate.
The second means: on that dataset “always pick the most common label” is a strong baseline and nothing exceeded it. In that situation a first-place rank proves nothing — you are measuring the label distribution, not the model.
Count malformed output as a miss. One benchmark explicitly does. That sounds pedantic until you notice it decides whether scores from different models can be compared at all — a model permitted to emit illegal values has a meaningless accuracy.
Statistical comparison: bigger is not better
Does 84.6% versus 85.1% mean one is better?
Not necessarily. At n=100 that is well inside noise. One test that has actually been used here is McNemar’s — it compares the cases where two models disagree on the same items, not two totals.
A replica replayed published answers on 256 judgments and got 231 against 238, McNemar p = 0.21 — no significant difference from the original model. “No significant difference” is a completely respectable conclusion, and far more useful than insisting one is closer.
A stricter variant is preregistration: fix the criteria, then run. One ranking benchmark did exactly that and reported “passes on 20 Newsgroups, fails four of six conditions on Amazon ESCI”. A half-and-half result like that is more credible than an all-pass — it shows the criteria are actually doing work.
Four traps in self-reported numbers
Most real projects run one dataset and report a number. Those numbers have four predictable traps:
① The shape of the question changes the answer.
One survey-style experiment produced a hard comparison: whether a yes/no question is
written as Noul or as Choice matters more than the gap between Jev and GPT-4.1.
Meaning: in many “A beats B” conclusions, what is really doing the work may be the
question’s phrasing. Changing the primitive requires re-measuring; results do not
transfer.
② One run, small sample, no variance. A self-reported “100% attack interception rate” may rest on ten items. Elsewhere in the same ecosystem, reproducible projects report with a denominator — “0 adversarial flips out of 50”. The denominator is the credibility.
③ Jitter between calls. A project that uses the model as a calculator (asking it, per character, to pick one of 13 options) deliberately ships a Rerun button to expose jitter between two calls on the same input. Jitter is not the scary part; not knowing about it is — it means the threshold you measured may not hold at another moment.
④ A specialist winning on home turf is not a general win. A 706,048-parameter (2.8 MB) specialist scores 99.7% on its own form-filling task where the general model gets 83.6%. The authors say plainly that this is “a specialist on home turf rather than a general win”. That sentence should be the default reading of every “beats Jev” entry.
In one sentence
An evaluation’s value is in its denominator, its variance and its subclasses, not in its headline. A 60 with sample sizes, an option-order test and a negative result beats a 95 without them.