Public criticism and results that did not reproduce
The previous page covered how to fix calibration. This page covers why it needs fixing — gathering the public contrary evidence, because it rarely makes it onto a project’s home page.
The right way to read this page is not “so the model is bad” but: these are known failure modes, and your system has to survive them.
The die experiment
The most direct one. Ask the model to guess a fair die, 400 times:
| Item | Result |
|---|---|
| Which face it picked | the same face all 400 times |
| Average self-reported probability | 82.9% |
| Actual hit rate | 19.0% |
These numbers say two things at once. First, a self-reported probability can carry no information at all — 82.9% told you nothing reliable about a task with a 19% hit rate. Second, the model does not express “I don’t know”. On a fair die the correct answer is “about 16.7% each”, and the output shape — one preferred option with a high probability — has no way to say that.
Another result from the same author is worth keeping: on a synthetic forecasting document,
a 30% shortage risk became 5.3% through Choice and 26.7% through Noul. Same input,
same meaning, different primitive — a five-fold difference. The choice of primitive
changes the answer, and there is harder evidence of that on the
evaluation page.
“Jev Can’t Be Calibrated”
A widely circulated critique argues at the statistical level that the calibration claim does not hold up on the available evidence. A highly upvoted discussion puts the same point more carefully:
The type-safety guarantee is real (verified from several directions), but the calibration claim has no published ECE or reliability curve behind it.
Note the precision: it does not say calibration is false, it says the public evidence supporting it is missing. Those are different statements, and for an engineering decision they lead to the same place — you cannot set a production threshold on something with no published reliability curve.
Reranking: a complete negative result
This class of result is the most valuable, because it is specific, reproducible, and negates exactly the “it can do anything” implication.
A reranking evaluation used 33,047 items, 164 real queries and 9,831 judged pairs, and concluded that pure decision-model reranking loses to vector retrieval.
The scale is the point: thirty-odd thousand items and over a hundred real queries is not a toy setting. And the use case it negates — semantic reranking — is exactly the one most likely to be assumed to follow for free from classification.
A split decision
Some results depend on which row you read. One scoring-versus-generation comparison:
On Japanese NLI, multi-dimensional scoring with locally fitted weights beat asking directly (0.9076 vs 0.8373). But it flagged about 25 times as many hard benign rows as attacks (37.2% vs 1.5%).
The average improved while one class’s false-positive rate went up 25×. If false-flagging benign input is expensive in your system, this improvement is negative for you — even though it looks positive.
There is also a preregistered experiment (criteria fixed before the run, results published either way) that produced inconsistent verdicts across datasets — passing on one, failing on another.
The vendor’s own docs admit this
The official “model jaggedness” page lists known failure modes of the current public version. That is not contrary evidence, but its existence says something: “this model performs badly on some tasks” is part of the official position, not a secret.
Turning this page into rules
Three engineering rules fall out:
- Never reason across tasks from the absolute value of a confidence. “0.9 means 90% correct” is simply false on a task like the die.
- The choice of primitive is an experimental variable, not a style choice. The same
question asked as
Choiceor asNoulcan differ five-fold. Measure that difference rather than picking one arbitrarily. - Look at subclasses, not totals. The 25× false-positive example is completely invisible in the aggregate.
In one sentence
None of this is a reason not to use these models — it is a reason not to trust the defaults. Every pattern elsewhere in this guide (code owns the threshold, gating, conformal abstention) is a way of working within this page’s constraints.