Decision model field guide
These pages answer three questions:
- How do people actually use it in real projects?
- Which practices keep showing up, and which of them only exist because someone got burned?
- What does the evidence say about the part the docs cannot settle — above all, whether the probabilities can be trusted?
Every figure on the pages below comes from a publicly checkable first-hand record: an independent recomputation, a preregistered experiment, or an artifact the project published itself. Which one it is gets said on the page — because those three carry very different weight.
The thread
Jev and Laya have a property that is both easy to understand and easy to oversell:
the shape of the output is guaranteed. Ask for a Choice and you get back one of
the options plus a probability; ask for a Noul and you get a number between 0 and 1.
There is no text to parse, so there is no “the model rambled and the program fell over”.
But that is type safety, not calibrated probability. The second is a different claim, and it is currently contested in public:
- An independent recomputation found that the
confidencereturned forScorequestions is frequently just the fractional part of the score (121 such outputs, tightly clustered) — which looks like an artifact of the wire format rather than a calibrated probability. - A public experiment asked Jev to guess a fair die 400 times. It picked the same face all 400 times, at an average self-reported probability of 82.9%; the actual hit rate was 19.0%.
- An independent out-of-distribution calibration test reported the sign of the error:
ChoiceandScoreare systematically overconfident,Booleansystematically underconfident.
So every page below keeps coming back to one sentence: treat confidence as a prior you have to calibrate, not as a guarantee. Every page below is about what that means in a specific setting.
How to read the numbers
The figures below come from three kinds of place, and the pages try to say which:
| Source | What it means | How to use it |
|---|---|---|
| Self-reported | The author’s own result, in their README or blog | A lead, not a conclusion |
| Independently recomputed | A third party re-ran it, or published something you can re-run | Citable, but still check the sample size |
| Preregistered | Criteria fixed before the run, results published either way | The most informative kind |
The distinction is not pedantry. The same model shows up both in self-reported “beats Jev” results and in an independent negative one — both are true, and the difference is who measured, and how much.