Open replicas and the systemone ecosystem
Other pages in this guide are about how to use these models. This one is about which one — and why that question is both easier to answer than a year ago and easier to get wrong.
One interface became the de facto standard
POST /v1/systemone is no longer just a detail of the official SDK. It became a standard
because the client-side migration cost is close to zero: switching backends means
changing a base URL environment variable, with request and response shapes unchanged.
That produced dozens of projects implementing the same interface — from “turn any open model into a decision model” (read the logits, generate no text), to fine-tunes on open bases, to pure local runtimes, to bridges that adapt other ecosystems.
Where the local alternatives sit
Open replicas span five orders of magnitude in parameter count, from a few hundred thousand to 27B:
| Implementation | Scale | Licence | Worth noting |
|---|---|---|---|
| Laya | non-autoregressive, ~35 ms per forward pass | Apache-2.0 | Non-autoregressive with RLCD-calibrated probabilities; on PyPI / Hugging Face |
| von | 395M | — | under 15 ms per forward pass |
| TinyJev | 596M | — | pointer head, returns calibrated confidence meant to be thresholded |
| NanoJev | 0.6B | — | full probability distribution in one pass, shipped with training pipeline, weights and dataset |
| Verdict | 118M | Apache-2.0 | temperature scaling + split-conformal abstain set; states it loses to Laya on typed decisions |
| jevos | 1B (MiniCPM5 cut to 17 layers) | MIT | GGUF q4_k_m at 619 MB; 54 ms short / 220 ms long on CPU |
| CLM | 8B | — | reports up to 9× lower latency |
| Jebadiah | 27B / 9B / 4B | Apache-2.0 | distributed as bf16 / GGUF / MLX |
A few numbers concrete enough to check:
- A 1B CPU implementation (GGUF, 619 MB) takes 54 ms on short requests and 220 ms on long ones, against 344 / 345 ms for the hosted service — but scores 0.815 versus 0.927, and answers only yes/no questions.
- Another open weight scores 33.1% versus 36.7% on 308 sealed decisions (about 13 ms).
- A specialist form-filling model (706,048 parameters, 2.8 MB) reaches 99.7% on its own task where the general model gets 83.6% — and the authors say explicitly that this is home-turf advantage, not a general win.
What this table should convey is magnitudes, not a ranking. “Six times faster but eleven points less accurate” and “comparable accuracy but only one question type” are completely different trade-offs, and both can be described as “a local replacement”.
Compliance: the number not to skip
When an ecosystem gets busy, the question to ask is whether everyone really honours the same interface.
One project turned the semantics of /v1/systemone into 48 checkable requirements
(for example: a Choice’s option probabilities sum to 1 by definition; a Score’s
expectation equals the probability-weighted sum; option counts run from 2 to 255; answers
do not change when question order changes) and ships a suite that tests any server claiming
compatibility.
Its measured result:
Of the eight highest-starred open ports, only two were fully compliant.
That cuts two ways. For clients, it means “just change the base URL” needs to be confirmed with the suite, not assumed. For anyone building a replacement, it points at a differentiation axis that is easier to reach than accuracy and far less crowded.
What this means for choosing
Putting it together, the centre of gravity has moved away from “which model scores higher”:
- The interface is shared and the model is replaceable. Those 48 requirements are the definition of replaceability — confirm them with the suite first, then talk about anything else.
- The moat moved outside the model. Once the model layer is a commodity, what is hard to copy is the interface contract + the threshold policy + calibration on your own data. Those are exactly what the other pages here are about, and all three live in your repository, not your vendor’s.
- Put accuracy gaps in the right coordinates. “83.6% vs 99.7%” is a specialist on home turf; “0.815 vs 0.927” is a small model running on CPU that only answers yes/no. Say both halves, or the table misleads.
The size of the ecosystem
The ecosystem’s public footprint is roughly 1,100 projects (September 2026).
Volume is not quality, and it is not maturity. That count includes
a great many single-commit projects sharing one scaffold: one author releasing a batch
of repositories on the same day, sharing an AGENTS.md, a single-commit history, and
considerably more prose than code. They may be entirely legitimate — they are simply
unproven. Filter on “is there something runnable”, not on “did it appear in a list”.
In one sentence
The model layer is becoming a consumable; the interface, the thresholds and the calibration are not. Spend your effort on the latter three — the first is replaceable at any time.