Docs

Open replicas and the systemone ecosystem

Other pages in this guide are about how to use these models. This one is about which one — and why that question is both easier to answer than a year ago and easier to get wrong.

One interface became the de facto standard

POST /v1/systemone is no longer just a detail of the official SDK. It became a standard because the client-side migration cost is close to zero: switching backends means changing a base URL environment variable, with request and response shapes unchanged.

That produced dozens of projects implementing the same interface — from “turn any open model into a decision model” (read the logits, generate no text), to fine-tunes on open bases, to pure local runtimes, to bridges that adapt other ecosystems.

Where the local alternatives sit

Open replicas span five orders of magnitude in parameter count, from a few hundred thousand to 27B:

Implementation Scale Licence Worth noting
Laya non-autoregressive, ~35 ms per forward pass Apache-2.0 Non-autoregressive with RLCD-calibrated probabilities; on PyPI / Hugging Face
von 395M — under 15 ms per forward pass
TinyJev 596M — pointer head, returns calibrated confidence meant to be thresholded
NanoJev 0.6B — full probability distribution in one pass, shipped with training pipeline, weights and dataset
Verdict 118M Apache-2.0 temperature scaling + split-conformal abstain set; states it loses to Laya on typed decisions
jevos 1B (MiniCPM5 cut to 17 layers) MIT GGUF q4_k_m at 619 MB; 54 ms short / 220 ms long on CPU
CLM 8B — reports up to 9× lower latency
Jebadiah 27B / 9B / 4B Apache-2.0 distributed as bf16 / GGUF / MLX

A few numbers concrete enough to check:

  • A 1B CPU implementation (GGUF, 619 MB) takes 54 ms on short requests and 220 ms on long ones, against 344 / 345 ms for the hosted service — but scores 0.815 versus 0.927, and answers only yes/no questions.
  • Another open weight scores 33.1% versus 36.7% on 308 sealed decisions (about 13 ms).
  • A specialist form-filling model (706,048 parameters, 2.8 MB) reaches 99.7% on its own task where the general model gets 83.6% — and the authors say explicitly that this is home-turf advantage, not a general win.

What this table should convey is magnitudes, not a ranking. “Six times faster but eleven points less accurate” and “comparable accuracy but only one question type” are completely different trade-offs, and both can be described as “a local replacement”.

Compliance: the number not to skip

When an ecosystem gets busy, the question to ask is whether everyone really honours the same interface.

One project turned the semantics of /v1/systemone into 48 checkable requirements (for example: a Choice’s option probabilities sum to 1 by definition; a Score’s expectation equals the probability-weighted sum; option counts run from 2 to 255; answers do not change when question order changes) and ships a suite that tests any server claiming compatibility.

Its measured result:

Of the eight highest-starred open ports, only two were fully compliant.

That cuts two ways. For clients, it means “just change the base URL” needs to be confirmed with the suite, not assumed. For anyone building a replacement, it points at a differentiation axis that is easier to reach than accuracy and far less crowded.

What this means for choosing

Putting it together, the centre of gravity has moved away from “which model scores higher”:

  1. The interface is shared and the model is replaceable. Those 48 requirements are the definition of replaceability — confirm them with the suite first, then talk about anything else.
  2. The moat moved outside the model. Once the model layer is a commodity, what is hard to copy is the interface contract + the threshold policy + calibration on your own data. Those are exactly what the other pages here are about, and all three live in your repository, not your vendor’s.
  3. Put accuracy gaps in the right coordinates. “83.6% vs 99.7%” is a specialist on home turf; “0.815 vs 0.927” is a small model running on CPU that only answers yes/no. Say both halves, or the table misleads.

The size of the ecosystem

The ecosystem’s public footprint is roughly 1,100 projects (September 2026).

Volume is not quality, and it is not maturity. That count includes a great many single-commit projects sharing one scaffold: one author releasing a batch of repositories on the same day, sharing an AGENTS.md, a single-commit history, and considerably more prose than code. They may be entirely legitimate — they are simply unproven. Filter on “is there something runnable”, not on “did it appear in a list”.

In one sentence

The model layer is becoming a consumable; the interface, the thresholds and the calibration are not. Spend your effort on the latter three — the first is replaceable at any time.