Project list
60 projects, organised by what you are trying to solve. One line each, with a star count and a snapshot date.
How to read a line
- The star count reflects attention and usage, not quality; the date beside it is when it was read.
- The evidence tier (reproducible / implemented) answers “how much can be checked”, not whether it is good. A few lines of code that publish raw results and a feature-complete project whose numbers all come from its author land on different tiers — and that has nothing to do with whether either is useful.
60 projects. Star counts are as of 2026-09-24.
- Reproducible
- Publishes raw artifacts, datasets or a re-runnable evaluation
- Implemented
- Readable code with tests; the numbers are the author’s own
Interfaces and browsers
Where every step is a click. Decisions are cheap enough to ask at every step.
Jev Ultrafast19,456★Implemented
Browser Use’s ultrafast agent: the decision model picks both the operation and the DOM element in one request, and a small LLM is called only when text must be typed.
jev-chat-jarvis5,561★Implemented
A phone-side conversation copilot: reads the other party in WeChat, QQ, X or Feishu, proposes replies and fills the input box — sending stays manual. It reads the screen only, hooking nothing.
typesafe-computer-use937★Implemented
Computer use on macOS: OCR the screen, classify the next action with the decision model, click. Reported at about $0.0002 a step.
hermes-jev-skills726★Implemented
A skill suite for Hermes agents covering model routing, memory, compaction, skill selection, and computer or browser use — also installs under Claude Code and Codex.
jev-browser-use453★Implemented
Lets the decision model pick the click while Codex thinks and verifies — reported 5–10× faster browser operations.
typesafe-mario383★Implemented
An agent that plays Super Mario Bros. from structured emulator state rather than pixels.
mobile-jev380★Implemented
A standalone Android agent where the decision model makes every decision, with a live studio showing decision telemetry.
jev-chat-jarvis-mac346★Implemented
A WeChat intent-detection overlay for macOS: screen reading plus a local small model to judge intent and risk, then reply candidates — strictly read-only, nothing injected into WeChat.
jev-voice-browser271★Implemented
Control a real browser by voice: intent and target are decided in about 300 ms per spoken word, then Playwright acts — often before you finish the sentence.
jev-browser257★Implemented
Browser use driven by the decision model.
unclutter227★Implemented
A browser extension that uses the decision model to find page clutter and remove it, with reusable template rules.
embodied-jev221★Implemented
A MuJoCo robot-decision workbench.
jev-drone170★Implemented
A camera-only autonomous drone in MuJoCo with the judgement step running at 2.5 Hz.
Putting brakes on an agent
The second layer of judgement for a coding agent: allow or not, is it done, what is still needed in context.
fast-jev-compaction6,653★Implemented
A Claude Code plugin replacing the compaction summary with decisions: every tool call and result is scored in one request, stale ones are dropped or truncated, and everything kept stays verbatim.
foreman540★Implemented
An agent supervisor that uses decisions to keep coding agents on task.
JevHarness191★Implemented
LLM-authored task-specific decision harnesses, with optional full-trajectory reward reflection and GEPA evolution.
jev-dsh-decision163★Implemented
A structured-decision plugin for agent harnesses: native on DeepSeek Harness, and through iPolloWork also OpenCode and Codex Harness.
jev-pruner144★Implemented
A Claude Code plugin that trims long Bash output before the model ever sees it.
pi-jev (y0usaf)144★Implemented
A decision layer for the Pi coding agent: a measured tool-call gate plus
jev_askfor typed, calibrated answers.pi-warden138★Implemented
Guardrails for Pi that steer rather than interrupt: the decision model judges irreversible and off-task tool calls, detects stuck loops, and checks unverified done claims.
building-with-jev-skill131★Implemented
A skill for writing and improving programs that call the decision model.
Routing, search and review
Deciding who handles it, or picking which candidates are worth looking at.
jev-review (devagrawal09)586★Implemented
A staged code-review workflow with a local dashboard.
jev-search446★Implemented
Web search with the decision model: source selection, query understanding and relevance ranking.
jev-router (gargpratyush)377★Implemented
Routes Claude Code tasks to the cheapest capable model by asking the decision model to choose among candidates.
pg-jev329★Implemented
A PostgreSQL extension for asking your tables questions in plain language.
jev-codex-router263★Implemented
Per-turn routing for Codex: every turn re-decides the model, the thinking depth and the speed mode.
jev-review (NiazMorshed2007)218★Implemented
A local-first MCP plugin for continuous software-quality review by coding agents.
notra215★Implemented
A marketing analytics platform that moved a brand-visibility classifier off an LLM and onto boolean decisions, switched by a feature flag.
JevRouter190★Implemented
A lightweight router for models, tools and subagents.
perch171★Implemented
A CLI that parses a repository into a method-level call graph, then asks the decision model to flag likely defects and check project rules written in plain English.
jev-semgrep134★Implemented
Grep by meaning rather than pattern: every line is scored against a meaning, and meanings combine with AND / OR / NOT.
neo4jev129★Implemented
Navigates a Neo4j graph hop by hop using a classifier over candidate outgoing relationships, with beam search over answer log-probabilities.
Local and open alternatives
Open weights and local runtimes behind the same HTTP interface.
SemIf4,162★Reproducible
Semantic ifs from open models, on a single 3090 at home. Independent; not affiliated with the vendor.
NanoJev2,159★Reproducible
A 0.6B replica: parallel decisions over dynamic candidates, no output-token decoding, with the end-to-end training pipeline.
jevlike1,279★Reproducible
Train a small model that chooses among a changing list of text options, one probability per option in a single pass — with Doom, chess and Wikispeedia demos.
von610★Reproducible
A 395M non-autoregressive model answering typed questions with calibrated probabilities in under 15 ms, positioned as a local drop-in.
simple-jev507★Reproducible
Turn any open model into a classifier, or into a compatible endpoint.
AnyJev440★Reproducible
Turns any LLM into a decision model: typed decisions and calibrated probabilities are read from next-token prefill distributions, with no training.
openjev386★Reproducible
An open, vendor-compatible decision server built on DiffusionGemma.
decider351★Reproducible
One-pass typed decisions with calibrated probabilities, fine-tuned from a 2B model.
LLM2Jev300★Reproducible
Adapts local language models into compatible structured decision engines, with all three primitives produced by prefill-only binary inference.
agent-jev286★Reproducible
A 0.6B replica aimed at agents: feed it unstructured state (diffs, traces, logs) and structured questions, get calibrated distributions back in one ~50 ms pass with no output tokens.
openJev-verdict-2.0283★Reproducible
A calibrated 151M non-autoregressive decision engine that its author reports as beating both the vendor model and Laya on a public benchmark: 77.10% accuracy, 0.0636 Brier, 0.0144 ECE.
jev-visual271★Reproducible
An educational inference experiment on Apple Silicon: shared context, direct candidate scoring, and local visual demos.
jeff238★Reproducible
A self-hosted drop-in replacement.
reflex133★Reproducible
A small open decision model: state plus typed questions, calibrated probabilities back.
SDKs and integration
Turning a judgement into ordinary control flow in your language, or into a framework you already use. Includes the official SDKs.
eve5,342★Implemented
Vercel’s open agent framework, which ships the decision model as the default evaluation model in its experimental evaluate path.
TypeSafe agent skills2,060★Implemented
The official agent skill for Claude Code, Codex and compatible agents: the primitives, the patterns, and how to structure evaluations.
ai-cli813★Implemented
A Vercel Labs terminal CLI that can run the decision model as the evaluation model for its
evaluatecommand.jev-mcp (jkudish)320★Implemented
A proof-of-concept MCP server for connecting the decision model to MCP clients.
typesafe-mcp292★Implemented
A single-binary Go MCP server exposing typed judgements to Claude Desktop, Claude Code and Codex.
System One adapter (Python)287★Implemented
The official drop-in client replacement backed by OpenAI, Anthropic or compatible LLM APIs — made for comparing the decision model against chat models in the same evaluation.
TypeSafe JavaScript SDK232★Implemented
The official TypeScript/JavaScript client, with inferred answer types.
TypeSafe Python SDK220★Implemented
The official sync and async Python client.
Applications and experiments
Things built for a person to use directly: trading, search, writing, killing an idea.
jev-trader2,240★Implemented
One trading decision every Monad block.
shapeshift491★Implemented
An input that becomes what you mean: one text box that morphs into the right UI as you type. Works offline.
JevRev304★Implemented
An LLM plus decision-model workflow. (The public description goes no further than this.)
jev-align (Sutro)284★Implemented
An active-learning CLI that uses human labels and GEPA to improve how decisions are defined.
killmyidea159★Implemented
Describe your startup idea and it decides: kill it, fix it, or ship it.
jev-trade131★Implemented
A live decision-driven trader on Hyperliquid.