Docs

Project list

60 projects, organised by what you are trying to solve. One line each, with a star count and a snapshot date.

How to read a line

  • The star count reflects attention and usage, not quality; the date beside it is when it was read.
  • The evidence tier (reproducible / implemented) answers “how much can be checked”, not whether it is good. A few lines of code that publish raw results and a feature-complete project whose numbers all come from its author land on different tiers — and that has nothing to do with whether either is useful.

60 projects. Star counts are as of 2026-09-24.

Reproducible
Publishes raw artifacts, datasets or a re-runnable evaluation
Implemented
Readable code with tests; the numbers are the author’s own

Interfaces and browsers

Where every step is a click. Decisions are cheap enough to ask at every step.

  • Jev Ultrafast19,456★Implemented

    Browser Use’s ultrafast agent: the decision model picks both the operation and the DOM element in one request, and a small LLM is called only when text must be typed.

  • jev-chat-jarvis5,561★Implemented

    A phone-side conversation copilot: reads the other party in WeChat, QQ, X or Feishu, proposes replies and fills the input box — sending stays manual. It reads the screen only, hooking nothing.

  • typesafe-computer-use937★Implemented

    Computer use on macOS: OCR the screen, classify the next action with the decision model, click. Reported at about $0.0002 a step.

  • hermes-jev-skills726★Implemented

    A skill suite for Hermes agents covering model routing, memory, compaction, skill selection, and computer or browser use — also installs under Claude Code and Codex.

  • jev-browser-use453★Implemented

    Lets the decision model pick the click while Codex thinks and verifies — reported 5–10× faster browser operations.

  • typesafe-mario383★Implemented

    An agent that plays Super Mario Bros. from structured emulator state rather than pixels.

  • mobile-jev380★Implemented

    A standalone Android agent where the decision model makes every decision, with a live studio showing decision telemetry.

  • jev-chat-jarvis-mac346★Implemented

    A WeChat intent-detection overlay for macOS: screen reading plus a local small model to judge intent and risk, then reply candidates — strictly read-only, nothing injected into WeChat.

  • jev-voice-browser271★Implemented

    Control a real browser by voice: intent and target are decided in about 300 ms per spoken word, then Playwright acts — often before you finish the sentence.

  • jev-browser257★Implemented

    Browser use driven by the decision model.

  • unclutter227★Implemented

    A browser extension that uses the decision model to find page clutter and remove it, with reusable template rules.

  • embodied-jev221★Implemented

    A MuJoCo robot-decision workbench.

  • jev-drone170★Implemented

    A camera-only autonomous drone in MuJoCo with the judgement step running at 2.5 Hz.

Putting brakes on an agent

The second layer of judgement for a coding agent: allow or not, is it done, what is still needed in context.

  • fast-jev-compaction6,653★Implemented

    A Claude Code plugin replacing the compaction summary with decisions: every tool call and result is scored in one request, stale ones are dropped or truncated, and everything kept stays verbatim.

  • foreman540★Implemented

    An agent supervisor that uses decisions to keep coding agents on task.

  • JevHarness191★Implemented

    LLM-authored task-specific decision harnesses, with optional full-trajectory reward reflection and GEPA evolution.

  • jev-dsh-decision163★Implemented

    A structured-decision plugin for agent harnesses: native on DeepSeek Harness, and through iPolloWork also OpenCode and Codex Harness.

  • jev-pruner144★Implemented

    A Claude Code plugin that trims long Bash output before the model ever sees it.

  • pi-jev (y0usaf)144★Implemented

    A decision layer for the Pi coding agent: a measured tool-call gate plus jev_ask for typed, calibrated answers.

  • pi-warden138★Implemented

    Guardrails for Pi that steer rather than interrupt: the decision model judges irreversible and off-task tool calls, detects stuck loops, and checks unverified done claims.

  • building-with-jev-skill131★Implemented

    A skill for writing and improving programs that call the decision model.

Routing, search and review

Deciding who handles it, or picking which candidates are worth looking at.

  • jev-review (devagrawal09)586★Implemented

    A staged code-review workflow with a local dashboard.

  • jev-search446★Implemented

    Web search with the decision model: source selection, query understanding and relevance ranking.

  • jev-router (gargpratyush)377★Implemented

    Routes Claude Code tasks to the cheapest capable model by asking the decision model to choose among candidates.

  • pg-jev329★Implemented

    A PostgreSQL extension for asking your tables questions in plain language.

  • jev-codex-router263★Implemented

    Per-turn routing for Codex: every turn re-decides the model, the thinking depth and the speed mode.

  • jev-review (NiazMorshed2007)218★Implemented

    A local-first MCP plugin for continuous software-quality review by coding agents.

  • notra215★Implemented

    A marketing analytics platform that moved a brand-visibility classifier off an LLM and onto boolean decisions, switched by a feature flag.

  • JevRouter190★Implemented

    A lightweight router for models, tools and subagents.

  • perch171★Implemented

    A CLI that parses a repository into a method-level call graph, then asks the decision model to flag likely defects and check project rules written in plain English.

  • jev-semgrep134★Implemented

    Grep by meaning rather than pattern: every line is scored against a meaning, and meanings combine with AND / OR / NOT.

  • neo4jev129★Implemented

    Navigates a Neo4j graph hop by hop using a classifier over candidate outgoing relationships, with beam search over answer log-probabilities.

Local and open alternatives

Open weights and local runtimes behind the same HTTP interface.

  • SemIf4,162★Reproducible

    Semantic ifs from open models, on a single 3090 at home. Independent; not affiliated with the vendor.

  • NanoJev2,159★Reproducible

    A 0.6B replica: parallel decisions over dynamic candidates, no output-token decoding, with the end-to-end training pipeline.

  • jevlike1,279★Reproducible

    Train a small model that chooses among a changing list of text options, one probability per option in a single pass — with Doom, chess and Wikispeedia demos.

  • von610★Reproducible

    A 395M non-autoregressive model answering typed questions with calibrated probabilities in under 15 ms, positioned as a local drop-in.

  • simple-jev507★Reproducible

    Turn any open model into a classifier, or into a compatible endpoint.

  • AnyJev440★Reproducible

    Turns any LLM into a decision model: typed decisions and calibrated probabilities are read from next-token prefill distributions, with no training.

  • openjev386★Reproducible

    An open, vendor-compatible decision server built on DiffusionGemma.

  • decider351★Reproducible

    One-pass typed decisions with calibrated probabilities, fine-tuned from a 2B model.

  • LLM2Jev300★Reproducible

    Adapts local language models into compatible structured decision engines, with all three primitives produced by prefill-only binary inference.

  • agent-jev286★Reproducible

    A 0.6B replica aimed at agents: feed it unstructured state (diffs, traces, logs) and structured questions, get calibrated distributions back in one ~50 ms pass with no output tokens.

  • openJev-verdict-2.0283★Reproducible

    A calibrated 151M non-autoregressive decision engine that its author reports as beating both the vendor model and Laya on a public benchmark: 77.10% accuracy, 0.0636 Brier, 0.0144 ECE.

  • jev-visual271★Reproducible

    An educational inference experiment on Apple Silicon: shared context, direct candidate scoring, and local visual demos.

  • jeff238★Reproducible

    A self-hosted drop-in replacement.

  • reflex133★Reproducible

    A small open decision model: state plus typed questions, calibrated probabilities back.

SDKs and integration

Turning a judgement into ordinary control flow in your language, or into a framework you already use. Includes the official SDKs.

  • eve5,342★Implemented

    Vercel’s open agent framework, which ships the decision model as the default evaluation model in its experimental evaluate path.

  • TypeSafe agent skills2,060★Implemented

    The official agent skill for Claude Code, Codex and compatible agents: the primitives, the patterns, and how to structure evaluations.

  • ai-cli813★Implemented

    A Vercel Labs terminal CLI that can run the decision model as the evaluation model for its evaluate command.

  • jev-mcp (jkudish)320★Implemented

    A proof-of-concept MCP server for connecting the decision model to MCP clients.

  • typesafe-mcp292★Implemented

    A single-binary Go MCP server exposing typed judgements to Claude Desktop, Claude Code and Codex.

  • System One adapter (Python)287★Implemented

    The official drop-in client replacement backed by OpenAI, Anthropic or compatible LLM APIs — made for comparing the decision model against chat models in the same evaluation.

  • TypeSafe JavaScript SDK232★Implemented

    The official TypeScript/JavaScript client, with inferred answer types.

  • TypeSafe Python SDK220★Implemented

    The official sync and async Python client.

Applications and experiments

Things built for a person to use directly: trading, search, writing, killing an idea.

  • jev-trader2,240★Implemented

    One trading decision every Monad block.

  • shapeshift491★Implemented

    An input that becomes what you mean: one text box that morphs into the right UI as you type. Works offline.

  • JevRev304★Implemented

    An LLM plus decision-model workflow. (The public description goes no further than this.)

  • jev-align (Sutro)284★Implemented

    An active-learning CLI that uses human labels and GEPA to improve how decisions are defined.

  • killmyidea159★Implemented

    Describe your startup idea and it decides: kill it, fix it, or ship it.

  • jev-trade131★Implemented

    A live decision-driven trader on Hyperliquid.