Skip to content
MrJev

reflex

A small open decision model for your own GPU: fixed answer options in, per-option percentages out, with no free text so it cannot answer off the list.

View on GitHub →

Hands-on review

An open re-creation of the System One idea that trains nothing: a frozen open-weights model, scored over your option set, behind one HTTP endpoint.

Good for

  • Nothing to fine-tune: it scores options against a frozen base model
  • A `stable` tag whose config file records what was selected and on what evidence
  • Bearer auth on the endpoint, and a state cache you can watch hit

Watch out for

  • Documented for Linux and an NVIDIA GPU; CPU works but is not a plan
  • Nine tests fail on a CPU-only box with the project's own dependencies
  • The README's benchmark tables are the author's; we verified none of them

Tested Sep 22, 2026 at dec50eb5a15e · python:3.12-slim in Docker on CPU, no GPU: its test suite, its engine and its HTTP server driven with Qwen3-0.6B

How we reviewed this: on CPU in python:3.12-slim, with no GPU — the test suite, then the engine and the HTTP server loaded with Qwen/Qwen3-0.6B instead of the 4B the README uses, because the method does not care which base model you hand it. This is an open model, so everything here is our own run of their code; we made no Jev calls and this review contains no measurements of Jev.

What it is

The idea in one line: do not train a decision model, make one. Take a normal open-weights model, freeze it, and instead of letting it write, score the options you supplied against its logits. Out comes a probability per option, for every question at once, and no free text — so, as the README puts it, “it can never make up an answer that isn’t on your list”.

The default is Qwen3.5-4B. Nothing is fine-tuned on the recommended path: serving/stable.json at the stable tag says "adapter": null, "calibration": null.

Any model will do

We never downloaded the 4B. We handed it Qwen/Qwen3-0.6B — a general chat model a fraction of the size, with no adapter — and asked the README’s own support-ticket questions on CPU:

"model": "Qwen/Qwen3-0.6B"
team      choice "billing"   billing 0.99029 · support 0.000573 · other 0.009137   confidence 0.948
escalate  noul   0.850154
urgency   score  0.432584    low 0.639 · medium 0.289 · high 0.072
usage     input_tokens 245 · state_tokens 71 · question_tokens 174 · state_cache_hit false

A 0.6B model that was never trained for this routed a double-charge complaint to billing at 0.99. That is the argument the repository is making, and it survives being run on the wrong model on the wrong hardware.

The usage block is worth its own sentence. State tokens and question tokens are counted separately, and state_cache_hit is reported per request — so “the second question about the same document is cheaper” is something the caller can read rather than assume.

The server

Started on CPU with REFLEX_API_KEY set:

request response
no Authorization header 401 {"error":{"message":"Invalid or missing API key","type":"invalid_api_key"}}
a wrong bearer key 401, same body
the right key 200, typed answers
"questions": {} 422 at least one question is required

Auth is off unless you set a key, and the docstring above it says what to do: it “makes /v1/systemone require Authorization: Bearer <key>, which you want on any endpoint reachable from the network”. Fair, and worth reading before you expose it.

The start-up line tells you exactly what you are running:

kernels: qwen3 | attn=sdpa | bfloat16 on cpu | no kernel-backed ops |
         installed: flash-linear-attention==0.5.2

“no kernel-backed ops” is the honest report of a CPU box: the fast paths are installed and inactive. And the state cache is visible in the log — the same 4,057-token state took 104,561 ms on the first request and 1,243 ms on the second, marked (hit). On a CPU with a 0.6B model, those numbers say nothing about the project’s own hardware, but they do show the cache doing what it claims.

The suite, and one rough edge

At the stable tag: 23 passed, 9 skipped. On main, which has moved a long way since — 55 passed, 18 skipped, and 9 failed, all in test_mps_delta_rule.py:

RuntimeError: 0 active drivers ([]). There should only be one.   triton/runtime/driver.py

We chased it rather than reporting it. That file compares the project’s hand-written gated delta rule against the one in transformers, and guards itself with skipif(torch.cuda.is_available()) because on CUDA the “reference” dispatches to a Triton kernel that refuses CPU tensors. On a CPU-only machine torch.cuda.is_available() is false, so the guard does not fire — but Triton is installed anyway, as a dependency of the project’s own flash-linear-attention>=0.5.2, and it explodes on a box with no driver. Remove Triton from the path and the same commit gives 64 passed, 18 skipped.

So: not a defect on the platform the README documents (Linux plus an NVIDIA GPU), and not worth an issue. Worth knowing if you clone this on a laptop and think you have found something.

What we did not check

The README leads with two benchmark tables — self-run JevBench figures and a placement from an external run of JevBench v1.2 by the benchmark’s author. This site does not republish other projects’ comparisons against Jev, and we did not attempt to reproduce either table; they are the author’s, labelled as the author’s, and you should read them as such.

What we can say about the surrounding process is that it is unusually disciplined for a project this young. stable is a git tag that moves with a configuration file, and that file names the model, the prompt style, the number of option orders, the release date and the external-set accuracies it was selected on — with a comment saying the tag and the file move together, “only after the external-set and public-item gates pass”. A project that writes down what a release was chosen on is a project you can argue with.

Verdict

The most approachable entry point in this category, because the answer to “where do I get the weights” is “you already have some”. If you have an NVIDIA box and a base model you trust, uv sync and the support-ticket example will tell you in five minutes whether typed scoring solves your problem.

Do not expect the CPU path to be a deployment, do not expose the endpoint without setting a key, and read serving/stable.json before you pin anything. For a project that trained its own weights instead, see decider and Dev-0.4B; for the same trick on images, Jev Visual.

See how it compares with other tools in Best Jev tools, tested hands-on.

Review updated Sep 22, 2026. Numbers quoted from the project are its author's own; we don't publish our own measurements of Jev.

More in Open Models & Reproductions

Laya

★ 27k▲ 12k

NandhaKishorM/laya

Non-autoregressive decision engine over 100+ languages: three checkpoints and a router that detects the script and dispatches per request. Its benchmarks end with a limits section naming the datasets it does not generalise to and the headline figure that came from a training split.

PythonReviewed

kev

★ 7.5k▲ 6.2k

jaredpalmer/kev

Jev-style decision models from 0.5B to 8B, built as LoRA adapters on Qwen and served behind a Jev-compatible /v1/systemone API.

PythonReviewed

SemIf

★ 4.5k▲ 1.9k

TheoLeeCJ/SemIf-OpenJev

Jev-style decisions from a frozen 4B model on a single RTX 3090, with a browser demo. Formerly OpenJev.

PythonReviewed

Get new Jev projects every week

New Jev releases, pricing changes, and the best new projects, once a week. No spam; unsubscribe anytime.

Powered by Buttondown. See our privacy policy.