How we reviewed this: on CPU in python:3.12-slim, with no GPU — the test suite, then the engine and the HTTP server loaded with Qwen/Qwen3-0.6B instead of the 4B the README uses, because the method does not care which base model you hand it. This is an open model, so everything here is our own run of their code; we made no Jev calls and this review contains no measurements of Jev.
What it is
The idea in one line: do not train a decision model, make one. Take a normal open-weights model, freeze it, and instead of letting it write, score the options you supplied against its logits. Out comes a probability per option, for every question at once, and no free text — so, as the README puts it, “it can never make up an answer that isn’t on your list”.
The default is Qwen3.5-4B. Nothing is fine-tuned on the recommended path: serving/stable.json at the stable tag says "adapter": null, "calibration": null.
Any model will do
We never downloaded the 4B. We handed it Qwen/Qwen3-0.6B — a general chat model a fraction of the size, with no adapter — and asked the README’s own support-ticket questions on CPU:
"model": "Qwen/Qwen3-0.6B"
team choice "billing" billing 0.99029 · support 0.000573 · other 0.009137 confidence 0.948
escalate noul 0.850154
urgency score 0.432584 low 0.639 · medium 0.289 · high 0.072
usage input_tokens 245 · state_tokens 71 · question_tokens 174 · state_cache_hit false
A 0.6B model that was never trained for this routed a double-charge complaint to billing at 0.99. That is the argument the repository is making, and it survives being run on the wrong model on the wrong hardware.
The usage block is worth its own sentence. State tokens and question tokens are counted separately, and state_cache_hit is reported per request — so “the second question about the same document is cheaper” is something the caller can read rather than assume.
The server
Started on CPU with REFLEX_API_KEY set:
| request |
response |
no Authorization header |
401 {"error":{"message":"Invalid or missing API key","type":"invalid_api_key"}} |
| a wrong bearer key |
401, same body |
| the right key |
200, typed answers |
"questions": {} |
422 at least one question is required |
Auth is off unless you set a key, and the docstring above it says what to do: it “makes /v1/systemone require Authorization: Bearer <key>, which you want on any endpoint reachable from the network”. Fair, and worth reading before you expose it.
The start-up line tells you exactly what you are running:
kernels: qwen3 | attn=sdpa | bfloat16 on cpu | no kernel-backed ops |
installed: flash-linear-attention==0.5.2
“no kernel-backed ops” is the honest report of a CPU box: the fast paths are installed and inactive. And the state cache is visible in the log — the same 4,057-token state took 104,561 ms on the first request and 1,243 ms on the second, marked (hit). On a CPU with a 0.6B model, those numbers say nothing about the project’s own hardware, but they do show the cache doing what it claims.
The suite, and one rough edge
At the stable tag: 23 passed, 9 skipped. On main, which has moved a long way since — 55 passed, 18 skipped, and 9 failed, all in test_mps_delta_rule.py:
RuntimeError: 0 active drivers ([]). There should only be one. triton/runtime/driver.py
We chased it rather than reporting it. That file compares the project’s hand-written gated delta rule against the one in transformers, and guards itself with skipif(torch.cuda.is_available()) because on CUDA the “reference” dispatches to a Triton kernel that refuses CPU tensors. On a CPU-only machine torch.cuda.is_available() is false, so the guard does not fire — but Triton is installed anyway, as a dependency of the project’s own flash-linear-attention>=0.5.2, and it explodes on a box with no driver. Remove Triton from the path and the same commit gives 64 passed, 18 skipped.
So: not a defect on the platform the README documents (Linux plus an NVIDIA GPU), and not worth an issue. Worth knowing if you clone this on a laptop and think you have found something.
What we did not check
The README leads with two benchmark tables — self-run JevBench figures and a placement from an external run of JevBench v1.2 by the benchmark’s author. This site does not republish other projects’ comparisons against Jev, and we did not attempt to reproduce either table; they are the author’s, labelled as the author’s, and you should read them as such.
What we can say about the surrounding process is that it is unusually disciplined for a project this young. stable is a git tag that moves with a configuration file, and that file names the model, the prompt style, the number of option orders, the release date and the external-set accuracies it was selected on — with a comment saying the tag and the file move together, “only after the external-set and public-item gates pass”. A project that writes down what a release was chosen on is a project you can argue with.
Verdict
The most approachable entry point in this category, because the answer to “where do I get the weights” is “you already have some”. If you have an NVIDIA box and a base model you trust, uv sync and the support-ticket example will tell you in five minutes whether typed scoring solves your problem.
Do not expect the CPU path to be a deployment, do not expose the endpoint without setting a key, and read serving/stable.json before you pin anything. For a project that trained its own weights instead, see decider and Dev-0.4B; for the same trick on images, Jev Visual.
See how it compares with other tools in Best Jev tools, tested hands-on.
Review updated Sep 22, 2026. Numbers quoted from the project are its author's own; we don't publish our own measurements of Jev.