Skip to content
MrJev

decider

A family of System One-style models that never generate text: one forward pass over a state and typed questions returns a probability distribution per question. Ships ten text games and a Super Mario Bros agent where each move is one typed decision over the legal actions.

View on GitHub →

Hands-on review

An Apache-2.0 reproduction of the System One idea: typed questions in, a probability per option out, and a limits section that lists what it cannot do.

Good for

  • Weights you can run today: the 2B and its server both run on plain CPU
  • A README example that reproduces against the released checkpoint
  • A limits section that names regressions, teacher bias and the weak axis

Watch out for

  • An empty question map is answered `200 {}`, by design — an empty result is not 'no findings'
  • The answer set is closed, so it always answers — even between two nonsense options
  • English only, and one pass means no arithmetic and no multi-hop chains

Tested Sep 23, 2026 at 104b844b4b8b · python:3.12-slim in Docker on CPU, no GPU: its own test suite, its quick-start example against the released 2B weights, and its HTTP server on each version through 1.1.3, the last of them installed from PyPI and run with --network none

How we reviewed this: on CPU in python:3.12-slim, with no GPU anywhere — its own pytest suite, then its quick-start example against the released Mapika/decider-2b weights downloaded from Hugging Face, then its HTTP server. This is an open model, so everything here is our own run of their model; we made no Jev calls and this review contains no measurements of Jev.

What it is

A language model that does not generate text. It reads a state and a set of typed questions and returns, in one forward pass, a probability distribution for each question — Choice over up to 255 options, Score over described levels, or Noul, the probability of yes.

Three sizes, built on Qwen3.5 bases: a 2B, a 4B released today, and a 35B mixture-of-experts. Apache-2.0, weights on Hugging Face, and the independence statement is in the second paragraph rather than the small print: “This is an independent project. It is not affiliated with or endorsed by TypeSafe AI… The training mixture is public datasets plus data labelled by a local Qwen3.5-27B teacher. Nothing was distilled from Jev.” The teacher data is in the repository.

The README’s example is real

The quick start shows a support ticket, three questions of three different types, and the exact answers it expects. We ran it on CPU against the released 2B:

"model": "decider-v10"
department        choice "billing"   billing 0.5786 · returns 0.4197 · other 0.0017
refund_requested  noul   0.986
frustration       score  0.76        0.3466 / 0.5471 / 0.1063

The README prints 0.56, 0.44, 0.99 and 0.76 for the same call. Every figure lands where it says, on a different device and dtype than the author used. A README example that reproduces against the shipped checkpoint is rarer than it should be, and it is the cheapest possible way to establish that a repository is what it claims.

It is not fast on a CPU, and nobody says it should be: the weights load, the kernels fall back with a warning naming what is missing (flash-linear-attention, causal_conv1d), and the answer arrives. For a laptop test of whether typed decisions fit your problem, that is enough.

The closed answer set, and what it costs

The best demonstration of the shape is a stupid question. We asked it to choose between two strings with no meaning:

d.decide("My card was charged twice.", [{"question": "Which team?", "options": ["zzzz", "qqqq"]}])
# [{"choice": "zzzz", "confidence": 0.63, "probs": {"zzzz": 0.63, "qqqq": 0.37}}]

There is no way for this model to say “neither”. The probabilities are over the options you supplied and they sum to one, so an option set that does not contain the right answer gets a confident-looking wrong one. Decider takes an abstain_below threshold for exactly this, and the limits section says the same thing from the other side: “A plain support next to other sends an in-scope complaint to other.” Design the option set, or the model will design it for you.

The whole thing runs without a GPU

$ python -m pytest tests
110 passed, 9 skipped

The skips are the CUDA and MPS paths, which is what the README promises.

The server used to be the exception. On 2026-09-22 it could not start without CUDA — Engine defaulted to device="cuda", no setting overrode it, and the failure was AssertionError: Torch not compiled with CUDA enabled a few lines under a Runs on section listing all three devices. We reported it and suggested a one-line doc fix; the author shipped considerably more than that within the hour, in 1.1.2. We re-ran it rather than take the word for it:

GET  /health        {"ok": true, "model": "Mapika/decider-2b", "device": "cpu"}
POST /v1/systemone  200   refund noul 0.9781 · team choice "billing" 0.9983
POST /decide        200   {"Which team should handle this?": {"choice": "billing", "confidence": 0.9387, …}}

The server now selects its device the way Decider does, DECIDER_DEVICE overrides it, and asking for a device this machine does not have stops start-up with a sentence that names the requirement rather than torch’s assertion:

DECIDER_DEVICE=cuda    but torch.cuda.is_available() is False on this machine.
                       Set DECIDER_DEVICE=mps or cpu (or leave it at auto).
DECIDER_DEVICE=bogus   expected auto, cuda, cuda:<index>, mps or cpu.

POST /v1/systemone speaks TypeSafe’s wire format, so a client written against Jev can be pointed at it with TYPESAFE_BASE_URL — now from a laptop, not only from a GPU box.

The rough edge that was left after that is now gone too. /decide used to read its body loosely: schema_: dict = None accepted a missing key and any dict, so {"context": "…"} and {"context": "…", "schema": {"Which team?": ["billing", "technical"]}} both reached the executor and came back 500 Internal Server Error with an AttributeError in the log. The second is the shape a reader might guess from the README’s d.decide(ctx, [{"question": …, "options": [...]}]); the real one is {question: {"type": "choice", "options": [...]}}. Reported, and fixed in 1.1.3 a day later — the schema is checked before any work, and the author found a third 500 of his own while doing it.

We ran the released 1.1.3 rather than the branch: pip install decider-ai[serve]==1.1.3 in python:3.12-slim, --network none, weights from a local cache, DECIDER_DEVICE=cpu.

POST /decide  {"context": …}                                   422 schema is required; schema: a map {question: field}, …
              schema {"Which team?": ["billing","technical"]}   422 schema['Which team?'] is a list, not a field object; …
              choice with "options": "billing"                  422 "options" must be a non-empty list of strings; …
              scale with no legend                              422 "legend" must be a non-empty list or {level number: description} map; …
              legend {"nan": "a", "1": "b"}                     422 the keys of a "legend" map must be finite numbers, …
              {"type": "mystery", …}                            422 unknown field type 'mystery'; …
              a well-formed two-question schema                 200 choice "billing" 0.9759 · noul 0.8368

Every message names the expected form, and the same messages come out of the library: decider.infer.Decider._check_schema is what both paths call, so decide_json raises ValueError: schema is required; … on the body that used to 500 over HTTP. /v1/systemone is unchanged, including its own limit — 300 options still gets 422 choice criteria: a map of 2..255 options.

One behaviour changed on purpose: options or a legend given as a bare string were split into single characters by 1.1.2 and are a 422 now. And one thing deliberately did not. /decide with {} and /v1/systemone with "questions": {} both answer 200 with an empty answer map, which the author is keeping because a caller that builds its question map at run time can legitimately send an empty one. Fair, and worth knowing at the other end: an empty answer map means nothing was asked, not that nothing was found.

The limits section

Most model repositories have a limitations heading with three hedges under it. This one has thirteen bullets, and they are specific enough to act on:

Rules written into the question are not followed at this size. On the form-filling probe a one-sentence question scores 0.67 and a paragraph of rules 0.24. A fixed convention has to be in the training data, not in the question.

Picking a record out of a long JSON array by position is the least accurate input shape. Address records by key.

Known regressions. TREC-fine with all 50 labels fell from 0.76 (v6) to 0.72 (v8). Held-out Freeway play fell to 0 and did not come back when the game data was replayed.

Teacher bias. The custom-question data is labelled by a 27B teacher that shares some of the biases it is meant to fix; it agreed with only 72% of its own generic-option labels.

Also: English only, no arithmetic and no multi-hop chains in one pass, and the vision variant is on older text weights and retraining. The demo captions carry the same instinct — the ten image questions in the README GIF were “chosen by a fixed rule that includes the confident wrong answers, not the ten best”, and docs/DEMOS.md gives the checkpoints, seeds and timers.

The repository carries a model card per size, a changelog, a history, a results document and the teacher data. It is a research repository written by someone who expects to be checked.

Verdict

The most complete open answer to “what is a System One model, actually” that we have run. Clone it, pip install -e ., run the quick start on whatever machine you have, and you will know within ten minutes whether typed decisions fit your problem — no account, no key, no card.

Take the limits section seriously before you put it in front of traffic: English only, and a closed answer set that always answers. For the same idea at a smaller scale, see reflex and Dev-0.4B; for the multilingual end of it, Laya.

See how it compares with other tools in Best Jev tools, tested hands-on.

Review updated Sep 23, 2026. Numbers quoted from the project are its author's own; we don't publish our own measurements of Jev.

More in Open Models & Reproductions

Laya

★ 27k▲ 12k

NandhaKishorM/laya

Non-autoregressive decision engine over 100+ languages: three checkpoints and a router that detects the script and dispatches per request. Its benchmarks end with a limits section naming the datasets it does not generalise to and the headline figure that came from a training split.

PythonReviewed

kev

★ 7.5k▲ 6.2k

jaredpalmer/kev

Jev-style decision models from 0.5B to 8B, built as LoRA adapters on Qwen and served behind a Jev-compatible /v1/systemone API.

PythonReviewed

SemIf

★ 4.5k▲ 1.9k

TheoLeeCJ/SemIf-OpenJev

Jev-style decisions from a frozen 4B model on a single RTX 3090, with a browser demo. Formerly OpenJev.

PythonReviewed

Get new Jev projects every week

New Jev releases, pricing changes, and the best new projects, once a week. No spam; unsubscribe anytime.

Powered by Buttondown. See our privacy policy.