Skip to content
MrJev

Dev-0.4B

A 399M bidirectional encoder with one universal choice head, answering Choice, Noul and Score in a single forward pass. Every README figure reconciles to an evaluation JSON shipped in the repository.

View on GitHub →

Hands-on review

A 399M encoder answering typed questions in one forward pass. Every published figure traces to a shipped artifact; the flagship example does not.

Good for

  • Every README number reconciles exactly to a JSON artifact in the repo
  • A real decision in about a quarter of a second on an ordinary CPU
  • Calibrated and uncalibrated runs both published, so you can see the cost

Watch out for

  • The model card's own choice example returns a different answer
  • `run.json` names a different checkpoint than the benchmark table
  • No LICENSE file, and the quickstart's clone URL is a 404

Tested Sep 22, 2026 at 2ce2563db938 · python:3.13 + CPU torch in Docker with --network none, running the published weights from a local cache

How we reviewed this: we downloaded the published weights, ran them offline on CPU with --network none, executed the model card’s own examples verbatim, and reconciled every figure in the README against the JSON artifacts in docs/. We made no Jev calls; this is an open model.

What it is

A 399M-parameter bidirectional encoder — ModernBERT-large with a single universal choice head — that answers choice, noul and score questions by reading pooled candidate spans in one forward pass. The argument is against constrained decoding on an 8B generative model: no token loop, no JSON to repair.

It loads, and it works. On an ordinary CPU, with nothing but torch:

[params] 396,880,897
[choice ]  314 ms
[noul   ]  266 ms   "ModernBERT … 8k context natively" -> probability_yes 0.9955
[noul-neg] 212 ms   "rolled back after three health checks failed" -> probability_yes 0.0012
[score  ]  251 ms   a mediocre restaurant review -> peak on "average", value 1.86/4

The model card’s noul example reproduces almost exactly — it says 0.9942, we got 0.9955 — and the negative control behaves. A clear billing ticket routes correctly at 0.9922 confidence.

The numbers reconcile

This is the part we most wanted to check, and it passed. Every headline figure in the README traces to a file in the repository:

README artifact
BoolQ 85.20% benchmark-uncalibrated.json → 0.8520
Banking77 91.33%, top-3 98.67% → 0.9133, 0.9867
Yelp 62.67%, MAE 0.4017 → 0.6267, 0.40169

Shipping the JSON beside the claim is what makes a benchmark checkable by a stranger, and we could do it in minutes rather than by re-running anything. Both a calibrated and an uncalibrated run are published, which is more honesty than the table strictly required.

Reading them side by side is instructive. Calibration leaves accuracy untouched on all three tasks — as temperature scaling must — and improves expected calibration error: BoolQ 0.1033 → 0.0771, Banking77 0.0754 → 0.0553, Yelp 0.3178 → 0.1545.

That Yelp figure is the thing to notice. For a model whose whole selling point is a usable probability rather than a token, an ECE of 0.32 on the ordinal head — 0.15 even after calibration — is the weakest part of the picture, and the project publishes it rather than hiding it. Calibration also moves Yelp MAE the wrong way, 0.4017 → 0.4287, and the README quotes the better one.

The demo does not reproduce

Running the model card’s flagship choice snippet verbatim against the published weights:

model card says:  criterion='email_infrastructure'  confidence=0.9987
what we get:      criterion='general_inquiry'       confidence=0.4786

Deterministic across runs. Rewriting the same four queues as descriptions rather than snake_case labels moves it to Login, passwords, account access and security at 0.61 — defensible, since the customer cannot log in, but still not the documented answer and still not confident.

There is a plausible explanation in the repository. The run.json shipped inside the model says the weights came from runs/jev-unified-large, while the README’s benchmark table comes from runs/dev-0.4b / runs/jev-unified-epoch2 — and those are measurably different models in the project’s own artifacts:

checkpoint BoolQ Banking77 Yelp
runs/dev-0.4b (the README) 0.8520 0.9133 0.6267
runs/jev-unified-large (run.json) 0.8340 0.9333 0.6700

Not cherry-picked — the published checkpoint is better on two of three — but it means the weights on the Hub may not be the ones the documentation describes. Reported, along with two smaller things: the quickstart’s git clone https://github.com/nikhilpujari/dev.git is a 404, and there is no LICENSE file despite both README and model card saying “Apache 2.0”.

On the comparison table

The README leads with a head-to-head against another project’s model, with trophies and “Decisive Win” in each row. We do not republish one project’s benchmark of another — those are the author’s numbers, measured on the author’s harness — and we have not attempted to reproduce them.

What we can say about the method is that it states its terms: held-out public benchmarks, raw out-of-the-box predictions, a standard 0.50 threshold, and “zero test set peeking”. The sample counts are in the JSON and they are small — 500 BoolQ, 300 Banking77, 300 Yelp — which the artifacts disclose and the README’s percentages do not.

Verdict

A real model that does a real thing, published with the evidence attached — and we would rather review ten projects that ship their evaluation JSON than one that ships a chart.

The gap is between the artifacts and the shipping. A flagship example that returns the wrong answer is the first thing a new user will run, and “which checkpoint is on the Hub” should not be a question a reviewer has to answer from a stray run.json. Both are a day’s work.

Wait for the checkpoint question to be resolved before building on it, and treat the missing licence as a real blocker if you intend to redistribute anything.

For a smaller open model read out of logits rather than a trained head, see JEV-CPU and SemIf; for an open reproduction that publishes its protocol first, von.

See how it compares with other tools in Best Jev tools, tested hands-on.

Review updated Sep 22, 2026. Numbers quoted from the project are its author's own; we don't publish our own measurements of Jev.

More in Open Models & Reproductions

Laya

★ 27k▲ 12k

NandhaKishorM/laya

Non-autoregressive decision engine over 100+ languages: three checkpoints and a router that detects the script and dispatches per request. Its benchmarks end with a limits section naming the datasets it does not generalise to and the headline figure that came from a training split.

PythonReviewed

kev

★ 7.5k▲ 6.2k

jaredpalmer/kev

Jev-style decision models from 0.5B to 8B, built as LoRA adapters on Qwen and served behind a Jev-compatible /v1/systemone API.

PythonReviewed

SemIf

★ 4.5k▲ 1.9k

TheoLeeCJ/SemIf-OpenJev

Jev-style decisions from a frozen 4B model on a single RTX 3090, with a browser demo. Formerly OpenJev.

PythonReviewed

Get new Jev projects every week

New Jev releases, pricing changes, and the best new projects, once a week. No spam; unsubscribe anytime.

Powered by Buttondown. See our privacy policy.