Skip to content
MrJev

Jev Visual

Multiple typed questions about one image in a single pass, on Qwen3.5-0.8B with MLX on Apple Silicon. States plainly that it explores the pattern and does not claim to reproduce Jev's architecture or training. Chinese and English.

View on GitHub →

Our review

Many typed questions about one image in a single pass on Apple Silicon — and every figure in its results file regenerates from the records it ships.

Good for

  • A published results table that regenerates byte-for-byte from its raw records
  • A scoring check against full-vocabulary forwards, with the errors committed
  • Disclaimers in the README rather than at the bottom of a docs page

Watch out for

  • Apple Silicon with Metal only; we could not run a single inference
  • Only the Qwen3.5 adapter is verified, at 1–64 questions and 2–26 options
  • Probabilities are relative to your options, and the project says so twice

How we reviewed this: we could not run it. It needs MLX on an Apple Silicon Mac, and we work on x86 Linux, so no image was ever scored here. What we could do is check everything the repository asserts about runs it already made: its model-free tests, its published results table regenerated from the raw records beside it, and both aggregates in its scoring-verification file recomputed by hand. We made no Jev calls.

What it is

One image, one shared prefill, then many typed questions answered off that same context: choose an option, judge yes/no, or score ordered levels. Qwen3.5-0.8B in 4-bit through MLX, about 596 MiB of weights, with a browser UI, a CLI and an HTTP API. Bilingual, English and Simplified Chinese.

The claim it is not making is in the second paragraph, before any demo:

This is an independent community implementation and does not claim to reproduce TypeSafe Jev’s proprietary model architecture, RLCD training, calibration, or serving system.

What we could check, and it all holds

The tests that need no model. The README says python -m pytest -q runs “Unit/API tests; no model inference”. With MLX absent entirely from the machine: 14 passed. Nothing in the schema, scoring or API-shape layer requires the accelerator.

The published table. benchmarks/RESULTS.md is the document the README links for results, and benchmarks/results.json is the raw record beside it. We ran the project’s own benchmarks.report over the committed records and diffed the output against the committed document:

$ python -m benchmarks.report benchmarks/results.json --output /py/regen.md
$ diff /py/regen.md benchmarks/RESULTS.md

One line differs, and it is the relative path of a link, rewritten by the script from wherever you put the output. All 26 table rows, every latency sample, every agreement figure: identical. The published table is the record, not a transcription of it.

The scoring verification. benchmarks/scoring-verification.json claims two things: that the shared-prefix path and the independent path agree, and that both match full-vocabulary model forwards. Both are aggregates over per-option scores stored in the same file, so we recomputed them:

aggregate published our recomputation
max_probability_delta, shared vs direct 1.75009471181653e-06 identical
max_oracle_score_error vs full forwards 3.906427846800398e-05 identical (3.9064278468e-05)

And the comparison is honest at the input end too: every answer carries a prompt_sha256, and the hashes match between the shared and direct runs, so the two paths were scored on the same prompt rather than on two prompts that happened to agree. The four questions cover both scoring modes — single-token candidate logits and full sequence log-probabilities including EOS — so the check is over the paths that actually differ.

This is the whole case for the optimisation the project exists to demonstrate, and it is checkable from a laptop with no Mac in it.

The disclaimers are load-bearing

Most projects put the caveats at the bottom of a docs page. Here they are in the paragraph that introduces the feature:

  • “Candidate probabilities are relative to supplied options, not correctness estimates.”
  • “Only the Qwen3.5 adapter is verified. Cache sharing is within one request and copies state; it is not zero-copy sharing. Supports 1–64 questions, 2–26 options each.”
  • The benchmark “is a scaling experiment, not an accuracy evaluation or Jev comparison.”
  • In the generated table itself: “Requested decisions/s is workload count divided by elapsed time, NOT completed throughput when generation fails.”

That last one is a rebuttal the author wrote against their own headline number. The comparison path is MLX-VLM’s ordinary generate() emitting JSON, and it failed the declared schema at several sizes; rather than repairing the output or dropping the row, the report keeps it, marks the failed runs, prints an example of the bad output and says the timing is not completion latency. The author’s own recorded figure for the thing being demonstrated — independent versus shared scoring at 64 decisions on an M4 — is 37.30 s against 2.40 s, medians, their machine, their weights.

The same instinct shows up in the game demos. The Breakout section opens by saying the task “exposes the limits of Qwen3.5-0.8B-4bit in our non-thinking, direct-scoring setup”, explains that asking it to follow the ball was unreliable, and then describes the three simplifications that made it work — bigger ball, five numbered regions, one question about which region. A demo that explains how it was made easy enough to pass is more useful than one that does not.

THIRD_PARTY.md reviews every borrowed idea with the revision it was reviewed at and the licence evidence, including a model card whose repository has no licence file. The dependencies are in a 57-line lock file.

What is unverified

Everything that needs the accelerator: the inference path itself, the browser UI, the three video demos, the latency figures, and whether any of it works on a Mac at all. There is no CI, so nothing runs those paths automatically either. If you have Apple Silicon, uv pip install -r requirements-lock.txt and jev-visual-download are the two commands between you and finding out; if you do not, this repository is a very well-documented read.

Verdict

The clearest piece of work in this directory on checking your own results. Two files ship the evidence for the claims the README makes, both regenerate exactly, and the project tells you what it has not verified before you ask.

Run it only if you have an Apple Silicon Mac. Read it regardless — particularly benchmarks/report.py, which is forty lines and whose docstring is the whole ethic: “Render recorded measurements; never replace failures with repaired answers.”

For the same scoring trick on text with weights of its own, see decider and reflex; for the project it credits, SemIf (formerly OpenJev).

See how it compares with other tools in Best Jev tools, tested hands-on.

Review updated Sep 22, 2026. Numbers quoted from the project are its author's own; we don't publish our own measurements of Jev.

More in Open Models & Reproductions

Laya

★ 27k▲ 12k

NandhaKishorM/laya

Non-autoregressive decision engine over 100+ languages: three checkpoints and a router that detects the script and dispatches per request. Its benchmarks end with a limits section naming the datasets it does not generalise to and the headline figure that came from a training split.

PythonReviewed

kev

★ 7.5k▲ 6.2k

jaredpalmer/kev

Jev-style decision models from 0.5B to 8B, built as LoRA adapters on Qwen and served behind a Jev-compatible /v1/systemone API.

PythonReviewed

SemIf

★ 4.5k▲ 1.9k

TheoLeeCJ/SemIf-OpenJev

Jev-style decisions from a frozen 4B model on a single RTX 3090, with a browser demo. Formerly OpenJev.

PythonReviewed

Get new Jev projects every week

New Jev releases, pricing changes, and the best new projects, once a week. No spam; unsubscribe anytime.

Powered by Buttondown. See our privacy policy.