How we reviewed this: we verified its checksums, recomputed every published metric from the per-row records committed beside it, ran its tests, then downloaded its runtime and weights and served it on four CPU cores — driving the Jev-compatible endpoint with the official typesafe-sdk and sending it a round of malformed bodies. We made no Jev calls.
What it is
A local server that turns Spark-X2.5-4B into typed decisions by scoring option tokens, with two front doors: its own /v1/decisions, and /v1/systemone speaking Jev’s shape.
The two surfaces are deliberately different, which is the interesting decision. Its native schema uses boolean rather than noul and adds a numeric type Jev does not have; the compatibility layer translates. It converges on the wire without pretending its own model is Jev’s — and that turns out to run right through the project.
The official SDK, one base URL
model: rizzo-spark-x2.5-4b-q8_0
is_urgent: type='noul' noul=0.99998
department: type='choice' choice='billing' confidence=0.99936
probabilities={'billing': 0.99968, 'technical': 0.00032}
severity: type='score' score=1.9917 confidence=0.98815
legend={0: 'minor', 1: 'annoying', 2: 'serious'}
probabilities={0: 0.00035, 1: 0.00755, 2: 0.99210}
All three types, the legend carried on the score answer, probabilities on both.
GET /v1/models works, which is where most of this category falls over — and what it returns is the best handling of the jev-latest alias we have seen:
jev-latest — “Compatibility alias: answered by rizzo-spark-x2.5-4b-q8_0, not by TypeSafe Jev.”
That sentence travels into any client that enumerates models. Compare OpenThai-SystemOne, which has no /v1/models at all and echoes back whatever model name the caller sent.
/health is the same instinct: it reports the model source, the requested revision, the GGUF repository and revision, the SHA-256 of the weights file actually loaded, and the precision, device, backend and runtime. A running server tells you exactly what is answering.
The artifacts are the review
We check published numbers against committed data as a matter of routine. This is the first project where the checking ran out of things to catch.
results/SHA256SUMS: 36 of 36 verify. And the set includes the source that produced the run — nine .py files snapshotted next to the results and checksummed with them.
Every dataset_sha256 matches a committed dataset. Each result file names its fixture set by hash; we recomputed those hashes the project’s own way and all twelve matched, with the -v1 results correctly pointing at the -v1 fixtures, both of which are still in the tree so the older runs stay checkable.
Every published metric recomputes from its own rows. Across all twelve result files — coverage, status accuracy, categorical accuracy and the 10-bin ECE — to within 1e-9:
| result file |
coverage |
status acc |
accuracy |
ECE |
spark-bf16-validation/perturbations |
0.5556 |
0.6667 |
0.6667 |
0.3710 |
spark-bf16-v2-validation/perturbations |
0.8889 |
0.7778 |
0.7778 |
0.1744 |
spark-bf16-final/perturbations |
0.7778 |
0.8889 |
0.8889 |
0.0657 |
spark-q8-final/perturbations |
0.8889 |
1.0000 |
1.0000 |
0.0413 |
Published and recomputed were identical in every cell, so the table above is both.
Note what those rows are. The first is a bad run — 56% coverage, an ECE of 0.37 — and it is committed next to the good ones, under a name that says which stage it was. The progression from validation through v2-validation to final is the work, left in the open. We have spent today reading projects that publish only the configuration where they win; this one shipped the ones where it lost, and warmup_excluded: true sits in every summary.
The perturbation fixtures test the right things, too: irrelevant text added to the state (paid-irrelevant), and options reordered (login-reordered, refund-reordered) — the property AnyJev is built around, here as a shipped benchmark rather than a README claim.
One more small honesty: every answer carries probability_status: "uncalibrated_conditional_option_scores". It declines to call its own numbers calibrated.
Where it pushes back
We pointed --model at a Qwen GGUF we had lying around:
rizzo: Only the Spark2.5 architecture is supported, not qwen2
Refused rather than served as nonsense. Which is why this review cost a 4.1 GB download — there is no substituting a smaller model, and that is the right trade.
What we found
Two things, both small.
NaN in the request body returned HTTP 500, where every other malformed input we sent returned a clean 422 with a detail array — no questions, an unknown question type, a one-option choice, an unknown top-level field, a bare number as state. A 500 tells a caller “my fault, retry” when the answer is “your body is malformed”.
Fixed two days later, and the diagnosis went past ours. It was not the parser: json.loads accepts NaN, validation rejected it correctly, and the 422 then died on the way out, because the validation error echoes the offending input and Starlette serialises the response with json.dumps and no allow_nan=False. A nested value had a second route in, through canonical()’s own ValueError landing in the error body. 6f40a48 adds a RequestValidationError handler that replaces anything json.dumps would refuse — non-finite floats and exception objects alike — with its text, plus a regression test. We re-ran it at bc38251 against their own fake backend:
bare NaN 422 detail=list
nested NaN 422 detail=list "Value error, Out of range float values are not JSON compliant"
nested Infinity 422 detail=list
no questions 422 detail=list
unknown type 422 detail=list
one-option choice 422 detail=list
bare number state 422 detail=list
Their service, decisions and compatibility suites are 23 passed on a machine with no weights and no network.
The second one stands, by the author’s decision. model: "jev-1.13.0" is accepted and answered locally. The response names the real model, which is most of what matters, but a caller who pinned a version did it to hold the answering model still. JevBERT refuses pinned jev-* names for that reason and documents it. Both reported; the author’s answer is that it lets a client written for the hosted API run unchanged against localhost while the response names the real model, that the compatibility notes now spell out the error shapes, and that JevBERT’s opposite choice is noted there.
What it costs
rizzo download pulls a pinned llama.cpp build and the Q8_0 weights: 112 MB and 4.1 GB. On four CPU cores a decision is seconds, not the milliseconds the README’s GPU figures describe — those are the author’s, on an RTX 5060 Ti, and we did not reproduce them.
pytest gives 58 passed, 11 skipped. MIT, a NOTICE, Italian docs, a self-contained playground page with no external calls.
Verdict
The best-evidenced project in this directory, and the gap is not close. Everything it publishes can be checked from what it commits, it committed the runs that made it look worse, and the one piece of Jev-compatibility theatre everyone else performs — a jev-latest alias that quietly implies Jev — it turned into a sentence telling you it is not.
Budget the 4.1 GB and a GPU if you want the latency. Read results/ either way; it is the model for how a project in this space should publish.
For the same API shape on a model you host, see JevBERT and the rest of the open models category.
See how it compares with other tools in Best Jev tools, tested hands-on.
Review updated Sep 24, 2026. Numbers quoted from the project are its author's own; we don't publish our own measurements of Jev.