How we reviewed this: we ran the suite in python:3.12 in Docker with no network, then patched the single endpoint constant to point at a local stand-in and ran the real router twice — once with a stand-in that could match on keywords, once with one that answered uniformly — to see both outcomes. We reverted and diffed. No Jev calls; we did not reproduce the author’s evaluation.
What it does
An agent with 24 specialist skills has to decide which one, if any, a request needs. This router asks Jev, and its answer space includes two things that aren’t skills: __no_skill__ and __review__.
The shape of the walk is visible from the other end. For a 24-skill catalogue it sent four requests:
- three
choice questions of nine options each, reducing the catalogue in batches;
- one final request combining the finalist
choice with a needs_specialist noul, a needs_review noul, and a fit_i noul for every surviving candidate.
That last part is the design worth copying. The finalists are re-checked one at a time, independently of the choice that produced them, so a winner that no independent check supports doesn’t get routed. The state in every request was just {"request": "…"} — 3,600 input tokens total for our sentence against a 24-skill catalogue.
Both outcomes reproduce:
stand-in matches on keywords → {"outcome": "route", "selected_skill": "accessibility-audit",
"reason": "confident_match", "confidence": 0.92}
stand-in answers uniformly → {"outcome": "review", "selected_skill": null,
"reason": "clarification_required", "confidence": 0.02}
And in an earlier run, where our crude stand-in said every finalist fitted, the router also returned review — contradictory evidence is treated as no answer rather than resolved by picking the top probability. The response carries an evidence array naming the candidates at each stage, so a surprising outcome is traceable.
Seven thresholds govern all this — confidence, winner probability, margin, specialist need, no-skill ceiling, review veto and candidate fit — and the config validates the relationships between them, not just their ranges.
The evaluation protocol is the best thing here
docs/evaluation.md was written as a protocol, and it reads like one:
Run the development split first… Any policy change must be documented, preserving the original complete results and rerunning the development split. Then freeze the policy and run the entire dataset once, reporting development and test separately. Never tune on test failures without explicitly declaring that the test set has been reused. A small synthetic holdout is not evidence of production performance.
The dataset is 72 synthetic requests in five deliberate categories — clear routes, near-neighbours, ordinary no-skill requests, ambiguous tasks and adversarial distractions — split 12 development / 60 test before any live call. The baseline is a documented lexical one (token-set cosine, fixed stop-words, thresholds stated) rather than a strawman.
Then the README declares its own violation of that protocol, in bold, above the numbers:
These are exploratory, reused-data results… The routing questions were revised after inspecting an earlier full run. The original policy scored 64/72 overall and 52/60 on its initially untouched test split. Neither run establishes production accuracy or calibrated probabilities. Typed output guarantees shape, not correctness.
The headline figures — 68 of 72 against 51 of 72 for the lexical baseline, median 1,287 ms — are the author’s, measured on 16 September with jev-1.13.0 pinned, and we did not reproduce them. What we can say is that the results directory keeps both runs, the earlier one included, which is what the protocol demands and what makes the disclosure checkable rather than decorative.
Engineering
101 tests pass in a tenth of a second (one needs git present, because it shells out to git ls-files to assert that no secrets or machine state are among the published files — a test we had not seen before and would steal). CI runs unit tests, ruff and mypy on Python 3.11 and 3.12, and was green on the last commit.
The transport is defensive in the same places as the best client in this batch: a no-redirect opener so a 302 cannot carry the bearer token elsewhere, a key rejected if it contains whitespace or non-ASCII, allow_nan=False on the outgoing body, a 2 MB response cap, and a single hardcoded endpoint. Two unrelated authors arriving at the same four precautions is a useful signal about which ones matter.
MIT, four commits, all on 16 September — the earliest work in this batch, and untouched since.
Verdict
A small, carefully argued demonstration rather than a product: 1,300 lines, no dependencies, and a jev-router route command you can point at your own catalogue. The two ideas worth taking are the independent per-finalist fit check and the protocol document, which is the only one in this directory that says in advance what would count as cheating and then admits to it.
Treat the seven thresholds as the starting values the docs say they are, and calibrate them on your own catalogue before routing anything that matters.
For routing between models rather than skills, see jev-router; for the capability-tree version of the same instinct, JCR.
See how it compares with other tools in Best Jev tools, tested hands-on.
Review updated Sep 21, 2026. Numbers quoted from the project are its author's own; we don't publish our own measurements of Jev.