jev-align
★ 297▲ 49sutro-sh/jev-align
CLI from Sutro that finds the examples a Jev function is least sure about, asks you to label them, and uses GEPA to improve the question.
Benchmark for typed decision models across several suites, with confidence cascades and committees reported separately.
View on GitHub →Hands-on review
Every ranked score recomputes from its published axes, and the datasets match their hashes. One constant needed to check the top axis is not written down.
Good for
Watch out for
Tested Sep 27, 2026 at d06ee95988da · No environment: the published artifacts recomputed from their own records, and the dataset files hashed against the manifest
How we reviewed this: no environment at all. This is a benchmark, so the useful check is whether its published figures follow from the records it ships — we recomputed them, and hashed its datasets against its own manifest. We ran no model and made no Jev calls. Every figure quoted below is JevBench’s own measurement, not ours.
Benchmark Heaven’s leaderboard for Jev-class decision models: state plus a bounded rubric in, a typed answer with probabilities out. Ninety-four systems, ninety ranked, scored on four axes — Intelligence, Calibration, Speed and Cost — combined into one number. It says in its second paragraph that it is Benchmark Heaven’s own benchmark, not affiliated with or endorsed by TypeSafe, whose model is one of the systems measured.
It matters to this directory beyond its own merits: it is the ruler other projects cite. When JevK5’s README says “second of 76 systems and first among open entrants”, this is the board it means.
Every ranked system’s headline score, exactly. The method states the composite as an equal-weight harmonic mean of the four axes, with a squared penalty when Intelligence, Speed or Cost falls below 50. We implemented that sentence and ran it over the artifact:
jevbench_score recomputed from axes: 90 match, 0 differ, 4 skipped
The four skipped are the unranked partials. Ranks follow the scores in order, and the top five in the README are the top five in the file.
The datasets match their hashes. datasets/manifest.json publishes a SHA-256 per split; the three public files match theirs and their item counts:
original.jsonl 72/72 items sha256 MATCH
easy.jsonl 48/48 items sha256 MATCH
hard.jsonl 111/111 items sha256 MATCH
The manifest also records frozen_at and, pointedly, labels_changed_after_inference: false.
The protocol is designed against the obvious failure. There is a sealed split of 308 items alongside the public ones, and the Intelligence axis subtracts a penalty for the gap between a system’s public and sealed accuracy once that gap passes 25 points. The hard tier was written half by one frontier model and half by another, each item reviewed blind and then against its gold by the other, frozen and hashed before any measured system saw it.
The Intelligence axis is published as a formula — 0.8 × a chance-corrected score on the frozen items, plus 0.2 × a chance-corrected sealed score, then multiplied by the gap penalty — and the file gives the sealed chance level (0.293) but not the chance level for the frozen items. Without it the axis cannot be checked at all.
Solving for it across the board, a single value of 0.319 reproduces the published Intelligence within 0.5 points for 83 of 90 ranked systems, mean error 0.35. So the axis is sound and the constant is simply missing from the artifact; publishing it, as the sealed chance already is, would make the whole axis checkable by anyone.
The seven that do not fit are mostly rerankers and label-only systems — mirror, bge-reranker-v2-m3, mxbai-rerank-base-v2, gte-reranker-modernbert-base, simplejev-qwen3.5-0.8b — with errors from 0.9 to 5.7 points. The method describes special handling for label-only systems on the Calibration axis (“contribute zero”), so a different treatment on Intelligence is the likely explanation rather than an error, but the artifact does not say which rule applied to which row.
Nothing runs, so nothing leaves. Each system’s row carries the endpoint condition it was measured under — production API, self-hosted, evaluator-owned GPU pod — and self-hosted or demo endpoints have their latency adjusted, which is stated in the method rather than buried.
Active and versioned with unusual discipline: a changelog, per-release notes, per-version result files kept side by side, superseded rows retained, and a revision_log. Prices behind the Cost axis are dated and corrected in place with a note when a tariff changes.
The most checkable leaderboard in this space, and now the one we will point at when a project cites a rank. Its arithmetic redoes itself from the files it ships, its datasets are what it says they are, and it was built to punish tuning to the public set.
Two things to keep in mind when you read a rank. Publish the frozen-items chance level and the Intelligence axis becomes verifiable too — it is one number away. And every figure on the board, including the ones about commercial models, is Benchmark Heaven’s own measurement under its own protocol; cite it that way, as its README does.
For the projects whose claims lean on it, see JevK5 and decider; for a calibration check you run on your own data instead, jev-calibrate.
See how it compares with other tools in Best Jev tools, tested hands-on.
Review updated Sep 27, 2026. Numbers quoted from the project are its author's own; we don't publish our own measurements of Jev.
sutro-sh/jev-align
CLI from Sutro that finds the examples a Jev function is least sure about, asks you to label them, and uses GEPA to improve the question.
openlayer-ai/jevals
Agent evals and guardrails as typed questions instead of an LLM judge, packing every eval for a trace into one request. From Openlayer, with a mock backend so the whole library runs without a key.
danielgshea/jev-as-a-judge
Uses Jev through langchain-typesafe as the judge in an eval suite, asking typed quality questions instead of asking a larger model to grade. No licence file.
New Jev releases, pricing changes, and the best new projects, once a week. No spam; unsubscribe anytime.
Powered by Buttondown. See our privacy policy.