Skip to content
MrJev

jev-calibrate

Checks a Jev question against your own labelled examples and grades it: act on it, only sort by it, or rewrite it. Refuses to grade a question whose classes have too few examples, however good the numbers look.

View on GitHub →

Hands-on review

Checks a Jev question against your own labels and grades it. We gave it a perfect score on three examples and it declined to bless the question.

Good for

  • A verdict ladder that separates 'act on this' from 'only sort by it'
  • Refusing to grade a question that has too few labelled examples
  • Every metric we recomputed by hand matched, including AUC and Brier

Watch out for

  • It cannot tell you whether your labels are right, only whether the model agrees
  • A `--split all` run marks held-out examples as seen, and says so
  • Thirty-one stars, one author, version 0.1.11

Tested Sep 22, 2026 at 28bc62065b95 · node:24 in Docker, its own suite, then the real CLI against a stand-in whose every probability we chose, with the metrics recomputed by hand

How we reviewed this: we ran its suite, then drove the real CLI against a stand-in whose probability for every example we chose ourselves — so we knew the confusion matrix before the tool did — and recomputed precision, recall, AUC and Brier by hand. Then we built three projects designed to earn three different verdicts. We made no Jev calls.

What it is

You write Jev questions. This tells you whether they work.

The README opens with the failure it exists for: on its own shipped example, a first draft of a frustration question got 18 of 26 labelled messages right, and four of its eight wrong answers came with a confidence of 0.94 or more. That is the whole argument. A question that looks right on the inputs you tried by hand can be confidently wrong on the ones you did not.

The verdict ladder

Every question ends in one of five verdicts, and the distinctions are the valuable part:

verdict what it means
gate act on the answer
gate-above-confidence act on the confident ones, review the rest
ranker the order is right, the cut is wrong — sort by it, do not threshold it
unusable rewrite the question or drop it
too-few-examples label more before trusting any number above

Most tools in this directory would have stopped at “good” and “bad”. The ranker rung is the one that matters in practice: a question can separate your classes perfectly and still be useless at the threshold you picked.

We chose the answers, then checked its arithmetic

We built a twenty-example project and a stand-in that returned a probability we fixed per example. Then we compared what it printed against what we computed independently:

metric our hand computation jev-calibrate
precision at 0.50 0.82 0.82
recall at 0.50 0.90 0.90
false positives 2 / 10 2 / 10
AUC 0.98 0.98
Brier 0.07 0.07
suggested threshold 0.75 → precision 1.00, recall 0.90 0.75 → precision 1.00, recall 0.90

Every field, and it named the three examples that missed — exactly the three our model predicts.

Worth recording how that went: our first run disagreed, and the tool was right. Our stand-in matched states by substring, so positive example 1 also matched positive example 10, and two examples got the wrong probability. The tool had faithfully reported what our broken harness actually returned. We have made this class of mistake twice before this month; the lesson keeps being the same one.

The verdict we wanted to see

The strongest thing here is what it does with a question that looks perfect:

q  [noul]  threshold 0.50
  verdict    too-few-examples
             3 positive and 10 negative examples; 5 of each are needed
             before the numbers mean much
  AUC        1.00
  at 0.50    precision 1.00, recall 1.00, false positives 0/10
  Brier      0.00

Perfect separation. Perfect precision, perfect recall, AUC 1.00, Brier 0.00 — and the verdict is not gate. It prints the flattering numbers and then declines to let you act on them, because three positive examples cannot support any of it.

That is the single behaviour we would most like to see copied across this directory. Plenty of projects here will show you an accuracy figure. This one shows you the figure and tells you it does not mean what you want it to.

For contrast, the two ends of the ladder behaved as designed: clean separation over twenty examples produced gate, and a question with no signal produced unusable — AUC 0.55 is below 0.85 — while still printing every number so you can see why.

When the provider misbehaves

Every case is the real CLI against a stand-in we controlled:

what the provider did what happened
HTTP 500 four attempts, then not checked per example, exit 2
a non-JSON body response is not JSON, exit 2
a 200 with an empty answers object answer is not an object, exit 2
a probability of 1.7 noul answer has no probability, exit 2
connection refused exit 2
no API key set TYPESAFE_API_KEY or OPENROUTER_API_KEY, exit 2

And the part that matters for a measuring instrument — with every request failing it reported 0 examples checked, 20 not checked and printed no verdict, no AUC, no accuracy. A calibration tool that quietly scored the examples that happened to succeed would be worse than no tool, and this one does not.

--require gate exits 1 and names the question that fell short, which makes it a CI gate. --runs 3 repeats every request, averages, and reports the spread and whether any verdict changed.

lint needs no key at all and catches the mistakes you make before you spend anything: a question no example carries a label for, a class with too few examples in a split.

Things to know

It grades agreement with your labels. If your labels are wrong, it will confidently tell you your question is good — it is a calibration tool, not an oracle, and nothing here pretends otherwise.

--split all is available for looking around and warns, every time, that it mixes tune and holdout and that the held-out examples it judges are recorded as seen. That is the right way to offer a footgun.

MIT, Node 20+, no runtime dependencies, TypeScript with erasable syntax. 83 tests pass and tsc --noEmit is clean; one of the tests checks that the version matches across the source, package.json and the lockfile. Thirty-one stars, one author, version 0.1.11.

Verdict

The most disciplined measuring instrument in this directory. Every number we could check was right, and more importantly it refuses to produce numbers it cannot stand behind — from a failed run, or from three examples.

If you are shipping a Jev question that gates anything, run this against labelled data first. And read the verdict table even if you never install it: ranker is a category most people building on confidence scores do not know they need.

For the eval side of the same argument, see jevals; for what publishing an honest benchmark looks like, Rizzo Flow and DocJev.

See how it compares with other tools in Best Jev tools, tested hands-on.

Review updated Sep 22, 2026. Numbers quoted from the project are its author's own; we don't publish our own measurements of Jev.

More in Evaluation & Benchmarks

jev-align

★ 297▲ 49

sutro-sh/jev-align

CLI from Sutro that finds the examples a Jev function is least sure about, asks you to label them, and uses GEPA to improve the question.

PythonReviewed

JevBench

★ 152▲ 120

fstandhartinger/jevbench

Benchmark for typed decision models across several suites, with confidence cascades and committees reported separately.

PythonReviewed

jevals

★ 97▲ 39

openlayer-ai/jevals

Agent evals and guardrails as typed questions instead of an LLM judge, packing every eval for a trace into one request. From Openlayer, with a mock backend so the whole library runs without a key.

PythonReviewed

Get new Jev projects every week

New Jev releases, pricing changes, and the best new projects, once a week. No spam; unsubscribe anytime.

Powered by Buttondown. See our privacy policy.