jev-align
★ 297▲ 49sutro-sh/jev-align
CLI from Sutro that finds the examples a Jev function is least sure about, asks you to label them, and uses GEPA to improve the question.
Checks a Jev question against your own labelled examples and grades it: act on it, only sort by it, or rewrite it. Refuses to grade a question whose classes have too few examples, however good the numbers look.
View on GitHub →Hands-on review
Checks a Jev question against your own labels and grades it. We gave it a perfect score on three examples and it declined to bless the question.
Good for
Watch out for
Tested Sep 22, 2026 at 28bc62065b95 · node:24 in Docker, its own suite, then the real CLI against a stand-in whose every probability we chose, with the metrics recomputed by hand
How we reviewed this: we ran its suite, then drove the real CLI against a stand-in whose probability for every example we chose ourselves — so we knew the confusion matrix before the tool did — and recomputed precision, recall, AUC and Brier by hand. Then we built three projects designed to earn three different verdicts. We made no Jev calls.
You write Jev questions. This tells you whether they work.
The README opens with the failure it exists for: on its own shipped example, a first draft of a frustration question got 18 of 26 labelled messages right, and four of its eight wrong answers came with a confidence of 0.94 or more. That is the whole argument. A question that looks right on the inputs you tried by hand can be confidently wrong on the ones you did not.
Every question ends in one of five verdicts, and the distinctions are the valuable part:
| verdict | what it means |
|---|---|
gate |
act on the answer |
gate-above-confidence |
act on the confident ones, review the rest |
ranker |
the order is right, the cut is wrong — sort by it, do not threshold it |
unusable |
rewrite the question or drop it |
too-few-examples |
label more before trusting any number above |
Most tools in this directory would have stopped at “good” and “bad”. The ranker rung is the one that matters in practice: a question can separate your classes perfectly and still be useless at the threshold you picked.
We built a twenty-example project and a stand-in that returned a probability we fixed per example. Then we compared what it printed against what we computed independently:
| metric | our hand computation | jev-calibrate |
|---|---|---|
| precision at 0.50 | 0.82 | 0.82 |
| recall at 0.50 | 0.90 | 0.90 |
| false positives | 2 / 10 | 2 / 10 |
| AUC | 0.98 | 0.98 |
| Brier | 0.07 | 0.07 |
| suggested threshold | 0.75 → precision 1.00, recall 0.90 | 0.75 → precision 1.00, recall 0.90 |
Every field, and it named the three examples that missed — exactly the three our model predicts.
Worth recording how that went: our first run disagreed, and the tool was right. Our stand-in matched states by substring, so positive example 1 also matched positive example 10, and two examples got the wrong probability. The tool had faithfully reported what our broken harness actually returned. We have made this class of mistake twice before this month; the lesson keeps being the same one.
The strongest thing here is what it does with a question that looks perfect:
q [noul] threshold 0.50
verdict too-few-examples
3 positive and 10 negative examples; 5 of each are needed
before the numbers mean much
AUC 1.00
at 0.50 precision 1.00, recall 1.00, false positives 0/10
Brier 0.00
Perfect separation. Perfect precision, perfect recall, AUC 1.00, Brier 0.00 — and the verdict is not gate. It prints the flattering numbers and then declines to let you act on them, because three positive examples cannot support any of it.
That is the single behaviour we would most like to see copied across this directory. Plenty of projects here will show you an accuracy figure. This one shows you the figure and tells you it does not mean what you want it to.
For contrast, the two ends of the ladder behaved as designed: clean separation over twenty examples produced gate, and a question with no signal produced unusable — AUC 0.55 is below 0.85 — while still printing every number so you can see why.
Every case is the real CLI against a stand-in we controlled:
| what the provider did | what happened |
|---|---|
| HTTP 500 | four attempts, then not checked per example, exit 2 |
| a non-JSON body | response is not JSON, exit 2 |
a 200 with an empty answers object |
answer is not an object, exit 2 |
| a probability of 1.7 | noul answer has no probability, exit 2 |
| connection refused | exit 2 |
| no API key | set TYPESAFE_API_KEY or OPENROUTER_API_KEY, exit 2 |
And the part that matters for a measuring instrument — with every request failing it reported 0 examples checked, 20 not checked and printed no verdict, no AUC, no accuracy. A calibration tool that quietly scored the examples that happened to succeed would be worse than no tool, and this one does not.
--require gate exits 1 and names the question that fell short, which makes it a CI gate. --runs 3 repeats every request, averages, and reports the spread and whether any verdict changed.
lint needs no key at all and catches the mistakes you make before you spend anything: a question no example carries a label for, a class with too few examples in a split.
It grades agreement with your labels. If your labels are wrong, it will confidently tell you your question is good — it is a calibration tool, not an oracle, and nothing here pretends otherwise.
--split all is available for looking around and warns, every time, that it mixes tune and holdout and that the held-out examples it judges are recorded as seen. That is the right way to offer a footgun.
MIT, Node 20+, no runtime dependencies, TypeScript with erasable syntax. 83 tests pass and tsc --noEmit is clean; one of the tests checks that the version matches across the source, package.json and the lockfile. Thirty-one stars, one author, version 0.1.11.
The most disciplined measuring instrument in this directory. Every number we could check was right, and more importantly it refuses to produce numbers it cannot stand behind — from a failed run, or from three examples.
If you are shipping a Jev question that gates anything, run this against labelled data first. And read the verdict table even if you never install it: ranker is a category most people building on confidence scores do not know they need.
For the eval side of the same argument, see jevals; for what publishing an honest benchmark looks like, Rizzo Flow and DocJev.
See how it compares with other tools in Best Jev tools, tested hands-on.
Review updated Sep 22, 2026. Numbers quoted from the project are its author's own; we don't publish our own measurements of Jev.
sutro-sh/jev-align
CLI from Sutro that finds the examples a Jev function is least sure about, asks you to label them, and uses GEPA to improve the question.
fstandhartinger/jevbench
Benchmark for typed decision models across several suites, with confidence cascades and committees reported separately.
openlayer-ai/jevals
Agent evals and guardrails as typed questions instead of an LLM judge, packing every eval for a trace into one request. From Openlayer, with a mock backend so the whole library runs without a key.
New Jev releases, pricing changes, and the best new projects, once a week. No spam; unsubscribe anytime.
Powered by Buttondown. See our privacy policy.