Skip to content
MrJev

jev-align

CLI from Sutro that finds the examples a Jev function is least sure about, asks you to label them, and uses GEPA to improve the question.

View on GitHub →

Hands-on review

Finds the rows Jev is least certain about, has you label them, and lets GEPA rewrite the question. It never measures the calibration in its name.

Good for

  • Turning a vague question into a written rubric, one labelling session at a time
  • A holdout that is frozen up front and provably never reaches the optimizer
  • An accept/reject gate a human has to press, every round

Watch out for

  • Your rows, your labels and your written rationales go to a second vendor
  • About 290 paid calls per human label on the documented defaults
  • A perfectly correct holdout round can be reported as F1 0.000

Tested Sep 20, 2026 at 49753df924d3 · Python 3.12 in Docker, offline against two local stand-ins; 127 tests, and two full optimisation rounds

How we reviewed this: we ran its 127 tests offline in Docker, then drove two complete optimisation rounds against two local stand-ins — one for the decision model, one for the reflection LLM — with every vendor host pointed at loopback. We made no Jev calls, so nothing here judges Jev’s own accuracy.

What it does

jev-align, by Sutro, is an active-learning loop around a written question. It scores your dataset with Jev, ranks rows by how close to 0.5 the answer came, asks you to label the most ambiguous handful, and then hands your labels to GEPA — a reflective optimizer that asks a second, generative LLM to rewrite the question’s text. It re-scores with the new wording, shows you both, and you accept or reject.

The thing being optimised is three strings: instructions, true_criteria, false_criteria. That’s the whole artifact, and it’s the right artifact — at the end you have a rubric you can read, not a fine-tuned weight file.

Two design decisions deserve credit. The holdout is frozen before anything runs: 20% of row ids are set aside, permanently excluded from sampling, and never passed to GEPA — there is a test named test_holdout_labels_are_added_separately_and_never_sent_to_gepa that asserts exactly that. And no proposal is ever auto-accepted; a better training score still needs a keystroke.

It doesn’t measure calibration

The repository’s one-line description is “Build calibrated AI classifiers”. Searching the entire codebase for calibrat|brier|ece|reliability|isotonic|platt returns three hits, all of them the same declaration:

uncertainty={"binary": "calibrated_probability", ...}

That is an assertion about Jev, made once, never checked. There is no Brier score, no expected calibration error, no reliability curve, no isotonic or Platt rescaling, and no threshold search. The decision threshold is the literal 0.5, twice:

return prediction.probability >= 0.5

What it optimises and reports is F1 on your labels. That is a perfectly good objective — it just isn’t calibration, and the README, to its credit, never says it is. The word only appears in the GitHub description.

What the interface puts at the top, in a progress bar, is “Certainty” — which is 1.0 - mean_ambiguity, the average distance of the pool’s probabilities from 0.5. That measures how confident the model is, not how often it’s right, and nothing in it touches a label. A rewrite that pushes every probability toward 0 or 1 raises this number without improving a single answer, and the UI colours a rising number green. The accept/reject decision is made while looking at it.

A correct holdout round can read as zero

precision = tp / (tp + fp) if tp + fp else 0.0
recall    = tp / (tp + fn) if tp + fn else 0.0
f1 = 2 * precision * recall / (precision + recall) if precision + recall else 0.0

Undefined is treated as zero. That matters here because the holdout batch is one row: max(1, round(batch_size * 0.2)) is 1 at the default batch size of 5. In our run, the single holdout row was a true negative — the model said no, the human said no, completely correct — and the report said:

Held-out F1   0.000 · P 0.000 · R 0.000

That 0.0 is then written into the run’s permanent history. Reported, with the rest of the findings below. The README’s “an optional 20% held-out evaluation set” also reads as 200 rows on a 1,000-row pool; what you actually get is one extra label per round, so ten rounds gives you a ten-example evaluation set. The wizard’s own phrasing — “add 20% extra held-out annotations each round” — is the accurate one.

The holdout is also off by default.

Your labels reach a second vendor

The reflection prompt we captured was 5,383 characters and contained, for each selected example: the row’s full contents, your label, the model’s prediction, and your written rationale, verbatim. Every one of the 24 reflection calls in our run carried them.

That is inherent to how GEPA works, and it isn’t hidden maliciously — but the README describes the reflection model only as “GEPA’s reflection model”, and never says that the data you are labelling, and the reasons you type, go to OpenAI, Anthropic, Google or whichever provider you picked. Nothing is redacted or truncated anywhere in the pipeline.

Two related things: the CLI phones PyPI on every interactive start and offers an upgrade with “Upgrade now” as the default, which the README doesn’t mention; and the labelling card pre-selects the model’s own answer as the default choice — including on holdout rows, which are the ones meant to judge it.

What it costs

One HTTP request per row per candidate. No batching, no cache inside GEPA (deliberately disabled, with a comment explaining why), concurrency 16, and no spend cap.

We measured 221 Jev requests and 24 reflection calls to collect ten human labels, and reproduced the per-round formula exactly. Applying it to the README’s own defaults — a 1,000-row pool, 5 labels a round, a 300-call GEPA budget, ten rounds — gives about 14,550 Jev requests for 50 labels, roughly 290 paid calls per label. Around 69% of that is re-scoring the whole pool each round to refresh the “Certainty” table, which happens whether you accept the proposal or reject it.

The budget is soft: we asked for 40 metric calls and got 55. The UI explains this as “one in-flight step past the budget”; for us it was 37.5%.

One round was worse than merely expensive. The reflection minibatch is chosen once and memoised for the whole run, and skip_perfect_score=True compares against a batch-wide F1 — so when those five fixed rows happen to be all correct, the round spends its entire budget, makes zero reflection calls, and proposes nothing. Our second round burned 40 of 40 calls for “No textual change.”

Finally: the TypeSafe SDK path retries on 429 and 5xx; the Cloudflare and Vercel gateway paths have no retry at all, and a single 429 aborts the round with a message telling you to “check credentials, access, and model availability”. At 16 concurrent requests, that is not a rare event.

Verdict

The loop is the right loop, and the parts that guard against self-deception — a frozen holdout that never reaches the optimizer, a human gate on every change — are better than most. The parts that report on it are not yet: the headline number rewards confidence rather than accuracy, a perfect holdout can print as zero, and the evaluation set is ten examples when it sounds like two hundred.

Worth knowing before you point it at anything sensitive: the rows you label and the reasons you write for them go to a second vendor, unredacted.

180 stars, 21 commits, one contributor, one day of history, and CI that only runs on a release. Treat the star count as interest in the idea.

For the same question without the labelling loop, see jevcal and Janus.

See how it compares with other tools in Best Jev tools, tested hands-on.

Review updated Sep 20, 2026. Numbers quoted from the project are its author's own; we don't publish our own measurements of Jev.

More in Evaluation & Benchmarks

JevBench

★ 152▲ 120

fstandhartinger/jevbench

Benchmark for typed decision models across several suites, with confidence cascades and committees reported separately.

PythonReviewed

jevals

★ 97▲ 39

openlayer-ai/jevals

Agent evals and guardrails as typed questions instead of an LLM judge, packing every eval for a trace into one request. From Openlayer, with a mock backend so the whole library runs without a key.

PythonReviewed

jev-as-a-judge

★ 93▲ 13

danielgshea/jev-as-a-judge

Uses Jev through langchain-typesafe as the judge in an eval suite, asking typed quality questions instead of asking a larger model to grade. No licence file.

Python

Get new Jev projects every week

New Jev releases, pricing changes, and the best new projects, once a week. No spam; unsubscribe anytime.

Powered by Buttondown. See our privacy policy.