jev-align
★ 297▲ 49sutro-sh/jev-align
CLI from Sutro that finds the examples a Jev function is least sure about, asks you to label them, and uses GEPA to improve the question.
Agent evals and guardrails as typed questions instead of an LLM judge, packing every eval for a trace into one request. From Openlayer, with a mock backend so the whole library runs without a key.
View on GitHub →Hands-on review
Agent evals as typed questions in one request instead of an LLM judge. Installing its PII extra stops phone numbers and SSNs being redacted.
Good for
Watch out for
Tested Sep 24, 2026 at 0457836c5696 · python:3.12-slim in Docker, its own mock backend counting requests; the suite run both the way CI installs it and with the optional extras
How we reviewed this: we ran its suite exactly as CI does and again with the optional extras, drove evaluate() against its own mock backend while counting requests, and compared PII detection with and without the pii extra on the same inputs in the same process. We made no Jev calls — the mock backend and --backend mock bench made that unnecessary, which is worth saying on its own.
An eval library from Openlayer that replaces the LLM judge with typed questions. You hand it a trace — the messages you sent, the tool schemas you gave — and a list of evals, and it turns all of them into one set of questions for a decision model:
r = evaluate(
{"messages": messages, "tools": tools},
[ToolChoice(), UsedToolResult(), Grounded(), StayedInScope(),
AnswerRelevancy(), Completeness(), IndirectInjection(), PHI()],
)
The argument in the README is a good one and it is about economics rather than accuracy: an LLM judge costs six to eleven round trips for four Ragas metrics, so teams sample 1% nightly and the results never reach the request path. Questions that come back in one forward pass can run on every trace.
The division of labour is the right one, too. Code does what code is good at — splitting sentences, matching tool calls, regexes for secrets, Presidio for entities — and typed questions handle only the judgment calls.
We counted at the backend rather than trusting the usage line. Seven evals, one trace:
HTTP requests for 7 evals: 1
call 1: 6 questions -> ['grounded.c0', 'indirect_injection.q', 'pii.personal',
'stayed_in_scope.q', 'tool_choice.q', 'used_tool_result.q']
One request, six questions, the seventh eval resolved in code without asking. That is the design working.
Everything CI runs is green: ruff check, ruff format --check, 56 tests passing, jevals validate on the shipped YAML evals, and jevals --backend mock bench. For a library at 0.1.4 that is a solid state.
pyproject.toml offers a pii extra, which brings in Presidio, and it is the one a team that cares about PII will install. Adding it turns the project’s own suite red:
$ uv sync --extra dev --extra pii && uv run pytest -q
FAILED tests/test_evals.py::test_pii_redact_and_not_personal
FAILED tests/test_runner.py::test_packs_all_evals_into_one_request
FAILED tests/test_runner.py::test_skipped_and_errors_are_reported
3 failed, 53 passed
CI runs uv sync --extra dev, so it has never seen this.
Same inputs, same call, the only difference being whether the extra is installed:
| input | --extra dev |
--extra dev --extra pii |
|---|---|---|
Contact support@acme.com or call 555-123-4567. |
<EMAIL_ADDRESS>, <PHONE_NUMBER> |
<EMAIL_ADDRESS>, 555-123-4567 |
Reach me on +1 (415) 555-0199 or at bob.smith@acme.co.uk |
both redacted | phone left in the clear |
card 4111 1111 1111 1111, ssn 123-45-6789 |
<CREDIT_CARD>, <US_SSN> |
<CREDIT_CARD>, 123-45-6789 |
security/_detect.py replaces the regex scan with Presidio’s rather than combining them:
ents = _presidio_scan(text, None, threshold) if use_presidio else None
if ents is None:
ents = _scan(text, _PII_PATTERNS)
else: # presidio lacks BR CPF and checksum-validated cards in some configs; add ours
ents += _scan(text, [p for p in _PII_PATTERNS if p[0] in ("BR_CPF",)])
Only BR_CPF is added back. Everything else Presidio misses is dropped — and detect_pii(..., use_presidio=False) finds all three in the same process, so the patterns are there and simply are not consulted.
The cause is a threshold boundary. Presidio scores that US phone at exactly 0.40 against jevals’ default cutoff of 0.5, and does not report the SSN at all:
EMAIL_ADDRESS score=1.00 'support@acme.com'
URL score=0.50 'acme.com'
PHONE_NUMBER score=0.40 '555-123-4567'
Lowering the cutoff is not the fix — 0.4 admits URL and still misses the SSN. Unioning the two scans is, and _dedupe already exists for the overlap that creates.
This is the finding that matters, because PII(action="redact") is a control rather than a metric. A control that gets weaker when you install the package named after it is the wrong way round, and nothing tells you: the output looks redacted, because the email still is.
The same seven-eval trace, with the extra installed:
HTTP requests for 7 evals: 2
call 1: 6 questions -> [... 'phi.phi' ...]
call 2: 1 questions -> ['pii.personal']
r.usage.requests correctly reports 2, so the library is not misreporting — it is the README’s first sentence that stops being true for the documented install. On the request path, where this is meant to run, a second round trip is the difference the whole design exists to avoid.
We reported both, with the two-line union and a suggestion to put --all-extras on one CI matrix entry.
a15f28f (2026-09-23) took the union — ents += _scan(text, _PII_PATTERNS) on the Presidio path, _dedupe on the way out — and went further than we suggested by restricting Presidio to a curated entity list, which is what keeps URL score=0.50 from being admitted when the cutoff is lowered. We installed the documented extras and re-ran our own canaries:
presidio available: True
Contact support@acme.com or call 555-123-4567. EMAIL_ADDRESS, PHONE_NUMBER (regex-only: the same)
Reach me on +1 (415) 555-0199 or at bob.smith@… EMAIL_ADDRESS, PHONE_NUMBER (regex-only: the same)
card 4111 1111 1111 1111, ssn 123-45-6789 CREDIT_CARD, US_SSN (regex-only: the same)
The extra is additive now, which is what installing it was supposed to mean. uv sync --extra dev --extra pii && uv run pytest -q is 58 passed where it was 3 failed, 53 passed. And the seven-eval trace that had become two requests is one again, counted at the mock backend with Presidio installed:
HTTP requests for 7 evals: 1 | usage.requests: 1
One thing did not change: .github/workflows/ci.yml still runs uv sync --extra dev, so the path that broke is covered by tests that CI does not install the dependencies for.
Everything in the trace. The evals send the user’s message, the assistant’s answer, the tool schemas and the tool results as state, because that is what they are judging. PII detection happens locally first, but the pii.personal question sends the detected entity to the model to ask whether it is a real person’s data or a support address — so a redaction decision involves a round trip carrying the thing being redacted.
That is defensible and is what makes the eval better than a regex alone. It is also worth knowing before you point this at production traffic, and the README does not spell it out.
The best-argued eval library we have looked at, and the only one whose central performance claim we could check by counting rather than believing. The mock backend and the bench command mean you can evaluate the whole thing without spending anything, which more projects should copy.
The pii extra is safe to install as of a15f28f; below that, pass use_presidio=False and keep the regex path you are already getting. Either way, read r.usage.requests rather than the README if the one-request property is what you are buying, and pin the version you audited — this is a control, and it changed behaviour once already.
For guardrails at the gate rather than the eval, see hermes-jev-approvals and jev-use.
See how it compares with other tools in Best Jev tools, tested hands-on.
Review updated Sep 24, 2026. Numbers quoted from the project are its author's own; we don't publish our own measurements of Jev.
sutro-sh/jev-align
CLI from Sutro that finds the examples a Jev function is least sure about, asks you to label them, and uses GEPA to improve the question.
fstandhartinger/jevbench
Benchmark for typed decision models across several suites, with confidence cascades and committees reported separately.
danielgshea/jev-as-a-judge
Uses Jev through langchain-typesafe as the judge in an eval suite, asking typed quality questions instead of asking a larger model to grade. No licence file.
New Jev releases, pricing changes, and the best new projects, once a week. No spam; unsubscribe anytime.
Powered by Buttondown. See our privacy policy.