Skip to content
MrJev

jevcal

Picks the confidence threshold that meets your accuracy target on your own data, and fails CI when a model update breaks it.

View on GitHub →

We haven't reviewed this project hands-on yet. The description above is our own summary of its README. Numbers it reports are the author's, not ours.

More in Evaluation & Benchmarks

iammrduncan/typesafe-ai-benchmark

Compares LLM structured output with Jev on latency, cost, and judgment quality.

TypeScript

AbdelStark/jev-benchmarks

Probability-aware evaluation: calibration, and how much work can be automated at a fixed error budget.

Python

Janus

★ 2

FirasSX914/Janus

Measures Jev's calibration and confidence-based routing on Banking77 and Web of Science.

Python

Get new Jev projects every week

New Jev releases, pricing changes, and the best new projects, once a week. No spam; unsubscribe anytime.

Powered by Buttondown. See our privacy policy.