Skip to content
MrJev

Jev Benchmarks and Evaluation Tools

TypeSafe's published speed, cost, and accuracy numbers come from its own evaluations. The projects here measure Jev independently, check whether its confidence scores are calibrated on real datasets, and help you pick a threshold for your own data.

If you plan to automate decisions based on Jev's confidence, start here.

4 projects

iammrduncan/typesafe-ai-benchmark

Compares LLM structured output with Jev on latency, cost, and judgment quality.

TypeScript

AbdelStark/jev-benchmarks

Probability-aware evaluation: calibration, and how much work can be automated at a fixed error budget.

Python

Janus

★ 2

FirasSX914/Janus

Measures Jev's calibration and confidence-based routing on Banking77 and Web of Science.

Python

jevcal

★ 1

abhixhek/jevcal

Picks the confidence threshold that meets your accuracy target on your own data, and fails CI when a model update breaks it.

Python

Other categories

Get new Jev projects every week

New Jev releases, pricing changes, and the best new projects, once a week. No spam; unsubscribe anytime.

Powered by Buttondown. See our privacy policy.