iammrduncan/typesafe-ai-benchmark
Compares LLM structured output with Jev on latency, cost, and judgment quality.
TypeSafe's published speed, cost, and accuracy numbers come from its own evaluations. The projects here measure Jev independently, check whether its confidence scores are calibrated on real datasets, and help you pick a threshold for your own data.
If you plan to automate decisions based on Jev's confidence, start here.
iammrduncan/typesafe-ai-benchmark
Compares LLM structured output with Jev on latency, cost, and judgment quality.
AbdelStark/jev-benchmarks
Probability-aware evaluation: calibration, and how much work can be automated at a fixed error budget.
FirasSX914/Janus
Measures Jev's calibration and confidence-based routing on Banking77 and Web of Science.
abhixhek/jevcal
Picks the confidence threshold that meets your accuracy target on your own data, and fails CI when a model update breaks it.
New Jev releases, pricing changes, and the best new projects, once a week. No spam; unsubscribe anytime.
Powered by Buttondown. See our privacy policy.