Skip to content
JevDirectory.org
CommunityPractices & Patterns10 starsVerified 2026-09-22

jevcal

CLI that fits a per-question confidence threshold to a target accuracy on your own labeled data, verifies it on a held-out split, estimates how much traffic still needs an LLM, and re-checks locked thresholds in CI. It publishes no Jev results of its own.

Category
Practices & Patterns
Published by
Community
Author
abhixhek
Added
2026-09-22
Tagscommunitypythoncalibrationconfidenceevaluation

Highlights

  • Fits a per-question confidence threshold to your accuracy target and verifies it on a held-out split.
  • Reports handled and accepted accuracy, ECE, and the share of traffic that still escalates.
  • Locks thresholds and their evidence in decisions.lock.json, then re-checks them in CI.
  • Publishes no Jev performance numbers, because TypeSafe's customer agreement restricts them.
  • jevcal lint flags negations, counting, dates, compound questions, and overlapping options without an API key.

Quickstart

bash
pip install "git+https://github.com/abhixhek/jevcal"
jevcal demo
open jevcal-demo/report.html

Watch out

MIT-licensed and not affiliated with TypeSafe. Needs Python 3.10+ and a TYPESAFE_API_KEY for real runs (the demo uses a built-in simulator), and thresholds fitted on fewer than about 100 labeled rows should not be trusted.

More like this

An independent calibration study of Jev over three public benchmarks and 900 rule-generated support tickets, publishing every raw Gateway response and the quantization limits of returned probabilities.
Practices & Patterns#community#python#calibration
74GitHub stars
Independent cross-model benchmark for Jev-class decision models, running 534 frozen cases per complete entrant with scoring code and a four-axis score of accuracy, calibration, latency and cost.
Practices & Patterns#community#python#benchmarks
Communityjevbench
An experiment comparing Jev with GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6 as evaluators of five frozen weather-agent runs, measuring pass-or-fail accuracy against human labels plus variance, cost, and latency.
Practices & Patterns#community#python#evaluation
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Beating Gemini Flash Lite on an eval

Browser Use Ultrafast, powered by Jev

A really smart switch statement

hype-free explanation of jev: jev does not replace gpt / claude jev is just a *really* smart switch statement like if 2016 ml classifiers got 2026 levels of intelligence it's a new* type of tool that will make a lot of workloads insanely fast, cheap, and accurate * = and by Show more

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

When a designer gets Jev

Full Jev video tutorial

The case against Jev-scored compaction

This is a terrible compaction strategy that fundamentally doesn't understand how compaction and context management work. Seems like a lot of people are confused so let's break this down. 1. Compaction isn't a filter The role of compaction is to clean up history to keep the Show more

tamara
tamara
@tamarajtran

found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant

Reply