Ran @typesafeai's Jev against an existing classifier eval that previously used Gemini 2.5 Flash Lite. It won both on quality (saturated the eval) and speed (6x)
jevcal
CLI that fits a per-question confidence threshold to a target accuracy on your own labeled data, verifies it on a held-out split, estimates how much traffic still needs an LLM, and re-checks locked thresholds in CI. It publishes no Jev results of its own.
- Category
- Practices & Patterns
- Published by
- Community
- Author
- abhixhek
- Added
- 2026-09-22
Highlights
- Fits a per-question confidence threshold to your accuracy target and verifies it on a held-out split.
- Reports handled and accepted accuracy, ECE, and the share of traffic that still escalates.
- Locks thresholds and their evidence in decisions.lock.json, then re-checks them in CI.
- Publishes no Jev performance numbers, because TypeSafe's customer agreement restricts them.
- jevcal lint flags negations, counting, dates, compound questions, and overlapping options without an API key.
Quickstart
pip install "git+https://github.com/abhixhek/jevcal"
jevcal demo
open jevcal-demo/report.htmlWatch out
MIT-licensed and not affiliated with TypeSafe. Needs Python 3.10+ and a TYPESAFE_API_KEY for real runs (the demo uses a built-in simulator), and thresholds fitted on fewer than about 100 labeled rows should not be trusted.
More like this
From the community
Posts from builders shipping with Jev right now.
Beating Gemini Flash Lite on an eval
Browser Use Ultrafast, powered by Jev
Breaking: Browser Use + Jev = Ultrafast ⚡ Findings flights took 7s and cost only $0.0039 🤯 > new action space every step > DOM state space > small LLM fallback to type (this video is at 1x speed btw) Built a tiny open source browser agent. try it below ↓
A really smart switch statement
hype-free explanation of jev: jev does not replace gpt / claude jev is just a *really* smart switch statement like if 2016 ml classifiers got 2026 levels of intelligence it's a new* type of tool that will make a lot of workloads insanely fast, cheap, and accurate * = and by Show more
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x
When a designer gets Jev
when a designer gets access to Jev
Full Jev video tutorial
Full Jev Tutorial What it is, how you can build with it and what new applications it can unlock → 0:00 Intro → 0:34 Jev explained → 4:06 API setup → 5:59 Demo 1: Voice-controlled browser → 11:33 Demo 2: AI memory → 17:27 Demo 3: YouTube predictor
The case against Jev-scored compaction
This is a terrible compaction strategy that fundamentally doesn't understand how compaction and context management work. Seems like a lot of people are confused so let's break this down. 1. Compaction isn't a filter The role of compaction is to clean up history to keep the Show more
found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant



