Ran @typesafeai's Jev against an existing classifier eval that previously used Gemini 2.5 Flash Lite. It won both on quality (saturated the eval) and speed (6x)
Jev Capability Atlas
A bilingual field map of where Jev fits and where it fails, separating the author's raw API suites from cited third-party results and TypeSafe's own claims.
- Category
- Practices & Patterns
- Published by
- Community
- Author
- Zaious
- Added
- 2026-09-22
Highlights
- Each case suite ships real API logs, methodology, and a report rather than a leaderboard, citing other benchmarks instead of re-running them.
- A citation-support suite separates paraphrase support from reversed-meaning overlap to show the judgments are not keyword rules.
- One history suite records a 0.90-confidence wrong answer caused by a typo in an option, corrected when the typo was fixed.
- It flags where calibration broke in cited third-party work, including an emotion task at 0.819 mean confidence and 48% accuracy.
- AGENTS.md defines scan criteria and a result-reporting protocol for agents asked to judge other projects.
Watch out
MIT-licensed for code and original content, while quoted third-party material keeps its own rights. Case studies frame task fit, not population-level accuracy or a reusable threshold.
More like this
From the community
Posts from builders shipping with Jev right now.
Beating Gemini Flash Lite on an eval
Browser Use Ultrafast, powered by Jev
Breaking: Browser Use + Jev = Ultrafast ⚡ Findings flights took 7s and cost only $0.0039 🤯 > new action space every step > DOM state space > small LLM fallback to type (this video is at 1x speed btw) Built a tiny open source browser agent. try it below ↓
A really smart switch statement
hype-free explanation of jev: jev does not replace gpt / claude jev is just a *really* smart switch statement like if 2016 ml classifiers got 2026 levels of intelligence it's a new* type of tool that will make a lot of workloads insanely fast, cheap, and accurate * = and by Show more
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x
When a designer gets Jev
when a designer gets access to Jev
Full Jev video tutorial
Full Jev Tutorial What it is, how you can build with it and what new applications it can unlock → 0:00 Intro → 0:34 Jev explained → 4:06 API setup → 5:59 Demo 1: Voice-controlled browser → 11:33 Demo 2: AI memory → 17:27 Demo 3: YouTube predictor
The case against Jev-scored compaction
This is a terrible compaction strategy that fundamentally doesn't understand how compaction and context management work. Seems like a lot of people are confused so let's break this down. 1. Compaction isn't a filter The role of compaction is to clean up history to keep the Show more
found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant



