Jev by @typesafeai is now on OpenRouter, in beta. Jev is a System One model. Instead of generating text, it takes your app's state plus a typed question and returns a typed decision with a probability attached. There is no JSON prompting, parsing layer, and nothing to validate Show more
Jev: one judge call, or twelve dimension scores? I measured both on three tasks
An independent measurement spanning 5,477 test rows and 34.1M input tokens, comparing one direct Jev question per row with 12-14 scored dimensions fitted to local labels.
- Category
- Practices & Patterns
- Published by
- Community
- Author
- —
- Added
- 2026-09-22
Highlights
- The full experiment cost $1.43: 25,174 Jev calls, 5,477 test rows, and 34.1M input tokens extracted in 9 minutes 14 seconds with zero failed calls.
- On Japanese NLI, 14 dimensions with fitted weights reached 0.908 accuracy versus 0.837 for the direct Jev call, a 7.03-point gain.
- A 12-option ledger classification broke the direct call at 0.400 accuracy; dimensions reached 0.911 and stacked n-gram probabilities 0.9695.
- On hard-benign guardrail text the direct question flagged 1.5% versus 37.2% for the dimension model, about 25x more false positives.
- The direct call reported confidence at or above 0.9 on 42% of rows where it was only 72.2% accurate, and four repair attempts all failed.
Watch out
The dimensions are hand-written by the author and output tokens were left unpriced, so the stated cost is input-only.
More like this
From the community
Posts from builders shipping with Jev right now.
Jev lands on OpenRouter
700 leads scored for $0.09
JEV is INSANE. We gave it 700 high-intent leads and personalised outreach messages. In 40 seconds, it predicted how each message would perform, assigned a confidence score and detected lead-message mismatches. All for just $0.09. JEV can also score leads, analyse buying Show more
Beating Gemini Flash Lite on an eval
Ran @typesafeai's Jev against an existing classifier eval that previously used Gemini 2.5 Flash Lite. It won both on quality (saturated the eval) and speed (6x)
Browser Use Ultrafast, powered by Jev
Breaking: Browser Use + Jev = Ultrafast ⚡ Findings flights took 7s and cost only $0.0039 🤯 > new action space every step > DOM state space > small LLM fallback to type (this video is at 1x speed btw) Built a tiny open source browser agent. try it below ↓
A really smart switch statement
hype-free explanation of jev: jev does not replace gpt / claude jev is just a *really* smart switch statement like if 2016 ml classifiers got 2026 levels of intelligence it's a new* type of tool that will make a lot of workloads insanely fast, cheap, and accurate * = and by Show more
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x
When a designer gets Jev
when a designer gets access to Jev



