found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant
The Bitterest Lesson
A TypeSafe blog essay arguing that in ML the order that matters is doing the right task, then data, then compute, then algorithms, using the InstructGPT result as its example.
- Category
- Sites & Guides
- Published by
- TypeSafe AI
- Author
- —
- Added
- 2026-09-22
Highlights
- Extends Sutton's compute-beats-algorithms lesson: doing the right task > data > compute > algorithms.
- Cites InstructGPT, where GPT-2-sized models over 100x smaller than GPT-3 beat it once trained on the right task.
- Says scaling pre-training would need roughly GPT-7 level to beat that baseline and GPT-9 to beat InstructGPT on GPT-3.
- Published September 10, 2026, with footnotes linking to the original version on The Complete Skeptic.
- Argues picking the right task often requires leaving ML to study users, products, and organizations.
Watch out
A persuasive essay rather than new experimental work; the InstructGPT comparison is read from the original paper's figure 31.
More like this
From the community
Posts from builders shipping with Jev right now.
Instant compaction with Jev
A Claude session from 1M to 86K tokens
This is actually insane. This uses @typesafeai Jev model, as a plugin in Claude to review all the un-nesseasary tool calls, and it takes 1s to run! Like, literally, 1 second to take my Claude session from nearly 1M to ... 86K tokens! 😮 Ask your claude to install it and be Show more
found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant
Vercel's fx safety reviewer, 18x faster
We're seeing extraordinary results from @typesafeai. Default mode in 𝚏𝚡 is auto, with a safety reviewer analyzing every command. That reviewer runs on GPT Luna today. Jev is up to 18x faster (p95) *and* more accurate. It's coming to @vercel AI Gateway and likely new default.
We benchmarked fx auto mode (safety) classifier with @typesafeai's Jev. tl;dr: ~5-18x faster and more accurate than 𝚐𝚙𝚝-𝟻.𝟼-𝚕𝚞𝚗𝚊, our current top choice
Jev lands on OpenRouter
Jev by @typesafeai is now on OpenRouter, in beta. Jev is a System One model. Instead of generating text, it takes your app's state plus a typed question and returns a typed decision with a probability attached. There is no JSON prompting, parsing layer, and nothing to validate Show more
700 leads scored for $0.09
JEV is INSANE. We gave it 700 high-intent leads and personalised outreach messages. In 40 seconds, it predicted how each message would perform, assigned a confidence score and detected lead-message mismatches. All for just $0.09. JEV can also score leads, analyse buying Show more
Beating Gemini Flash Lite on an eval
Ran @typesafeai's Jev against an existing classifier eval that previously used Gemini 2.5 Flash Lite. It won both on quality (saturated the eval) and speed (6x)

