I just open sourced Foreman: a software factory foreman built with @typesafeai's Jev. Coding agents work the factory floor. Foreman watches them, continuously assessing progress, completeness, tests, drift, and verification, and intervenes when needed. GitHub: Show more
Cookbook
An end-to-end recipe in the TypeSafe docs that shows a real problem, with its dataset, latencies, and measured numbers.
- Category
- Ecosystem
- Also known as
- —
- Related terms
- 3
- Directory entries
- 21
- Docs
- docs.typesafe.ai
- Added
- 2026-09-24
Definition
The official cookbooks run from a few questions to full pipelines: parallel questions, re-ranking, line-by-line search, function calling, skill suggestion, entity alignment, date extraction, guardrails, and more. Each publishes the numbers behind its claims rather than only the code.
They are the best source of thresholds and patterns because the caveats come with them — tuned datasets, historical cost sweeps, and the failure cases each recipe leaves open.
Related terms
Definitions that connect to this one.
From the directory
15 more matching entries in the full directory.
From the community
Posts from builders shipping with Jev right now.
Foreman keeps coding agents on task
Unclutter: an ad and slop blocker that runs on Jev
introducing Unclutter: a smart ad + slop blocker with Jev 🤓 it auto cleans up pages from slop elements: ⬖ ads ⬖ cookie banners ⬖ upsells ⬖ bs dialogs BYOK. open source + free, download below 👇
An open-source BS meter for debates and investor calls
🚨 Open Source Jev BS meter you can use this to analyze any debate / investor call / interview / sales pitch / podcast video fact check live , for example this dario interview cost 60 Jev calls / 111K tokens / $0.0047 github.com/ChetasLua/jevm…
🚨 I gave the Trump vs Kamala debate a live BS meter using Jev every sentence, both candidates, 5 yes/no questions each 1,191 Jev calls / 1.18M tokens / 415 ms median total cost : $0.0497 same questions for both, clips picked by one fixed rule, not a fact-check
A Jev-shaped model on Cerebras and Qwen
Built an alternative version of @typesafeai but on @cerebras with Qwen 3.8 27b. Similar quality, similar performance, but vastly different cost. TypeSafe was way cheaper, and did beat Qwen on performance. Closest we can get using LLMs I think. Source: github.com/iammrduncan/ty…
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x
Jev benchmarked on two public safety corpora
1/ Benchmarked TypeSafe's Jev on two public safety corpora. It doesn't generate text, it returns calibrated probabilities you threshold in code. 96.5% on prompt injection, all 662 messages in deepset/prompt-injections. No tuning. 325ms p50.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x
An open 151M-parameter decision engine
TypeSafe AI came out of stealth with Jev, and access is behind a waitlist. I built an open source version Verdict (Open-jev) you can run right now in a browser tab: And its a real post trained model..(link in comments) It is a post trained 151M parameters model. ModernBERT-base Show more
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x
