Skip to content

Benchmarks & Evaluations

2 curated entries. Benchmarks, calibration studies, independent evaluations, and observability write-ups that measure how Jev actually behaves.

By format

  • Repos & sites2

2 resources

A 770-post benchmark that scores Jev, Sonnet 5, GPT-5 nano, and two local models on Reddit verdicts with weighted Brier scores.
Benchmarks & Evaluations#community#python#evaluation
A sentiment study of whether Jev's stated confidence matches accuracy, comparing raw scores with Platt scaling and isotonic regression.
Benchmarks & Evaluations#community#python#evaluation

Watch: Jev in action

Video walkthroughs and demos from the community.

YouTube: Sam Witteveen on System 1 thinking, then Choice, Score, and Noul demos, and chained actions.

YouTube: Nate Herk compares speed and cost against ordinary models, then builds an X feed classifier and a paper-trading prototype.

YouTube: Ray Amjad puts Jev inside an agentic coding loop and looks at what it costs to run.

Discussions

Threads from Reddit and Hacker News about building with Jev.

Reddit: A skeptical read of the calibration claim: no ECE or reliability curves published.

Hacker News: Jev-Leftpad: the joke that measures how cheap decisions are

fka233 points · 87 commentsJev-LeftpadRead the thread on Hacker News

Hacker News: A product-taxonomy migration from a gpt-5.2 agent loop to one Jev Choice per level.

sammy_rulez2 pointsReplacing an agentic classification loop with Jev: 7x fasterRead the thread on Hacker News

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

700 leads scored for $0.09

Beating Gemini Flash Lite on an eval

Browser Use Ultrafast, powered by Jev

A really smart switch statement

hype-free explanation of jev: jev does not replace gpt / claude jev is just a *really* smart switch statement like if 2016 ml classifiers got 2026 levels of intelligence it's a new* type of tool that will make a lot of workloads insanely fast, cheap, and accurate * = and by Show more

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

When a designer gets Jev

Full Jev video tutorial