Skip to content

Typed judgments or agentic loops?

Sam Reghenzi's product-taxonomy benchmark: a Jev Choice at every level versus a gpt-5.2 agent that descends the tree and can be sent back by a judge.

Category
Guides & Articles
Format
Article
Published by
Community
Author
Sam Reghenzi
Added
2026-09-28
Last verified
2026-09-28

Highlights

  • Fifty products, same taxonomies and Postgres. Mean latency fell from 9.62 s to 1.38 s. Model calls fell from 7.22 to 3.18.
  • The slowest Jev run, 3.68 s, beat the fastest old run, 5.50 s. Fifty sequential items finished in 68.8 s.
  • One request can ask the current level and the children of each candidate. On eight products, 1 to 7 questions stayed about 0.32–0.39 s.
  • Those eight paths matched with fan-out on or off. The agent still won some close branches, such as whey powder versus exercise.
  • Twelve of 50 old paths stopped mid-descent when the budget ran out. Four cases the author refused to score.

Watch out

The old pipeline used 5 threads and the new one ran sequentially, so the 7x headline is the author's generous reading. Per-call 1.30 s versus 0.43 s does not depend on that. Output tokens rose because each Choice returns a probability per criterion.

Reactions & coverage

Posts, threads, and videos about this entry from around the web.

Hacker News: A product-taxonomy migration from a gpt-5.2 agent loop to one Jev Choice per level.

sammy_rulez2 pointsReplacing an agentic classification loop with Jev: 7x fasterRead the thread on Hacker News

More like this

Datadog's walkthrough of one Jev rubric that scores every criterion in a single request, then runs as online evals on live spans and as offline experiments.
Guides & ArticlesArticle#community#article#evaluation
Community
A release guide with eval tables, pricing, curl and Python examples, and fit boundaries for Jev's early access.
Guides & ArticlesArticle#community#article#pricing
Community
A heavily footnoted claims audit that separates documented Jev facts from marketing, including the seed round, third-party tests, and missing transparency.
Guides & ArticlesArticle#community#article#analysis
Community
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Computer use at 155× cheaper than Opus 5

i built computer use using @typesafeai ! it is 155x cheaper than opus 5, ~20x faster, and generalizes across OS's more on how it works in the vid & thread below:

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

Foreman keeps coding agents on task

Unclutter: an ad and slop blocker that runs on Jev

An open-source BS meter for debates and investor calls

🚨 Open Source Jev BS meter you can use this to analyze any debate / investor call / interview / sales pitch / podcast video fact check live , for example this dario interview cost 60 Jev calls / 111K tokens / $0.0047 github.com/ChetasLua/jevm…

Chetaslua
Chetaslua
@chetaslua

🚨 I gave the Trump vs Kamala debate a live BS meter using Jev every sentence, both candidates, 5 yes/no questions each 1,191 Jev calls / 1.18M tokens / 415 ms median total cost : $0.0497 same questions for both, clips picked by one fixed rule, not a fact-check

Reply

A Jev-shaped model on Cerebras and Qwen

Built an alternative version of @typesafeai but on @cerebras with Qwen 3.8 27b. Similar quality, similar performance, but vastly different cost. TypeSafe was way cheaper, and did beat Qwen on performance. Closest we can get using LLMs I think. Source: github.com/iammrduncan/ty…

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

Jev benchmarked on two public safety corpora

1/ Benchmarked TypeSafe's Jev on two public safety corpora. It doesn't generate text, it returns calibrated probabilities you threshold in code. 96.5% on prompt injection, all 662 messages in deepset/prompt-injections. No tuning. 325ms p50.

Terminal dashboard showing 96.5% accuracy on prompt injection and 89.0% pairwise on vulnerable code, with a context ablation and calibration plot.
Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply