Skip to content
JevDirectory.org
CommunityPractices & PatternsArticleVerified 2026-09-22

Mini-Vibe Check: TypeSafe's Jev Judged Everything I've Written in 0.7 Seconds

Every's head of evals ran Jev across 27 of his articles plus 10 AI-styled fakes, asking 21 AI-tell questions about each and reporting what the judgments caught and where he would not trust them yet.

Category
Practices & Patterns
Published by
Community
Author
Added
2026-09-22
Tagscommunityevaluationwritingbenchmarks

Highlights

  • Mike Taylor, Every's head of evals, sent 27 of his articles plus 10 deliberately AI-styled counterparts through Jev as 37 documents.
  • The same 21 AI-tell questions ran concurrently over all 37 documents: 777 judgments in under 0.7 seconds for about a quarter of a cent.
  • Questions came from Every's own AI-tells skill, including whether text repeats an idea without evidence or forces a symmetrical both-sides argument.
  • Taylor says Jev correctly flagged pieces that leaned more on AI, but he wants a fuller accuracy check before putting it in production.
  • In a small defect test, Jev caught six of seven planted problems at 0.35 seconds per passage; Claude Fable 5.1 caught seven at 8.8 seconds.

Watch out

A personal afternoon-scale experiment rather than a benchmark, and Taylor explicitly hedges the accuracy verdict.

More like this

208GitHub stars
A self-hosted implementation of TypeSafe's Jev System One API powered by the 400M-parameter GLiFormer encoder: it serves choice, score, and noul and drops into the official typesafe-sdk via TYPESAFE_BASE_URL, but trails Jev on reasoning-heavy tasks.
Practices & Patterns#community#python#open-models
Communityjeff
104GitHub stars
A personal-assistant agent with 100 mocked tools that measures how many steps a task takes when the LLM picks the tool versus when Jev picks it before every model step.
Practices & Patterns#community#typescript#agents
91GitHub stars
A comparison arena that runs the same batch of review comments through Jev and DeepSeek, showing processing time, cost, and per-label results with CSV/Excel import, replay, and offline reports.
Practices & Patterns#community#javascript#benchmarks
Communityjev-arena
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

LLM-as-a-judge, sped up

Jev has spoken. It picked which model is AGI. 20–200x faster. 40–400x cheaper. This could make things like LLM-as-a-judge insanely fast and nearly free. (I tried a bunch of prompts and still didn’t burn through $0.10.)

Image
Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

Instant compaction with Jev

A Claude session from 1M to 86K tokens

This is actually insane. This uses @typesafeai Jev model, as a plugin in Claude to review all the un-nesseasary tool calls, and it takes 1s to run! Like, literally, 1 second to take my Claude session from nearly 1M to ... 86K tokens! 😮 Ask your claude to install it and be  Show more

Image
Image
tamara
tamara
@tamarajtran

found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant

Reply

Vercel's fx safety reviewer, 18x faster

We're seeing extraordinary results from @typesafeai. Default mode in 𝚏𝚡 is auto, with a safety reviewer analyzing every command. That reviewer runs on GPT Luna today. Jev is up to 18x faster (p95) *and* more accurate. It's coming to @vercel AI Gateway and likely new default.

Pranit
Pranit
Vercel
@fazxes

We benchmarked fx auto mode (safety) classifier with @typesafeai's Jev. tl;dr: ~5-18x faster and more accurate than 𝚐𝚙𝚝-𝟻.𝟼-𝚕𝚞𝚗𝚊, our current top choice

Image
Reply

Jev lands on OpenRouter

700 leads scored for $0.09