Skip to content
JevDirectory.org
CommunityPractices & PatternsArticleVerified 2026-09-22

Testing Jev on public and private data: classifier or filter?

Sixteen thousand calls against gpt-5.4-mini and gpt-5.6-luna on four public datasets plus production pipeline decisions, with a threshold procedure and an explicit filter-not-replacement verdict.

Category
Practices & Patterns
Published by
Community
Author
Added
2026-09-22
Tagscommunitybenchmarksevaluationcalibration

Highlights

  • Jev led on Enron spam (98.7), SST-2 (95.7) and AG News (91.3) but lost Banking77, 76.0 against 81.7 for gpt-5.6-luna.
  • On production page gates it agreed with outcomes 96-98% and sent 25-60% of pages past the classifier with nothing lost.
  • Threshold rule: put the drop line at half the lowest P(yes) any known positive received, set it on part of the data and check it on another fold.
  • At 100 in-flight calls a few per thousand took 10 to 35 seconds; a whole-document read scored 68.6% and no prompt fixed it.
  • Cost gaps: 5-15x versus gpt-5.6-luna, 23-56x versus gpt-5.4-mini, and about 800x versus an agent session for the same email triage.

Watch out

Public data and scoring code are published, but the production numbers come from the author's private pipeline and cannot be checked.

More like this

74GitHub stars
Independent cross-model benchmark for Jev-class decision models, running 534 frozen cases per complete entrant with scoring code and a four-axis score of accuracy, calibration, latency and cost.
Practices & Patterns#community#python#benchmarks
Communityjevbench
An experiment comparing Jev with GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6 as evaluators of five frozen weather-agent runs, measuring pass-or-fail accuracy against human labels plus variance, cost, and latency.
Practices & Patterns#community#python#evaluation
38GitHub stars
Side-by-side benchmark of TypeSafe Jev, Qwen 3.8 27B on Cerebras, and a local Needle 3 across seven synthetic workloads, recording validated outputs, mistakes, latency, tokens, and estimated cost. Raw exports and per-scene limitations are published.
Practices & Patterns#community#typescript#benchmarks
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

700 leads scored for $0.09

Beating Gemini Flash Lite on an eval

Browser Use Ultrafast, powered by Jev

A really smart switch statement

hype-free explanation of jev: jev does not replace gpt / claude jev is just a *really* smart switch statement like if 2016 ml classifiers got 2026 levels of intelligence it's a new* type of tool that will make a lot of workloads insanely fast, cheap, and accurate * = and by Show more

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

When a designer gets Jev

Full Jev video tutorial