Skip to content
JevDirectory.org
CommunityPractices & PatternsArticleVerified 2026-09-22

Jev: one judge call, or twelve dimension scores? I measured both on three tasks

An independent measurement spanning 5,477 test rows and 34.1M input tokens, comparing one direct Jev question per row with 12-14 scored dimensions fitted to local labels.

Category
Practices & Patterns
Published by
Community
Author
Added
2026-09-22
Tagscommunitybenchmarksevaluationcalibration

Highlights

  • The full experiment cost $1.43: 25,174 Jev calls, 5,477 test rows, and 34.1M input tokens extracted in 9 minutes 14 seconds with zero failed calls.
  • On Japanese NLI, 14 dimensions with fitted weights reached 0.908 accuracy versus 0.837 for the direct Jev call, a 7.03-point gain.
  • A 12-option ledger classification broke the direct call at 0.400 accuracy; dimensions reached 0.911 and stacked n-gram probabilities 0.9695.
  • On hard-benign guardrail text the direct question flagged 1.5% versus 37.2% for the dimension model, about 25x more false positives.
  • The direct call reported confidence at or above 0.9 on 42% of rows where it was only 72.2% accurate, and four repair attempts all failed.

Watch out

The dimensions are hand-written by the author and output tokens were left unpriced, so the stated cost is input-only.

More like this

74GitHub stars
Independent cross-model benchmark for Jev-class decision models, running 534 frozen cases per complete entrant with scoring code and a four-axis score of accuracy, calibration, latency and cost.
Practices & Patterns#community#python#benchmarks
Communityjevbench
An experiment comparing Jev with GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6 as evaluators of five frozen weather-agent runs, measuring pass-or-fail accuracy against human labels plus variance, cost, and latency.
Practices & Patterns#community#python#evaluation
38GitHub stars
Side-by-side benchmark of TypeSafe Jev, Qwen 3.8 27B on Cerebras, and a local Needle 3 across seven synthetic workloads, recording validated outputs, mistakes, latency, tokens, and estimated cost. Raw exports and per-scene limitations are published.
Practices & Patterns#community#typescript#benchmarks
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Jev lands on OpenRouter

700 leads scored for $0.09

Beating Gemini Flash Lite on an eval

Browser Use Ultrafast, powered by Jev

A really smart switch statement

hype-free explanation of jev: jev does not replace gpt / claude jev is just a *really* smart switch statement like if 2016 ml classifiers got 2026 levels of intelligence it's a new* type of tool that will make a lot of workloads insanely fast, cheap, and accurate * = and by Show more

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

When a designer gets Jev