Skip to content
JevDirectory.org
CommunityPractices & PatternsArticleVerified 2026-09-22

An early-access test of TypeSafe's Jev: calibrated judgments for half a cent

An independent trial running Jev on 24 Norwegian resource-tax hearing responses with eleven typed questions each, then comparing agreement, calibration, latency and cost against DeepSeek V4.1 Flash.

Category
Practices & Patterns
Published by
Community
Author
Added
2026-09-22
Tagscommunityevaluationcalibrationpricing

Highlights

  • Jev answered eleven questions per document for 24 documents at a half cent total, median 0.32s versus 2.7s for DeepSeek with reasoning off.
  • On the ordered substance Score it beat DeepSeek 19 of 24 against 14, the one clear accuracy difference in the trial.
  • Calibration on 192 argument judgments was directionally right: 0% yes in the 0.0-0.1 bin and 98% in the 0.9-1.0 bin, slightly underconfident.
  • Rewriting short draft questions with careful qualifiers cut agreement 0.89 to 0.86 and raised calibration error from 0.040 to 0.116.
  • Norwegian runs about 2.06 characters per token, making Jev's 32k-token state limit roughly 64,000 characters.

Watch out

Twenty-four documents is a first look, not a benchmark, and reference labels come from Claude Fable 5.1, so results measure agreement with a frontier model.

More like this

74GitHub stars
Independent cross-model benchmark for Jev-class decision models, running 534 frozen cases per complete entrant with scoring code and a four-axis score of accuracy, calibration, latency and cost.
Practices & Patterns#community#python#benchmarks
Communityjevbench
An experiment comparing Jev with GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6 as evaluators of five frozen weather-agent runs, measuring pass-or-fail accuracy against human labels plus variance, cost, and latency.
Practices & Patterns#community#python#evaluation
38GitHub stars
Side-by-side benchmark of TypeSafe Jev, Qwen 3.8 27B on Cerebras, and a local Needle 3 across seven synthetic workloads, recording validated outputs, mistakes, latency, tokens, and estimated cost. Raw exports and per-scene limitations are published.
Practices & Patterns#community#typescript#benchmarks
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Beating Gemini Flash Lite on an eval

Browser Use Ultrafast, powered by Jev

A really smart switch statement

hype-free explanation of jev: jev does not replace gpt / claude jev is just a *really* smart switch statement like if 2016 ml classifiers got 2026 levels of intelligence it's a new* type of tool that will make a lot of workloads insanely fast, cheap, and accurate * = and by Show more

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

When a designer gets Jev

Full Jev video tutorial

The case against Jev-scored compaction

This is a terrible compaction strategy that fundamentally doesn't understand how compaction and context management work. Seems like a lot of people are confused so let's break this down. 1. Compaction isn't a filter The role of compaction is to clean up history to keep the Show more

tamara
tamara
@tamarajtran

found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant

Reply