Skip to content
JevDirectory.org
CommunityPractices & Patterns24 starsVerified 2026-09-22

Jev Capability Atlas

A bilingual field map of where Jev fits and where it fails, separating the author's raw API suites from cited third-party results and TypeSafe's own claims.

Category
Practices & Patterns
Published by
Community
Author
Zaious
Added
2026-09-22
Tagscommunitypythonevaluationcalibrationresearch

Highlights

  • Each case suite ships real API logs, methodology, and a report rather than a leaderboard, citing other benchmarks instead of re-running them.
  • A citation-support suite separates paraphrase support from reversed-meaning overlap to show the judgments are not keyword rules.
  • One history suite records a 0.90-confidence wrong answer caused by a typo in an option, corrected when the typo was fixed.
  • It flags where calibration broke in cited third-party work, including an emotion task at 0.819 mean confidence and 48% accuracy.
  • AGENTS.md defines scan criteria and a result-reporting protocol for agents asked to judge other projects.

Watch out

MIT-licensed for code and original content, while quoted third-party material keeps its own rights. Case studies frame task fit, not population-level accuracy or a reusable threshold.

More like this

74GitHub stars
Independent cross-model benchmark for Jev-class decision models, running 534 frozen cases per complete entrant with scoring code and a four-axis score of accuracy, calibration, latency and cost.
Practices & Patterns#community#python#benchmarks
Communityjevbench
An experiment comparing Jev with GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6 as evaluators of five frozen weather-agent runs, measuring pass-or-fail accuracy against human labels plus variance, cost, and latency.
Practices & Patterns#community#python#evaluation
17GitHub stars
A probability-aware evaluation harness that compares TypeSafe Jev with GLiNER2.5 on zero-shot single-label text classification, measuring calibration, coverage at a fixed error budget, latency, and token cost.
Practices & Patterns#community#python#benchmarks
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Beating Gemini Flash Lite on an eval

Browser Use Ultrafast, powered by Jev

A really smart switch statement

hype-free explanation of jev: jev does not replace gpt / claude jev is just a *really* smart switch statement like if 2016 ml classifiers got 2026 levels of intelligence it's a new* type of tool that will make a lot of workloads insanely fast, cheap, and accurate * = and by Show more

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

When a designer gets Jev

Full Jev video tutorial

The case against Jev-scored compaction

This is a terrible compaction strategy that fundamentally doesn't understand how compaction and context management work. Seems like a lot of people are confused so let's break this down. 1. Compaction isn't a filter The role of compaction is to clean up history to keep the Show more

tamara
tamara
@tamarajtran

found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant

Reply