Skip to content
JevDirectory.org
CommunityPractices & Patterns0 starsVerified 2026-09-22

ASSAY-001

A pre-registered, independently rescored audit of Jev's calibration and type safety: calibrated on CLINC150 (ECE 0.0204), systematically overconfident on Banking77 (ECE 0.0936), and zero type errors across 8,576 responses.

Category
Practices & Patterns
Published by
Community
Author
jourdanlabs
Added
2026-09-22
Tagscommunitypythoncalibrationevaluationbenchmarks

Highlights

  • On CLINC150 Jev's chosen-option probabilities were calibrated at ECE 0.0204; on Banking77 they were not, at ECE 0.0936 and systematically overconfident.
  • Across 8,576 responses there were zero type errors.
  • The protocol was frozen on 2026-09-17 before any query, with corpus hashes for Banking77 and CLINC150 sealed before the first API call.
  • Every request and response is logged verbatim and sealed with SHA-256; smoke-test connectivity runs are excluded from scoring.
  • A blind independent re-score written from a spec on a different base model matched every field of the original scoring.

Quickstart

bash
python3 harness/controls.py
python3 harness/score.py banking77 --sum-tol 0.02
python3 harness/score.py clinc150 --sum-tol 0.02

Watch out

No license file, so reuse terms are unclear; re-running the model needs a TypeSafe key at ~/.config/typesafe/api_key and produces a new run, not the recorded one.

More like this

74GitHub stars
Independent cross-model benchmark for Jev-class decision models, running 534 frozen cases per complete entrant with scoring code and a four-axis score of accuracy, calibration, latency and cost.
Practices & Patterns#community#python#benchmarks
Communityjevbench
An experiment comparing Jev with GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6 as evaluators of five frozen weather-agent runs, measuring pass-or-fail accuracy against human labels plus variance, cost, and latency.
Practices & Patterns#community#python#evaluation
17GitHub stars
A probability-aware evaluation harness that compares TypeSafe Jev with GLiNER2.5 on zero-shot single-label text classification, measuring calibration, coverage at a fixed error budget, latency, and token cost.
Practices & Patterns#community#python#benchmarks
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Jev lands on OpenRouter

700 leads scored for $0.09

Beating Gemini Flash Lite on an eval

Browser Use Ultrafast, powered by Jev

A really smart switch statement

hype-free explanation of jev: jev does not replace gpt / claude jev is just a *really* smart switch statement like if 2016 ml classifiers got 2026 levels of intelligence it's a new* type of tool that will make a lot of workloads insanely fast, cheap, and accurate * = and by Show more

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

When a designer gets Jev