Skip to content
JevDirectory.org
Evaluation & Training

Repeatability

How stable answers are across repeated evaluations, measured separately from accuracy and reported as probability standard deviation.

Category
Evaluation & Training
Also known as
—
Related terms
4
Directory entries
2
Docs
docs.typesafe.ai
Added
2026-09-24

Definition

The consistency cookbooks repeat the same rubric 15 times per condition and report a mean probability standard deviation of about 0.01 for jev-1.13.0, against 0.0245 to 0.0543 for five LLM distributions.

Repeatability is not correctness: Jev still flipped its top label on two of its own questions, and a deterministic temperature-0 baseline can score perfect agreement by never abstaining. The docs keep the two measurements distinct.

Tagsevaluationconsistency

From the directory

A repeatability study: an 8-question moderation rubric run 15 times per condition shows a mean probability standard deviation of 0.0098, and a 0.60 uncertainty gate lifts agreement to 99.2%.
Cookbooks & DemosDocs#official#cookbook#consistency
Official
A 14-Noul claims-triage rubric over one insurance claim, repeated 15 times: mean probability standard deviation of 0.0102 at 111 ms per call, with a 0.30 to 0.70 band for human review.
Cookbooks & DemosDocs#official#cookbook#consistency
Official

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Fraud detection with Jev and Kimi K3

Jev Detector scans ~10,000 words for slop in ~2 s

Computer use without screenshots

Okay so Jev can actually do computer use really well Without any screenshots, or LLMs and no Pixels leave my mac I dont even read the Dom elements A local CoreML model segments every button and UI element on screen. On-device OCR reads the labels. That text is all Jev gets. Show more

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

Jev plays Subway Surfers

Agentic browsing in Chrome

I built a Chrome extension for agentic browsing using Jev by @typesafeai, fx.sh including AI Gateway by @vercel. Now agents can browse, click, and interact with websites directly in your browser. Cost effective and fassst. Decision-making by Jev.

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

724 competitor ads, broken down