Skip to content
Evaluation & Training

Platt scaling

A two-parameter sigmoid fit on labeled data that maps a model's raw probability onto a better-calibrated one.

Category
Evaluation & Training
Also known as
—
Related terms
4
Directory entries
2
Docs
—
Added
2026-09-28

Definition

On 8,801 sentiment examples, Jev's raw expected calibration error for P(positive) was 0.117. A Platt curve fit on a calibration split cut that to 0.052 on held-out data. Isotonic regression on the same split reached 0.008, because the error was not sigmoid-shaped.

The AITA study used a related idea: a logistic regression learned from Jev's answers on 300 older posts, then adjusted probabilities on the 770 test posts. That chart is not a like-for-like win unless every model is adjusted the same way.

Tagscalibrationevaluation

From the directory

A sentiment study of whether Jev's stated confidence matches accuracy, comparing raw scores with Platt scaling and isotonic regression.
Benchmarks & Evaluations#community#python#evaluation
A 770-post benchmark that scores Jev, Sonnet 5, GPT-5 nano, and two local models on Reddit verdicts with weighted Brier scores.
Benchmarks & Evaluations#community#python#evaluation

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Jev plays Subway Surfers

Agentic browsing in Chrome

I built a Chrome extension for agentic browsing using Jev by @typesafeai, fx.sh including AI Gateway by @vercel. Now agents can browse, click, and interact with websites directly in your browser. Cost effective and fassst. Decision-making by Jev.

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

724 competitor ads, broken down

A chat bot with no LLM

3,282 posts, eight questions each

Jev repositories worth a look, in Japanese

やぁ!兄弟たち! Jevに関するGitHubの実用性と発展性がありそうなリポジトリをまとめたよ! やはり、高速判断を要するComputerUseや完全自動トレードなんかに対しての活用が多い印象だね! Jevは公式のウェイトリストも1日ほどで承認されるけど、待たなくてもVercel AI GatewayからModel: Show more

Reply