Skip to content
JevDirectory.org
Evaluation & TrainingReinforcement learning for calibrated decisions

RLCD

TypeSafe's post-training method: reinforcement learning that optimizes for calibrated decisions and probabilities instead of generated text.

Category
Evaluation & Training
Also known as
Reinforcement learning for calibrated decisions
Related terms
5
Directory entries
5
Docs
docs.typesafe.ai
Added
2026-09-24

Definition

RLCD is the third post-training path the AI primer describes, alongside RLHF and RLVR. It optimizes a different output contract: the model does not generate text, it returns decisions and probabilities, and higher probability should correspond to a greater chance that the answer is correct.

Calibration is the measurable result: outcomes assigned 0.2 should occur about 20% of the time across many predictions. That property is what makes Jev's confidence usable by code.

Tagstrainingcalibration

From the directory

Why TypeSafe trains decision models with RLCD instead of RLHF: calibrated probabilities where 0.2 outcomes happen about 20% of the time, and the case for machine-to-machine automation.
Sites & GuidesDocs#official#docs#evaluation
Official
TypeSafe AI's product site for Jev, its first System One Model: typed decisions with calibrated confidence, performance and pricing claims, a FAQ, and links to the docs, console, workflow evals, and launch post.
Sites & GuidesDocs#official#docs#models
Official
127GitHub stars
Together AI's open recipe and weights for a Jev-inspired decision model fine-tuned on Qwen3.5-4B, with the full data pipeline, training config, and saved benchmark results.
Repos & SDKs#community#python#open-models
A TypeSafe blog essay arguing that in ML the order that matters is doing the right task, then data, then compute, then algorithms, using the InstructGPT result as its example.
Sites & GuidesArticle#official#article#models
Official
1.3kGitHub stars
Train a small model that chooses among a changing list of text options, one probability per option in a single pass. Includes Doom, chess, and Wikispeedia demos.
Repos & SDKs#community#training#research

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Computer use without screenshots

Okay so Jev can actually do computer use really well Without any screenshots, or LLMs and no Pixels leave my mac I dont even read the Dom elements A local CoreML model segments every button and UI element on screen. On-device OCR reads the labels. That text is all Jev gets. Show more

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

Jev plays Subway Surfers

Agentic browsing in Chrome

I built a Chrome extension for agentic browsing using Jev by @typesafeai, fx.sh including AI Gateway by @vercel. Now agents can browse, click, and interact with websites directly in your browser. Cost effective and fassst. Decision-making by Jev.

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

724 competitor ads, broken down

A chat bot with no LLM

3,282 posts, eight questions each