Skip to content
JevDirectory.org
Evaluation & TrainingReinforcement learning from human feedback

RLHF

The human-preference training method that turned pretrained models into chatbots, and the method TypeSafe deliberately did not use.

Category
Evaluation & Training
Also known as
Reinforcement learning from human feedback
Related terms
5
Directory entries
1
Docs
docs.typesafe.ai
Added
2026-09-24

Definition

RLHF trains a model to produce responses people prefer. It powered InstructGPT and ChatGPT and was co-invented by Diogo Almeida, now a TypeSafe cofounder, but the primer argues preference optimization can reward sycophancy and confident-sounding hallucinations.

It remains a good fit for conversational models. TypeSafe's position is that production automation needs a different objective — constrained decisions and calibrated uncertainty.

Tagstraining

From the directory

Why TypeSafe trains decision models with RLCD instead of RLHF: calibrated probabilities where 0.2 outcomes happen about 20% of the time, and the case for machine-to-machine automation.
Sites & GuidesDocs#official#docs#evaluation
Official

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Jev plays Subway Surfers

Agentic browsing in Chrome

I built a Chrome extension for agentic browsing using Jev by @typesafeai, fx.sh including AI Gateway by @vercel. Now agents can browse, click, and interact with websites directly in your browser. Cost effective and fassst. Decision-making by Jev.

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

724 competitor ads, broken down

A chat bot with no LLM

3,282 posts, eight questions each

Jev repositories worth a look, in Japanese

やぁ!兄弟たち! Jevに関するGitHubの実用性と発展性がありそうなリポジトリをまとめたよ! やはり、高速判断を要するComputerUseや完全自動トレードなんかに対しての活用が多い印象だね! Jevは公式のウェイトリストも1日ほどで承認されるけど、待たなくてもVercel AI GatewayからModel: Show more

Reply