typesafe's jev is fun! live demo you can play with: typesafe-demo.val.run
Brier score
A score for a full probability vector against one observed label. Lower is better, and a confident wrong answer costs more than a hedge.
- Category
- Evaluation & Training
- Also known as
- —
- Related terms
- 4
- Directory entries
- 2
- Docs
- —
- Added
- 2026-09-28
Definition
The r/AmItheAsshole benchmark grades every verdict probability, from 0 for a perfect call to 2 for the worst. On 770 posts from 2025, unadjusted Jev scored 0.369 and Sonnet 5 scored 0.344. Top-1 accuracy is reported beside it and is a different number.
Rizzo Flow reports Brier on typed-decisions as well: 0.205 after its fine-tune, against 0.480 for the untuned Spark-X2.5-4B and 0.148 on Jev's dataset card. Compare Brier only when the label set and the weighting match.
Related terms
Definitions that connect to this one.
From the directory
From the community
Posts from builders shipping with Jev right now.
A playable 16-judgment demo
AI multiple choice, not essay writing
WTF is Jev by @typesafeai? Here’s the tl;dr ELI5: Think AI multiple choice, not AI essay writing. It doesn’t chat. It makes decisions your software can act on: “Spam or not?” “Which tool should this agent use?” “Does this need a human?” The exciting part: roughly 200x faster Show more
Screening agent actions with Jev
Tested TypeSafe’s Jev (no-text, probability-only model) as an AI agent safety monitor. Checking each action first worked well caught most attacks with almost no false blocks, and much faster than Gemini.
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x
Cua's small System One models
1/ Introducing CUA-S1: a family of System One Models, small, specialized, and built for computer use. Today we're open-sourcing CUA-S1-FORMS, the first in the family: github.com/trycua/cua
A 706K-parameter form filler
cua open sourced a 706k param model that fills a whole form in one 50ms pass the llm agent doing the same form took 23 turns and 39.6 seconds the specialists are going to eat the generalists from the bottom
1/ Introducing CUA-S1: a family of System One Models, small, specialized, and built for computer use. Today we're open-sourcing CUA-S1-FORMS, the first in the family: github.com/trycua/cua
Navigating Neo4j with Jev
Jev 这个 waitlist 还是很给力的,昨天申请,今天就能用上。 给已经拿到 API、但还不知道怎么玩的人整理了一份 Awesome Jev,目前我能确认到的 Jev 项目基本都在这里: 1. jev-ultrafast Browser Use 做的高速浏览器 Agent。Jev Show more
前 OpenAI 研究员 Diogo Almeida 创办的 TypeSafe AI 推出新模型 Jev。它有点像一个能读懂自然语言的超级分类器,不生成文本,只返回选项、分数和概率,专门给软件做判断。 普通大模型需要一个 token 一个 token 往外生成,Jev 则可以并行给出多个结果。TypeSafe 还用新的 RLCD
