Skip to content
JevDirectory.org
Evaluation & Training

Benchmark claims

TypeSafe's headline multipliers — 193.6x faster, 444.6x cheaper, a 238x lower input price — come from its own workflow comparisons.

Category
Evaluation & Training
Also known as
—
Related terms
4
Directory entries
2
Docs
typesafe.ai
Added
2026-09-24

Definition

The product page lists Jev at $42 per billion input tokens, described as a 238x lower input price than Claude Fable 5.1, and claims 193.6x faster and 444.6x cheaper than LLMs on its own System One workflows.

The caveat belongs with the number: these are vendor benchmarks from TypeSafe's own comparison, not independent testing. The independent picture lives in community benchmarks, field reports, and the recurring finding that question design dominates outcomes.

Tagsbenchmarksevaluation

From the directory

TypeSafe AI's product site for Jev, its first System One Model: typed decisions with calibrated confidence, performance and pricing claims, a FAQ, and links to the docs, console, workflow evals, and launch post.
Sites & GuidesDocs#official#docs#models
Official
Published evaluations of four automation workflows, security incidents, agent trace observability, invoice processing, and customer service, comparing Jev and frontier LLMs as structured workflows versus single prompts.
Practices & PatternsDocs#official#benchmarks#evaluation
Official

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

The case against Jev-scored compaction

This is a terrible compaction strategy that fundamentally doesn't understand how compaction and context management work. Seems like a lot of people are confused so let's break this down. 1. Compaction isn't a filter The role of compaction is to clean up history to keep the Show more

tamara
tamara
@tamarajtran

found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant

Reply

Classifying rows in DuckDB

A playable 16-judgment demo

AI multiple choice, not essay writing

Screening agent actions with Jev

Tested TypeSafe’s Jev (no-text, probability-only model) as an AI agent safety monitor. Checking each action first worked well caught most attacks with almost no false blocks, and much faster than Gemini.

Image
Image
Image
Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

Cua's small System One models