Skip to content
JevDirectory.org
Evaluation & Training

Oracle floor

The best error rate achievable with perfect choices from the available options, used to judge how close a system gets.

Category
Evaluation & Training
Also known as
—
Related terms
3
Directory entries
1
Docs
docs.typesafe.ai
Added
2026-09-24

Definition

The skill suggestion cookbook measures wrong skill loads at 16.8% before and 7.3% after, against an oracle floor of 2.5%, and needless loads at 9.8% to 4.0% against a floor of 1.2%. The floor is what a perfect chooser would score given the same catalog and inputs.

Floors make improvement legible: halving the gap to the oracle is progress, while a raw percentage alone cannot say whether a system is near the ceiling.

Tagsevaluationbenchmarks

From the directory

Rank 182 agent skills in one request and re-read the top three in a second: over 488 requests, wrong skill loads fell from 16.8% to 7.3% and needless loads from 9.8% to 4.0%.
Cookbooks & DemosDocs#official#cookbook#agents
Official

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Custom Jev-style models for agent workflows

Prediction: millionaires will be made using custom Jev style models (parallel constrained decoding) to make the agent systems companies already run more token efficient. Let me explain with a scenario: Imagine a company already has an agent workflow running where an llm reviews Show more

Harsha Gundala
Harsha Gundala
@harshagundal

They were building in stealth for 2 years, I was building in stealth for 2 hours… Happy to open source Qwen-2.5-1B-RLCD, 5x faster on-device inference for JSON workloads that need to be type-safe. ⚡️Demo below on a M4 MacBook⚡️ every LLM has the ability to efficiently batch

Reply

Live viral post analyzer

Jev plays Tetris: 134 lines in two minutes

Support answers in a Mac app

A local Jev build with room to get faster

jev-review: a local-first score loop for coding agents

built `jev-review` @typesafeai it's an experimental, local-first MCP plugin that gives coding agents a score quality feedback loop across different metrics. agents call jev while they work, get scored, make improvements, and repeat the loop try below 👇

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply