Skip to content
JevDirectory.org

Jev within 5 points of a fine-tune

An internal benchmark report: Jev zero-shot came within about 5 points of recall of a fine-tuned model at matched precision, at roughly $70 a month.

Category
Practices & Patterns
Format
Post
Published by
Community
Author
iden
Added
2026-09-25
Last verified
2026-09-25

Highlights

  • Within about 5 points of recall of the team's fine-tune at matched precision.
  • Full volume estimated at roughly $70 a month with sub-second latency.
  • Concludes a fine-tuned Qwen still beats it on their task.

Watch out

One company's internal benchmark with no dataset or threshold details.

Reactions & coverage

Posts, threads, and videos about this entry from around the web.

X: 100,000 viral posts in 20.4 seconds

X: SEO and GEO fixes, 90% cheaper

X: Hard-coded rules moved to Jev

Related terms

Glossary definitions related to this entry.

More like this

Vercel's Pranit reports benchmarking the fx auto-mode safety classifier with Jev: about 5 to 18 times faster and more accurate than GPT-5.6 Luna.
Practices & PatternsPost#community#x#benchmark
Community
Vercel CTO Malte Ubl reports that Jev beat an existing classifier eval previously run on Gemini 2.5 Flash Lite, saturating the eval on quality and running 6x faster.
Practices & PatternsPost#community#x#evaluation
Community
The WebMCP benchmark author reports that Jev paired with the small Mercury 2.5 model solved 100% of tasks at roughly 112x lower model cost than GPT-6 Astra with computer use.
Practices & PatternsPost#community#x#benchmark
Community
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Custom Jev-style models for agent workflows

Prediction: millionaires will be made using custom Jev style models (parallel constrained decoding) to make the agent systems companies already run more token efficient. Let me explain with a scenario: Imagine a company already has an agent workflow running where an llm reviews Show more

Harsha Gundala
Harsha Gundala
@harshagundal

They were building in stealth for 2 years, I was building in stealth for 2 hours… Happy to open source Qwen-2.5-1B-RLCD, 5x faster on-device inference for JSON workloads that need to be type-safe. ⚡️Demo below on a M4 MacBook⚡️ every LLM has the ability to efficiently batch

Reply

Live viral post analyzer

Jev plays Tetris: 134 lines in two minutes

Support answers in a Mac app

A local Jev build with room to get faster

jev-review: a local-first score loop for coding agents

built `jev-review` @typesafeai it's an experimental, local-first MCP plugin that gives coding agents a score quality feedback loop across different metrics. agents call jev while they work, get scored, make improvements, and repeat the loop try below 👇

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply