We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost LLM) using WebMCP solved 100% of the tasks at roughly 112× lower model cost than GPT-6 Astra using computer use with code execution. Compared to Astra using Show more
Jev + WebMCP Breaks a Benchmark
The WebMCP benchmark author reports that Jev paired with the small Mercury 2.5 model solved 100% of tasks at roughly 112x lower model cost than GPT-6 Astra with computer use.
- Category
- Practices & Patterns
- Published by
- Community
- Author
- idan levin
- Added
- 2026-09-20
The post
Highlights
- Jev + Mercury 2.5 solved 100% of the WebMCP tasks in the run.
- Reported roughly 112x lower model cost than GPT-6 Astra with computer use.
- Shared by the benchmark's own authors rather than a third party.
Watch out
A single benchmark configuration; reproduce it before quoting the multiple.
More like this
From the community
Posts from builders shipping with Jev right now.
Full Jev video tutorial
Full Jev Tutorial What it is, how you can build with it and what new applications it can unlock → 0:00 Intro → 0:34 Jev explained → 4:06 API setup → 5:59 Demo 1: Voice-controlled browser → 11:33 Demo 2: AI memory → 17:27 Demo 3: YouTube predictor
The case against Jev-scored compaction
This is a terrible compaction strategy that fundamentally doesn't understand how compaction and context management work. Seems like a lot of people are confused so let's break this down. 1. Compaction isn't a filter The role of compaction is to clean up history to keep the Show more
found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant
