Jevの内部アーキテクチャを推測している技術記事(Jev’s Architecture Unmasked)からメモ。 ・本記事はJevのAPIを約1万回の呼び出して、内部構造を推測したもの ・従来の言語モデルを用いた分類やルーティングでは、トークンを1文字ずつ逐次生成するために膨大な無駄な計算コストが発生していた。 Show more
jev-agent-failure-benchmark
A benchmark that runs Jev on all 6,257 text traces of Who&When Pro to attribute agent failures, scoring it with the official pinned scorer against the paper's GPT-5.4, Claude Sonnet 4.6, GLM-5, and Qwen3.5-122B baselines.
- Category
- Practices & Patterns
- Published by
- Community
- Author
- TokenTrim
- Added
- 2026-09-22
Highlights
- Jev outperformed GPT-5.4 on every axis at about $1.28 total: Who 73.4, When 76.4, and error-type F1 23.7, the best in the field.
- Jev answers three typed Choice questions per trace (responsible agent, decisive step, error mode), each with a calibrated probability distribution.
- The official whowhen_eval scorer is pinned at commit 14369dcb so rows are directly comparable with the paper's numbers.
- Who and When are constrained-choice for Jev while the LLM baselines free-generate, so the like-for-like axis is What, the error type over a shared taxonomy.
- Ground-truth labels never enter Jev's input, and a leakage test enforces it; runs are resumable and record failures rather than dropping them.
Quickstart
uv venv && uv pip install -e ".[dev]"
mkdir -p data
REV=0bd196c8a040841c4ae167ab33cc8151de246f1f
curl -L "https://huggingface.co/datasets/Leoxx/whowhen_pro/resolve/$REV/data/text.jsonl" -o data/text.jsonl
curl -L "https://huggingface.co/datasets/Leoxx/whowhen_pro/resolve/$REV/taxonomy.yaml" -o data/taxonomy.yaml
jevbench sample --n 300 --out results/run/sample.json
jevbench run --sample results/run/sample.json
jevbench reportWatch out
Apache-2.0 code; the CC-BY-4.0 Who&When Pro dataset is not included and must be fetched, failures are injected by a controlled pipeline rather than being natural incidents, and running needs Python with uv plus a TYPESAFE_API_KEY.
More like this
From the community
Posts from builders shipping with Jev right now.
Inferring Jev's internals from 1,000 calls
The open System One roundup
Jev 发布没几天,开源社区已经开始疯狂复刻了🔥 最值得推荐的五个模型: 1、Laya 421M:原生决策模型,支持 Mac 2、Decider-2B:最像 Jev,基于 Qwen3.5 3、NanoJev 0.6B:专门的 Decision Head 4、Reflex:Qwen3.5 + Direct Logits 5、System-One 4B:专门做概率校准 Show more
Jev 刚发布没几天,开源社区就出现了同款🔥 Decider-2B模型,是基于 Qwen3.5-2B 做了特殊调整 它和 Jev 模型是一样的 只做选择 评分和判断 不是文本类的 LLM 模型 但两者还是有几个明显区别: 1、模型 Jev:闭源 System One Model Decider:Qwen3.5-2B,约 1.9B 参数,Apache 2.0 开源 2、价格
The launch post
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x Show more
Trading bot, one decision per block
I built a trading bot with Jev! Jev decides if it should "buy" or "sell", given the price feed of an asset pair, and executes real trades. It uses Monad to place the orders on Kuru's on-chain order book in every 300ms block. Demo link → jev-trader.vercel.app
Classifying 1,500 real emails
this model is actually insane at email classification i tested it on 1500 of my own emails to see how well it works and I am blown away
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x
Fast browser use with Stagehand
we built blazing fast computer/browser use with Jev + @Stagehanddev. this task cost $0.001 and executed at near instant speed (in a remote browser btw) the loop: observe the page, send a11y tree as state + actions as questions, Jev decides the next action, then Stagehand Show more
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x



