The target is the post's verdict flair, not a moral ground truth. 770 posts are from 2025.
Unadjusted, Jev's weighted Brier was 0.369 against Sonnet 5 at 0.344. Lower is better.
Jev's median call was 6.3 times faster than Sonnet's and 62 times cheaper. The author paid about $2.75.
Eight question designs were tried first. None beat asking which verdict the subreddit would reach.
A logistic adjustment fit on 300 older posts moved Jev to 0.337. Sonnet was not adjusted the same way in that chart.
Quickstart
bash
python -m jevbench.bench --help
Watch out
No license file. Calls go through OpenRouter. The author is unaffiliated with TypeSafe. A later shared adjustment still left Sonnet ahead by 0.007, which the README calls inconclusive.
Independent cross-model benchmark for Jev-class decision models, running 534 frozen cases per complete entrant with scoring code and a four-axis score of accuracy, calibration, latency and cost.
An experiment comparing Jev with GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6 as evaluators of five frozen weather-agent runs, measuring pass-or-fail accuracy against human labels plus variance, cost, and latency.
cua open sourced a 706k param model that fills a whole form in one 50ms pass
the llm agent doing the same form took 23 turns and 39.6 seconds
the specialists are going to eat the generalists from the bottom
Cua
@trycua
1/ Introducing CUA-S1: a family of System One Models, small, specialized, and built for computer use.
Today we're open-sourcing CUA-S1-FORMS, the first in the family: github.com/trycua/cua
I've been using Jev by @typesafeai Here's the six things i've tried and am confident I'll still use Jev for 60 days from now.
There's many more experiments, ideas, and things I think I will use it for. It's a big deal (more on why in next post).
But I am only sharing thingsShow more
This made me rethink where AI actually fits into security engineering.
For purely engineering work, forget about ChatGPT or Claude.
TypeSafe AI just released Jev, and I think it’s going to change how we build AI into security workflows.
Instead of asking an LLM to “investigateShow more
TypeSafe AI
@typesafeai
we are officially out of stealth! join the frontier and get access to Jev on our website (link on profile)