On 8,801 unique sentiment examples, raw expected calibration error for P(positive) was 0.117.
Platt scaling cut that to 0.052. Isotonic regression cut it to 0.008 on the held-out split.
Isotonic matched or beat Platt at every calibration size from 20 to 5,280 examples.
Noul accuracy rose from 47.6% at 50–60% confidence to 100% above 95% (2,104 of 2,104).
For a two-option Choice, the confidence field matched 2 times p_top minus 1, within rounding.
Quickstart
bash
# results live in results/variants.json and results/raw_confidence_bands.json
Watch out
No license file. One constructed sentiment set and jev-1.13.0. Neutral-tier labels are arbitrary by design. The whole dataset was about 330 input tokens per request and a median latency of 0.24 seconds.
Independent cross-model benchmark for Jev-class decision models, running 534 frozen cases per complete entrant with scoring code and a four-axis score of accuracy, calibration, latency and cost.
An experiment comparing Jev with GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6 as evaluators of five frozen weather-agent runs, measuring pass-or-fail accuracy against human labels plus variance, cost, and latency.
I've been using Jev by @typesafeai Here's the six things i've tried and am confident I'll still use Jev for 60 days from now.
There's many more experiments, ideas, and things I think I will use it for. It's a big deal (more on why in next post).
But I am only sharing thingsShow more
This made me rethink where AI actually fits into security engineering.
For purely engineering work, forget about ChatGPT or Claude.
TypeSafe AI just released Jev, and I think it’s going to change how we build AI into security workflows.
Instead of asking an LLM to “investigateShow more
TypeSafe AI
@typesafeai
we are officially out of stealth! join the frontier and get access to Jev on our website (link on profile)