JEV is INSANE. We gave it 700 high-intent leads and personalised outreach messages. In 40 seconds, it predicted how each message would perform, assigned a confidence score and detected lead-message mismatches. All for just $0.09. JEV can also score leads, analyse buying Show more
Cookbook: Self-Consistency for Choices
A repeatability study: an 8-question moderation rubric run 15 times per condition shows a mean probability standard deviation of 0.0098, and a 0.60 uncertainty gate lifts agreement to 99.2%.
- Category
- Cookbooks & Demos
- Published by
- TypeSafe AI
- Author
- —
- Added
- 2026-09-20
Tagsofficialcookbookconsistencymoderationconfidence
Highlights
- 15 repeats across 9 conditions over 8 Choice questions, with jev-1.13.0 returning in 114 ms at $0.000046 per call.
- TypeSafe's mean probability standard deviation was 0.0098 versus 0.0245 to 0.0543 for five LLM distributions.
- With top probability below 0.60 routed to uncertain, policy agreement reached 99.2% while 74.2% of labels stayed automatic.
- TypeSafe still flipped its top label on primary_risk and link_handling across its own repeats.
Watch out
These are repeatability numbers, not accuracy numbers, and a deterministic temperature-0 baseline can score 100% agreement by never abstaining.
More like this
A 14-Noul claims-triage rubric over one insurance claim, repeated 15 times: mean probability standard deviation of 0.0102 at 111 ms per call, with a 0.30 to 0.70 band for human review.
Cookbooks & Demos#official#cookbook#consistency
Official
Semantic search over GitHub's Terms of Service: one request ranks all 218 lines with a Choice while a Noul checks whether the document contains an answer at all, including when it should say no.
Cookbooks & Demos#official#cookbook#search
Official
A two-stage extraction cascade: a mini model extracts, a Noul battery verifies each field in one request, and a 0.7 gate escalates to a reasoning model, sitting on the cost/quality frontier.
Cookbooks & Demos#official#cookbook#extraction
Official
From the community
Posts from builders shipping with Jev right now.
700 leads scored for $0.09
Beating Gemini Flash Lite on an eval
Ran @typesafeai's Jev against an existing classifier eval that previously used Gemini 2.5 Flash Lite. It won both on quality (saturated the eval) and speed (6x)
