Skip to content
JevDirectory.org
OfficialCookbooks & DemosVerified 2026-09-20

Cookbook: Self-Consistency for Choices

A repeatability study: an 8-question moderation rubric run 15 times per condition shows a mean probability standard deviation of 0.0098, and a 0.60 uncertainty gate lifts agreement to 99.2%.

Category
Cookbooks & Demos
Published by
TypeSafe AI
Author
Added
2026-09-20
Tagsofficialcookbookconsistencymoderationconfidence

Highlights

  • 15 repeats across 9 conditions over 8 Choice questions, with jev-1.13.0 returning in 114 ms at $0.000046 per call.
  • TypeSafe's mean probability standard deviation was 0.0098 versus 0.0245 to 0.0543 for five LLM distributions.
  • With top probability below 0.60 routed to uncertain, policy agreement reached 99.2% while 74.2% of labels stayed automatic.
  • TypeSafe still flipped its top label on primary_risk and link_handling across its own repeats.

Watch out

These are repeatability numbers, not accuracy numbers, and a deterministic temperature-0 baseline can score 100% agreement by never abstaining.

More like this

A 14-Noul claims-triage rubric over one insurance claim, repeated 15 times: mean probability standard deviation of 0.0102 at 111 ms per call, with a 0.30 to 0.70 band for human review.
Cookbooks & Demos#official#cookbook#consistency
Official
Semantic search over GitHub's Terms of Service: one request ranks all 218 lines with a Choice while a Noul checks whether the document contains an answer at all, including when it should say no.
Cookbooks & Demos#official#cookbook#search
Official
A two-stage extraction cascade: a mini model extracts, a Noul battery verifies each field in one request, and a 0.7 gate escalates to a reasoning model, sitting on the cost/quality frontier.
Cookbooks & Demos#official#cookbook#extraction
Official
Back to all resources

From the community

Posts from builders shipping with Jev right now.

700 leads scored for $0.09

Beating Gemini Flash Lite on an eval