Skip to content
JevDirectory.org
CommunityPractices & Patterns3 starsVerified 2026-09-22

Jev vs LLM: a phishing decision benchmark with a calibration audit

A public benchmark comparing Jev with Claude Haiku 4.5 on whether an email agent should click a link, over 2,000 PhishNChips emails, with calibration, latency, cost, and five signal questions.

Category
Practices & Patterns
Published by
Community
Author
anisselbd
Added
2026-09-22
Tagscommunitypythonsecuritybenchmarksevaluationcalibration

Highlights

  • Both systems saw the same 2,000 PhishNChips v5.2 emails in seeded order, one call per email, no concurrency.
  • Jev was 62.6% accurate with AUROC 0.689 and ECE 0.154; Claude Haiku 4.5 was 81.3% accurate with AUROC 0.837.
  • Jev won on cost and speed: $0.038 per 1,000 emails at list price and 239 ms p50 from France versus $0.462 and 687 ms for Haiku.
  • A cross-validated logistic regression on Jev's five signals reached 95.0% on held-out half B, against 91.8% for a two-feature regex regression.
  • The signal result was re-controlled after review: Jev's best single signal lost to both the regex (89.4% vs 91.8%) and Haiku asked the same question (94.2%).

Quickstart

bash
cp .env.example .env            # fill in the TypeSafe and Anthropic keys
uv run run_jev.py --limit 10    # smoke test, prints raw answers
uv run run_llm.py --limit 10

Watch out

No license file, so reuse terms are unclear. Needs Python 3.13, uv, and both TypeSafe and Anthropic keys; published recall on the dataset varies with the system prompt, and one pass per arm limits stability claims.

More like this

74GitHub stars
Independent cross-model benchmark for Jev-class decision models, running 534 frozen cases per complete entrant with scoring code and a four-axis score of accuracy, calibration, latency and cost.
Practices & Patterns#community#python#benchmarks
Communityjevbench
An experiment comparing Jev with GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6 as evaluators of five frozen weather-agent runs, measuring pass-or-fail accuracy against human labels plus variance, cost, and latency.
Practices & Patterns#community#python#evaluation
17GitHub stars
A probability-aware evaluation harness that compares TypeSafe Jev with GLiNER2.5 on zero-shot single-label text classification, measuring calibration, coverage at a fixed error budget, latency, and token cost.
Practices & Patterns#community#python#benchmarks
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Navigating Neo4j with Jev

Jev 这个 waitlist 还是很给力的,昨天申请,今天就能用上。 给已经拿到 API、但还不知道怎么玩的人整理了一份 Awesome Jev,目前我能确认到的 Jev 项目基本都在这里: 1. jev-ultrafast Browser Use 做的高速浏览器 Agent。Jev Show more

Image
思维怪怪
思维怪怪
@0xLogicrw

前 OpenAI 研究员 Diogo Almeida 创办的 TypeSafe AI 推出新模型 Jev。它有点像一个能读懂自然语言的超级分类器,不生成文本,只返回选项、分数和概率,专门给软件做判断。 普通大模型需要一个 token 一个 token 往外生成,Jev 则可以并行给出多个结果。TypeSafe 还用新的 RLCD

Reply

Reranking 33,047 catalog entries

拿 Jev 做搜索重排,我先泼一盆冷水:单独用,它没打赢向量检索 TypeSafe 的 Jev 这阵子很火,一堆项目拿它做重排。我们在 Agent Skills Hub 的 33,047 条目录上认真测了一次,164 条中英文真实查询,9,831 对分级标注,整套只花了 2.6 美元 三个结论 01|单独重排,约等于没赢 Jev 重排 bge-m3 Show more

Jason Zhu
Jason Zhu
@GoSailGlobal

有美团、阿里的老哥嘛? 试试加一路召回、重排(离线、近实时实现),我觉得有奇效 他在文本理解上 跟之前机器学习、llm很不一样 还能自动打标签做特征

Reply

Six uses that stuck after 60 days

Security decisions that fit Jev

This made me rethink where AI actually fits into security engineering. For purely engineering work, forget about ChatGPT or Claude. TypeSafe AI just released Jev, and I think it’s going to change how we build AI into security workflows. Instead of asking an LLM to “investigate Show more

TypeSafe AI
TypeSafe AI
@typesafeai

we are officially out of stealth! join the frontier and get access to Jev on our website (link on profile)

Reply

A million judged questions

Inferring Jev's internals from 1,000 calls

Jevの内部アーキテクチャを推測している技術記事(Jev’s Architecture Unmasked)からメモ。 ・本記事はJevのAPIを約1万回の呼び出して、内部構造を推測したもの ・従来の言語モデルを用いた分類やルーティングでは、トークンを1文字ずつ逐次生成するために膨大な無駄な計算コストが発生していた。 Show more

Reply