Skip to content
JevDirectory.org

Jev in Search: Three Practical Evaluations

An independent evaluation of Jev in three retrieval stacks, with recalculated recall, latency, and cost against DeepSeek, GPT-4o-mini, and GPT-5-mini baselines.

Category
Practices & Patterns
Format
Article
Published by
Community
Author
Cheney Zhang
Added
2026-09-25
Last verified
2026-09-25

Highlights

  • In DeepSearcher both Jev and DeepSeek V4 Flash hit 93.25% Recall@5, while decision latency fell from 2.23 s to 0.55 s.
  • On 2,172 MemSearch questions Jev lifted Recall@5 from 74.71% to 79.41%, behind Voyage rerank-3 at 81.87%.
  • On Vector Graph RAG Jev beat GPT-4o-mini but trailed GPT-5-mini by 1 point on MuSiQue and 4.13 on HotpotQA.
  • Jev cost about $3.15 per 1,000 questions against $6 to $9 for GPT-5-mini.

Watch out

The author works on the evaluated projects, and baselines reuse published 1,000-question results against Jev's 500-question samples, so the comparisons are not paired tests.

Reactions & coverage

Posts, threads, and videos about this entry from around the web.

X: Search and tagging on keep.md

X: Jev reranks station suggestions in TrainLCD

YouTube: Jev in the reranking step of a RAG pipeline: why cosine similarity fails and how a steerable reranker compares. A Colab notebook comes with it.

X: Natural-language search over Zillow

Related terms

Glossary definitions related to this entry.

More like this

A graded relevance evaluation of Jev as a reranker: 9,831 labelled pairs from 164 queries over a 33,047-item skills catalog, comparing Jev score reranks with BM25, bge-m3, and rank fusion.
Practices & Patterns#community#python#search
8GitHub stars
Reranking benchmark that gave Jev, Cohere Rerank 4 Pro, ZeroEntropy zerank-2, and DeepSeek the same thirty BM25 candidates across eight English datasets, publishing saved responses, scoring code, and paired-bootstrap intervals. Jev's rubric scored 0.692 nDCG@10 against Cohere Pro's 0.691.
Practices & Patterns#community#python#search
A measured search reranking run over 33,047 catalog entries, 164 real queries, and 9,831 labelled pairs, reporting how Jev compares with BM25 and bge-m3.
Practices & PatternsPost#community#x#search
Community
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Security decisions that fit Jev

This made me rethink where AI actually fits into security engineering. For purely engineering work, forget about ChatGPT or Claude. TypeSafe AI just released Jev, and I think it’s going to change how we build AI into security workflows. Instead of asking an LLM to “investigate Show more

TypeSafe AI
TypeSafe AI
@typesafeai

we are officially out of stealth! join the frontier and get access to Jev on our website (link on profile)

Reply

A million judged questions

Inferring Jev's internals from 1,000 calls

Jevの内部アーキテクチャを推測している技術記事(Jev’s Architecture Unmasked)からメモ。 ・本記事はJevのAPIを約1万回の呼び出して、内部構造を推測したもの ・従来の言語モデルを用いた分類やルーティングでは、トークンを1文字ずつ逐次生成するために膨大な無駄な計算コストが発生していた。 Show more

Reply

The open System One roundup

Jev 发布没几天,开源社区已经开始疯狂复刻了🔥 最值得推荐的五个模型: 1、Laya 421M:原生决策模型,支持 Mac 2、Decider-2B:最像 Jev,基于 Qwen3.5 3、NanoJev 0.6B:专门的 Decision Head 4、Reflex:Qwen3.5 + Direct Logits 5、System-One 4B:专门做概率校准 Show more

小墨同学
小墨同学
@xiaomovps

Jev 刚发布没几天,开源社区就出现了同款🔥 Decider-2B模型,是基于 Qwen3.5-2B 做了特殊调整 它和 Jev 模型是一样的 只做选择 评分和判断 不是文本类的 LLM 模型 但两者还是有几个明显区别: 1、模型 Jev:闭源 System One Model Decider:Qwen3.5-2B,约 1.9B 参数,Apache 2.0 开源 2、价格

Image
Reply

Jev lands on the Vercel AI Gateway

Vercel ships the AI SDK provider for Jev