Skip to content
JevDirectory.org
CommunityRepos & SDKs2 starsVerified 2026-09-22

pytest-jev

A pytest plugin for semantic assertions: jev.expect sends many holds and lacks claims about one text to Jev in a single request and fails visibly when a probability is uncertain.

Category
Repos & SDKs
Published by
Community
Author
allebee
Added
2026-09-22
Tagscommunitypythonevaluationguardrails

Highlights

  • All claims about one text go to Jev in one request; holds needs p >= 0.8 and lacks needs p <= 0.2, and unsure claims fail.
  • Answers cache in .pytest_cache, so rerunning unchanged tests makes no requests and gives identical verdicts.
  • On 12 example tests it matched Claude Sonnet 5's verdicts in 5.3 s versus 27.1 s, at $0.00017 versus $0.0192 per run.
  • The jev fixture exposes holds, lacks, expect, choice, and score methods; each sends one request and returns an assertable result.
  • Without a key, tests that use the jev fixture are skipped by default.

Quickstart

python
def test_refund_reply(jev):
    reply = support_bot("I was charged twice for order #1042.")
    jev.expect(
        reply,
        holds=["apologizes to the customer", "says the duplicate payment was refunded"],
        lacks=["blames the customer", "asks for a password or a full card number"],
    )

Watch out

MIT-licensed and not affiliated with TypeSafe; needs Python 3.10+ and pytest 7.4+, plus a TYPESAFE_API_KEY or OPENROUTER_API_KEY, and asserted text is sent to the provider.

More like this

14.3kGitHub stars
Open multilingual System 1 decision models with published checkpoints for choice, score and noul questions, plus a router that dispatches each request to the right checkpoint in one forward pass.
Repos & SDKs#community#python#open-models
Communitylaya
1.9kGitHub stars
An open 0.6B replica of Jev that turns states and questions into full probability distributions without decoding answer tokens, trained and evaluated on Maze, Snake, and ViZDoom.
Repos & SDKs#community#python#open-models
CommunityNanoJev
1kGitHub stars
Train a small model that chooses among a changing list of text options, one probability per option in a single pass. Includes Doom, chess, and Wikispeedia demos.
Repos & SDKs#community#training#research
Communityjevlike
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Reranking 33,047 catalog entries

拿 Jev 做搜索重排,我先泼一盆冷水:单独用,它没打赢向量检索 TypeSafe 的 Jev 这阵子很火,一堆项目拿它做重排。我们在 Agent Skills Hub 的 33,047 条目录上认真测了一次,164 条中英文真实查询,9,831 对分级标注,整套只花了 2.6 美元 三个结论 01|单独重排,约等于没赢 Jev 重排 bge-m3 Show more

Jason Zhu
Jason Zhu
@GoSailGlobal

有美团、阿里的老哥嘛? 试试加一路召回、重排(离线、近实时实现),我觉得有奇效 他在文本理解上 跟之前机器学习、llm很不一样 还能自动打标签做特征

Reply

Six uses that stuck after 60 days

Security decisions that fit Jev

This made me rethink where AI actually fits into security engineering. For purely engineering work, forget about ChatGPT or Claude. TypeSafe AI just released Jev, and I think it’s going to change how we build AI into security workflows. Instead of asking an LLM to “investigate Show more

TypeSafe AI
TypeSafe AI
@typesafeai

we are officially out of stealth! join the frontier and get access to Jev on our website (link on profile)

Reply

A million judged questions

Inferring Jev's internals from 1,000 calls

Jevの内部アーキテクチャを推測している技術記事(Jev’s Architecture Unmasked)からメモ。 ・本記事はJevのAPIを約1万回の呼び出して、内部構造を推測したもの ・従来の言語モデルを用いた分類やルーティングでは、トークンを1文字ずつ逐次生成するために膨大な無駄な計算コストが発生していた。 Show more

Reply

The open System One roundup

Jev 发布没几天,开源社区已经开始疯狂复刻了🔥 最值得推荐的五个模型: 1、Laya 421M:原生决策模型,支持 Mac 2、Decider-2B:最像 Jev,基于 Qwen3.5 3、NanoJev 0.6B:专门的 Decision Head 4、Reflex:Qwen3.5 + Direct Logits 5、System-One 4B:专门做概率校准 Show more

小墨同学
小墨同学
@xiaomovps

Jev 刚发布没几天,开源社区就出现了同款🔥 Decider-2B模型,是基于 Qwen3.5-2B 做了特殊调整 它和 Jev 模型是一样的 只做选择 评分和判断 不是文本类的 LLM 模型 但两者还是有几个明显区别: 1、模型 Jev:闭源 System One Model Decider:Qwen3.5-2B,约 1.9B 参数,Apache 2.0 开源 2、价格

Image
Reply