Jev is now available on the @vercel AI Gateway vercel.com/ai-gateway/mod…
Using Jev for evals in Datadog
Datadog's walkthrough of one Jev rubric that scores every criterion in a single request, then runs as online evals on live spans and as offline experiments.
- Category
- Guides & Articles
- Format
- Article
- Published by
- Community
- Author
- —
- Added
- 2026-09-28
- Last verified
- 2026-09-28
Highlights
- One rubric scores every criterion in a single Jev request, for both online evals and offline experiments.
- Online evals use LLMObs.submit_evaluation. Offline evals use LLMObs.experiment and BaseEvaluator.
- The post requires ddtrace 4.5.0 or newer. It says there are no preview builds or private endpoints.
- Notebooks in DataDog/llm-observability: 1-jev-rubric, 2-online-evals, and 3-experiments.
- The rubric notebook needs a TypeSafe key. The agent trace also needs OpenAI and Datadog keys.
Quickstart
git clone https://github.com/DataDog/llm-observability
cd llm-observability/typesafe-jev
pip install -r requirements.txtWatch out
Datadog says to measure agreement with human reviewers and repeatability on your own traffic before relying on the scores.
More like this
From the community
Posts from builders shipping with Jev right now.
Jev lands on the Vercel AI Gateway
Vercel ships the AI SDK provider for Jev
Jev from @typesafeai is on AI Gateway. Build agents that decide, route, score, and stop in milliseconds: 𝚊𝚠𝚊𝚒𝚝 𝚎𝚟𝚊𝚕𝚞𝚊𝚝𝚎({ 𝚖𝚘𝚍𝚎𝚕: '𝚝𝚢𝚙𝚎𝚜𝚊𝚏𝚎-𝚊𝚒/𝚓𝚎𝚟', 𝚜𝚝𝚊𝚝𝚎, 𝚚𝚞𝚎𝚜𝚝𝚒𝚘𝚗𝚜, }); vercel.com/changelog/type…
Computer use at 155× cheaper than Opus 5
i built computer use using @typesafeai ! it is 155x cheaper than opus 5, ~20x faster, and generalizes across OS's more on how it works in the vid & thread below:
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x
Foreman keeps coding agents on task
I just open sourced Foreman: a software factory foreman built with @typesafeai's Jev. Coding agents work the factory floor. Foreman watches them, continuously assessing progress, completeness, tests, drift, and verification, and intervenes when needed. GitHub: Show more
Unclutter: an ad and slop blocker that runs on Jev
introducing Unclutter: a smart ad + slop blocker with Jev 🤓 it auto cleans up pages from slop elements: ⬖ ads ⬖ cookie banners ⬖ upsells ⬖ bs dialogs BYOK. open source + free, download below 👇
An open-source BS meter for debates and investor calls
🚨 Open Source Jev BS meter you can use this to analyze any debate / investor call / interview / sales pitch / podcast video fact check live , for example this dario interview cost 60 Jev calls / 111K tokens / $0.0047 github.com/ChetasLua/jevm…
🚨 I gave the Trump vs Kamala debate a live BS meter using Jev every sentence, both candidates, 5 yes/no questions each 1,191 Jev calls / 1.18M tokens / 415 ms median total cost : $0.0497 same questions for both, clips picked by one fixed rule, not a fact-check


