Skip to content

Using Jev for evals in Datadog

Datadog's walkthrough of one Jev rubric that scores every criterion in a single request, then runs as online evals on live spans and as offline experiments.

Category
Guides & Articles
Format
Article
Published by
Community
Author
—
Added
2026-09-28
Last verified
2026-09-28

Highlights

  • One rubric scores every criterion in a single Jev request, for both online evals and offline experiments.
  • Online evals use LLMObs.submit_evaluation. Offline evals use LLMObs.experiment and BaseEvaluator.
  • The post requires ddtrace 4.5.0 or newer. It says there are no preview builds or private endpoints.
  • Notebooks in DataDog/llm-observability: 1-jev-rubric, 2-online-evals, and 3-experiments.
  • The rubric notebook needs a TypeSafe key. The agent trace also needs OpenAI and Datadog keys.

Quickstart

bash
git clone https://github.com/DataDog/llm-observability
cd llm-observability/typesafe-jev
pip install -r requirements.txt

Watch out

Datadog says to measure agreement with human reviewers and repeatability on your own traffic before relying on the scores.

More like this

Sam Reghenzi's product-taxonomy benchmark: a Jev Choice at every level versus a gpt-5.2 agent that descends the tree and can be sent back by a judge.
Guides & ArticlesArticle#community#article#evaluation
Community
A release guide with eval tables, pricing, curl and Python examples, and fit boundaries for Jev's early access.
Guides & ArticlesArticle#community#article#pricing
Community
A heavily footnoted claims audit that separates documented Jev facts from marketing, including the seed round, third-party tests, and missing transparency.
Guides & ArticlesArticle#community#article#analysis
Community
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Jev lands on the Vercel AI Gateway

Vercel ships the AI SDK provider for Jev

Computer use at 155× cheaper than Opus 5

i built computer use using @typesafeai ! it is 155x cheaper than opus 5, ~20x faster, and generalizes across OS's more on how it works in the vid & thread below:

Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply

Foreman keeps coding agents on task

Unclutter: an ad and slop blocker that runs on Jev

An open-source BS meter for debates and investor calls

🚨 Open Source Jev BS meter you can use this to analyze any debate / investor call / interview / sales pitch / podcast video fact check live , for example this dario interview cost 60 Jev calls / 111K tokens / $0.0047 github.com/ChetasLua/jevm…

Chetaslua
Chetaslua
@chetaslua

🚨 I gave the Trump vs Kamala debate a live BS meter using Jev every sentence, both candidates, 5 yes/no questions each 1,191 Jev calls / 1.18M tokens / 415 ms median total cost : $0.0497 same questions for both, clips picked by one fixed rule, not a fact-check

Reply