Skip to content
JevDirectory.org
Evaluation & Training

Workflow evals

TypeSafe's published evaluations of four automation workflows, comparing Jev and frontier LLMs as structured workflows versus single prompts.

Category
Evaluation & Training
Also known as
—
Related terms
4
Directory entries
48
Docs
evals.typesafe.ai
Added
2026-09-24

Definition

evals.typesafe.ai covers security incidents, agent trace observability, invoice processing, and customer service. Models are scored as workflows of Noul, Choice, and Score questions against consensus labels from GPT-6 Astra and Claude Fable 5.1 at high thinking; Jev averaged 67.8% accuracy at $0.0004 and 0.4 seconds per case.

The headline finding is not about Jev alone: every model was more accurate, cheaper, and faster in a workflow than with the same policy as a single prompt. The evals are vendor-run, and the harness and reference labels are assumed rather than independently audited.

Tagsevaluationbenchmarks

From the directory

Published evaluations of four automation workflows, security incidents, agent trace observability, invoice processing, and customer service, comparing Jev and frontier LLMs as structured workflows versus single prompts.
Practices & PatternsDocs#official#benchmarks#evaluation
Official
The index of TypeSafe's worked examples, from parallel questions and reranking to guardrails, date extraction, and self-consistency, each with datasets and measured results.
Sites & GuidesDocs#official#docs#cookbook
Official
Scott Chacon runs hosted Jev against two local decision models, laya and kev, in a Tetris match and a GitHub Settings filter, with a public test repo.
Practices & PatternsVideo#community#video#open-models
Community
A head-to-head benchmark that prices 200 products from descriptions with Jev against GPT-5.6 Luna and GPT-4.1-Nano.
Practices & PatternsVideo#community#video#benchmarks
Community
A release guide with eval tables, pricing, curl and Python examples, and fit boundaries for Jev's early access.
Sites & GuidesArticle#community#article#pricing
Community
A heavily footnoted claims audit that separates documented Jev facts from marketing, including the seed round, third-party tests, and missing transparency.
Sites & GuidesArticle#community#article#analysis
Community

42 more matching entries in the full directory.

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Game levels generated in real time

900 images in 40 seconds

Intent-based search in Gmail

Jev is really good at intent-based search! How it looks in Gmail: (for a huge inbox you'd prob let semantic search / embeddings pull first but still much better experience)

nader dabit
nader dabit
Cognition
@dabit3

Another crazy @typesafeai Jev example: Predictive spreadsheets Spreadsheets recalculate numbers, not meaning. Jev reads intent. Type "Urgency" at the top of a column and, as you type, it figures out you want each row rated from "no follow-up needed" to "urgent" in ~100 ms.

Reply

End-to-end tests run by agents

An always-on assistant with no wake word

400 companies matched to one candidate