Skip to content
JevDirectory.org
CommunityPractices & Patterns38 starsVerified 2026-09-22

typesafe-ai-benchmark

Side-by-side benchmark of TypeSafe Jev, Qwen 3.8 27B on Cerebras, and a local Needle 3 across seven synthetic workloads, recording validated outputs, mistakes, latency, tokens, and estimated cost. Raw exports and per-scene limitations are published.

Category
Practices & Patterns
Published by
Community
Author
iammrduncan
Added
2026-09-22
Tagscommunitytypescriptbenchmarksevaluationcalibration

Highlights

  • Runs Jev, Qwen 3.8 27B on Cerebras, and a local Needle 3 side by side across seven synthetic workloads.
  • Jev logged 176 ms p50 and $0.011919 estimated cost against Qwen's 215 ms and $0.310581.
  • All three approaches validate outputs before applying simulated actions.
  • Needle 3 was measured separately on an Apple M4 Pro at 864 tok/s median native decode.
  • Raw exports, source hashes, and per-scene limitations are published under docs/benchmarks.

Quickstart

bash
test -f .env || cp .env.example .env
npm run setup:needle   # optional local Needle 3 lane
npm run check          # type checks, lint, offline tests, both builds

Watch out

MIT-licensed. Requires Node 22 and npm 10, API keys for the cloud lanes, and the Hugging Face CLI for the optional local Needle lane; the Needle timing is not a controlled speed ranking.

More like this

104GitHub stars
A personal-assistant agent with 100 mocked tools that measures how many steps a task takes when the LLM picks the tool versus when Jev picks it before every model step.
Practices & Patterns#community#typescript#agents
74GitHub stars
Independent cross-model benchmark for Jev-class decision models, running 534 frozen cases per complete entrant with scoring code and a four-axis score of accuracy, calibration, latency and cost.
Practices & Patterns#community#python#benchmarks
Communityjevbench
An experiment comparing Jev with GPT-5.6 Luna, GPT-5.6 Terra, and Claude Sonnet 4.6 as evaluators of five frozen weather-agent runs, measuring pass-or-fail accuracy against human labels plus variance, cost, and latency.
Practices & Patterns#community#python#evaluation
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Instant compaction with Jev

A Claude session from 1M to 86K tokens

This is actually insane. This uses @typesafeai Jev model, as a plugin in Claude to review all the un-nesseasary tool calls, and it takes 1s to run! Like, literally, 1 second to take my Claude session from nearly 1M to ... 86K tokens! 😮 Ask your claude to install it and be  Show more

Image
Image
tamara
tamara
@tamarajtran

found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant

Reply

Vercel's fx safety reviewer, 18x faster

We're seeing extraordinary results from @typesafeai. Default mode in 𝚏𝚡 is auto, with a safety reviewer analyzing every command. That reviewer runs on GPT Luna today. Jev is up to 18x faster (p95) *and* more accurate. It's coming to @vercel AI Gateway and likely new default.

Pranit
Pranit
Vercel
@fazxes

We benchmarked fx auto mode (safety) classifier with @typesafeai's Jev. tl;dr: ~5-18x faster and more accurate than 𝚐𝚙𝚝-𝟻.𝟼-𝚕𝚞𝚗𝚊, our current top choice

Image
Reply

Jev lands on OpenRouter

700 leads scored for $0.09

Beating Gemini Flash Lite on an eval