Skip to content

Rizzo Flow

A local server that reads next-token probabilities from a fine-tuned Spark model and returns Choice, Score, and Noul answers with zero generated tokens.

Category
Models & Reimplementations
Format
—
Published by
Community
Author
Rizzo-AI-Academy
Added
2026-09-28
Last verified
2026-09-28

Rizzo-AI-Academy/rizzo-flow

Organization repository on GitHub

View on GitHub

The open, local take on Jev: typed decisions from an LLM, without generating a single token

GitHub stars
689
Forks
41
Primary language
Python
License
apache-2.0
Last pushed
Updated Sep 2026

Repo stats from the GitHub API, cached Sep 2026.Homepage

Highlights

  • Serves POST /v1/systemone on llama.cpp across Metal, CUDA, Vulkan, ROCm, SYCL, and CPU.
  • Default weights are a LoRA on Spark-X2.5-4B. On typed-decisions at Q8_0 the README lists accuracy 0.648, Brier 0.205, and ECE 0.112.
  • The same table lists the untuned 4B at 0.574 accuracy and Jev 1.13.0's dataset-card accuracy at 0.727.
  • A Snake demo records about 150 ms per decision and zero generated tokens.
  • The README states probabilities are uncalibrated unless you fit them on your own data.

Quickstart

bash
uv sync --locked
uv run rizzo download
uv run rizzo serve

Watch out

Apache-2.0. Independent of TypeSafe; it does not claim to reproduce Jev's architecture or RLCD. The 4B Q8_0 download is about 4.4 GB.

Reactions & coverage

Posts, threads, and videos about this entry from around the web.

X: deepseek-v4.1-flash-jev

More like this

39GitHub stars
A local /v1/systemone server that reads label probabilities from a GGUF chat model, with prompt repetition and a confidence shrink toward uniform.
Models & Reimplementations#community#python#open-models
7GitHub stars
A Python server in front of llama-server that reads one-token label probabilities and exposes them as POST /v1/systemone.
Models & Reimplementations#community#python#open-models
1.1kGitHub stars
A local yes/no decision model that speaks Jev's wire format on a laptop CPU, refusing Choice and Score until those question types ship.
Models & Reimplementations#community#python#open-models

From the guides

Original write-ups that draw on this entry.

A base URL is the whole client change. What still works when Ollaya, Lichen, verdict, Rizzo Flow, or jevos answers POST /v1/systemone on your machine.
Models & Reimplementations#open-models#local#api
Read article
Back to all resources

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Instant compaction with Jev

A Claude session from 1M to 86K tokens

This is actually insane. This uses @typesafeai Jev model, as a plugin in Claude to review all the un-nesseasary tool calls, and it takes 1s to run! Like, literally, 1 second to take my Claude session from nearly 1M to ... 86K tokens! 😮 Ask your claude to install it and be  Show more

Image
Image
tamara
tamara
@tamarajtran

found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant

Reply

Vercel's fx safety reviewer, 18x faster

We're seeing extraordinary results from @typesafeai. Default mode in 𝚏𝚡 is auto, with a safety reviewer analyzing every command. That reviewer runs on GPT Luna today. Jev is up to 18x faster (p95) *and* more accurate. It's coming to @vercel AI Gateway and likely new default.

Pranit
Pranit
Vercel
@fazxes

We benchmarked fx auto mode (safety) classifier with @typesafeai's Jev. tl;dr: ~5-18x faster and more accurate than 𝚐𝚙𝚝-𝟻.𝟼-𝚕𝚞𝚗𝚊, our current top choice

Image
Reply

Jev lands on OpenRouter

700 leads scored for $0.09

Beating Gemini Flash Lite on an eval