Skip to content
Patterns

Logit readout

Reading the next-token probability of each declared label, then renormalising, so no answer text is generated.

Category
Patterns
Also known as
—
Related terms
4
Directory entries
3
Docs
—
Added
2026-09-28

Definition

Lichen, verdict, and Rizzo Flow all answer by scoring label tokens in one forward pass. Lichen lists options twice in rotated order and shrinks confidence when the two readings disagree. Verdict keeps the probability mass that landed on the labels before renormalising, because a tiny mass can still rank options after it is scaled to 1.

The mechanism is not Jev's. These projects say so. Output tokens in their sample responses are zero. Accuracy and calibration still depend on the GGUF you loaded and on any temperature or shrink you fitted.

Tagspatternsopen-models

From the directory

39GitHub stars
A local /v1/systemone server that reads label probabilities from a GGUF chat model, with prompt repetition and a confidence shrink toward uniform.
Models & Reimplementations#community#python#open-models
7GitHub stars
A Python server in front of llama-server that reads one-token label probabilities and exposes them as POST /v1/systemone.
Models & Reimplementations#community#python#open-models
689GitHub stars
A local server that reads next-token probabilities from a fine-tuned Spark model and returns Choice, Score, and Noul answers with zero generated tokens.
Models & Reimplementations#community#python#open-models

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

How Jev makes agents faster and cheaper

500 emails for 3.5 cents

Headless Chromium agent

Custom Jev-style models for agent workflows

Prediction: millionaires will be made using custom Jev style models (parallel constrained decoding) to make the agent systems companies already run more token efficient. Let me explain with a scenario: Imagine a company already has an agent workflow running where an llm reviews Show more

Harsha Gundala
Harsha Gundala
@harshagundal

They were building in stealth for 2 years, I was building in stealth for 2 hours… Happy to open source Qwen-2.5-1B-RLCD, 5x faster on-device inference for JSON workloads that need to be type-safe. ⚡️Demo below on a M4 MacBook⚡️ every LLM has the ability to efficiently batch

Reply

Live viral post analyzer

Jev plays Tetris: 134 lines in two minutes