Skip to content

Records at Rest: Documents, Inboxes, and Risk Queues

Jev labels a record that already exists. Code moves, flags, or holds it. How the document, mail, and transfer tools keep the action out of the model.

Published
Category
Tools & CLIs
Resources cited
6
Reading time
3 min
Tags#classification#email#security

A support ticket, a PDF packet, and a withdrawal are the same shape of work: a record already exists, and something in code has to move, flag, or hold it. Jev's job in these tools is the label. The action stays in ordinary code, where it can be unit-tested and turned off.

Documents

DocJev takes a PDF, DOCX, or PPTX, extracts page text locally with LiteParse, and asks Jev either for one category or for the boundaries inside a packet. On its 40-document pilot both engines classified every original correctly, and Jev split 7 of 8 packets exactly. The measured packet keeps a ten-page statistical release together and separates two adjacent Treasury results that share a category. DOCX and PPTX need LibreOffice. Local OCR does not make the model call offline: TYPESAFE_API_KEY is still required.

The design detail worth copying is the rule file. Categories have stable ids and descriptions of purpose, and other is added if you leave it out. Splitting instructions say what counts as a new document rather than a continuation page. That text is the question. The page ranges come back for code to export.

Inboxes

jev-mail-classifier connects over IMAP and turns each category you configure into one Noul inside a single call per message, so a mail can match more than one label. Each category has its own threshold and its own actions: tag, move, flag, mark read, or call a webhook. Processed mail is marked with the IMAP keyword $JevProcessed, not the Seen flag, and a run defaults to 25 unprocessed messages, newest first. run --dry-run prints the actions without touching the mailbox.

Jevmail is the read-only cousin. It sorts Gmail into five trays with the AI SDK's evaluate API and stores trays in local SQLite. The Gmail scope cannot send, label, or archive. The README's figure is about 1,000 emails in a minute for around 3 cents. Classifying 1,500 real emails is a single person's screen recording of the same idea, useful as a field note and not as a method.

The shared rule is the threshold per label. Urgent at 0.7 and spam at 0.9 is a policy, and it belongs next to the action, not inside the question.

Risk queues

jev-guard screens a deposit or withdrawal. Blacklist, sanctioned counterparty, and amounts over $500k are hard rules and never reach the model. Jev then answers risk level, pattern, freeze probability, and a suggested route. The route is written to the audit log and never executed. Timeouts and server errors become manual review, not a pass. On 100 author-labeled synthetic events the rules-only baseline scored 68% and the Jev gate scored 50%, with zero missed freezes. Route and the composed action agreed 72% of the time. The README treats that 28% gap as the reason the matrix stays in Java.

jev-suite applies the same split to four checks: whether a sponsored video matched the brief, where a candidate misses the job, whether an edit preserved the source, and what a rental listing still needs confirmed. Missing evidence fails the check closed. A missing model escalates. Calibration runners refuse to score mock answers unless you pass a flag that watermarks the file.

What to copy

  • Put arithmetic, caps, and vetoes in code. Send the digested record as state.
  • Log the model's suggestion beside the action you actually took.
  • Degrade toward a person. A timeout that auto-passes is the failure mode these projects are written to avoid.
  • Dry-run before the first live mailbox or the first live transfer.

The tools category has the CLIs. The thresholds worth trusting are the ones you fit on labels from your own queue, which is the subject of confidence gates.

From the directory

The resources behind this article.

468GitHub stars
Classifies and splits PDFs, DOCX, and PPTX with local page text from LiteParse and category decisions from Jev.
Tools & CLIs#community#python#document
19GitHub stars
An IMAP client that turns each inbox category into a Noul, then tags, moves, flags, or calls a webhook when the probability clears that category's threshold.
Tools & CLIs#community#python#email
86GitHub stars
A local, read-only Gmail triage app that sorts mail into five trays with Jev: each message becomes one experimental_evaluate call answering tray, urgency, and human-written questions, and 1,000 emails sort in about a minute for around 3 cents.
Cookbooks#community#typescript#privacy
vogel ran Jev over 1,500 of his own emails to test classification quality and posted the results as a video, calling it the most impressive model he has tried for the task.
CookbooksVideo#community#x#video
Community
A Java gateway that asks Jev four questions about a deposit or withdrawal, then applies a direction-aware routing matrix in code.
Tools & CLIs#community#java#security
33GitHub stars
Four Java checkers — sponsored video, hiring fit, edit fidelity, and rental listings — that keep thresholds and vetoes in unit-tested code.
Tools & CLIs#community#java#verification
Back to all articles

More articles

A base URL is the whole client change. What still works when Ollaya, Lichen, verdict, Rizzo Flow, or jevos answers POST /v1/systemone on your machine.
Models & Reimplementations#open-models#local#api
Read article
A game tick is a state plus a legal action list. How Clash Royale, Minecraft, and the earlier Mario, Pokémon, StarCraft, and drone loops ask Jev one step at a time.
Demos & Experiments#gaming#agents#control
Read article

From the community

Posts from builders shipping with Jev right now.

Follow @typesafeai

Full Jev video tutorial

The case against Jev-scored compaction

This is a terrible compaction strategy that fundamentally doesn't understand how compaction and context management work. Seems like a lot of people are confused so let's break this down. 1. Compaction isn't a filter The role of compaction is to clean up history to keep the Show more

tamara
tamara
@tamarajtran

found the perfect use case for @typesafeai Jev: instant compaction in 2026, why is compaction still a summarization prompt? Jev can make it instant by scoring every tool call and dropping what’s irrelevant

Reply

Classifying rows in DuckDB

A playable 16-judgment demo

AI multiple choice, not essay writing

Screening agent actions with Jev

Tested TypeSafe’s Jev (no-text, probability-only model) as an AI agent safety monitor. Checking each action first worked well caught most attacks with almost no false blocks, and much faster than Gemini.

Image
Image
Image
Diogo Almeida
Diogo Almeida
TypeSafe AI
@CompleteSkeptic

After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x

Reply