A support ticket, a PDF packet, and a withdrawal are the same shape of work: a record already exists, and something in code has to move, flag, or hold it. Jev's job in these tools is the label. The action stays in ordinary code, where it can be unit-tested and turned off.
Documents
DocJev takes a PDF, DOCX, or PPTX, extracts page text locally with LiteParse, and asks Jev either for one category or for the boundaries inside a packet. On its 40-document pilot both engines classified every original correctly, and Jev split 7 of 8 packets exactly. The measured packet keeps a ten-page statistical release together and separates two adjacent Treasury results that share a category. DOCX and PPTX need LibreOffice. Local OCR does not make the model call offline: TYPESAFE_API_KEY is still required.
The design detail worth copying is the rule file. Categories have stable ids and descriptions of purpose, and other is added if you leave it out. Splitting instructions say what counts as a new document rather than a continuation page. That text is the question. The page ranges come back for code to export.
Inboxes
jev-mail-classifier connects over IMAP and turns each category you configure into one Noul inside a single call per message, so a mail can match more than one label. Each category has its own threshold and its own actions: tag, move, flag, mark read, or call a webhook. Processed mail is marked with the IMAP keyword $JevProcessed, not the Seen flag, and a run defaults to 25 unprocessed messages, newest first. run --dry-run prints the actions without touching the mailbox.
Jevmail is the read-only cousin. It sorts Gmail into five trays with the AI SDK's evaluate API and stores trays in local SQLite. The Gmail scope cannot send, label, or archive. The README's figure is about 1,000 emails in a minute for around 3 cents. Classifying 1,500 real emails is a single person's screen recording of the same idea, useful as a field note and not as a method.
The shared rule is the threshold per label. Urgent at 0.7 and spam at 0.9 is a policy, and it belongs next to the action, not inside the question.
Risk queues
jev-guard screens a deposit or withdrawal. Blacklist, sanctioned counterparty, and amounts over $500k are hard rules and never reach the model. Jev then answers risk level, pattern, freeze probability, and a suggested route. The route is written to the audit log and never executed. Timeouts and server errors become manual review, not a pass. On 100 author-labeled synthetic events the rules-only baseline scored 68% and the Jev gate scored 50%, with zero missed freezes. Route and the composed action agreed 72% of the time. The README treats that 28% gap as the reason the matrix stays in Java.
jev-suite applies the same split to four checks: whether a sponsored video matched the brief, where a candidate misses the job, whether an edit preserved the source, and what a rental listing still needs confirmed. Missing evidence fails the check closed. A missing model escalates. Calibration runners refuse to score mock answers unless you pass a flag that watermarks the file.
What to copy
- Put arithmetic, caps, and vetoes in code. Send the digested record as state.
- Log the model's suggestion beside the action you actually took.
- Degrade toward a person. A timeout that auto-passes is the failure mode these projects are written to avoid.
- Dry-run before the first live mailbox or the first live transfer.
The tools category has the CLIs. The thresholds worth trusting are the ones you fit on labels from your own queue, which is the subject of confidence gates.
