Local servers trade the hosted API's operations for a model you run. Ollaya pulls several open checkpoints. Rizzo Flow downloads a fine-tuned Spark model. Lichen and verdict read a GGUF through llama.cpp. jevos is a yes/no model aimed at a laptop CPU.
The state stays on the machine only if nothing else forwards it. Jeview is the opposite pattern: a local proxy that still calls TypeSafe. A local server also drops hosted rate limits and picks up your GPU, RAM, and license constraints.
A local server that reads next-token probabilities from a fine-tuned Spark model and returns Choice, Score, and Noul answers with zero generated tokens.
Jev is cool not because it re-invented classification, but because it makes ARBITRARY classification into a type-safe programmable primitive. A general purpose zero shot decision model whose native interface is RUNTIME-DEFINED typed decisions, optimized for that exact interface
cocktail peanut
@cocktailpeanut
If you called Yann LeCun an idiot for saying we need to move beyond LLMs and build something new, you are banned from using Jev.
Jev solved local harness/model routing
I use a combination of Claude Code, Codex and Opencode as my local agentic stack and routing to other harnesses was always enforced in the system prompt/rules
With a deterministic hook that Claude Code can decide before delegation, JevShow more
Acabo de terminar la implementación de @typesafeai + Chromium Headless para que mis agentes puedan navegar por internet a una buena velocidad!
En este ejemplo le pido al agente que entre a la página del término "Café" en Wikipedia y navegue por los hipervínculos hasta terminarShow more
Prediction: millionaires will be made using custom Jev style models (parallel constrained decoding) to make the agent systems companies already run more token efficient.
Let me explain with a scenario:
Imagine a company already has an agent workflow running where an llm reviewsShow more
Harsha Gundala
@harshagundal
They were building in stealth for 2 years, I was building in stealth for 2 hours…
Happy to open source Qwen-2.5-1B-RLCD, 5x faster on-device inference for JSON workloads that need to be type-safe.
⚡️Demo below on a M4 MacBook⚡️
every LLM has the ability to efficiently batch
Just created this with Jev by @typesafeai. A live viral post analyzer. As soon as you stop typing for .5 seconds it analyzes the viral potential.
Going to try and actually make this good, will need to scrape a lot of twitter data...
Notice how it also categorizes the tweetShow more