Hosted Jev is one endpoint. A growing set of local servers speak the same POST /v1/systemone shape, so the change in a client is the base URL. Weights, training, and how to read their benchmarks are covered in the open reimplementations. This is the serving question: what still works when the model is on your machine.
What "compatible" means
Ollaya is the broadest host. It pulls open decision models by name and serves /v1/systemone and /v1/models. Setting TYPESAFE_BASE_URL=http://localhost:11435 is the documented switch for the official SDK. Its recommended winnow:e4b is listed at 0.722 accuracy on typed decisions against Jev's 0.738, at 89 ms for five questions on an RTX 4090. Without an NVIDIA GPU the README points you at laya, which is the CPU path.
Lichen and verdict are narrower: each sits on a GGUF you already have and reads the probability of label tokens instead of generating text. Lichen's own table on 231 public JevBench items has gemma-4-26B-A4B at 207 correct against 200 for Jev 1.13, with a median 93 ms on one laptop GPU. The README also says the edge is not significant (p = 0.09), and the server does not check API keys. Verdict's conformance check is an unmodified jev-ultrafast client pointed at TYPESAFE_BASE_URL. It records option mass before renormalising, and it refuses a model that cannot use its alphabet. It also says a small Gemma build reported confidence 1.000 on wrong answers, so a local gate on raw confidence is not safe until you fit one.
Rizzo Flow fine-tunes Spark-X2.5 and serves the same endpoint on llama.cpp. On the typed-decisions set at Q8_0 its README lists accuracy 0.648, Brier 0.205, and ECE 0.112, against 0.574 for the untuned 4B and 0.727 for Jev's dataset-card accuracy. It says the probabilities are uncalibrated unless you fit them yourself.
What is not a full swap
jevos speaks the wire format and then refuses anything that is not a Noul, with HTTP 422. On 2,000 unseen yes/no items it reports accuracy 0.815 against 0.927 for Jev. Three questions that share one state take about 165 ms on CPU, and optional criteria are accepted and ignored. Use it when every question is yes or no. Do not point a Choice or Score client at it and expect a soft fallback.
@receptron/laya is not a server. It is a Node package that runs Convai's Laya checkpoint through ONNX Runtime, with answers the README says match the Python reference to four decimal places. Three questions take about 140 ms on a warm Apple-silicon CPU. The state truncates at 512 tokens on the English checkpoint. Call systemOne from your process, or serve Laya through Ollaya if you want the HTTP shape.
System One Connector is the agent-side piece: one evaluate tool that can target hosted Jev or TYPESAFE_BASE_URL for a local Laya or CLM server. A non-empty TYPESAFE_API_KEY selects that route even when the local server does not check keys.
How to read a local number
- Separate wire compatibility from quality. A client that parses the JSON is not evidence the answers match Jev.
- Read the hardware column. Lichen's 93 ms is a local GPU. Jev's 665 ms in that table includes the network.
- Treat calibration as unfitted unless the README shows a held-out fit. Rizzo Flow and verdict both say so.
- Check the question types. jevos is Noul only. Lichen caps a choice at 62 labels. Laya's Node build recommends about 20 options.
The models category collects these servers next to the weight projects. If the decision has to leave the machine, stay on the hosted API and keep the base URL pointed at TypeSafe.
