Open Jev reproductions beyond Ollama
Written by Jakub Rusinowski · Last updated
Ollama ships three decision models. The community has published about 70 open reproductions of Jev, and several score far higher on the Decision Index — Winnow-12B (50.0) fits a 16 GB card at Q8_0, and AutoJev-27B (56.4) comes within 1.5 points of Jev (57.9) but needs about 49 GiB for its BF16 weights. Almost all of them expose the same POST /v1/systemone request shape, so code you wrote against Ollama keeps working — you change the URL and port.
When to leave Ollama's three models
Stay with nimble / tev1 if they're accurate enough on 50–100 of your own labelled examples. Move on when:
- your questions are harder than routing — multi-step judgments, knowledge questions, long policies — where the Index gap between Nimble (39.6) and the leaders (50–57) shows up;
- you need more than 26 options in a
choice(Ollama's limit; Decider and Jev go to 255); - you need non-English input (Laya's multilingual router covers 100+ languages; most others are English-first);
- you need images in the state (AutoJev, Intern-Decision and Winnow accept images; Ollama's endpoint doesn't).
The shortlist
Only models with open weights, a permissive licence, and a documented local runtime made the cut. Decision Index 0.2.1 snapshot of 28 September 2026 (chance-corrected; 0 = random, 100 = perfect).
| Model | Index | Base | Weights you run | Runtime | Licence |
|---|---|---|---|---|---|
| AutoJev-27B | 56.4 | Qwen3.8-27B | ~49 GiB BF16 | autojev-serve (CUDA) | Apache-2.0 weights, MIT code |
| Winnow-12B | 50.0 | Gemma 4 12B | 12.7 GB Q8_0 · 23.8 GB BF16 GGUF | Winnow llama.cpp server (CUDA, Metal) | Apache-2.0 |
| Decider 35B-A3B (NVFP4) | 47.1 | Qwen3.5-35B-A3B | 19.6 GB NVFP4 | vLLM / TensorRT-LLM on Blackwell | Apache-2.0 |
| Decider 4B | 40.7 | Qwen3.5-4B | 8.4 GB BF16 | decider Python package + server | Apache-2.0 |
| Winnow-E4B | 39.9 | Gemma 4 E4B | 8.0 GB Q8_0 GGUF | Winnow llama.cpp server | Apache-2.0 |
| JevK5 | 38.8 | Qwen3.5-4B | ~9 GB BF16, or GGUF Q4_K_M/Q8_0 | jevk5-serve, or llama.cpp via JevK5-GGUF | Apache-2.0 |
| Intern-Decision-4B | 37.8 | Qwen3.5-4B | Safetensors (Transformers) | InternLM reference code | Apache-2.0 |
| Decider 2B | 29.0 | Qwen3.5-2B | 3.8 GB BF16 | decider package + server | Apache-2.0 |
| Decider 0.8B | not on the board | Qwen3.5-0.8B | 1.4 GB BF16 | decider package + server | Apache-2.0 |
| Laya (EN + multilingual) | 6.0 ² | ModernBERT-large / mmBERT | 421M / 322M params | pip install laya | Apache-2.0 |
| Jev 1.13 (reference) | 57.9 | closed | hosted only | TypeSafe API | proprietary |
² Laya is an encoder built for fast routing and guardrails in many languages, not a general reasoner — its Index score reflects that. It's on the list because nothing else covers 100+ languages locally.
Left off on purpose: the top-ranked Surogate Rune 26B-A4B (57.4) because its Hugging Face repo listed no GGUF file when we checked on 30 September; the JPT family because it is CC-BY-NC-4.0 (non-commercial); and every "inference technique" entry on the Index, because those are decoding methods applied to stock models rather than weights you download. We'll add Rune once its files are published.
| Model | Runs with | Published size | Decision Index 0.2.1 |
|---|---|---|---|
| Decider 0.8B preview Mapika | decider.serve | 1.4 GB (BF16) | — |
| Decider 2B Mapika | decider.serve | 3.8 GB (BF16) | 28.97 |
| Winnow E4B EldanRing | winnow-inference :8091 | 8.01 GB (Q8_0) | 39.89 |
| Decider 4B Mapika | decider.serve | 8.4 GB (BF16) | 40.7 |
| JevK5 4B alibiserikbay | jevk5-serve :8090 | 9 GB (BF16) | 38.81 |
| Winnow 12B EldanRing | winnow-inference :8091 | 12.67 GB (Q8_0) | 50.02 |
| Decider 35B-A3B (NVFP4) Mapika | decider.serve | 19.6 GB (NVFP4) | 47.11 |
| AutoJev 27B denis-pplx | autojev-serve :8000 | 52.6 GB (BF16 (49 GiB)) | 56.4 |
| Intern-Decision 4B InternLM | reference Python code | — | 37.81 |
| Laya ConvAI Innovations | laya-serve / Python Router | — | 6.04 |
Fit on your hardware: set your GPU or unified memory in the hardware analyzer and this page shows whether each model fits it. Check your hardware
Pattern 1 — GGUF + llama.cpp-based server (Winnow)
The most hardware-flexible path. Winnow's own server is built on llama.cpp and exposes /v1/systemone and a normal /v1/chat/completions from the same loaded model — one download gives you a decision model and a chat/vision model.
Tested setups, per its README: Linux with an NVIDIA GPU (RTX 5070 Ti, CUDA 12.8+), Apple Silicon with 24 GB or more of unified memory, and an NVIDIA Docker container. The Q8_0 build was run with a 64K context on a 16 GB RTX 5070 Ti.
# bash — Linux (NVIDIA) or macOS (Apple Silicon)
git clone https://github.com/EldanRing/winnow-inference
cd winnow-inference
python3 scripts/setup.py # downloads the model (~12.9 GB)
python3 scripts/serve.py --profile 5070ti-64k # NVIDIA
# python3 scripts/serve.py --profile apple-silicon # MacIt listens on http://127.0.0.1:8091. Your Ollama request works unchanged apart from the URL:
curl http://127.0.0.1:8091/v1/systemone \
-H 'Content-Type: application/json' \
-d '{"state":"The store accepts returns within 30 days. This item was bought 12 days ago.",
"questions":{"eligible":{"type":"noul","instructions":"Is the item eligible for return?"}}}'Check the repo's quickstart for whether model is required and for the current profile names — this project is two weeks old and moves quickly.
Pattern 2 — Python package with a built-in server (Decider, JevK5, AutoJev)
These ship their own inference code and a small HTTP server that speaks TypeSafe's format. You need a working PyTorch install with GPU support first.
Decider — the most downloaded open decision model (over 255,000 downloads of the 2B in its first two weeks). Four sizes share one interface.
# Python 3.11+; needs torch, transformers>=5 and flash-linear-attention
# (it runs without flash-linear-attention, but several times slower)
from decider.infer import Decider # the decider/ package ships inside the model repo
d = Decider("Mapika/decider-2b")
d.decide(
"My card was charged twice for the same purchase.",
[{"question": "Which department should handle this?", "options": ["billing", "technical support", "sales"]},
{"question": "Does this need a refund action?", "options": ["no", "yes"]}],
)For an HTTP endpoint, run decider.serve with DECIDER_MODEL pointing at the downloaded model folder; the official typesafe-sdk works against it unchanged with TYPESAFE_BASE_URL set to that server. Decider allows 2–255 options per question, and abstain_below=t makes it return None instead of an answer when confidence is under t.
JevK5 — one of the best-calibrated 4B models on the Index (expected calibration error 0.027). English only; inputs over 16,384 tokens are refused rather than cut.
# bash — needs a CUDA GPU with ~9 GB free for BF16
jevk5-serve --model alibiserikbay/JevK5 --port 8090No NVIDIA card? The alibiserikbay/JevK5-GGUF builds run under llama.cpp on NVIDIA, AMD, Intel and Apple GPUs or the CPU. Per its card, the Q8_0 file matches the full-precision answer on 229 of 231 public JevBench items, and Q4_K_M on 221.
AutoJev-27B — the strongest open model you can actually download today. Its card asks for a GPU with room for about 49 GiB of BF16 weights plus runtime overhead — in practice a 64 GB-class or larger workstation card.
# bash — Python 3.12+ and uv
git clone https://github.com/denis-pplx/autojev.git
cd autojev
uv sync --frozen --python 3.12
uv run hf download denis-pplx/autojev-27b --local-dir checkpoints/selected
AUTOJEV_CHECKPOINT=checkpoints/selected uv run autojev-serveIt serves a browser playground at http://localhost:8000 and POST /v1/systemone with optional base64 images. Set AUTOJEV_API_KEY to turn on authentication — do that before exposing it to a network.
Pattern 3 — A tiny encoder for routing in any language (Laya)
pip install laya # extras: laya[serve], laya[mcp], laya[langchain], laya[onnx]from laya import Router
router = Router() # downloads a checkpoint on first use
result = router.predict(
"Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan.",
{"department": {"type": "choice", "instructions": "Which department should handle this?",
"criteria": {"billing": "invoices, payments, refunds",
"technical": "bugs, outages, system errors",
"other": "everything else"}}},
)
print(result["answers"]["department"]["choice"], result["routing"]["model"])The router detects the script and language and sends non-English text to the multilingual checkpoint automatically. It's the fastest thing on this page and the least capable at reasoning — use it as a first-pass router, not a judge.
Which one for your hardware?
| Your machine | Start with | Why |
|---|---|---|
| CPU only, 8–16 GB RAM | tev1:0.8b (Ollama) or Laya | Smallest files; Laya is an encoder and cheap on CPU |
| 8 GB GPU | Decider 2B, or JevK5-GGUF Q4_K_M | Small BF16 or quantised 4B |
| 12 GB GPU | nimble (Ollama) or Decider 4B | Best accuracy that fits comfortably |
| 16 GB GPU (e.g. RTX 5070 Ti, 4080) | Winnow-12B Q8_0 | Highest Index score that fits; tested on exactly this class of card |
| Apple Silicon, 24–48 GB | Winnow-12B (Metal profile) | Its README lists 24 GB+ unified memory for Apple Silicon |
| 24–32 GB Blackwell card | Decider 35B-A3B NVFP4 | 19.6 GB, needs vLLM/TensorRT-LLM with NVFP4 support |
| 64 GB-class or larger workstation GPU | AutoJev-27B | Within 1.5 Index points of Jev |
This table is guidance from the published file sizes and each project's tested setups. Your analyzer result is the verdict for your exact machine. If the model you need doesn't fit what you have, renting a GPU by the hour is often cheaper than buying for an occasional batch job — compare in the cost calculator.
Train your own
Two of these projects publish their full recipe:
- Together AI's Tev1 — data recipe and training code at github.com/togethercomputer/tev1, with a Together blog post on training your own classifier for about $17.
- Decider — code, data registry and training scripts at github.com/Mapika/decider.
A decision model fine-tuned on a few thousand of your own labelled decisions usually beats a bigger general one on your task. Our fine-tuning guides cover the basics of LoRA on a single GPU.
Frequently asked questions
Are any of these actually Jev?
No. TypeSafe hasn't released Jev's weights or its RLCD training method. Every model here is an independent reproduction of Jev's interface — state plus typed questions in, probabilities out — trained with its own data and method.
Why does a 27B model score like a hosted model?
Most open reproductions fine-tune a strong base model (Qwen3.8-27B, Gemma 4) to read the answer directly off the logits of each option in one forward pass. A capable base plus a narrow task goes a long way. The Index measures accuracy; it does not measure latency under your load or cost at your volume.
Can I load these GGUFs into Ollama?
Not for /v1/systemone today. Ollama's endpoint requires a model packaged for its System One runner; it rejects models it can't score. Use each project's own server.
How do I compare two models fairly?
Label 100 of your own real decisions, run both models on the same requests, and compare accuracy and how often a high-probability answer was wrong. The thresholds guide walks through it. --- Index figures from data/index.json of the Jev Decision Index Space (edition 0.2.1, generated 28 September 2026). Runtimes and sizes from each model's Hugging Face card and GitHub README, read 30 September 2026. The Index is community-maintained and not affiliated with TypeSafe AI.
Verified against: Decision Index 0.2.1 data (generated 28 Sep 2026); Hugging Face model cards and GitHub READMEs read 30 Sep 2026.