Decision models: run Jev-style System One models on your own hardware
Written by Jakub Rusinowski · Last updated
A decision model reads some text (the "state") and a set of typed questions, and returns a probability for every allowed answer in a single scoring pass. It never generates text. TypeSafe AI launched the category with its hosted model Jev on 15 September 2026. Since Ollama 0.35 (29 September 2026) you can run three decision models locally — nimble, tev1 and tev1:0.8b — through a new /v1/systemone endpoint. More than 70 open reproductions of Jev are now tracked on a community leaderboard, and the strongest of them score close to Jev on it.
This hub is the "can I run it and how" guide for the whole category.
Pick your starting point
| If you are… | Start here |
|---|---|
| New to the idea — "what is this and why should I care?" | Decision models explained, with infographics |
| Ready to run one tonight on a laptop | Set up Nimble or Tev1 with Ollama |
| A developer who needs more accuracy than Ollama's three models | Open Jev reproductions beyond Ollama: Winnow, Decider, JevK5, Laya |
| Wiring probabilities into production code | Confidence, thresholds and calibration |
| Writing your first request | Decision request builder — generates curl, Python and JS for you |
| Stuck on an error | Ollama /v1/systemone errors and fixes |
What a decision model is good for
Think of it as a smart if-statement you can drop into ordinary code. Your program stays in control of the logic; the model only answers narrow questions about messy input:
- Route a support ticket to billing, technical or other.
- Gate an action: "does the customer explicitly ask for a refund?" → probability of yes.
- Score urgency, quality or risk on a rubric you define.
- Moderate or guard an LLM's input or output before it reaches a user.
- Choose a tool or a skill from a fixed list for an agent.
It is the wrong tool for anything that needs words out: summaries, drafts, explanations, chat. For those, you still want a regular local LLM — see the model library.
Which decision model can your hardware run?
| Model | Runs with | Published size | Decision Index 0.2.1 |
|---|---|---|---|
| Tev1 0.8B Together AI | ollama pull tev1:0.8b | 812 MB (Ollama build) | 12.85 |
| Decider 0.8B preview Mapika | decider.serve | 1.4 GB (BF16) | — |
| Decider 2B Mapika | decider.serve | 3.8 GB (BF16) | 28.97 |
| Tev1 4B Together AI | ollama pull tev1 | 4.5 GB (Ollama build) | 29.24 |
| Winnow E4B EldanRing | winnow-inference :8091 | 8.01 GB (Q8_0) | 39.89 |
| Decider 4B Mapika | decider.serve | 8.4 GB (BF16) | 40.7 |
| JevK5 4B alibiserikbay | jevk5-serve :8090 | 9 GB (BF16) | 38.81 |
| Nimble 9B Bespoke Labs | ollama pull nimble | 9.5 GB (Ollama build) | 39.57 |
| Winnow 12B EldanRing | winnow-inference :8091 | 12.67 GB (Q8_0) | 50.02 |
| Decider 35B-A3B (NVFP4) Mapika | decider.serve | 19.6 GB (NVFP4) | 47.11 |
| AutoJev 27B denis-pplx | autojev-serve :8000 | 52.6 GB (BF16 (49 GiB)) | 56.4 |
| Intern-Decision 4B InternLM | reference Python code | — | 37.81 |
| Laya ConvAI Innovations | laya-serve / Python Router | — | 6.04 |
Fit on your hardware: set your GPU or unified memory in the hardware analyzer and this page shows whether each model fits it. Check your hardware
The short version, by published download size of the build you would actually run:
| Model | Runs in Ollama? | Published size | Decision Index 0.2.1 |
|---|---|---|---|
| Tev1 0.8B (Together AI, experimental) | ✅ tev1:0.8b | 812 MB | 12.9 |
| Decider 2B (Mapika) | Own server | 3.8 GB (BF16) | 29.0 |
| Tev1 4B (Together AI, experimental) | ✅ tev1 | 4.5 GB | 29.2 |
| Winnow-E4B (EldanRing) | Own llama.cpp server | 8.0 GB (Q8_0) | 39.9 |
| Decider 4B (Mapika) | Own server | 8.4 GB (BF16) | 40.7 |
| Nimble 9B (Bespoke Labs) | ✅ nimble | 9.5 GB | 39.6 ¹ |
| Winnow-12B (EldanRing) | Own llama.cpp server | 12.7 GB (Q8_0) | 50.0 |
| AutoJev-27B | Own server | ~49 GiB (BF16) | 56.4 |
| Jev 1.13 (TypeSafe, hosted, closed weights) | — | — | 57.9 |
¹ The Index scores the Bespoke-Nimble-9B-v2 repository; which checkpoint the Ollama nimble tag packages is not stated in one place. Decision Index snapshot: 28 September 2026. Scores are chance-corrected (0 = random guessing, 100 = perfect). Download size is a floor, not a verdict — use the hardware analyzer for a fit check that includes runtime overhead.
The wider field: the top 25 downloadable models on the Decision Index
| Index rank | Model | Params | Base | Licence | Decision Index 0.2.1 | ECE | Links |
|---|---|---|---|---|---|---|---|
| — | Jev 1.13 (TypeSafe AI) hosted, closed weights | — | — | Proprietary API | 57.91 | 0.074 | reference only |
| 1 | Surogate Rune 26B-A4B v3 | 25.8B | google/gemma-4-26B-A4B-it | Apache-2.0 | 57.44 | 0.1196 | weights ↗ · code ↗ |
| 3 | AutoJev-27B | 27.8B | Qwen/Qwen3.8-27B | Apache-2.0 | 56.4 | 0.0178 | weights ↗ · code ↗ |
| 5 | Jebadiah 27B | 27.8B | Qwen/Qwen3.8-27B | Apache-2.0 | 54.67 | 0.014 | weights ↗ · code ↗ |
| 6 | Eikos-27B | 27.8B | Qwen/Qwen3.8-27B | MIT | 53.13 | 0.0574 | weights ↗ · code ↗ |
| 9 | Winnow-12B | 12B | google/gemma-4-12B | Apache-2.0 | 50.02 | 0.1679 | weights ↗ · code ↗ |
| 12 | Decider 35B-A3B | 36B | Qwen/Qwen3.5-35B-A3B-Base | Apache-2.0 | 47.11 | 0.0226 | weights ↗ |
| 13 | JPT-9B | 9.7B | Qwen/Qwen3.5-9B-Base | CC-BY-NC-4.0 NC | 46.89 | 0.0773 | weights ↗ · code ↗ |
| 14 | Decision 1.0 Lux | 9.7B | Qwen/Qwen3.5-9B-Base | Apache-2.0 | 43.49 | 0.076 | weights ↗ |
| 15 | JPT-4B | 4.7B | Qwen/Qwen3.5-4B-Base | CC-BY-NC-4.0 NC | 43.04 | 0.0793 | weights ↗ · code ↗ |
| 16 | Jet v6.2 | 4.7B | Qwen/Qwen3.5-4B-Base | Apache-2.0 | 42.6 | 0.1411 | weights ↗ · code ↗ |
| 17 | Xor | 36B | Qwen/Qwen3.6-35B-A3B | Apache-2.0 | 41.48 | 0.0149 | weights ↗ |
| 18 | Hopper (G) 1.2 | 4.7B | Qwen/Qwen3.5-4B-Base | Other | 40.77 | 0.0927 | weights ↗ · code ↗ |
| 19 | Decider 4B | 4.7B | Qwen/Qwen3.5-4B-Base | Apache-2.0 | 40.7 | 0.0837 | weights ↗ · code ↗ |
| 20 | Jev-Omni | 12B | google/gemma-4-12B | Apache-2.0 | 40.53 | 0.161 | weights ↗ |
| 22 | Winnow-E4B | 8B | google/gemma-4-E4B | Apache-2.0 | 39.89 | 0.058 | weights ↗ · code ↗ |
| 23 | Bespoke Nimble 9B v2 | 9.7B | Qwen/Qwen3.5-9B-Base | Apache-2.0 | 39.57 | 0.024 | weights ↗ · code ↗ |
| 24 | JevK5 | 4.7B | Qwen/Qwen3.5-4B-Base | Apache-2.0 | 38.81 | 0.0268 | weights ↗ · code ↗ |
| 25 | lev | 4.7B | Qwen/Qwen3.5-4B-Base | Apache-2.0 | 38.54 | 0.0638 | weights ↗ · code ↗ |
| 26 | Kev 9B | 9.7B | Qwen/Qwen3.5-9B-Base | Apache-2.0 | 38.48 | 0.1378 | weights ↗ · code ↗ |
| 27 | Intern-Decision-4B | 4.7B | Qwen/Qwen3.5-4B-Base | Apache-2.0 | 37.81 | 0.0278 | weights ↗ · code ↗ |
| 29 | NeoHorse-Jev-4B | 4.7B | Qwen/Qwen3.5-4B-Base | Apache-2.0 | 36.75 | 0.1039 | weights ↗ · code ↗ |
| 30 | Solomon v1.1 | 27.8B | Qwen/Qwen3.8-27B | Apache-2.0 | 36.43 | 0.0807 | weights ↗ |
| 31 | Kev 4B | 4.7B | Qwen/Qwen3.5-4B-Base | Apache-2.0 | 34.64 | 0.1756 | weights ↗ · code ↗ |
| 32 | Decision 1.0 Nox | 4.7B | Qwen/Qwen3.5-4B-Base | Not stated | 34.36 | 0.1399 | weights ↗ |
| 35 | open-jev (pngwn) | 4.7B | Qwen/Qwen3.5-4B-Base | CC-BY-NC-4.0 NC | 29.91 | 0.0589 | weights ↗ |
Index scores are chance-corrected (0 = random guessing, 100 = perfect); ECE is expected calibration error, lower is better. These are not the model authors' own benchmark numbers and are never plotted with them.
The honest caveats, up front
- "Can't hallucinate" means "can't answer off-schema." The answer is always one of your options. It can still be the wrong option.
- Ollama's
confidenceis not a correctness score. Ollama's own API reference defines it as how concentrated the probability distribution is, and says it is not calibrated correctness. Test thresholds on your own data before you trust them — here's how. - Two benchmarks, two stories. On Bespoke Labs' 13 human-labelled datasets, Nimble scores 75.7% against Jev's 76.0%. On the broader community Decision Index, Nimble v2 scores 39.6 against Jev's 57.9. Both are fair; they measure different things. We never put them on one chart.
- Score is the weakest primitive. Nimble's public-benchmark accuracy is 81.6% on choice, 80.2% on yes/no and 54.6% on score questions.
- Everything here is weeks old. Jev launched on 15 September, Ollama support on 29 September. Expect model versions, limits and leaderboard positions to move. Every page in this cluster shows its last-verified date.
Keep going
- The model library — every decision model has its own page with hardware fit and sources
- Installing Ollama
- Check your hardware
Frequently asked questions
Is Jev open source?
No. Jev is a hosted API from TypeSafe AI with closed weights, and its training method (RLCD, Reinforcement Learning for Calibrated Decisions) is described only at blog-post level. Everything you can run locally is an independent open reproduction of the interface, not Jev itself.
Do I need a GPU?
Not for the smallest models. tev1:0.8b is an 812 MB download, and decision requests are short, so CPU inference is usable for low volumes. For interactive, sub-100 ms decisions you want the model fully in GPU or unified memory — Nimble 9B averaged 91 ms per decision on an M5 Max MacBook Pro in Ollama's own demo.
Can I use a normal chat model like Llama or Qwen on /v1/systemone?
No. Ollama requires a local model trained for System One with compatible GGUF weights and a scoring-capable runner. Chat models, cloud models and MLX/Safetensors models are rejected.
Why not just ask a chat model for JSON?
You can, and for low volumes it works. A decision model returns a probability for every option, guarantees the schema, answers all your questions in one request, and is much faster per decision. Chat models give you one sampled string that you still have to parse and validate.
Which one should I start with?
nimble (9.5 GB download) if the analyzer says it fits your GPU or unified memory with room to spare; tev1 (4.5 GB) if it doesn't; tev1:0.8b (812 MB) for a CPU-only box or a quick test. Then measure on your own examples before deciding whether you need a larger open reproduction.
In this cluster
- Run Jev-style decision models locally with Ollama (Nimble, Tev1)Set up Nimble or Tev1 in Ollama 0.35 and send your first /v1/systemone request with curl, PowerShell, Python or JavaScript. macOS, Windows and Linux.
- Open Jev reproductions beyond Ollama: Winnow, Decider, JevK5, AutoJev, LayaThe open decision models closest to Jev on the Decision Index, what hardware each needs, and how to serve them on a /v1/systemone endpoint your code already speaks.
- Decision-model confidence, thresholds and calibration — use the probabilities safelyOllama's confidence score isn't a correctness score. How to measure calibration on your own data and set accept, escalate and abstain thresholds for local decision models.
- Ollama /v1/systemone errors and fixesWhat a 400, 404 or 413 from the decision endpoint means, and how to fix each.
- Decision models explained, with infographicsWhat Jev started, how the open reproductions compare, and the two benchmarks.
- Decision request builderGenerates a valid curl, PowerShell, Python or JavaScript request and flags limit violations before you send it.