Decision models: run Jev-style System One models on your own hardware

Written by Jakub Rusinowski · Last updated

A decision model reads some text (the "state") and a set of typed questions, and returns a probability for every allowed answer in a single scoring pass. It never generates text. TypeSafe AI launched the category with its hosted model Jev on 15 September 2026. Since Ollama 0.35 (29 September 2026) you can run three decision models locally — nimble, tev1 and tev1:0.8b — through a new /v1/systemone endpoint. More than 70 open reproductions of Jev are now tracked on a community leaderboard, and the strongest of them score close to Jev on it.

This hub is the "can I run it and how" guide for the whole category.

Pick your starting point

If you are…Start here
New to the idea — "what is this and why should I care?"Decision models explained, with infographics
Ready to run one tonight on a laptopSet up Nimble or Tev1 with Ollama
A developer who needs more accuracy than Ollama's three modelsOpen Jev reproductions beyond Ollama: Winnow, Decider, JevK5, Laya
Wiring probabilities into production codeConfidence, thresholds and calibration
Writing your first requestDecision request builder — generates curl, Python and JS for you
Stuck on an errorOllama /v1/systemone errors and fixes

What a decision model is good for

Think of it as a smart if-statement you can drop into ordinary code. Your program stays in control of the logic; the model only answers narrow questions about messy input:

  • Route a support ticket to billing, technical or other.
  • Gate an action: "does the customer explicitly ask for a refund?" → probability of yes.
  • Score urgency, quality or risk on a rubric you define.
  • Moderate or guard an LLM's input or output before it reaches a user.
  • Choose a tool or a skill from a fixed list for an agent.

It is the wrong tool for anything that needs words out: summaries, drafts, explanations, chat. For those, you still want a regular local LLM — see the model library.

Which decision model can your hardware run?

ModelRuns withPublished sizeDecision Index 0.2.1
Tev1 0.8B
Together AI
ollama pull tev1:0.8b812 MB (Ollama build)12.85
Decider 0.8B preview
Mapika
decider.serve1.4 GB (BF16)—
Decider 2B
Mapika
decider.serve3.8 GB (BF16)28.97
Tev1 4B
Together AI
ollama pull tev14.5 GB (Ollama build)29.24
Winnow E4B
EldanRing
winnow-inference :80918.01 GB (Q8_0)39.89
Decider 4B
Mapika
decider.serve8.4 GB (BF16)40.7
JevK5 4B
alibiserikbay
jevk5-serve :80909 GB (BF16)38.81
Nimble 9B
Bespoke Labs
ollama pull nimble9.5 GB (Ollama build)39.57
Winnow 12B
EldanRing
winnow-inference :809112.67 GB (Q8_0)50.02
Decider 35B-A3B (NVFP4)
Mapika
decider.serve19.6 GB (NVFP4)47.11
AutoJev 27B
denis-pplx
autojev-serve :800052.6 GB (BF16 (49 GiB))56.4
Intern-Decision 4B
InternLM
reference Python code—37.81
Laya
ConvAI Innovations
laya-serve / Python Router—6.04

Fit on your hardware: set your GPU or unified memory in the hardware analyzer and this page shows whether each model fits it. Check your hardware

The short version, by published download size of the build you would actually run:

ModelRuns in Ollama?Published sizeDecision Index 0.2.1
Tev1 0.8B (Together AI, experimental)✅ tev1:0.8b812 MB12.9
Decider 2B (Mapika)Own server3.8 GB (BF16)29.0
Tev1 4B (Together AI, experimental)✅ tev14.5 GB29.2
Winnow-E4B (EldanRing)Own llama.cpp server8.0 GB (Q8_0)39.9
Decider 4B (Mapika)Own server8.4 GB (BF16)40.7
Nimble 9B (Bespoke Labs)✅ nimble9.5 GB39.6 ¹
Winnow-12B (EldanRing)Own llama.cpp server12.7 GB (Q8_0)50.0
AutoJev-27BOwn server~49 GiB (BF16)56.4
Jev 1.13 (TypeSafe, hosted, closed weights)——57.9

¹ The Index scores the Bespoke-Nimble-9B-v2 repository; which checkpoint the Ollama nimble tag packages is not stated in one place. Decision Index snapshot: 28 September 2026. Scores are chance-corrected (0 = random guessing, 100 = perfect). Download size is a floor, not a verdict — use the hardware analyzer for a fit check that includes runtime overhead.

The wider field: the top 25 downloadable models on the Decision Index

Decision Index 0.2.1 · snapshot 28 Sep 2026 · community-maintained, not affiliated with TypeSafe AI. Top 25 of 54 entrants with downloadable weights (70 entrants in all; inference-technique entries excluded). Full leaderboard ↗
Index rankModelParamsBaseLicenceDecision Index 0.2.1ECELinks
—Jev 1.13 (TypeSafe AI) hosted, closed weights——Proprietary API57.910.074reference only
1Surogate Rune 26B-A4B v325.8Bgoogle/gemma-4-26B-A4B-itApache-2.057.440.1196weights ↗ · code ↗
3AutoJev-27B27.8BQwen/Qwen3.8-27BApache-2.056.40.0178weights ↗ · code ↗
5Jebadiah 27B27.8BQwen/Qwen3.8-27BApache-2.054.670.014weights ↗ · code ↗
6Eikos-27B27.8BQwen/Qwen3.8-27BMIT53.130.0574weights ↗ · code ↗
9Winnow-12B12Bgoogle/gemma-4-12BApache-2.050.020.1679weights ↗ · code ↗
12Decider 35B-A3B36BQwen/Qwen3.5-35B-A3B-BaseApache-2.047.110.0226weights ↗
13JPT-9B9.7BQwen/Qwen3.5-9B-BaseCC-BY-NC-4.0 NC46.890.0773weights ↗ · code ↗
14Decision 1.0 Lux9.7BQwen/Qwen3.5-9B-BaseApache-2.043.490.076weights ↗
15JPT-4B4.7BQwen/Qwen3.5-4B-BaseCC-BY-NC-4.0 NC43.040.0793weights ↗ · code ↗
16Jet v6.24.7BQwen/Qwen3.5-4B-BaseApache-2.042.60.1411weights ↗ · code ↗
17Xor36BQwen/Qwen3.6-35B-A3BApache-2.041.480.0149weights ↗
18Hopper (G) 1.24.7BQwen/Qwen3.5-4B-BaseOther40.770.0927weights ↗ · code ↗
19Decider 4B4.7BQwen/Qwen3.5-4B-BaseApache-2.040.70.0837weights ↗ · code ↗
20Jev-Omni12Bgoogle/gemma-4-12BApache-2.040.530.161weights ↗
22Winnow-E4B8Bgoogle/gemma-4-E4BApache-2.039.890.058weights ↗ · code ↗
23Bespoke Nimble 9B v29.7BQwen/Qwen3.5-9B-BaseApache-2.039.570.024weights ↗ · code ↗
24JevK54.7BQwen/Qwen3.5-4B-BaseApache-2.038.810.0268weights ↗ · code ↗
25lev4.7BQwen/Qwen3.5-4B-BaseApache-2.038.540.0638weights ↗ · code ↗
26Kev 9B9.7BQwen/Qwen3.5-9B-BaseApache-2.038.480.1378weights ↗ · code ↗
27Intern-Decision-4B4.7BQwen/Qwen3.5-4B-BaseApache-2.037.810.0278weights ↗ · code ↗
29NeoHorse-Jev-4B4.7BQwen/Qwen3.5-4B-BaseApache-2.036.750.1039weights ↗ · code ↗
30Solomon v1.127.8BQwen/Qwen3.8-27BApache-2.036.430.0807weights ↗
31Kev 4B4.7BQwen/Qwen3.5-4B-BaseApache-2.034.640.1756weights ↗ · code ↗
32Decision 1.0 Nox4.7BQwen/Qwen3.5-4B-BaseNot stated34.360.1399weights ↗
35open-jev (pngwn)4.7BQwen/Qwen3.5-4B-BaseCC-BY-NC-4.0 NC29.910.0589weights ↗

Index scores are chance-corrected (0 = random guessing, 100 = perfect); ECE is expected calibration error, lower is better. These are not the model authors' own benchmark numbers and are never plotted with them.

The honest caveats, up front

  1. "Can't hallucinate" means "can't answer off-schema." The answer is always one of your options. It can still be the wrong option.
  2. Ollama's confidence is not a correctness score. Ollama's own API reference defines it as how concentrated the probability distribution is, and says it is not calibrated correctness. Test thresholds on your own data before you trust them — here's how.
  3. Two benchmarks, two stories. On Bespoke Labs' 13 human-labelled datasets, Nimble scores 75.7% against Jev's 76.0%. On the broader community Decision Index, Nimble v2 scores 39.6 against Jev's 57.9. Both are fair; they measure different things. We never put them on one chart.
  4. Score is the weakest primitive. Nimble's public-benchmark accuracy is 81.6% on choice, 80.2% on yes/no and 54.6% on score questions.
  5. Everything here is weeks old. Jev launched on 15 September, Ollama support on 29 September. Expect model versions, limits and leaderboard positions to move. Every page in this cluster shows its last-verified date.

Keep going

Frequently asked questions

Is Jev open source?

No. Jev is a hosted API from TypeSafe AI with closed weights, and its training method (RLCD, Reinforcement Learning for Calibrated Decisions) is described only at blog-post level. Everything you can run locally is an independent open reproduction of the interface, not Jev itself.

Do I need a GPU?

Not for the smallest models. tev1:0.8b is an 812 MB download, and decision requests are short, so CPU inference is usable for low volumes. For interactive, sub-100 ms decisions you want the model fully in GPU or unified memory — Nimble 9B averaged 91 ms per decision on an M5 Max MacBook Pro in Ollama's own demo.

Can I use a normal chat model like Llama or Qwen on /v1/systemone?

No. Ollama requires a local model trained for System One with compatible GGUF weights and a scoring-capable runner. Chat models, cloud models and MLX/Safetensors models are rejected.

Why not just ask a chat model for JSON?

You can, and for low volumes it works. A decision model returns a probability for every option, guarantees the schema, answers all your questions in one request, and is much faster per decision. Chat models give you one sampled string that you still have to parse and validate.

Which one should I start with?

nimble (9.5 GB download) if the analyzer says it fits your GPU or unified memory with room to spare; tev1 (4.5 GB) if it doesn't; tev1:0.8b (812 MB) for a CPU-only box or a quick test. Then measure on your own examples before deciding whether you need a larger open reproduction.

In this cluster