Winnow 12B — VRAM & /v1/systemone setup
Written by Jakub Rusinowski · Last updated
EldanRing's 12B decision model on Gemma 4 12B, shipped as GGUF and served by the llama.cpp-based winnow-inference server, which also handles chat and vision.
Winnow 12B needs about 8 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.
Call Winnow 12B
Served by winnow-inference (default port 8091). This model does not run in Ollama.
http://localhost:8091/v1/systemone
Setup for servers other than Ollama →
Build a request for this model Decision models guide
Other decision models: Nimble 9B · Tev1 4B · Tev1 0.8B · Winnow E4B · Decider 2B · Decider 4B · Decider 35B-A3B (NVFP4) · JevK5 4B · Intern-Decision 4B · AutoJev 27B · Laya
Hardware fit
Weights plus overhead plus the KV cache at a 8,192-token prompt, on NVIDIA RTX 4090 (24 GB). A publisher build is sized from its file; the other rows are modelled at a standard quant. Decision requests are short, so no long-context figure is shown.
| Quant | Memory | VRAM | Fit |
|---|---|---|---|
| BF16 Publisher build · 23.83 GB file | 25.9 GB | Offload | |
| Q4_K_M 4.83 bpw · modelled quant | 9.3 GB | Fits | |
| Q6_K 6.56 bpw · modelled quant | 11.9 GB | Fits | |
| Q8_0 8.50 bpw · modelled quant | 14.8 GB | Fits |
Published file size: Q8_0 12.67 GB · BF16 23.83 GB. A download size from the model publisher — not a VRAM requirement.
Runs in the Winnow llama.cpp-based server (default port 8091): /v1/systemone + /v1/chat/completions + vision. Tested: RTX 5070 Ti 16 GB (Q8_0, 64K ctx); Apple Silicon 24 GB+. Not usable on Ollama's /v1/systemone.
How it was scored
Two different suites on two different scales. Never compare the numbers across the two cards.
Decision Index 0.2.1 (snapshot 2026-09-28)
Decision Index 0.2.1- ECE (lower is better)
- 0.1679
- Median compute
- 72.5 ms on 1x NVIDIA RTX PRO 6000 (96 GB)
Community-maintained; not affiliated with the model authors.
Source ↗Decision models are not ranked on chat, creative or coding scores.
Specifications
Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source.
- Parameters
- 11.96B
- Context window
- Not published
- Architecture
- Fine-tune of google/gemma-4-12B
- Provider
- EldanRing
- Licence
- Apache-2.0
- Specified at
- Q4_K_M
- System RAM
- 18 GB
- Record updated
- 2026-09-30
Commercial use permitted. No usage restrictions beyond attribution.
Other Winnow sizes
Winnow 12B — frequently asked questions
What is Winnow 12B?
Winnow 12B is a decision model: you send it a state and typed questions (choice, yes/no/unknown, or a score) and it returns one answer per question with a probability for every option. It is not a chat model.
How do I run Winnow 12B locally?
Winnow 12B does not run in Ollama. Served by winnow-inference (default port 8091). This model does not run in Ollama. See the setup guide for servers other than Ollama.
How much memory does Winnow 12B need?
About 8 GB for the weights plus overhead at Q4_K_M, before the prompt's KV cache. Decision prompts are short, so the cache stays small.
How accurate is Winnow 12B?
It has two separate published scores on two different suites — the author's own benchmark and the community Decision Index 0.2.1. They are not comparable with each other, and neither is a calibration guarantee.