Decider 35B-A3B (NVFP4) — VRAM & /v1/systemone setup
Written by Jakub Rusinowski · Last updated
Mapika's 35B mixture-of-experts decision model (3B active) on Qwen3.5-35B-A3B-Base, in an NVFP4 build for NVIDIA Blackwell GPUs via vLLM or TensorRT-LLM.
Decider 35B-A3B (NVFP4) needs about 23 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.
Call Decider 35B-A3B (NVFP4)
Served by decider.serve. This model does not run in Ollama.
Setup for servers other than Ollama →
Build a request for this model Decision models guide
Other decision models: Nimble 9B · Tev1 4B · Tev1 0.8B · Winnow 12B · Winnow E4B · Decider 2B · Decider 4B · JevK5 4B · Intern-Decision 4B · AutoJev 27B · Laya
Affiliate links — we may earn a commission if you sign up, at no extra cost to you.
Hardware fit
Weights plus overhead plus the KV cache at a 8,192-token prompt, on NVIDIA RTX 4090 (24 GB). A publisher build is sized from its file; the other rows are modelled at a standard quant. Decision requests are short, so no long-context figure is shown.
| Quant | Memory | VRAM | Fit |
|---|---|---|---|
| NVFP4 Publisher build · 19.6 GB file | 22.3 GB | Tight | |
| BF16 Publisher build · 65 GB file | 67.7 GB | Too big | |
| Q4_K_M 4.83 bpw · modelled quant | 24.4 GB | Offload | |
| Q6_K 6.56 bpw · modelled quant | 32.2 GB | Offload | |
| Q8_0 8.50 bpw · modelled quant | 40.9 GB | Offload |
Published file size: NVFP4 19.6 GB · BF16 65 GB. A download size from the model publisher — not a VRAM requirement.
NVFP4 build targets NVIDIA Blackwell via vLLM or TensorRT-LLM. BF16 sibling repo Mapika/decider-35b-a3b is 65 GB.
How it was scored
Two different suites on two different scales. Never compare the numbers across the two cards.
Decision Index 0.2.1 (snapshot 2026-09-28)
Decision Index 0.2.1- ECE (lower is better)
- 0.0226
- Median compute
- 101.4 ms on 1x NVIDIA RTX PRO 6000 (96 GB)
Community-maintained; not affiliated with the model authors.
Source ↗Decision models are not ranked on chat, creative or coding scores.
Specifications
Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source.
- Parameters
- 35.95B
- Context window
- Not published
- Architecture
- Fine-tune of Qwen/Qwen3.5-35B-A3B-Base (mixture-of-experts)
- Provider
- Mapika
- Licence
- Apache-2.0
- Specified at
- Q4_K_M
- System RAM
- 46 GB
- Record updated
- 2026-09-30
Commercial use permitted. No usage restrictions beyond attribution.
Other Decider sizes
Decider 35B-A3B (NVFP4) — frequently asked questions
What is Decider 35B-A3B (NVFP4)?
Decider 35B-A3B (NVFP4) is a decision model: you send it a state and typed questions (choice, yes/no/unknown, or a score) and it returns one answer per question with a probability for every option. It is not a chat model.
How do I run Decider 35B-A3B (NVFP4) locally?
Decider 35B-A3B (NVFP4) does not run in Ollama. Served by decider.serve. This model does not run in Ollama. See the setup guide for servers other than Ollama.
How much memory does Decider 35B-A3B (NVFP4) need?
About 23 GB for the weights plus overhead at Q4_K_M, before the prompt's KV cache. Decision prompts are short, so the cache stays small.
How accurate is Decider 35B-A3B (NVFP4)?
It has two separate published scores on two different suites — the author's own benchmark and the community Decision Index 0.2.1. They are not comparable with each other, and neither is a calibration guarantee.