Mapika2B~2 GB VRAM at Q4_K_M

Decider 2B — VRAM & /v1/systemone setup

Written by Jakub Rusinowski · Last updated

Mapika's 2B decision model on Qwen3.5-2B-Base, served by decider.serve with a TypeSafe-compatible /v1/systemone.

Decider 2B needs about 2 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.

Call Decider 2B

Served by decider.serve. This model does not run in Ollama.

Hardware fit

Weights plus overhead plus the KV cache at a 8,192-token prompt, on NVIDIA RTX 4090 (24 GB). A publisher build is sized from its file; the other rows are modelled at a standard quant. Decision requests are short, so no long-context figure is shown.

QuantVRAMFit
BF16
Publisher build · 3.8 GB file
5.4 GBFits
Q4_K_M
4.83 bpw · modelled quant
2.9 GBFits
Q6_K
6.56 bpw · modelled quant
3.4 GBFits
Q8_0
8.50 bpw · modelled quant
4 GBFits

Published file size: BF16 3.8 GB. A download size from the model publisher — not a VRAM requirement.

Needs torch, transformers>=5, flash-linear-attention (optional but several times faster). typesafe-sdk works against decider.serve via TYPESAFE_BASE_URL. Current revision v11; v10 available as a tag.

How it was scored

Two different suites on two different scales. Never compare the numbers across the two cards.

Decision Index 0.2.1 (snapshot 2026-09-28)

Decision Index 0.2.1
Balanced skill28.97
ECE (lower is better)
0.0771
Median compute
8.1 ms on 1x NVIDIA RTX PRO 6000 (96 GB)

Community-maintained; not affiliated with the model authors.

Source ↗

Decision models are not ranked on chat, creative or coding scores.

Specifications

Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source.

Parameters
2.27B
Context window
32K
Architecture
Fine-tune of Qwen/Qwen3.5-2B-Base
Provider
Mapika
Licence
Apache-2.0
Specified at
Q4_K_M
System RAM
8 GB
Record updated
2026-09-30
LicenceApache-2.0Commercial use permitted

Commercial use permitted. No usage restrictions beyond attribution.

Other Decider sizes

Decider 2B — frequently asked questions

What is Decider 2B?

Decider 2B is a decision model: you send it a state and typed questions (choice, yes/no/unknown, or a score) and it returns one answer per question with a probability for every option. It is not a chat model.

How do I run Decider 2B locally?

Decider 2B does not run in Ollama. Served by decider.serve. This model does not run in Ollama. See the setup guide for servers other than Ollama.

How much memory does Decider 2B need?

About 2 GB for the weights plus overhead at Q4_K_M, before the prompt's KV cache. Decision prompts are short, so the cache stays small.

How accurate is Decider 2B?

It has two separate published scores on two different suites — the author's own benchmark and the community Decision Index 0.2.1. They are not comparable with each other, and neither is a calibration guarantee.