Bespoke Labs9B~7 GB VRAM at Q4_K_M

Nimble 9B — VRAM & /v1/systemone setup

Written by Jakub Rusinowski · Last updated

Bespoke Labs' 9B decision model: a LoRA on Qwen3.5-9B. The default decision model in Ollama's library, called at /v1/systemone.

Nimble 9B needs about 7 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.

Call Nimble 9B

Pull it with Ollama 0.35 or newer, then POST to /v1/systemone. Do not use ollama chat commands — this is not a chat model.

ollama pull nimble
curl http://localhost:11434/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "nimble",
    "state": {
      "ticket": "I was charged twice. Please refund the extra payment."
    },
    "questions": {
      "team": {
        "type": "choice",
        "instructions": "Which team should handle this ticket?",
        "criteria": {
          "billing": "Payments and refunds",
          "technical": "Bugs and integrations",
          "other": "None of the above"
        }
      }
    }
  }'
Weights on Hugging Face: bespokelabs/Bespoke-Nimble-9B ↗

Build a request for this model Decision models guide

Other decision models: Tev1 4B · Tev1 0.8B · Winnow 12B · Winnow E4B · Decider 2B · Decider 4B · Decider 35B-A3B (NVFP4) · JevK5 4B · Intern-Decision 4B · AutoJev 27B · Laya

Hardware fit

Weights plus overhead plus the KV cache at a 8,192-token prompt, on NVIDIA RTX 4090 (24 GB). A publisher build is sized from its file; the other rows are modelled at a standard quant. Decision requests are short, so no long-context figure is shown.

QuantVRAMFit
Ollama build
Publisher build · 9.5 GB file
11.5 GBFits
Q4_K_M
4.83 bpw · modelled quant
7.9 GBFits
Q6_K
6.56 bpw · modelled quant
10 GBFits
Q8_0
8.50 bpw · modelled quant
12.3 GBFits

Published file size: Ollama build 9.5 GB. A download size from the model publisher — not a VRAM requirement.

Ollama build is 9.5 GB. Which Bespoke checkpoint the tag packages (Sep 24 update vs original-2676) is not stated in one place; do not print a version.

How it was scored

Two different suites on two different scales. Never compare the numbers across the two cards.

Author's public benchmark

Bespoke Labs public benchmarks (13 datasets, 3,880 decisions)
Accuracy75.7%
choice questions
81.6%
noul questions
80.2%
score questions
54.6%
Source ↗

Decision Index 0.2.1 (snapshot 2026-09-28)

Decision Index 0.2.1
Balanced skill39.57
ECE (lower is better)
0.024
Median compute
77.3 ms on 1x NVIDIA RTX PRO 6000 (96 GB)

Community-maintained; not affiliated with the model authors.

Source ↗

The Index scored the repository bespokelabs/Bespoke-Nimble-9B-v2, not the checkpoint this page describes.

Decision models are not ranked on chat, creative or coding scores.

Specifications

Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source.

Parameters
9.65B
Context window
8K (prompt limit)
Architecture
Fine-tune of Qwen/Qwen3.5-9B
Provider
Bespoke Labs
Licence
Apache-2.0
Specified at
Q4_K_M
System RAM
14 GB
Record updated
2026-09-30
LicenceApache-2.0Commercial use permitted

Commercial use permitted. No usage restrictions beyond attribution.

Nimble 9B — frequently asked questions

What is Nimble 9B?

Nimble 9B is a decision model: you send it a state and typed questions (choice, yes/no/unknown, or a score) and it returns one answer per question with a probability for every option. It is not a chat model.

How do I run Nimble 9B locally?

Install Ollama 0.35 or newer, run `ollama pull nimble`, then POST your state and questions to http://localhost:11434/v1/systemone. It is called through the API, not a chat session.

How much memory does Nimble 9B need?

About 7 GB for the weights plus overhead at Q4_K_M, before the prompt's KV cache. Decision prompts are short, so the cache stays small.

How accurate is Nimble 9B?

It has two separate published scores on two different suites — the author's own benchmark and the community Decision Index 0.2.1. They are not comparable with each other, and neither is a calibration guarantee.