Written by Jakub Rusinowski · Last updated July 15, 2026
Model library → Bonsai 27B → 1-bit Bonsai 27B
The phone build. Binary {−1, +1} weights at ~1.125 bits/weight bring a full 27B multimodal model down to 3.9 GB — small enough to run on an iPhone 17 Pro at ~11 tok/s. PrismML reports it keeps more than 90% of the FP16 baseline (MATH 91.66, coding 81.88); tool-calling is the weak spot, dropping from ~80 to ~66. Ships as GGUF (llama.cpp / LM Studio) and MLX (Apple). Specs from launch coverage — verify on the Hugging Face model card.
1-bit Bonsai 27B needs about 17 GB of VRAM at 1-bit (binary, ~1.125 bpw) — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.
| Parameters | 27 Billion |
| Context window | 262,144 |
| Architecture | Qwen3.6-27B (multimodal), 1-bit binary weights |
| Provider | PrismML |
| Licence | Apache 2.0 |
| Specified at | 1-bit (binary, ~1.125 bpw) |
| System RAM | 8 GB |
| Record updated | 2026-07-15 |
Apache-2.0 — commercial use permitted. Commercial use permitted. No usage restrictions beyond attribution.
Modelled on a reference NVIDIA RTX 4090 (24 GB), with no KV cache (this record has no published architecture). Speed figures are ESTIMATES from the memory-bandwidth roofline described on the methodology page, not benchmarks we ran — rows marked measured come from published or reader-submitted runs. VRAM here includes the KV cache, so it reads higher than the headline figure above, which does not.
| Quant | Weights | VRAM needed | Est. speed | Fit on 24 GB |
|---|---|---|---|---|
| Q2_K | 8.9 GB | 9.7 GB | ~66 tok/s (est.) | Fits comfortably |
| Q3_K_M | 11.5 GB | 12.3 GB | ~54 tok/s (est.) | Fits comfortably |
| Q4_K_M | 16.3 GB | 17.1 GB | ~40 tok/s (est.) | Fits comfortably |
| Q5_K_M | 19.1 GB | 19.9 GB | ~35 tok/s (est.) | Fits comfortably |
| Q6_K | 22.1 GB | 22.9 GB | ~31 tok/s (est.) | Tight fit |
| Q8_0 | 28.7 GB | 29.5 GB | ~4 tok/s (est.) | Offloads to system RAM (slow) |
| F16 | 54.0 GB | 54.8 GB | ~2 tok/s (est.) | Offloads to system RAM (slow) |
Want the memory numbers alone, at every quantization level and your own context length? Use the 1-bit Bonsai 27B VRAM calculator.
or compare on Vast.ai from $0.35/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
The cheapest catalogued GPU that runs 1-bit Bonsai 27B is the AMD Radeon RX 7900 XT (20 GB).
Install Ollama, then run:
ollama run bonsai-27b
Weights on Hugging Face: prism-ml/Bonsai-27B-gguf.
Quality scores as published by the model's authors or an independent evaluator — not throughput, and not measured by us.
| Benchmark | Score | Provenance |
|---|---|---|
| MATH | 91.66 % | vendor-claimed · PrismML (launch) |
| Coding (HumanEval+ class) | 81.88 % | vendor-claimed · PrismML (launch) |
| Tool-calling (BFCL class) | 66 % | vendor-claimed · PrismML (launch) |
Best for: on device, phone, offline, multimodal, reasoning.
← All Bonsai 27B models | VRAM calculator | Check your own hardware