Written by Jakub Rusinowski · Last updated July 15, 2026
Model library → Bonsai 27B → Ternary Bonsai 27B
The laptop build and recommended pick when you have the memory. Ternary {−1, 0, +1} weights at ~1.71 bits/weight land at 5.9 GB while keeping more than 95% of the FP16 baseline. PrismML reports MATH-500 99.20, AIME 2025 90.84, HumanEval+ 93.90 and LiveCodeBench 82.8, with tool-use and agentic ability largely intact (~74). Multimodal, 262K context, Apache 2.0. Ships as GGUF and MLX; also hosted on Together AI. Specs from launch coverage — verify on the Hugging Face model card.
Ternary Bonsai 27B needs about 17 GB of VRAM at Ternary (1.58-bit, ~1.71 bpw) — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.
| Parameters | 27 Billion |
| Context window | 262,144 |
| Architecture | Qwen3.6-27B (multimodal), ternary weights |
| Provider | PrismML |
| Licence | Apache 2.0 |
| Specified at | Ternary (1.58-bit, ~1.71 bpw) |
| System RAM | 16 GB |
| Record updated | 2026-07-15 |
Apache-2.0 — commercial use permitted. Commercial use permitted. No usage restrictions beyond attribution.
Modelled on a reference NVIDIA RTX 4090 (24 GB), with no KV cache (this record has no published architecture). Speed figures are ESTIMATES from the memory-bandwidth roofline described on the methodology page, not benchmarks we ran — rows marked measured come from published or reader-submitted runs. VRAM here includes the KV cache, so it reads higher than the headline figure above, which does not.
| Quant | Weights | VRAM needed | Est. speed | Fit on 24 GB |
|---|---|---|---|---|
| Q2_K | 8.9 GB | 9.7 GB | ~66 tok/s (est.) | Fits comfortably |
| Q3_K_M | 11.5 GB | 12.3 GB | ~54 tok/s (est.) | Fits comfortably |
| Q4_K_M | 16.3 GB | 17.1 GB | ~40 tok/s (est.) | Fits comfortably |
| Q5_K_M | 19.1 GB | 19.9 GB | ~35 tok/s (est.) | Fits comfortably |
| Q6_K | 22.1 GB | 22.9 GB | ~31 tok/s (est.) | Tight fit |
| Q8_0 | 28.7 GB | 29.5 GB | ~4 tok/s (est.) | Offloads to system RAM (slow) |
| F16 | 54.0 GB | 54.8 GB | ~2 tok/s (est.) | Offloads to system RAM (slow) |
Want the memory numbers alone, at every quantization level and your own context length? Use the Ternary Bonsai 27B VRAM calculator.
or compare on Vast.ai from $0.35/hr (typical low · varies)
As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.
The cheapest catalogued GPU that runs Ternary Bonsai 27B is the AMD Radeon RX 7900 XT (20 GB).
Install Ollama, then run:
ollama run bonsai-27b
Weights on Hugging Face: prism-ml/Ternary-Bonsai-27B-gguf.
Quality scores as published by the model's authors or an independent evaluator — not throughput, and not measured by us.
| Benchmark | Score | Provenance |
|---|---|---|
| MATH-500 | 99.2 % | vendor-claimed · PrismML (launch) |
| AIME 2025 | 90.84 % | vendor-claimed · PrismML (launch) |
| HumanEval+ | 93.9 % | vendor-claimed · PrismML (launch) |
| LiveCodeBench | 82.8 % | vendor-claimed · PrismML (launch) |
| Tool-use (agentic) | 74.01 % | vendor-claimed · PrismML (launch) |
| Vision | 65.19 % | vendor-claimed · PrismML (launch) |
Best for: laptop, on device, multimodal, reasoning, coding, agents.
← All Bonsai 27B models | VRAM calculator | Build a PC for this model | Check your own hardware