Qwen 3.7 — local AI model by Alibaba Cloud
Written by Jakub Rusinowski · Last updated
Qwen 3.7 shipped as API-only Max and Plus tiers in May and June 2026. Alibaba skipped the generation for open weights and went straight to Qwen 3.8, so nothing in this family is locally runnable.
Variants
The smallest Qwen 3.7 variant needs about 22 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache.
| Model | VRAM at Q4 | VRAM | Context | Run it |
|---|---|---|---|---|
| Qwen 3.7 35B-A3B 35B (3B active) | ~21.9 GB | 262,144 | ollama run qwen3-7 |
Memory is quantized weights plus overhead at Q4_K_M, from the same engine as the GPU & VRAM checker.
How to run Qwen 3.7 locally
Install Ollama, then pull the tag.
ollama run qwen3-7Pick a size above for its own VRAM figure, speed estimate and install command.
Licence
Commercial use permitted. No usage restrictions beyond attribution.
Applies to: Qwen 3.7 35B-A3BRecommended GPU
The cheapest catalogued GPU that runs Qwen 3.7 locally (min 22 GB VRAM) is the AMD Radeon RX 7900 XTX (24 GB).
Can I run Qwen 3.7 on my GPU?
- Qwen 3.7 on AMD Radeon RX 7800 XT
- Qwen 3.7 on AMD Radeon RX 7900 XT
- Qwen 3.7 on AMD Radeon RX 7900 XTX
- Qwen 3.7 on AMD Radeon RX 9060 XT 16GB
- Qwen 3.7 on AMD Radeon RX 9070
- Qwen 3.7 on AMD Radeon RX 9070 XT
- Qwen 3.7 on Apple M1
- Qwen 3.7 on Apple M1 Pro
- Qwen 3.7 on Apple M2
- Qwen 3.7 on Apple M2 Pro
- Qwen 3.7 on Apple M3
- Qwen 3.7 on Apple M3 Pro
- Qwen 3.7 on Apple M4
- Qwen 3.7 on Apple M4 Pro
- Qwen 3.7 on Apple M5
- Qwen 3.7 on Intel Arc B570
- Qwen 3.7 on Intel Arc B580
- Qwen 3.7 on NVIDIA GeForce RTX 3060 (12GB)
- Qwen 3.7 on NVIDIA GeForce RTX 3080 (10GB)
- Qwen 3.7 on NVIDIA GeForce RTX 3090
- Qwen 3.7 on NVIDIA GeForce RTX 4060 Ti 16GB
- Qwen 3.7 on NVIDIA GeForce RTX 4070
- Qwen 3.7 on NVIDIA GeForce RTX 4070 Super
- Qwen 3.7 on NVIDIA GeForce RTX 4070 Ti
- Qwen 3.7 on NVIDIA GeForce RTX 4070 Ti Super
- Qwen 3.7 on NVIDIA GeForce RTX 4080
- Qwen 3.7 on NVIDIA GeForce RTX 4080 Super
- Qwen 3.7 on NVIDIA GeForce RTX 4090
- Qwen 3.7 on NVIDIA GeForce RTX 5060 Ti 16GB
- Qwen 3.7 on NVIDIA GeForce RTX 5070
- Qwen 3.7 on NVIDIA GeForce RTX 5070 Ti
- Qwen 3.7 on NVIDIA GeForce RTX 5080
- Qwen 3.7 on NVIDIA GeForce RTX 5090
- Qwen 3.7 on NVIDIA L40S
- Qwen 3.7 on NVIDIA RTX 6000 Ada Generation
Qwen 3.7 — frequently asked questions
How much VRAM does Qwen 3.7 need?
Qwen 3.7 needs about 22 GB VRAM at Q4_K_M quantization for its smallest variant. Variants: Qwen 3.7 35B-A3B (22 GB, Q4_K_M). On Apple Silicon, unified memory counts toward this requirement.
Can I run Qwen 3.7 on an RTX 4090 (24 GB)?
Yes — Qwen 3.7 runs on an RTX 4090 (24 GB) and other 24 GB cards such as the RTX 3090. Smaller variants also fit comfortably on 8–16 GB GPUs at Q4_K_M.
What quantization should I use for Qwen 3.7?
Q4_K_M is the best balance of quality and VRAM for Qwen 3.7 in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.
How do I run Qwen 3.7 with Ollama?
Qwen 3.7 has no local Ollama tag — the published tag is cloud-hosted, so running it sends your prompts to a hosted GPU rather than your own machine.