Qwen 3.7 — local AI model by Alibaba Cloud

Written by Jakub Rusinowski · Last updated

Qwen 3.7 shipped as API-only Max and Plus tiers in May and June 2026. Alibaba skipped the generation for open weights and went straight to Qwen 3.8, so nothing in this family is locally runnable.

Variants

The smallest Qwen 3.7 variant needs about 22 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache.

ModelVRAM
Qwen 3.7 35B-A3B
35B (3B active)
~21.9 GB

Memory is quantized weights plus overhead at Q4_K_M, from the same engine as the GPU & VRAM checker.

How to run Qwen 3.7 locally

Install Ollama, then pull the tag.

ollama run qwen3-7

Pick a size above for its own VRAM figure, speed estimate and install command.

Licence

Apache-2.0Commercial use permitted

Commercial use permitted. No usage restrictions beyond attribution.

Applies to: Qwen 3.7 35B-A3B

Recommended GPU

The cheapest catalogued GPU that runs Qwen 3.7 locally (min 22 GB VRAM) is the AMD Radeon RX 7900 XTX (24 GB).

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
AMD Radeon RX 7900 XTX 24GB
24 GB VRAM · 355 W board power
2026 prices are volatile — check the current listing.

Can I run Qwen 3.7 on my GPU?

Qwen 3.7 — frequently asked questions

How much VRAM does Qwen 3.7 need?

Qwen 3.7 needs about 22 GB VRAM at Q4_K_M quantization for its smallest variant. Variants: Qwen 3.7 35B-A3B (22 GB, Q4_K_M). On Apple Silicon, unified memory counts toward this requirement.

Can I run Qwen 3.7 on an RTX 4090 (24 GB)?

Yes — Qwen 3.7 runs on an RTX 4090 (24 GB) and other 24 GB cards such as the RTX 3090. Smaller variants also fit comfortably on 8–16 GB GPUs at Q4_K_M.

What quantization should I use for Qwen 3.7?

Q4_K_M is the best balance of quality and VRAM for Qwen 3.7 in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.

How do I run Qwen 3.7 with Ollama?

Qwen 3.7 has no local Ollama tag — the published tag is cloud-hosted, so running it sends your prompts to a hosted GPU rather than your own machine.