Best Local LLMs for Vision

Written by Jakub Rusinowski · Last updated June 26, 2026

Reading images, screenshots, charts and diagrams alongside text.

Top pick: Qwen 3.7 35B-A3B

Scores 95.5/100 for vision and image understanding. 35B parameters, needing about 21.9 GB at Q4_K_M, 256K context, Apache-2.0.

Ranked for vision and image understanding

ModelScoreParamsContextLicenceQuality index
1. Qwen 3.7 35B-A3B95.535B256KApache-2.0— (estimated)
2. Gemma 4 31B95.231B250KApache-2.0— (estimated)
3. Qwen 3.5 14B9514B125KApache 2.0— (estimated)
4. Qwen 3.6 35B-A3B94.835B256KApache-2.0— (estimated)
5. Gemma 4 27B ⭐94.227B125KGemma License (commercial OK)— (estimated)
6. Kimi K2.69432B125KKimi License (research)— (estimated)

Best pick for your memory budget

The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.

MemoryTypical hardwareRecommended models
8 GBRTX 4060, RTX 3070, base MacBook AirGemma 4 E4B (91.6)
Llama 3.2 11B Vision Instruct (91.6)
Qwen 2.5 VL 7B Instruct (90.7)
12 GBRTX 3060 12 GB, RTX 5070Qwen 3.5 14B (95.7)
Gemma 4 12B (92.3)
Gemma 4 12B (Unified) (92)
16 GBRTX 5080, RTX 4080, RX 9070 XTQwen 3.5 14B (95.4)
Mistral Small 3.1 24B (95.2)
Gemma 4 12B (91.2)
24 GBRTX 4090, RTX 3090, RX 7900 XTXQwen 3.7 35B-A3B (98.1)
Gemma 4 31B (97.8)
Qwen 3.6 35B-A3B (97.4)
48 GBRTX 6000 Ada, MacBook Pro M4 Max 48 GBQwen 3.5 72B (99.7)
Qwen 3.7 35B-A3B (96.5)
Qwen 3.6 35B-A3B (95.8)
128 GB+Mac Studio, DGX Spark, multi-GPUQwen 3.5 72B (96.9)
Llama 3.2 90B Vision Instruct (95.2)
Mistral Small 4 119B-A6.5B (93.1)

How this ranking works

Vision capability is a hard gate — text-only models are excluded outright rather than penalised, so this ranking is drawn from a genuinely different candidate pool than every other workload here.

Worked example — Qwen 3.7 35B-A3B: capability 92.7 × 0.422, quality 92.7 × 0.248, context 100 × 0.161, license 100 × 0.039, accessibility 80 × 0.13 + 3 tag bonus (multimodal).

Requirements applied: context floor 8,192 tokens (ideal 131,072), quality floor 45, licence weight 0.35, latency weight 0.5. Vision capability is required.

Running vision and image understanding locally

FAQ

What is the best local LLM for vision and image understanding?

Qwen 3.7 35B-A3B, scoring 95.5/100 against this workload's published requirements. 31 models qualified.

What hardware do I need for vision and image understanding?

A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.

How were these models ranked?

Vision capability is a hard gate — text-only models are excluded outright rather than penalised, so this ranking is drawn from a genuinely different candidate pool than every other workload here.

Hardware for This Workload

Related Workloads

Top Pick

Tools

← All workloads | Check your hardware