Written by Jakub Rusinowski · Last updated August 15, 2026
Turning scanned pages, screenshots and photographed documents into accurate text.
Top pick: Gemma 4 27B ⭐
Scores 93.7/100 for OCR and document scanning. 27B parameters, needing about 17.1 GB at Q4_K_M, 125K context, Gemma License (commercial OK).
| Model | Score | Params | Context | Licence | Quality index |
|---|---|---|---|---|---|
| 1. Gemma 4 27B ⭐ | 93.7 | 27B | 125K | Gemma License (commercial OK) | — (estimated) |
| 2. Mistral Small 3.1 24B | 93.5 | 24B | 125K | Apache 2.0 | — (estimated) |
| 3. Qwen 3.5 14B | 92.5 | 14B | 125K | Apache 2.0 | — (estimated) |
| 4. Qwen 3.7 35B-A3B | 92.2 | 35B | 256K | Apache-2.0 | — (estimated) |
| 5. Llama 3.2 11B Vision Instruct | 91.9 | 11B | 125K | Llama Community | — (estimated) |
| 6. Qwen 2.5 VL 7B Instruct | 91.8 | 8B | 125K | Apache 2.0 | — (estimated) |
The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.
| Memory | Typical hardware | Recommended models |
|---|---|---|
| 8 GB | RTX 4060, RTX 3070, base MacBook Air | Llama 3.2 11B Vision Instruct (91.9) Qwen 2.5 VL 7B Instruct (91.8) Gemma 4 E4B (89.9) |
| 12 GB | RTX 3060 12 GB, RTX 5070 | Qwen 3.5 14B (93.3) Gemma 4 12B (92.9) Llama 3.2 11B Vision Instruct (91.9) |
| 16 GB | RTX 5080, RTX 4080, RX 9070 XT | Mistral Small 3.1 24B (95.5) Qwen 3.5 14B (93) Gemma 4 12B (91.5) |
| 24 GB | RTX 4090, RTX 3090, RX 7900 XTX | Gemma 4 27B ⭐ (97.2) Qwen 3.7 35B-A3B (95.6) Mistral Small 3.1 24B (95.5) |
| 48 GB | RTX 6000 Ada, MacBook Pro M4 Max 48 GB | Qwen 2.5 VL 72B Instruct (98.5) Qwen 3.5 72B (97.1) Gemma 4 27B ⭐ (93.6) |
| 128 GB+ | Mac Studio, DGX Spark, multi-GPU | Qwen 2.5 VL 72B Instruct (94.8) Qwen 3.5 72B (93.3) Llama 3.2 90B Vision Instruct (91.5) |
Shares the vision hard gate but weights structured-output ability higher (25% coding) and latency higher than general vision, because OCR is typically run in batch over many pages where throughput compounds.
Worked example — Gemma 4 27B ⭐: capability 93.5 × 0.393, quality 92.7 × 0.231, context 100 × 0.15, license 70 × 0.052, accessibility 80 × 0.173 + 3 tag bonus (vision).
Requirements applied: context floor 8,192 tokens (ideal 32,768), quality floor 45, licence weight 0.5, latency weight 0.8. Vision capability is required.
Gemma 4 27B ⭐, scoring 93.7/100 against this workload's published requirements. 31 models qualified.
A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.
Shares the vision hard gate but weights structured-output ability higher (25% coding) and latency higher than general vision, because OCR is typically run in batch over many pages where throughput compounds.