Best Local LLMs for Document analysis

Written by Jakub Rusinowski · Last updated July 16, 2026

Pulling structured facts, tables and clauses out of long documents such as contracts and reports.

Top pick: Gemma 4 31B

Scores 95.2/100 for document analysis and extraction. 31B parameters, needing about 19.5 GB at Q4_K_M, 250K context, Apache-2.0.

Ranked for document analysis and extraction

ModelScoreParamsContextLicenceQuality index
1. Gemma 4 31B95.231B250KApache-2.0— (estimated)
2. Qwen 3.6 27B94.628B256KApache-2.0— (estimated)
3. Qwen 3.7 35B-A3B93.235B256KApache-2.0— (estimated)
4. Inkling (NVFP4)91.21000B977KApache 2.0— (estimated)
5. MiniMax M3 230B-A10B91.1230B977KModified MIT (attribution required)— (estimated)
6. MiniMax M2.5 230B90.9230B977KModified MIT (attribution required)— (estimated)

Best pick for your memory budget

The strongest model overall is rarely the right answer — what matters is the strongest model that fits the memory you have. These picks are re-ranked per tier, so each one uses its budget rather than simply being small.

MemoryTypical hardwareRecommended models
8 GBRTX 4060, RTX 3070, base MacBook AirIBM Granite 4.1 Granite 4.1 8B (82.2)
GLM-6 9B (81.2)
Qwen3-Coder 8B (80.7)
12 GBRTX 3060 12 GB, RTX 5070Qwen 3 14B (83.4)
Qwen 3.5 14B (83.2)
Gemma 4 12B (82.8)
16 GBRTX 5080, RTX 4080, RX 9070 XTQwen 3 14B (83.4)
Qwen 3.5 14B (83)
Mistral Small 3.1 24B (82.6)
24 GBRTX 4090, RTX 3090, RX 7900 XTXGemma 4 31B (97.4)
Qwen 3.6 27B (96.8)
Qwen 3.7 35B-A3B (95.4)
48 GBRTX 6000 Ada, MacBook Pro M4 Max 48 GBGemma 4 31B (95.6)
Qwen 3.6 27B (94.6)
Qwen 3.7 35B-A3B (94)
128 GB+Mac Studio, DGX Spark, multi-GPUGemma 4 31B (92.8)
Llama 4.5 Scout (92.4)
Qwen 3.6 27B (92)

How this ranking works

Same context profile as RAG but with a higher coding weight (20%), since extraction almost always ends in structured output — JSON, a table, or a schema the model must not break. Licence importance is 0.6: the documents are usually commercial.

Worked example — Gemma 4 31B: capability 92.5 × 0.418, quality 91.7 × 0.246, context 97.3 × 0.16, license 100 × 0.066, accessibility 80 × 0.111 + 3 tag bonus (long-context).

Requirements applied: context floor 32,768 tokens (ideal 262,144), quality floor 55, licence weight 0.6, latency weight 0.4.

Running document analysis and extraction locally

FAQ

What is the best local LLM for document analysis and extraction?

Gemma 4 31B, scoring 95.2/100 against this workload's published requirements. 110 models qualified.

What hardware do I need for document analysis and extraction?

A credible answer starts at 8 GB of memory. Larger budgets unlock materially stronger models — the table above lists the best pick at each tier.

How were these models ranked?

Same context profile as RAG but with a higher coding weight (20%), since extraction almost always ends in structured output — JSON, a table, or a schema the model must not break. Licence importance is 0.6: the documents are usually commercial.

Hardware for This Workload

Related Workloads

Top Pick

Tools

← All workloads | Check your hardware