DeepSeek-OCR — Local AI Model by DeepSeek

Written by Jakub Rusinowski · Last updated September 11, 2026

A 3B vision-language model built for one job — turning document pages into structured text — and a demonstration of "contexts optical compression": rendering text as an image and decoding it costs roughly 10x fewer tokens at near-lossless fidelity, or 20x at about 60% precision. A DeepEncoder compresses each page into vision tokens which a DeepSeek-3B MoE decoder reads, activating only ~570M parameters. A second generation, DeepSeek-OCR-2, followed in January 2026.

Licence

LicenceWhat it permitsApplies to
MITCommercial use permitted
Commercial use permitted. No usage restrictions beyond attribution.
DeepSeek-OCR 3B

Hardware Requirements

DeepSeek-OCR 3BMin 3 GB VRAM · Q4_K_M · 8,192 ctx ·

Recommended GPU

The cheapest GPU that runs DeepSeek-OCR locally (min 3 GB VRAM) is the Intel Arc B570 (10 GB).

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
Intel Arc B570 10GB
10 GB VRAM · 150 W board power
2026 prices are volatile — check the current listing.
Check price on Amazon

How to Run Locally

Install Ollama then run: ollama run

Minimum VRAM: 3 GB. For best results use Q4_K_M quantization.

DeepSeek-OCR — Frequently Asked Questions

How much VRAM does DeepSeek-OCR need?

DeepSeek-OCR needs about 3 GB VRAM at Q4_K_M quantization for its smallest variant. Variants: DeepSeek-OCR 3B (3 GB, Q4_K_M). On Apple Silicon, unified memory counts toward this requirement.

Can I run DeepSeek-OCR on an RTX 4090 (24 GB)?

Yes — DeepSeek-OCR runs on an RTX 4090 (24 GB) and other 24 GB cards such as the RTX 3090. Smaller variants also fit comfortably on 8–16 GB GPUs at Q4_K_M.

What quantization should I use for DeepSeek-OCR?

Q4_K_M is the best balance of quality and VRAM for DeepSeek-OCR in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.

How do I run DeepSeek-OCR with Ollama?

Install Ollama, then run: ollama run . This downloads DeepSeek-OCR and starts a local, OpenAI-compatible endpoint — no internet connection is needed after the initial download.

Can I Run DeepSeek-OCR on My GPU?