The latest generation of local LLMs can see images, process documents, and even understand audio — not just generate text. This guide explains multimodal AI, which models support it, and how to run vision models locally.
What Is a Multimodal LLM?
A multimodal model accepts multiple types of input: text, images, PDFs, audio, or video. The output is still typically text (descriptions, answers, analysis), though some models can also generate images.
Why it matters: You can now:
- Drop an image into your local AI and ask "What's wrong with this code screenshot?"
- Upload a PDF and ask "Summarize the key points of this contract"
- Describe a chart and get an analysis — without the data ever leaving your machine
Supported Modalities in 2026
| Modality | Description | Local Support |
|---|---|---|
| Text → Text | Classic LLM inference | ✅ Universal |
| Image → Text | Describe, analyze, OCR images | ✅ Many models |
| PDF → Text | Extract and reason over document content | ✅ Via image chunking |
| Audio → Text | Transcription + understanding | ⚠️ Limited (Whisper + LLM pipeline) |
| Video → Text | Frame-by-frame analysis | ⚠️ Experimental |
| Text → Image | Generate images from prompts | ✅ Separate models (Stable Diffusion) |
Best Local Vision Models (2026)
Llama 3.2 Vision (11B and 90B)
Meta's vision model integrates seamlessly with Ollama:
- 11B Vision: 8GB VRAM (Q4), excellent for most image tasks
- 90B Vision: Requires 50+ GB VRAM (impractical without multi-GPU)
- Strengths: Strong OCR, document understanding, chart analysis
- Weaknesses: Slower than text-only 8B equivalent
# Pull and run Llama 3.2 Vision
ollama run llama3.2-vision:11b
# Then send an image in your chat
# (Works in Open WebUI with drag-and-drop)Gemma 3 (All sizes support vision)
Google's Gemma 3 family is natively multimodal across all sizes:
- Gemma 3 4B Vision: 4GB VRAM — fastest vision model for low-end hardware
- Gemma 3 12B Vision: 8GB VRAM — best quality/speed balance
- Gemma 3 27B Vision: 16GB VRAM — near-GPT-4V quality
ollama run gemma3:12b # Vision enabled by default
ollama run gemma3:27b # Best quality for complex imagesQwen VL (Qwen 2.5-VL)
Alibaba's Qwen VL series excels at document and chart understanding:
- Qwen 2.5-VL 7B: Strong OCR, excellent for receipts, tables, and scanned documents
- Qwen 2.5-VL 72B: Flagship vision model, competitive with GPT-4V
- Best-in-class for dense text in images (menus, invoices, presentations)
ollama run qwen2.5-vl:7bMistral Small 3.1 (Vision Capable)
- 24B parameters with vision, fits in 16GB VRAM at Q4
- Strong at instruction following with image context
- Good balance of size and multimodal quality
VRAM Requirements for Vision Models
Vision models are heavier than their parameter count suggests, which matters most on a phone: storage and download planning on iOS covers what a multi-gigabyte model actually costs there.
Vision models use more VRAM than text-only models of the same size because they include a vision encoder (typically CLIP or SigLIP):
| Model | VRAM (Q4_K_M) | VRAM (Q8) | Notes |
|---|---|---|---|
| Gemma 3 4B Vision | ~4 GB | ~7 GB | Fastest, entry-level |
| Llama 3.2 Vision 11B | ~7.5 GB | ~12 GB | Good all-rounder |
| Gemma 3 12B Vision | ~8.5 GB | ~14 GB | Best mid-range |
| Qwen 2.5-VL 7B | ~5.5 GB | ~8 GB | Best for documents |
| Gemma 3 27B Vision | ~17 GB | ~28 GB | Near GPT-4V quality |
| Qwen 2.5-VL 72B | ~42 GB | — | Multi-GPU required |
Setting Up Vision Models with Open WebUI
Open WebUI provides the easiest interface for vision tasks:
# Start Open WebUI with Ollama backend
docker run -d -p 3000:8080 \
--add-host=host.docker.internal:host-gateway \
-v open-webui:/app/backend/data \
--name open-webui \
ghcr.io/open-webui/open-webui:main
# Access at http://localhost:3000
# Drag and drop images directly into the chatFor programmatic use via the API:
import ollama
import base64
# Load and encode image
with open("screenshot.png", "rb") as f:
image_data = base64.b64encode(f.read()).decode()
response = ollama.chat(
model="llama3.2-vision:11b",
messages=[{
"role": "user",
"content": "What error is shown in this screenshot?",
"images": [image_data]
}]
)
print(response["message"]["content"])Real-World Use Cases
Code Screenshot Analysis
Drop a screenshot of an error message and ask: "What's causing this error and how do I fix it?"
Document Processing
Upload a PDF scan: "Extract all dates and amounts from this invoice."
Chart and Graph Understanding
Upload a business chart: "Describe the trend shown in this revenue chart."
UI/UX Feedback
Upload a website screenshot: "What UX improvements would you suggest for this form?"
Medical Imaging (Research Only)
With proper HIPAA setup: "Describe what you observe in this X-ray" — for research assistance, not diagnosis.
Audio: Local Transcription + LLM Pipeline
For audio understanding, combine Whisper (transcription) with any LLM:
# Install Whisper
pip install openai-whisper
# Transcribe audio
whisper meeting_recording.mp3 --model medium --output_format txt
# Then pipe transcript to Ollama
cat meeting_recording.txt | ollama run llama3.1 "Summarize the key decisions from this meeting transcript:"Performance Tips for Vision Models
- Resize images before sending — models process at fixed resolution internally (typically 336×336 or 448×448 tiles). Very large images get tiled, which increases VRAM and time.
- Use Q4_K_M quantization — vision quality is less affected by quantization than text quality
- Batch processing — for multiple images, use the API in a loop rather than reloading the model
- GPU offload — ensure
num_gpuis set to maximum in Ollama for vision models
Frequently Asked Questions
Can local vision models read text in images (OCR)? Yes — Qwen 2.5-VL 7B is particularly strong at OCR. It can read text from menus, receipts, whiteboards, and documents accurately.
How does local vision quality compare to GPT-4V? Gemma 3 27B and Qwen 2.5-VL 72B are competitive with GPT-4V for most tasks. For simple image description, even the 7–12B models are often sufficient.
Can I run video analysis locally? Not efficiently yet. Most tools process video as individual frames, which is slow. Frame-by-frame analysis at 1 frame/second is possible but impractical for long videos.
Do vision models work with PDF files? Indirectly — you need to convert PDF pages to images first, then send them to the vision model. Tools like pdf2image automate this.