Local AI decision engine
Find the right model, hardware & setup for your workload.
Compare models, GPUs and local AI setups to find what you can actually run, what you should use, and what it will cost.
169 Model variants · 113 Guides · 70 GPUs & chips · Free — Always open
Written by Jakub Rusinowski · Last updated September 19, 2026
Everything you need to make a local AI decision.
01 What can I run?
GPU & VRAM checker — Check any GPU or Mac against every model we track and see what fits in its memory.
Check my hardware →02 What should I run?
Model library — Browse open-weight models by size, licence, context and hardware fit.
Explore models →03 What should I buy?
Build generator — Give a budget and a workload; get a GPU, RAM and storage recommendation.
Find a build →04 How fast & expensive?
Benchmarks & cost calculator — Estimated tokens per second by GPU and model, and local-versus-cloud running costs.
See benchmarks → · Calculate costs →05 How do I deploy?
Setup guides — Step-by-step guides for Ollama, LM Studio, vLLM, RAG and fine-tuning.
Read the guides →
What can your hardware actually run?
Pick a GPU or Mac and every model that fits its memory is ranked at Q4_K_M, with the VRAM it needs and an estimated speed. Detection runs in your browser; nothing is uploaded.
Or browse ready-made answers:
Local AI tools
Free calculators and checkers that answer one question each — and share one engine, so their numbers agree.
GPU & VRAM Checker
Find which LLMs your GPU can run — Llama, DeepSeek, Gemma
- VRAM & Fine-Tuning Calculator — Exact VRAM and fine-tuning memory for any model and quant
- Upgrade Advisor — Name the model you want to run and get ranked routes there — including the free ones
- Best GPU by Workload — The strongest GPU for coding, RAG, agents, vision and more
- Local vs Cloud Cost Calculator — See when local AI beats ChatGPT pricing
- GPU Benchmark Leaderboard — Tokens/sec rankings for RTX 4090, Apple M4, and more
- Cloud AI Directory — Rent GPUs, free LLM API tiers, and hosted models — plus local vs cloud comparisons
Explore the model ecosystem
Open-weight models with their real memory requirements, ranked for your hardware and your workload.
- LLM Model Library — Browse 196 open-weight model variants: Llama 4, Qwen 3, Gemma 3, DeepSeek V3
- Best Models for Your GPU — The best models that fit each specific GPU
- Best Models by Workload — Top local models for coding, RAG, agents, writing and more
- Can I Run It? — Straight yes/no verdicts for every model on your GPU
- Best LLMs by VRAM — The strongest models that fit in 8, 12, 16, 24 GB and up
- Local LLM Compare Tool — Side-by-side VRAM, speed, and quant comparison for 2–4 local models
- Local Model Explorer — Scatter plot and table of every local LLM with hardware-aware estimates
Recently updated models
- Apertus 70B — Parameters: 70 Billion · Architecture: Dense · Context: 65,536 · Memory (Q4): 43.1 GB · Licence: Apache 2.0 (Specs updated 2026-09-19)
- Apertus 8B — Parameters: 8 Billion · Architecture: Dense · Context: 65,536 · Memory (Q4): 5.6 GB · Licence: Apache 2.0 (Specs updated 2026-09-19)
- Bielik PL 11B v3.0 Instruct — Parameters: 11 Billion · Architecture: Dense · Context: 32,768 · Memory (Q4): 7.4 GB · Licence: Apache 2.0 (Specs updated 2026-09-19)
- EuroLLM 22B — Parameters: 22 Billion · Architecture: Dense · Context: 4,096 · Memory (Q4): 14.1 GB · Licence: Apache 2.0 (Specs updated 2026-09-19)
- EuroLLM 9B — Parameters: 9 Billion · Architecture: Dense · Context: 4,096 · Memory (Q4): 6.2 GB · Licence: Apache 2.0 (Specs updated 2026-09-19)
- GLM-4.6V-Flash 9B — Parameters: 9 Billion · Architecture: Dense · Context: 65,536 · Memory (Q4): 6.2 GB · Licence: MIT (Specs updated 2026-09-19)
Memory: Q4_K_M weights plus runtime overhead, before context (KV cache).
Find hardware for local AI
GPUs, Macs, laptops and AI stations compared by the two numbers that decide local AI: memory and bandwidth.
- GPU Database
- Hardware Database
- Laptops
- AI Stations & Mini-PCs
- Mac Buying Guide
- Best AI Build Generator
- GPU Upgrade Calculator
| Device | Memory | Bandwidth | Models (Q4_K_M) |
|---|---|---|---|
| NVIDIA GeForce RTX 4060 | 8 GB | 272 GB/s | 53 fit in VRAM |
| NVIDIA GeForce RTX 5090 | 32 GB | 1792 GB/s | 109 fit in VRAM |
| AMD Radeon RX 7900 XTX | 24 GB | 960 GB/s | 108 fit in VRAM |
| Apple M4 Max | 128 GB unified | 546 GB/s | 125 fit in VRAM |
| Intel Arc B580 | 12 GB | 456 GB/s | 67 fit in VRAM |
| NVIDIA GeForce RTX 5080 | 16 GB | 960 GB/s | 74 fit in VRAM |
Model counts come from the same check as the hardware monitor above, at Q4_K_M.
Know before you run it.
Published measurements, credited to their source — next to the estimates our model makes for every other pairing.
| Model | Hardware | Speed | Evidence |
|---|---|---|---|
| Qwen 3 30B-A3B (MoE) | NVIDIA GeForce RTX 4090 | 195.8 tok/s at Q4_K_XL | Measured · Hardware Corner |
| Qwen 3 8B | NVIDIA GeForce RTX 4090 | 141.3 tok/s at Q4_K_M | Measured · Hardware Corner |
| Llama 3.1 8B Instruct | NVIDIA GeForce RTX 4090 | 131.0 tok/s at Q4_K_XL | Measured · Hardware Corner |
| Qwen 3 14B | NVIDIA GeForce RTX 4090 | 82.8 tok/s at Q4_K_XL | Measured · Hardware Corner |
| Qwen 2.5 7B Instruct | NVIDIA GeForce RTX 4090 | 135.0 tok/s at Q4_K_M | Measured · Mustafa.net |
Measured rows are third-party runs we cite, not our own lab. Every other figure on the site is an estimate, and labelled as one.
What will local AI actually cost?
The cost calculator weighs the same terms for your hardware, your electricity price and your hours of use.
Hardware (purchase price ÷ the years you depreciate it over) + Electricity (power draw × hours of use a day × your price per kWh) vs Cloud (the equivalent API or rented-GPU spend)
Build a machine around your models.
Start from what you want to run, not from a parts list.
- Workload
- Models
- Budget
- Hardware
- Compatibility
- Performance
- Cost
Learn local AI
Step-by-step guides, from your first model to production serving.
- Beginner's Guide to Local AI — Start here: run your first local model step by step (Updated 2026-07-12)
- VRAM Requirements 2026 — How much GPU memory each model size really needs (Updated 2026-07-12)
- Quantization Explained — What Q4, Q5, Q8 and GGUF mean for memory and quality (Updated 2026-07-12)
- Local RAG Guide — Chat with your own documents, fully offline (Updated 2026-07-12)
- Fine-Tuning Guide (LoRA/DPO) — Train a model on your own data with LoRA and DPO (Updated 2026-07-12)
- Building AI Agents — Build autonomous agents on top of local models (Updated 2026-07-12)
More topics
- Setup Guides & Tutorials
- Coding Agents
- Install Ollama
- Ollama vs LM Studio
- vLLM Setup Guide
- Troubleshooting
- LLM Glossary
- Local AI by Operating System
- Run LLMs on Phone
What is a Private LLM?
A Private LLM is an artificial intelligence model that runs entirely on your own infrastructure—whether that's a powerful gaming laptop, an on-premise server, or a private cloud instance.
Unlike using public APIs like ChatGPT or Claude where your data is sent to external servers, a private LLM ensures your data never leaves your device. This allows businesses to process sensitive documents, proprietary code, and personal information with zero risk of data leakage or model training on your intellectual property.
Cloud vs. Local: The Breakdown
| Feature | Public Cloud API | Private Local LLM |
|---|---|---|
| Cost Structure | Recurring monthly fees + Cost per token (Expensive at scale) | Free* — Just electricity & hardware |
| Data Privacy | Data sent to 3rd party servers. Risk of retention/training. | 100% Private — Data stays on your disk |
| Latency | Network dependent. Unpredictable spikes. | Zero Network Latency — Speed depends on GPU |
| Control | Censored / Guardrailed. Models can change anytime. | Full Control — Uncensored / Fine-tunable |
Local AI data, benchmarks & research
Reports, open datasets and the methodology behind every number.
Local AI Report #5 — Europe's AI Compute Gap and Sovereignty Published 2026-09-14
The EU runs just 2 GW of AI compute to America's 35 GW. Inside Europe's 19 AI Factories, the gigafactory tender, and Luxembourg's 20-exaflop sovereignty bet.
- Methodology — How every VRAM and speed figure on this site is calculated
- Local AI Reports — Biweekly digest: new models, GPU price moves, tooling updates
- Efficiency Leaderboard — Tokens per watt, ranked across GPUs
- Local vs Cloud Comparisons — When local hardware beats the OpenAI and Anthropic APIs
- Training Dataset Hub — 47 curated datasets for fine-tuning and pretraining
- Local AI Blog — Guides, comparisons, and insights on running AI locally
What's new in local AI
The newest dated records in our catalogue: model releases, published measurements, reports and guide updates.
- New report ·
Local AI Report #5 — Europe's AI Compute Gap and Sovereignty — New issue of the Local AI Report. - Guide updated ·
"App is damaged and can't be opened" — Gatekeeper on macOS — Troubleshooting fix revised. - Guide updated ·
"Failed to initialize NVML: Driver/library version mismatch" — Troubleshooting fix revised. - Guide updated ·
"ollama is not recognized as an internal or external command" — Troubleshooting fix revised. - New report ·
Local AI Report #4 — Best Laptops for Local AI — New issue of the Local AI Report. - New model ·
GLM-5.3-Flash 320B-A18B — 320 Billion (18B active) mixture-of-experts model from Zhipu AI (Z.ai). - New model ·
Qwen3.8-Flash-Next — 180 Billion (6B active) mixture-of-experts model from Alibaba Cloud. - New model ·
Granite 4.2 30B — 30 Billion dense model from IBM.
Local AI for organizations
Planning private AI for a team or a company? Start with the questions that decide the architecture.
- Private, on-premise inference — Deployment options for running models on infrastructure you control.
- Data sovereignty — Keeping weights, compute and data inside one jurisdiction.
- Infrastructure cost — Hardware and running costs weighed against cloud APIs.
- Model selection — The strongest local models for each business workload.
Frequently asked questions
What is LLM Configurator?
LLM Configurator is a free decision engine for local AI. It checks which of 169 open-weight model variants fit your GPU or Mac, estimates how fast they run, compares hardware and costs, and links 113 setup guides. Every VRAM figure and speed estimate comes from one shared calculation, so the same question gets the same answer on every page. Start with the GPU & VRAM checker.
What can I run locally with my GPU?
It depends on memory first: a model runs well only when its weights, context cache and runtime overhead fit in your GPU's VRAM or your Mac's usable unified memory. Pick your card and we rank every model that fits, with the quantization used and an estimated speed. Models that only fit by spilling into system RAM are listed separately, because they run much slower. Browse ready answers on the Can I Run It? pages.
How much VRAM do I need for a local LLM?
Roughly the model's weights plus a context cache plus a little runtime overhead. Weights scale with total parameters and quantization: Llama 3.1 8B needs about 16 GB of weights at FP16 but only 4.8 GB at Q4_K_M. Mixture-of-experts models need memory for ALL their parameters, not just the active ones. Longer context adds more. Our VRAM requirements guide works through the numbers by model size.
How accurate are the compatibility results?
Memory fit is calculated from each model's published parameter count, the quantization's bits per weight and a context cache, then compared with the usable memory of your card; it is accurate to within the runtime overhead the estimate assumes. Speed is harder: it is a calibrated estimate shown as a range, not a promise. The methodology page publishes every formula and constant we use.
Are the benchmark results measured or estimated?
Mostly estimated, and always labelled. We cite 14 measured runs across 3 GPUs, each credited to the third party that published it; everything else is an estimate from a latency model calibrated against those runs, shown as a range rather than a single number. When reader-submitted runs are reviewed, they appear labelled as community results. See the benchmark database for which is which.
Which GPU is best for local LLMs in 2026?
Choose by memory first, then bandwidth. Memory decides which models fit at all, and nothing else can compensate when a model does not fit. Among cards with enough memory, bandwidth decides how fast tokens come out. Price per gigabyte of VRAM matters more than gaming performance. Compare every card on the GPU comparison table.
Can I run LLMs on a Mac?
Yes. Apple Silicon Macs share one pool of unified memory between the CPU and GPU, so a Mac with a lot of memory can load models no single consumer graphics card can hold. Only part of that pool is available to the GPU by default; we size models against about 75% of it. Bandwidth varies a lot between chips, and it sets the speed. Our Mac buying guide compares chips and memory sizes.
How do I choose the right local model?
Start from the job, not the leaderboard. Coding, agents, document search, writing and translation reward different models. Then filter by what fits your memory at a sensible quantization, and prefer the strongest model that still leaves room for the context length you need. A smaller model that fits comfortably usually beats a bigger one squeezed into RAM. Our best models by workload pages do this ranking for you.
What is the difference between VRAM and unified memory?
VRAM is memory soldered to a graphics card; only the GPU uses it, and it is very fast. Unified memory, on Apple Silicon and some AI mini-PCs, is one pool shared by the CPU and GPU, so the GPU can use far more of it, but the operating system and apps need a share too. For local AI, compare usable memory and bandwidth rather than the label. The hardware database lists both for every device.
Can LLM Configurator help me build a local AI PC?
Yes. The build generator starts from what you want to run and your budget, then recommends a GPU, RAM and storage combination, with the models each configuration can actually load. It also covers compact AI stations and laptops if a tower is not what you want. Each recommendation links to its compatibility results, so you can check the fit before you buy. Try the build generator.
Can I run AI models completely offline?
Yes. Once a model is downloaded, runtimes such as Ollama, LM Studio and llama.cpp run it with no internet connection at all, and your prompts never leave the machine. That also works on phones with offline apps, within much tighter memory limits. For a fully air-gapped setup, download models and installers on a connected machine first, then move them across. Our offline and air-gapped guide covers the steps.
Which runtimes are supported?
Our guides cover the runtimes most people use to run models locally: Ollama, LM Studio, llama.cpp, vLLM for high-throughput serving, and MLX on Apple Silicon. Memory and speed estimates assume GGUF quantizations as these runtimes load them; vLLM and MLX use their own formats, so treat those figures as a guide. Our Ollama vs LM Studio comparison is a good place to start.
How much VRAM does Llama 4 Scout need?
About 66.6 GB at Q4_K_M, before context. Scout is a mixture-of-experts model with 17B active and 109B total parameters: only the active experts are read per token, which keeps it fast, but all 109B must sit in memory. That is why it needs a large unified-memory machine or a data-centre GPU, and why Maverick needs about 242.3 GB. See the Llama 4 Scout model page for every quantization.
Can I run DeepSeek R1 locally?
The full DeepSeek R1 is a 671B mixture-of-experts model that needs about 405.9 GB at Q4_K_M, which means multi-GPU servers or very large unified-memory machines. The R1 Distill models are different, much smaller models trained to imitate it: the Llama 8B distill needs about 5.6 GB and the Qwen 32B distill about 20.1 GB. Compare them on the DeepSeek R1 model page.
What is quantization and why does it matter?
Quantization stores model weights at lower precision, which shrinks the memory they need. Llama 3.1 8B needs about 16 GB of weights at FP16 and about 4.8 GB at Q4_K_M, with a small quality cost. Q4_K_M is the usual balance of size and quality; Q8_0 is close to lossless; below Q4 quality drops noticeably. Read quantization explained for the trade-offs.
Can I run a local LLM on CPU only?
Yes. llama.cpp, and the tools built on it such as Ollama and LM Studio, run models on the CPU with no graphics card at all. It is much slower, because system RAM has far less bandwidth than VRAM, so small models are the practical choice, and you need enough RAM to hold the whole model. The checker handles a machine with no GPU. Our beginner's guide walks through a first install.
Is it cheaper to run AI locally than use ChatGPT?
It depends on how much you use it. Local AI costs the hardware, spread over its useful life, plus electricity for the hours it runs; cloud AI costs per token or per subscription. Heavy, steady use tends to favour local hardware, and light use tends to favour the cloud. Privacy and offline use can matter more than price. Put in your own numbers with the local vs cloud cost calculator.
Stop guessing. Know what your hardware can run.
Compare models, GPUs and local AI setups to find what you can actually run, what you should use, and what it will cost.