Find the right model, hardware & setup for your workload.
Compare models, GPUs and local AI setups to find what you can actually run, what you should use, and what it will cost.
Everything you need to make a local AI decision.
What can I run?
GPU & VRAM checker
Check any GPU or Mac against every model we track and see what fits in its memory.
Check my hardware →What should I run?
Model library
Browse open-weight models by size, licence, context and hardware fit.
Explore models →What should I buy?
Build generator
Give a budget and a workload; get a GPU, RAM and storage recommendation.
Find a build →How fast & expensive?
Benchmarks & cost calculator
Estimated tokens per second by GPU and model, and local-versus-cloud running costs.
See benchmarks →Calculate costs →How do I deploy?
Setup guides
Step-by-step guides for Ollama, LM Studio, vLLM, RAG and fine-tuning.
Read the guides →
What can your hardware actually run?
Pick a GPU or Mac and every model that fits its memory is ranked at Q4_K_M, with the VRAM it needs and an estimated speed. Detection runs in your browser; nothing is uploaded.
Local AI tools
Free calculators and checkers that answer one question each — and share one engine, so their numbers agree.
- VRAM & Fine-Tuning CalculatorExact VRAM and fine-tuning memory for any model and quant
- Upgrade AdvisorName the model you want to run and get ranked routes there — including the free ones
- Best GPU by WorkloadThe strongest GPU for coding, RAG, agents, vision and more
- Decision Request BuilderBuild a /v1/systemone request and export curl, Python or JS — runs in your browser
- Local vs Cloud Cost CalculatorSee when local AI beats ChatGPT pricing
- GPU Benchmark LeaderboardTokens/sec rankings for RTX 4090, Apple M4, and more
- Cloud AI DirectoryRent GPUs, free LLM API tiers, and hosted models — plus local vs cloud comparisons
Explore the model ecosystem
Open-weight models with their real memory requirements, ranked for your hardware and your workload.
- LLM Model Library
Browse 209 open-weight model variants: Llama 4, Qwen 3, Gemma 3, DeepSeek V3
- Best Models for Your GPU
The best models that fit each specific GPU
- Best Models by Workload
Top local models for coding, RAG, agents, writing and more
- Can I Run It?
Straight yes/no verdicts for every model on your GPU
- Best LLMs by VRAM
The strongest models that fit in 8, 12, 16, 24 GB and up
- Local LLM Compare Tool
Side-by-side VRAM, speed, and quant comparison for 2–4 local models
- Local Model Explorer
Scatter plot and table of every local LLM with hardware-aware estimates
Recently updated models
- Specs updated 2026-09-29Qwen3-Coder-Next (80B-A3B MoE)
- Parameters
- 80 Billion (3B active)
- Architecture
- Mixture of experts
- Context
- 262,144
- Memory (Q4)
- 49.1 GB
- Licence
- Apache 2.0
- Specs updated 2026-09-19Apertus 70B
- Parameters
- 70 Billion
- Architecture
- Dense
- Context
- 65,536
- Memory (Q4)
- 43.1 GB
- Licence
- Apache 2.0
- Specs updated 2026-09-19Apertus 8B
- Parameters
- 8 Billion
- Architecture
- Dense
- Context
- 65,536
- Memory (Q4)
- 5.6 GB
- Licence
- Apache 2.0
- Specs updated 2026-09-19Bielik PL 11B v3.0 Instruct
- Parameters
- 11 Billion
- Architecture
- Dense
- Context
- 32,768
- Memory (Q4)
- 7.4 GB
- Licence
- Apache 2.0
- Specs updated 2026-09-19EuroLLM 22B
- Parameters
- 22 Billion
- Architecture
- Dense
- Context
- 4,096
- Memory (Q4)
- 14.1 GB
- Licence
- Apache 2.0
- Specs updated 2026-09-19EuroLLM 9B
- Parameters
- 9 Billion
- Architecture
- Dense
- Context
- 4,096
- Memory (Q4)
- 6.2 GB
- Licence
- Apache 2.0
Memory: Q4_K_M weights plus runtime overhead, before context (KV cache).
Find hardware for local AI
GPUs, Macs, laptops and AI stations compared by the two numbers that decide local AI: memory and bandwidth.
| Device | Memory | Bandwidth | Models (Q4_K_M) |
|---|---|---|---|
| NVIDIA GeForce RTX 4060 | 8 GB | 272 GB/s | 53 fit in VRAM |
| NVIDIA GeForce RTX 5090 | 32 GB | 1792 GB/s | 109 fit in VRAM |
| AMD Radeon RX 7900 XTX | 24 GB | 960 GB/s | 108 fit in VRAM |
| Apple M4 Max | 128 GB unified | 546 GB/s | 125 fit in VRAM |
| Intel Arc B580 | 12 GB | 456 GB/s | 67 fit in VRAM |
| NVIDIA GeForce RTX 5080 | 16 GB | 960 GB/s | 74 fit in VRAM |
Model counts come from the same check as the hardware monitor above, at Q4_K_M.
Know before you run it.
Published measurements, credited to their source — next to the estimates our model makes for every other pairing.
| Model | Hardware | Speed | Evidence |
|---|---|---|---|
| Qwen 3 30B-A3B (MoE) | NVIDIA GeForce RTX 4090 | 195.8 tok/s at Q4_K_XL | Measured · Hardware Corner |
| Qwen 3 8B | NVIDIA GeForce RTX 4090 | 141.3 tok/s at Q4_K_M | Measured · Hardware Corner |
| Llama 3.1 8B Instruct | NVIDIA GeForce RTX 4090 | 131.0 tok/s at Q4_K_XL | Measured · Hardware Corner |
| Qwen 3 14B | NVIDIA GeForce RTX 4090 | 82.8 tok/s at Q4_K_XL | Measured · Hardware Corner |
| Qwen 2.5 7B Instruct | NVIDIA GeForce RTX 4090 | 135.0 tok/s at Q4_K_M | Measured · Mustafa.net |
Measured rows are third-party runs we cite, not our own lab. Every other figure on the site is an estimate, and labelled as one.
What will local AI actually cost?
The cost calculator weighs the same terms for your hardware, your electricity price and your hours of use.
Build a machine around your models.
Start from what you want to run, not from a parts list.
- Workload
- Models
- Budget
- Hardware
- Compatibility
- Performance
- Cost
Learn local AI
Step-by-step guides, from your first model to production serving.
- Updated 2026-07-12Beginner's Guide to Local AIStart here: run your first local model step by step
- Updated 2026-09-29VRAM Requirements 2026How much GPU memory each model size really needs
- Updated 2026-07-12Quantization ExplainedWhat Q4, Q5, Q8 and GGUF mean for memory and quality
- Updated 2026-07-12Local RAG GuideChat with your own documents, fully offline
- Updated 2026-07-12Fine-Tuning Guide (LoRA/DPO)Train a model on your own data with LoRA and DPO
- Updated 2026-07-12Building AI AgentsBuild autonomous agents on top of local models
More topics
What is a Private LLM?
A Private LLM is an artificial intelligence model that runs entirely on your own infrastructure—whether that's a powerful gaming laptop, an on-premise server, or a private cloud instance.
Unlike using public APIs like ChatGPT or Claude where your data is sent to external servers, a private LLM ensures your data never leaves your device. This allows businesses to process sensitive documents, proprietary code, and personal information with zero risk of data leakage or model training on your intellectual property.
| Feature | Public Cloud API | Private Local LLM |
|---|---|---|
| Cost Structure | Recurring monthly fees + Cost per token (Expensive at scale) | Free*Just electricity & hardware |
| Data Privacy | Data sent to 3rd party servers. Risk of retention/training. | 100% PrivateData stays on your disk |
| Latency | Network dependent. Unpredictable spikes. | Zero Network LatencySpeed depends on GPU |
| Control | Censored / Guardrailed. Models can change anytime. | Full ControlUncensored / Fine-tunable |
Local AI data, benchmarks & research
Reports, open datasets and the methodology behind every number.
- Methodology
How every VRAM and speed figure on this site is calculated
- Local AI Reports
Biweekly digest: new models, GPU price moves, tooling updates
- Efficiency Leaderboard
Tokens per watt, ranked across GPUs
- Local vs Cloud Comparisons
When local hardware beats the OpenAI and Anthropic APIs
- Training Dataset Hub
47 curated datasets for fine-tuning and pretraining
- Local AI Blog
Guides, comparisons, and insights on running AI locally
What's new in local AI
The newest dated records in our catalogue: model releases, new hardware, blog posts, published measurements, reports and guide updates.
- New hardwareApple M4 Max (32-core GPU)Apple hardware added to the catalogue.
- New hardwareApple M5 Max (32-core GPU)Apple hardware added to the catalogue.
- New hardwareApple M6Apple hardware added to the catalogue.
- New blog postJev and the decision-model wave, explained — and how to run one on your own machineNew post on the blog.
- Guide updatedBest Mac for Local LLMsSetup guide revised.
- Guide updatedBest GPU Buyer's Guide 2026Setup guide revised.
- Guide updatedVRAM Requirements GuideSetup guide revised.
- New reportLocal AI Report #5 — Europe's AI Compute Gap and SovereigntyNew issue of the Local AI Report.
Local AI for organizations
Planning private AI for a team or a company? Start with the questions that decide the architecture.
- Private, on-premise inferenceDeployment options for running models on infrastructure you control.
- Data sovereigntyKeeping weights, compute and data inside one jurisdiction.
- Infrastructure costHardware and running costs weighed against cloud APIs.
- Model selectionThe strongest local models for each business workload.
Frequently asked questions
Straight answers about running AI models on your own hardware.
What is LLM Configurator?
LLM Configurator is a free decision engine for local AI. It checks which of 169 open-weight model variants fit your GPU or Mac, estimates how fast they run, compares hardware and costs, and links 114 setup guides. Every VRAM figure and speed estimate comes from one shared calculation, so the same question gets the same answer on every page. Start with the GPU & VRAM checker.
What can I run locally with my GPU?
It depends on memory first: a model runs well only when its weights, context cache and runtime overhead fit in your GPU's VRAM or your Mac's usable unified memory. Pick your card and we rank every model that fits, with the quantization used and an estimated speed. Models that only fit by spilling into system RAM are listed separately, because they run much slower. Browse ready answers on the Can I Run It? pages.
How much VRAM do I need for a local LLM?
Roughly the model's weights plus a context cache plus a little runtime overhead. Weights scale with total parameters and quantization: Llama 3.1 8B needs about 16 GB of weights at FP16 but only 4.8 GB at Q4_K_M. Mixture-of-experts models need memory for ALL their parameters, not just the active ones. Longer context adds more. Our VRAM requirements guide works through the numbers by model size.
How accurate are the compatibility results?
Memory fit is calculated from each model's published parameter count, the quantization's bits per weight and a context cache, then compared with the usable memory of your card; it is accurate to within the runtime overhead the estimate assumes. Speed is harder: it is a calibrated estimate shown as a range, not a promise. The methodology page publishes every formula and constant we use.
Are the benchmark results measured or estimated?
Mostly estimated, and always labelled. We cite 14 measured runs across 3 GPUs, each credited to the third party that published it; everything else is an estimate from a latency model calibrated against those runs, shown as a range rather than a single number. When reader-submitted runs are reviewed, they appear labelled as community results. See the benchmark database for which is which.
Which GPU is best for local LLMs in 2026?
Choose by memory first, then bandwidth. Memory decides which models fit at all, and nothing else can compensate when a model does not fit. Among cards with enough memory, bandwidth decides how fast tokens come out. Price per gigabyte of VRAM matters more than gaming performance. Compare every card on the GPU comparison table.
Can I run LLMs on a Mac?
Yes. Apple Silicon Macs share one pool of unified memory between the CPU and GPU, so a Mac with a lot of memory can load models no single consumer graphics card can hold. Only part of that pool is available to the GPU by default; we size models against about 75% of it. Bandwidth varies a lot between chips, and it sets the speed. Our Mac buying guide compares chips and memory sizes.
How do I choose the right local model?
Start from the job, not the leaderboard. Coding, agents, document search, writing and translation reward different models. Then filter by what fits your memory at a sensible quantization, and prefer the strongest model that still leaves room for the context length you need. A smaller model that fits comfortably usually beats a bigger one squeezed into RAM. Our best models by workload pages do this ranking for you.
What is the difference between VRAM and unified memory?
VRAM is memory soldered to a graphics card; only the GPU uses it, and it is very fast. Unified memory, on Apple Silicon and some AI mini-PCs, is one pool shared by the CPU and GPU, so the GPU can use far more of it, but the operating system and apps need a share too. For local AI, compare usable memory and bandwidth rather than the label. The hardware database lists both for every device.
Can LLM Configurator help me build a local AI PC?
Yes. The build generator starts from what you want to run and your budget, then recommends a GPU, RAM and storage combination, with the models each configuration can actually load. It also covers compact AI stations and laptops if a tower is not what you want. Each recommendation links to its compatibility results, so you can check the fit before you buy. Try the build generator.
Can I run AI models completely offline?
Yes. Once a model is downloaded, runtimes such as Ollama, LM Studio and llama.cpp run it with no internet connection at all, and your prompts never leave the machine. That also works on phones with offline apps, within much tighter memory limits. For a fully air-gapped setup, download models and installers on a connected machine first, then move them across. Our offline and air-gapped guide covers the steps.
Which runtimes are supported?
Our guides cover the runtimes most people use to run models locally: Ollama, LM Studio, llama.cpp, vLLM for high-throughput serving, and MLX on Apple Silicon. Memory and speed estimates assume GGUF quantizations as these runtimes load them; vLLM and MLX use their own formats, so treat those figures as a guide. Our Ollama vs LM Studio comparison is a good place to start.
How much VRAM does Llama 4 Scout need?
About 66.6 GB at Q4_K_M, before context. Scout is a mixture-of-experts model with 17B active and 109B total parameters: only the active experts are read per token, which keeps it fast, but all 109B must sit in memory. That is why it needs a large unified-memory machine or a data-centre GPU, and why Maverick needs about 242.3 GB. See the Llama 4 Scout model page for every quantization.
Can I run DeepSeek R1 locally?
The full DeepSeek R1 is a 671B mixture-of-experts model that needs about 405.9 GB at Q4_K_M, which means multi-GPU servers or very large unified-memory machines. The R1 Distill models are different, much smaller models trained to imitate it: the Llama 8B distill needs about 5.6 GB and the Qwen 32B distill about 20.1 GB. Compare them on the DeepSeek R1 model page.
What is quantization and why does it matter?
Quantization stores model weights at lower precision, which shrinks the memory they need. Llama 3.1 8B needs about 16 GB of weights at FP16 and about 4.8 GB at Q4_K_M, with a small quality cost. Q4_K_M is the usual balance of size and quality; Q8_0 is close to lossless; below Q4 quality drops noticeably. Read quantization explained for the trade-offs.
Can I run a local LLM on CPU only?
Yes. llama.cpp, and the tools built on it such as Ollama and LM Studio, run models on the CPU with no graphics card at all. It is much slower, because system RAM has far less bandwidth than VRAM, so small models are the practical choice, and you need enough RAM to hold the whole model. The checker handles a machine with no GPU. Our beginner's guide walks through a first install.
Is it cheaper to run AI locally than use ChatGPT?
It depends on how much you use it. Local AI costs the hardware, spread over its useful life, plus electricity for the hours it runs; cloud AI costs per token or per subscription. Heavy, steady use tends to favour local hardware, and light use tends to favour the cloud. Privacy and offline use can matter more than price. Put in your own numbers with the local vs cloud cost calculator.
Written by Jakub Rusinowski · Last updated September 30, 2026
Stop guessing. Know what your hardware can run.
Compare models, GPUs and local AI setups to find what you can actually run, what you should use, and what it will cost.