Local AI decision engine

Find the right model, hardware & setup for your workload.

Compare models, GPUs and local AI setups to find what you can actually run, what you should use, and what it will cost.

169 Model variants · 113 Guides · 70 GPUs & chips · Free — Always open

Written by Jakub Rusinowski · Last updated September 19, 2026

Everything you need to make a local AI decision.

  1. 01 What can I run?

    GPU & VRAM checker — Check any GPU or Mac against every model we track and see what fits in its memory.

    Check my hardware →
  2. 02 What should I run?

    Model library — Browse open-weight models by size, licence, context and hardware fit.

    Explore models →
  3. 03 What should I buy?

    Build generator — Give a budget and a workload; get a GPU, RAM and storage recommendation.

    Find a build →
  4. 04 How fast & expensive?

    Benchmarks & cost calculator — Estimated tokens per second by GPU and model, and local-versus-cloud running costs.

    See benchmarks → · Calculate costs →
  5. 05 How do I deploy?

    Setup guides — Step-by-step guides for Ollama, LM Studio, vLLM, RAG and fine-tuning.

    Read the guides →

What can your hardware actually run?

Pick a GPU or Mac and every model that fits its memory is ranked at Q4_K_M, with the VRAM it needs and an estimated speed. Detection runs in your browser; nothing is uploaded.

Open the GPU & VRAM checker →

Or browse ready-made answers:

Local AI tools

Free calculators and checkers that answer one question each — and share one engine, so their numbers agree.

GPU & VRAM Checker
Find which LLMs your GPU can run — Llama, DeepSeek, Gemma

Explore the model ecosystem

Open-weight models with their real memory requirements, ranked for your hardware and your workload.

Recently updated models

  • Apertus 70B — Parameters: 70 Billion · Architecture: Dense · Context: 65,536 · Memory (Q4): 43.1 GB · Licence: Apache 2.0 (Specs updated 2026-09-19)
  • Apertus 8B — Parameters: 8 Billion · Architecture: Dense · Context: 65,536 · Memory (Q4): 5.6 GB · Licence: Apache 2.0 (Specs updated 2026-09-19)
  • Bielik PL 11B v3.0 Instruct — Parameters: 11 Billion · Architecture: Dense · Context: 32,768 · Memory (Q4): 7.4 GB · Licence: Apache 2.0 (Specs updated 2026-09-19)
  • EuroLLM 22B — Parameters: 22 Billion · Architecture: Dense · Context: 4,096 · Memory (Q4): 14.1 GB · Licence: Apache 2.0 (Specs updated 2026-09-19)
  • EuroLLM 9B — Parameters: 9 Billion · Architecture: Dense · Context: 4,096 · Memory (Q4): 6.2 GB · Licence: Apache 2.0 (Specs updated 2026-09-19)
  • GLM-4.6V-Flash 9B — Parameters: 9 Billion · Architecture: Dense · Context: 65,536 · Memory (Q4): 6.2 GB · Licence: MIT (Specs updated 2026-09-19)

Memory: Q4_K_M weights plus runtime overhead, before context (KV cache).

Find hardware for local AI

GPUs, Macs, laptops and AI stations compared by the two numbers that decide local AI: memory and bandwidth.

DeviceMemoryBandwidthModels (Q4_K_M)
NVIDIA GeForce RTX 40608 GB272 GB/s53 fit in VRAM
NVIDIA GeForce RTX 509032 GB1792 GB/s109 fit in VRAM
AMD Radeon RX 7900 XTX24 GB960 GB/s108 fit in VRAM
Apple M4 Max128 GB unified546 GB/s125 fit in VRAM
Intel Arc B58012 GB456 GB/s67 fit in VRAM
NVIDIA GeForce RTX 508016 GB960 GB/s74 fit in VRAM

Model counts come from the same check as the hardware monitor above, at Q4_K_M.

Know before you run it.

Published measurements, credited to their source — next to the estimates our model makes for every other pairing.

ModelHardwareSpeedEvidence
Qwen 3 30B-A3B (MoE)NVIDIA GeForce RTX 4090195.8 tok/s at Q4_K_XLMeasured · Hardware Corner
Qwen 3 8BNVIDIA GeForce RTX 4090141.3 tok/s at Q4_K_MMeasured · Hardware Corner
Llama 3.1 8B InstructNVIDIA GeForce RTX 4090131.0 tok/s at Q4_K_XLMeasured · Hardware Corner
Qwen 3 14BNVIDIA GeForce RTX 409082.8 tok/s at Q4_K_XLMeasured · Hardware Corner
Qwen 2.5 7B InstructNVIDIA GeForce RTX 4090135.0 tok/s at Q4_K_MMeasured · Mustafa.net

Measured rows are third-party runs we cite, not our own lab. Every other figure on the site is an estimate, and labelled as one.

GPU Benchmark Leaderboard · Methodology

What will local AI actually cost?

The cost calculator weighs the same terms for your hardware, your electricity price and your hours of use.

Hardware (purchase price ÷ the years you depreciate it over) + Electricity (power draw × hours of use a day × your price per kWh) vs Cloud (the equivalent API or rented-GPU spend)

Local vs Cloud Cost Calculator · Compare local vs cloud →

Build a machine around your models.

Start from what you want to run, not from a parts list.

  1. Workload
  2. Models
  3. Budget
  4. Hardware
  5. Compatibility
  6. Performance
  7. Cost

Open the build generator →

Learn local AI

Step-by-step guides, from your first model to production serving.

More topics

What is a Private LLM?

A Private LLM is an artificial intelligence model that runs entirely on your own infrastructure—whether that's a powerful gaming laptop, an on-premise server, or a private cloud instance.

Unlike using public APIs like ChatGPT or Claude where your data is sent to external servers, a private LLM ensures your data never leaves your device. This allows businesses to process sensitive documents, proprietary code, and personal information with zero risk of data leakage or model training on your intellectual property.

Cloud vs. Local: The Breakdown

Feature Public Cloud API Private Local LLM
Cost StructureRecurring monthly fees + Cost per token (Expensive at scale)Free* — Just electricity & hardware
Data PrivacyData sent to 3rd party servers. Risk of retention/training.100% Private — Data stays on your disk
LatencyNetwork dependent. Unpredictable spikes.Zero Network Latency — Speed depends on GPU
ControlCensored / Guardrailed. Models can change anytime.Full Control — Uncensored / Fine-tunable

Local AI data, benchmarks & research

Reports, open datasets and the methodology behind every number.

Local AI Report #5 — Europe's AI Compute Gap and Sovereignty Published 2026-09-14
The EU runs just 2 GW of AI compute to America's 35 GW. Inside Europe's 19 AI Factories, the gigafactory tender, and Luxembourg's 20-exaflop sovereignty bet.

What's new in local AI

The newest dated records in our catalogue: model releases, published measurements, reports and guide updates.

Local AI for organizations

Planning private AI for a team or a company? Start with the questions that decide the architecture.

AI Compliance & Readiness

Frequently asked questions

What is LLM Configurator?

LLM Configurator is a free decision engine for local AI. It checks which of 169 open-weight model variants fit your GPU or Mac, estimates how fast they run, compares hardware and costs, and links 113 setup guides. Every VRAM figure and speed estimate comes from one shared calculation, so the same question gets the same answer on every page. Start with the GPU & VRAM checker.

What can I run locally with my GPU?

It depends on memory first: a model runs well only when its weights, context cache and runtime overhead fit in your GPU's VRAM or your Mac's usable unified memory. Pick your card and we rank every model that fits, with the quantization used and an estimated speed. Models that only fit by spilling into system RAM are listed separately, because they run much slower. Browse ready answers on the Can I Run It? pages.

How much VRAM do I need for a local LLM?

Roughly the model's weights plus a context cache plus a little runtime overhead. Weights scale with total parameters and quantization: Llama 3.1 8B needs about 16 GB of weights at FP16 but only 4.8 GB at Q4_K_M. Mixture-of-experts models need memory for ALL their parameters, not just the active ones. Longer context adds more. Our VRAM requirements guide works through the numbers by model size.

How accurate are the compatibility results?

Memory fit is calculated from each model's published parameter count, the quantization's bits per weight and a context cache, then compared with the usable memory of your card; it is accurate to within the runtime overhead the estimate assumes. Speed is harder: it is a calibrated estimate shown as a range, not a promise. The methodology page publishes every formula and constant we use.

Are the benchmark results measured or estimated?

Mostly estimated, and always labelled. We cite 14 measured runs across 3 GPUs, each credited to the third party that published it; everything else is an estimate from a latency model calibrated against those runs, shown as a range rather than a single number. When reader-submitted runs are reviewed, they appear labelled as community results. See the benchmark database for which is which.

Which GPU is best for local LLMs in 2026?

Choose by memory first, then bandwidth. Memory decides which models fit at all, and nothing else can compensate when a model does not fit. Among cards with enough memory, bandwidth decides how fast tokens come out. Price per gigabyte of VRAM matters more than gaming performance. Compare every card on the GPU comparison table.

Can I run LLMs on a Mac?

Yes. Apple Silicon Macs share one pool of unified memory between the CPU and GPU, so a Mac with a lot of memory can load models no single consumer graphics card can hold. Only part of that pool is available to the GPU by default; we size models against about 75% of it. Bandwidth varies a lot between chips, and it sets the speed. Our Mac buying guide compares chips and memory sizes.

How do I choose the right local model?

Start from the job, not the leaderboard. Coding, agents, document search, writing and translation reward different models. Then filter by what fits your memory at a sensible quantization, and prefer the strongest model that still leaves room for the context length you need. A smaller model that fits comfortably usually beats a bigger one squeezed into RAM. Our best models by workload pages do this ranking for you.

What is the difference between VRAM and unified memory?

VRAM is memory soldered to a graphics card; only the GPU uses it, and it is very fast. Unified memory, on Apple Silicon and some AI mini-PCs, is one pool shared by the CPU and GPU, so the GPU can use far more of it, but the operating system and apps need a share too. For local AI, compare usable memory and bandwidth rather than the label. The hardware database lists both for every device.

Can LLM Configurator help me build a local AI PC?

Yes. The build generator starts from what you want to run and your budget, then recommends a GPU, RAM and storage combination, with the models each configuration can actually load. It also covers compact AI stations and laptops if a tower is not what you want. Each recommendation links to its compatibility results, so you can check the fit before you buy. Try the build generator.

Can I run AI models completely offline?

Yes. Once a model is downloaded, runtimes such as Ollama, LM Studio and llama.cpp run it with no internet connection at all, and your prompts never leave the machine. That also works on phones with offline apps, within much tighter memory limits. For a fully air-gapped setup, download models and installers on a connected machine first, then move them across. Our offline and air-gapped guide covers the steps.

Which runtimes are supported?

Our guides cover the runtimes most people use to run models locally: Ollama, LM Studio, llama.cpp, vLLM for high-throughput serving, and MLX on Apple Silicon. Memory and speed estimates assume GGUF quantizations as these runtimes load them; vLLM and MLX use their own formats, so treat those figures as a guide. Our Ollama vs LM Studio comparison is a good place to start.

How much VRAM does Llama 4 Scout need?

About 66.6 GB at Q4_K_M, before context. Scout is a mixture-of-experts model with 17B active and 109B total parameters: only the active experts are read per token, which keeps it fast, but all 109B must sit in memory. That is why it needs a large unified-memory machine or a data-centre GPU, and why Maverick needs about 242.3 GB. See the Llama 4 Scout model page for every quantization.

Can I run DeepSeek R1 locally?

The full DeepSeek R1 is a 671B mixture-of-experts model that needs about 405.9 GB at Q4_K_M, which means multi-GPU servers or very large unified-memory machines. The R1 Distill models are different, much smaller models trained to imitate it: the Llama 8B distill needs about 5.6 GB and the Qwen 32B distill about 20.1 GB. Compare them on the DeepSeek R1 model page.

What is quantization and why does it matter?

Quantization stores model weights at lower precision, which shrinks the memory they need. Llama 3.1 8B needs about 16 GB of weights at FP16 and about 4.8 GB at Q4_K_M, with a small quality cost. Q4_K_M is the usual balance of size and quality; Q8_0 is close to lossless; below Q4 quality drops noticeably. Read quantization explained for the trade-offs.

Can I run a local LLM on CPU only?

Yes. llama.cpp, and the tools built on it such as Ollama and LM Studio, run models on the CPU with no graphics card at all. It is much slower, because system RAM has far less bandwidth than VRAM, so small models are the practical choice, and you need enough RAM to hold the whole model. The checker handles a machine with no GPU. Our beginner's guide walks through a first install.

Is it cheaper to run AI locally than use ChatGPT?

It depends on how much you use it. Local AI costs the hardware, spread over its useful life, plus electricity for the hours it runs; cloud AI costs per token or per subscription. Heavy, steady use tends to favour local hardware, and light use tends to favour the cloud. Privacy and offline use can matter more than price. Put in your own numbers with the local vs cloud cost calculator.

Stop guessing. Know what your hardware can run.

Compare models, GPUs and local AI setups to find what you can actually run, what you should use, and what it will cost.

Find My Setup → · Explore the model library →