Local AI on Windows
Written by Jakub Rusinowski · Last updated October 7, 2026
The widest hardware choice and the easiest NVIDIA setup, at the cost of the highest idle memory overhead of the three platforms.
Runtimes that work on Windows
| Runtime | How to install | Acceleration | Watch out for |
|---|---|---|---|
| Ollama | Official Windows installer (.exe), or winget.winget install Ollama.Ollama | CUDA (NVIDIA), ROCm (select AMD), Vulkan, CPU | ROCm support on Windows covers far fewer AMD cards than on Linux; check your specific model before relying on it. |
| LM Studio | Official Windows installer, GUI-first with a built-in model browser. | CUDA (NVIDIA), Vulkan, CPU | — |
| llama.cpp | Prebuilt release binaries, or build from source with CUDA enabled.winget install llama.cpp | CUDA (NVIDIA), Vulkan, SYCL (Intel), CPU | Prebuilt Windows binaries are split by backend — download the CUDA build, not the CPU one, or you will silently run on CPU. |
| vLLM | Runs under WSL2, not natively on Windows.wsl --install | CUDA (NVIDIA, via WSL2) | WSL2 splits memory with Windows and adds a virtualisation layer; for single-user local inference Ollama or llama.cpp is the simpler choice. |
What Windows is good at
- The broadest GPU compatibility of the three platforms — NVIDIA, AMD and Intel Arc all have working paths.
- NVIDIA setup is genuinely one installer; no driver or kernel-module work.
- Easiest platform on which to reuse an existing gaming GPU for local AI.
Windows limitations
- Windows itself reserves noticeably more VRAM and system RAM than Linux, so usable memory for a model is lower on identical hardware.
- AMD ROCm coverage is narrower on Windows than on Linux; Vulkan is the reliable fallback but is slower than ROCm where both work.
- The desktop compositor holds VRAM even when idle, which matters most on 8–12 GB cards where every gigabyte counts.
When it breaks on Windows
- Ollama is running on the CPU even though you have an NVIDIA card
- CUDA works in PowerShell but not inside WSL2
- The model crawls instead of crashing — shared GPU memory
- The terminal cannot find the `ollama` command
- Moving the model store off a full C: drive
- Nothing on the LAN can reach the API — OLLAMA_HOST and the firewall
- A hybrid laptop picked the integrated GPU
- SmartScreen or Defender blocked the install or ate a GGUF
Hardware and models on Windows
Windows supports NVIDIA, AMD, Intel accelerators — 86 of the 109 in our hardware database, up to 128 GB.
| Memory | Hardware on this platform | Models to run |
|---|---|---|
| 8 GB | NVIDIA GeForce RTX 3080 (10GB) Intel Arc B570 NVIDIA GeForce RTX 4060 Ti 8GB | Qwen 3 8B GLM-6 9B GLM-4 9B |
| 16 GB | AMD Radeon RX 7900 XT NVIDIA GeForce RTX 4090 Laptop GPU NVIDIA GeForce RTX 5080 | Mistral Small 3.1 24B Qwen 3.5 14B Qwen 3.5 14B |
| 24 GB | NVIDIA GeForce RTX 5090 Laptop GPU NVIDIA GeForce RTX 3090 Ti NVIDIA GeForce RTX 4090 | Gemma 4 27B ⭐ Mistral Small 3.1 24B Qwen3.8 27B |
| 48 GB | NVIDIA RTX 6000 Ada Generation NVIDIA L40S NVIDIA RTX PRO 5000 Blackwell (48 GB) | GLM-5.1 72B Qwen 3.5 72B Gemma 4 27B ⭐ |
| 96 GB | NVIDIA RTX PRO 6000 Blackwell NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition) | GPT-OSS 120B Qwen 3.5 122B-A10B (MoE) GLM-5.1 72B |
Machines that run Windows
| Machine | Memory for models | Form factor | Upgradeable |
|---|---|---|---|
| RTX 4090 Desktop (24 GB VRAM, 64 GB RAM) | 24 GB VRAM | Desktop | Yes |
| RTX 5090 Desktop (32 GB VRAM, 64 GB RAM) | 32 GB VRAM | Desktop | Yes |
| RTX 3090 Desktop (24 GB VRAM, 64 GB RAM) | 24 GB VRAM | Desktop | Yes |
| RTX 5080 Desktop (16 GB VRAM, 32 GB RAM) | 16 GB VRAM | Desktop | Yes |
| RTX 3060 12 GB Desktop (12 GB VRAM, 32 GB RAM) | 12 GB VRAM | Desktop | Yes |
| RTX 5060 Ti 16 GB Desktop (16 GB VRAM, 32 GB RAM) | 16 GB VRAM | Desktop | Yes |
| Framework Desktop (Ryzen AI Max+ 395, 128 GB) | 128 GB unified | Mini PC | No |
| RTX 4060 Laptop (8 GB VRAM, 16 GB RAM) | 8 GB VRAM | Laptop | Yes |
FAQ
What is the best way to run an LLM locally on Windows?
Ollama — Official Windows installer (.exe), or winget. It reaches CUDA (NVIDIA), ROCm (select AMD), Vulkan, CPU.
Which GPUs work for local AI on Windows?
NVIDIA, AMD, Intel hardware, up to 128 GB in our database.
What are the downsides of running local AI on Windows?
Windows itself reserves noticeably more VRAM and system RAM than Linux, so usable memory for a model is lower on identical hardware.
Machines Running This Platform
- Run BitNet b1.58 on RTX 4090 desktop
- Run BitNet b1.58 on RTX 5090 desktop
- Run BitNet b1.58 on RTX 3090 desktop
- Run BitNet b1.58 on RTX 5080 desktop
Hardware for This Platform
- Best models for the NVIDIA GB10 Grace Blackwell
- Best models for the NVIDIA DGX Spark
- Best models for the AMD Ryzen AI Max+ 395
- Best models for the NVIDIA RTX PRO 6000 Blackwell