Local AI on Linux
Written by Jakub Rusinowski · Last updated October 7, 2026
The fastest and most memory-efficient platform, and the only one where every serving stack runs natively — at the cost of doing your own driver work.
Runtimes that work on Linux
| Runtime | How to install | Acceleration | Watch out for |
|---|---|---|---|
| Ollama | Official install script.curl -fsSL https://ollama.com/install.sh | sh | CUDA (NVIDIA), ROCm (AMD), Vulkan, CPU | — |
| llama.cpp | Build from source with the backend flag for your GPU.cmake -B build -DGGML_CUDA=ON && cmake --build build --config Release | CUDA (NVIDIA), ROCm (AMD), Vulkan, SYCL (Intel), CPU | Build flags decide the backend — a default build is CPU-only and will be an order of magnitude slower than it should be. |
| vLLM | pip, ideally into a clean virtual environment.pip install vllm | CUDA (NVIDIA), ROCm (AMD) | Built for concurrent serving, and it pre-allocates most of the GPU's memory on startup. For a single user it uses far more memory than llama.cpp for no benefit. |
| LM Studio | AppImage. | CUDA (NVIDIA), Vulkan, CPU | — |
What Linux is good at
- The lowest OS memory overhead of the three, so more of the card is available to the model.
- The only platform where vLLM, SGLang and the fine-tuning stacks all run natively.
- Best AMD support: ROCm coverage is materially wider here than on Windows.
Linux limitations
- NVIDIA driver and CUDA toolkit installation is on you, and a version mismatch between driver, toolkit and runtime is the most common cause of a silent fall back to CPU.
- AMD ROCm supports a specific list of cards; a supported-on-paper GPU can still need an environment override to work.
- No first-party GUI in most distributions — the workflow assumes a terminal.
When it breaks on Linux
- nvidia-smi stopped working after a kernel update
- NVML driver/library version mismatch
- Secure Boot is refusing to load the driver module
- The systemd service ignores your environment variables
- A container cannot see the GPU
- ROCm does not target your Radeon — overrides and the Vulkan path
- Permission denied on /dev/kfd
- Headless server: the disk fills and the API only answers on localhost
Hardware and models on Linux
Linux supports NVIDIA, AMD, Intel accelerators — 86 of the 109 in our hardware database, up to 128 GB.
| Memory | Hardware on this platform | Models to run |
|---|---|---|
| 8 GB | NVIDIA GeForce RTX 3080 (10GB) Intel Arc B570 NVIDIA GeForce RTX 4060 Ti 8GB | Qwen 3 8B GLM-6 9B GLM-4 9B |
| 16 GB | AMD Radeon RX 7900 XT NVIDIA GeForce RTX 4090 Laptop GPU NVIDIA GeForce RTX 5080 | Mistral Small 3.1 24B Qwen 3.5 14B Qwen 3.5 14B |
| 24 GB | NVIDIA GeForce RTX 5090 Laptop GPU NVIDIA GeForce RTX 3090 Ti NVIDIA GeForce RTX 4090 | Gemma 4 27B ⭐ Mistral Small 3.1 24B Qwen3.8 27B |
| 48 GB | NVIDIA RTX 6000 Ada Generation NVIDIA L40S NVIDIA RTX PRO 5000 Blackwell (48 GB) | GLM-5.1 72B Qwen 3.5 72B Gemma 4 27B ⭐ |
| 96 GB | NVIDIA RTX PRO 6000 Blackwell NVIDIA RTX PRO 6000 Blackwell (Max-Q Edition) | GPT-OSS 120B Qwen 3.5 122B-A10B (MoE) GLM-5.1 72B |
Machines that run Linux
| Machine | Memory for models | Form factor | Upgradeable |
|---|---|---|---|
| RTX 4090 Desktop (24 GB VRAM, 64 GB RAM) | 24 GB VRAM | Desktop | Yes |
| RTX 5090 Desktop (32 GB VRAM, 64 GB RAM) | 32 GB VRAM | Desktop | Yes |
| RTX 3090 Desktop (24 GB VRAM, 64 GB RAM) | 24 GB VRAM | Desktop | Yes |
| NVIDIA DGX Spark (128 GB) | 128 GB unified | Mini PC | No |
| RTX 5080 Desktop (16 GB VRAM, 32 GB RAM) | 16 GB VRAM | Desktop | Yes |
| RTX 3060 12 GB Desktop (12 GB VRAM, 32 GB RAM) | 12 GB VRAM | Desktop | Yes |
| RTX 5060 Ti 16 GB Desktop (16 GB VRAM, 32 GB RAM) | 16 GB VRAM | Desktop | Yes |
| Framework Desktop (Ryzen AI Max+ 395, 128 GB) | 128 GB unified | Mini PC | No |
FAQ
What is the best way to run an LLM locally on Linux?
Ollama — Official install script. It reaches CUDA (NVIDIA), ROCm (AMD), Vulkan, CPU.
Which GPUs work for local AI on Linux?
NVIDIA, AMD, Intel hardware, up to 128 GB in our database.
What are the downsides of running local AI on Linux?
NVIDIA driver and CUDA toolkit installation is on you, and a version mismatch between driver, toolkit and runtime is the most common cause of a silent fall back to CPU.
Machines Running This Platform
- Run BitNet b1.58 on RTX 4090 desktop
- Run BitNet b1.58 on RTX 5090 desktop
- Run BitNet b1.58 on RTX 3090 desktop
Hardware for This Platform
- Best models for the NVIDIA GB10 Grace Blackwell
- Best models for the NVIDIA DGX Spark
- Best models for the AMD Ryzen AI Max+ 395
- Best models for the NVIDIA RTX PRO 6000 Blackwell