Best Local LLMs for the NVIDIA DGX Spark

Written by Jakub Rusinowski · Last updated July 21, 2026

Best all-round pick: GPT-OSS 120B

The NVIDIA DGX Spark has 128 GB of VRAM, of which about 128 GB is available to a model. The largest model it can hold is Qwen3.8-Flash-Next (180B, 109.5 GB at Q4_K_M).

Best models overall on the NVIDIA DGX Spark

ModelScoreMemory at Q4_K_MContextLicence
GPT-OSS 120B95.871.3 GB128KApache-2.0
Qwen 3.5 122B-A10B (MoE)94.274.5 GB125KApache 2.0
Qwen 3.5 122B-A10B92.574.5 GB128KApache-2.0
GLM-5.1 72B91.244.3 GB125KMIT
Command R+ (104B)91.263.6 GB125KCC-BY-NC
Llama 4 Scout 17B91.266.6 GB10MLlama Community

Best model by what you are doing

WorkloadRecommended on this GPU
CodingDevstral-2 123B (99)
Qwen3-Coder-Next (80B-A3B MoE) (96.6)
Qwen 3.5 122B-A10B (MoE) (96.5)
General assistantGPT-OSS 120B (95.8)
Qwen 3.5 122B-A10B (MoE) (94.2)
Qwen 3.5 122B-A10B (92.5)
ReasoningGLM-5.1 72B (98.4)
Qwen 3.5 122B-A10B (MoE) (97)
Gemma 4 27B ⭐ (95.9)
RAGLlama 4 Scout 17B (93.6)
Qwen3.8-Flash-Next (93)
Gemma 4 31B (92.8)
AgentsDevstral-2 123B (95.2)
Qwen3.8-Flash-Next (93.9)
Qwen3-Coder-Next (80B-A3B MoE) (93.1)
VisionQwen 3.5 72B (96.9)
Llama 3.2 90B Vision Instruct (95.2)
Qwen3.8-Flash-Next (93.7)

How these numbers are calculated

FAQ

What is the best LLM for the NVIDIA DGX Spark?

GPT-OSS 120B is the strongest all-round pick that fits its 128 GB.

What is the largest model the NVIDIA DGX Spark can run?

Qwen3.8-Flash-Next — 180B parameters, needing 109.5 GB at Q4_K_M.

How much of the NVIDIA DGX Spark's memory can a model actually use?

About 128 GB of its 128 GB, before the desktop and runtime overhead are accounted for.

Compatibility Checks for the NVIDIA DGX Spark

Similar GPUs

By Workload

More

← All GPUs | NVIDIA DGX Spark specs