NVIDIA GB10 Grace Blackwell — Local LLM Performance & Compatibility

Written by Jakub Rusinowski · Last updated September 19, 2026

The silicon inside DGX Spark and the OEM boxes built on it — ASUS Ascent GX10, Dell Pro Max, HP ZGX Nano, Lenovo ThinkStation PGX. A 20-core Grace ARM CPU and a Blackwell GPU sharing 128 GB of LPDDR5X at 273 GB/s, running the full CUDA stack. Recorded as a CHIP so those machines can share one record instead of restating it.

Technical Specifications

VRAM128 GB
Memory Bandwidth273 GB/s
TDP140 W
ArchitectureGB10 Grace Blackwell Superchip
Release Year2025
MSRP at Launch$0
Inference Speed (Llama 3.1 8B Q4_K_M)26–53 tok/s (estimated)
Inference Speed (Llama 3.3 70B Q4_K_M)3.4–7.1 tok/s (estimated)
Buy This HardwareRyzen AI Max+ 395 Laptop (Strix Halo, up to 128GB) — 128 GB VRAM · 120 W board powerDeploy in the Cloud NowRunPod

or compare on Vast.ai

As an Amazon Associate we earn from qualifying purchases. Cloud GPU links are referral links — we may earn a commission at no extra cost to you.

LLMs Compatible with 128 GB VRAM

All models below run comfortably in 128 GB VRAM with Q4_K_M quantization.

Qwen3.8Qwen3.8-Flash-Next · 109 GB VRAM · Q4_K_M · qwen3-8
Qwen 3.5Qwen 3.5 122B-A10B · 74 GB VRAM · Q4_K_M · ollama run qwen3.5:122b
Mistral Small 4Mistral Small 4 119B-A6.5B · 73 GB VRAM · Q4_K_M · ollama run mistral-small
Llama 4Llama 4 Scout 17B · 67 GB VRAM · Q4_K_M · ollama run llama4:scout
Command R FamilyCommand R+ (104B) · 64 GB VRAM · Q4_K_M · ollama run command-r-plus
Llama 3.2 FamilyLlama 3.2 90B Vision Instruct · 54 GB VRAM · Q4_K_M · llama-3-2
Llama 3.2 VisionLlama 3.2 Vision 90B · 54 GB VRAM · Q4_K_M · ollama run llama3.2-vision:90b
Qwen3-CoderQwen3-Coder 80B-A3B (MoE) · 49 GB VRAM · Q4_K_M · ollama run qwen3-coder:80b-a3b-q4

62 more families also fit 128 GB — browse the full model library.

Best Use Cases

FAQ

Can the NVIDIA GB10 Grace Blackwell run local LLMs?

Yes — the NVIDIA GB10 Grace Blackwell has 128 GB VRAM and runs The silicon inside DGX Spark and the OEM boxes built on it — ASUS Ascent GX10, Dell Pro Max, HP ZGX Nano, Lenovo ThinkSt

How fast is the NVIDIA GB10 Grace Blackwell for AI inference?

The NVIDIA GB10 Grace Blackwell is estimated to run Llama 3.1 8B at 26–53 tok/s with Q4_K_M quantization. For Llama 3.3 70B the estimate is 3.4–7.1 tok/s. These are modelled estimates, not measurements — see /en/methodology.

What LLMs can I run on 128 GB VRAM?

With 128 GB you can run: Qwen3.8, Qwen 3.5, Mistral Small 4, Llama 4, Command R Family. Use Ollama for the easiest setup: ollama run llama3.1:8b.

Compare Similar GPUs

VRAM Tier

Buying Guide

← All GPU Reviews | All Hardware | Check Your Hardware | Full Benchmarks | Can I Run It?