IntelDesktop GPUCurrent

Intel Arc Pro B60 for local LLMs

Written by Jakub Rusinowski · Last updated

With 24 GB of GDDR6 at 456 GB/s, the Intel Arc Pro B60 runs 109 catalogued models at Q4_K_M with 8K context. The largest that fits is Qwen 3 32B (~22.8 GB), and the top pick is Laguna XS 2.1 33B-A3B at 37–78 tok/s.

24 GB of GDDR6 on a 192-bit bus, 456 GB/s. Board power is a range by partner design (120-200 W); 200 W is the upper bound. Not the 32 GB Arc Pro B65 or B70. Runs through SYCL or Vulkan builds of llama.cpp.

Models that run on the Intel Arc Pro B60

Q4_K_M, 8K context, 24 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
Laguna XS 2.1 33B-A3B
Poolside Laguna XS 2.1
20.7 GB37–78 tok/s
Granite 4.0 Small-H 32B-A9B
IBM Granite 4.0
20.1 GB21–44 tok/s
Olmo 3 32B
Olmo 3
20.1 GB8.0–17 tok/s
Nemotron-Cascade 2 30B-A3B
Nemotron Cascade 2
19.9 GB35–74 tok/s
Gemma 4 31B
Gemma 4
19.5 GB8.2–17 tok/s
Qwen 3 30B-A3B (MoE)
Qwen 3
20 GB46–95 tok/s
Qwen3-Coder 30B-A3B (MoE)
Qwen3-Coder
20 GB46–95 tok/s
GLM-4.7-Flash 30B-A3B
GLM-4.7 / GLM-Z1
18.9 GB38–79 tok/s
Nemotron 3 Nano Omni 30B-A3B
Nemotron 3 Nano Omni
18.9 GB38–79 tok/s
North Mini Code 1.0 30B-A3B
North Mini Code
18.9 GB38–79 tok/s
Nemotron 3.5 Lightning 30B-A3B
Nemotron 3.5
18.9 GB38–79 tok/s
Granite 4.1 30B
IBM Granite 4.1
18.9 GB8.5–18 tok/s
Showing 12 of 109

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-10-07.

Memory
24 GB GDDR6
Memory bandwidth
456 GB/s
Memory bus
192-bit
Architecture
Xe2 Battlemage BMG-G21
Series
Arc Pro B-series
Board power
200 W
Release year
2025
Compute backends
SYCL, VULKAN
Usable for models
24 GB
Best for24 GB workstationSYCL / Vulkan

Similar GPUs

Frequently asked questions

Can the Intel Arc Pro B60 run local LLMs?

Yes. With 24 GB (24 GB usable by a model) it runs 109 of the catalogued models at Q4_K_M with 8K context; the largest is Qwen 3 32B, needing about 22.8 GB.

How fast is the Intel Arc Pro B60 for AI inference?

It is estimated to run Llama 3.1 8B at 28–57 tok/s at Q4_K_M. Llama 3.3 70B does not fit: it needs about 44 GB against 24 GB usable. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.

What LLMs can I run on 24 GB?

Among the best that fit: Laguna XS 2.1 33B-A3B, Granite 4.0 Small-H 32B-A9B, Olmo 3 32B, Nemotron-Cascade 2 30B-A3B, Gemma 4 31B.