IntelDesktop GPUCurrent

Intel Arc Pro B50 for local LLMs

Written by Jakub Rusinowski · Last updated

With 16 GB of GDDR6 at 224 GB/s, the Intel Arc Pro B50 runs 74 catalogued models at Q4_K_M with 8K context. The largest that fits is OLMo 2 13B Instruct (~15.8 GB), and the top pick is GPT-OSS 20B at 28–58 tok/s.

16 GB of GDDR6 on a 128-bit bus at 70 W, powered from the slot. 224 GB/s, so it holds 14B-class models but runs them slowly.

Models that run on the Intel Arc Pro B50

Q4_K_M, 8K context, 16 GB usable. Ranked by quality and speed.

ModelVRAMSpeed
GPT-OSS 20B
GPT-OSS
13.8 GB28–58 tok/s
Cosmos 3 Nano
Cosmos 3
10.5 GB13–27 tok/s
Phi-4 (14B)
Phi-4 Family
10.9 GB8.2–17 tok/s
Gemma 4 12B (Unified)
Gemma 4
8 GB9.6–20 tok/s
Bielik PL 11B v3.0 Instruct
Bielik
7.4 GB10–21 tok/s
Llama 3.2 Vision 11B
Llama 3.2 Vision
8.5 GB11–22 tok/s
Falcon 3 10B Instruct
Falcon 3
8.4 GB11–22 tok/s
Qwen 3.5 9B
Qwen 3.5
6.2 GB12–25 tok/s
Ministral 3 14B
Ministral 3
9.3 GB8.4–18 tok/s
Cogito v1 14B
Cogito v1
10.9 GB8.2–17 tok/s
Gemma 4 E2B
Gemma 4
3.9 GB19–39 tok/s
Gemma 4 E4B
Gemma 4
5.6 GB13–28 tok/s
Showing 12 of 74

Buy it or rent the same memory

Buy the card, or rent a GPU with the same memory by the hour to try models first.

Speed vs other GPUs

Llama 3.1 8B, Q4_K_M. Estimated ranges. How this is calculated

Specifications

Specs last updated 2026-10-07.

Memory
16 GB GDDR6
Memory bandwidth
224 GB/s
Memory bus
128-bit
Architecture
Xe2 Battlemage BMG-G21
Series
Arc Pro B-series
Board power
70 W
Release year
2025
Launch price
$349
Compute backends
SYCL, VULKAN
Usable for models
16 GB
Best for16 GB workstationSYCL / Vulkan

Similar GPUs

Frequently asked questions

Can the Intel Arc Pro B50 run local LLMs?

Yes. With 16 GB (16 GB usable by a model) it runs 74 of the catalogued models at Q4_K_M with 8K context; the largest is OLMo 2 13B Instruct, needing about 15.8 GB.

How fast is the Intel Arc Pro B50 for AI inference?

It is estimated to run Llama 3.1 8B at 15–31 tok/s at Q4_K_M. Llama 3.3 70B does not fit: it needs about 44 GB against 16 GB usable. These are modelled estimates from memory bandwidth, not measurements; the methodology page shows the formula.

What LLMs can I run on 16 GB?

Among the best that fit: GPT-OSS 20B, Cosmos 3 Nano, Phi-4 (14B), Gemma 4 12B (Unified), Bielik PL 11B v3.0 Instruct.