Can I Run Poolside Laguna XS 2.1 on Beelink SER9 (Ryzen AI 9, 32 GB)?

Written by Jakub Rusinowski · Last updated August 15, 2026

Yes, but it is tight

Yes, but it is tight — Laguna XS 2.1 33B-A3B at Q6_K needs about 29.7 GB of the 32 GB usable on Beelink SER9 (Ryzen AI 9, 32 GB), leaving only ~2.3 GB before the runtime starts swapping. Expect ~50.3 tok/s (estimated), with room for about 16,384 tokens of context.

Confidence: low · Recommended quantization: Q6_K · Estimated speed: ~50.3 tok/s

Beelink SER9 (Ryzen AI 9, 32 GB) — what it gives a model

Usable memory for models32 GB
Memory bandwidth256 GB/s
Form factorMini PC
Operating systemWindows or Linux
Memory upgradeableYes
Price$859 (lib/data/ai-stations.ts (street price), checked 2026-07-06)

Poolside Laguna XS 2.1 on Beelink SER9 (Ryzen AI 9, 32 GB): memory by quantization

QuantMemory neededFits 32 GB?Max contextEst. speedDownload
F1668.6 GB✗ No66 GB
Q8_037.7 GB✗ No35.1 GB
Q6_K29.7 GB✓ Yes16K~50.3 tok/s27.1 GB
Q5_K_M26 GB✓ Yes32K~55.2 tok/s23.4 GB
Q4_K_M22.6 GB✓ Yes32K~60.7 tok/s19.9 GB
Q3_K_M16.7 GB✓ Yes64K~72.9 tok/s14.1 GB
Q2_K13.5 GB✓ Yes64K~82 tok/s10.8 GB

What to watch out for

Beelink SER9 32 GB limitations

Recommended setup

Ollama or llama.cpp (CUDA/ROCm) — vLLM if you need concurrent serving

How these numbers are calculated

FAQ

Can I run Poolside Laguna XS 2.1 on Beelink SER9 (Ryzen AI 9, 32 GB)?

Yes, but it is tight — Laguna XS 2.1 33B-A3B at Q6_K needs about 29.7 GB of the 32 GB usable on Beelink SER9 (Ryzen AI 9, 32 GB), leaving only ~2.3 GB before the runtime starts swapping. Expect ~50.3 tok/s (estimated), with room for about 16,384 tokens of context.

Which quantization of Poolside Laguna XS 2.1 should I use on Beelink SER9 (Ryzen AI 9, 32 GB)?

Q6_K — it needs about 29.7 GB of the 32 GB available, downloads as roughly 27.1 GB, and runs at an estimated 50.3 tokens/sec with up to 16K of context.

What limits Poolside Laguna XS 2.1 on Beelink SER9 (Ryzen AI 9, 32 GB)?

Nothing binding — the model fits with headroom and generates at a usable speed on this hardware.

Which runtime should I use?

Ollama or llama.cpp (CUDA/ROCm) — vLLM if you need concurrent serving

Other Computers

Other Models on Beelink SER9 (Ryzen AI 9, 32 GB)

Poolside Laguna XS 2.1 on GPUs

What This Model Is Good At

Model & Tools

← Can I Run It? | Poolside Laguna XS 2.1 model page | Check your hardware