ByteDance Seed36B~23 GB VRAM at Q4_K_M
Licence: Apache-2.0Local route: Community quantized routeHardware: WorkstationMaturity: Legacy but supportedEvidence: Moderate

Evidence: benchmark results vendor-reported

Seed-OSS-36B-Instruct — VRAM, speed & local setup

Written by Jakub Rusinowski · Last updated

ByteDance Seed's dense 36B open-weight model with a controllable thinking budget, native 512K context and an Apache-2.0 licence. Upstream llama.cpp implements the architecture, vLLM 0.10 or newer serves it with a dedicated tool-call parser, and community GGUFs exist. It dates from August 2025, so it is listed as supported legacy. A 512K context is not a single-GPU promise: its KV cache is very large because all 64 layers use full attention.

Seed-OSS-36B-Instruct needs about 23 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache. On Apple Silicon that figure comes out of unified memory.

VRAM and speed by quantization

Quoted against NVIDIA RTX 4090 (24 GB). Includes the KV cache at 8K context, so it reads higher than the headline figure.

QuantVRAMSpeed (est.)Fit
Q2_K
2.63 bpw
14.8 GB~52 tok/sFits
Q3_K_M
3.41 bpw
18.4 GB~42 tok/sFits
Q4_K_M
4.83 bpw
24.7 GB~5 tok/sOffload
Q5_K_M
5.67 bpw
28.6 GB~4 tok/sOffload
Q6_K
6.56 bpw
32.6 GB~3 tok/sOffload
Q8_0
8.50 bpw
41.4 GB~3 tok/sOffload
F16
16.00 bpw
75.2 GB—Too big

Black marker = usable memory on the NVIDIA RTX 4090 (24 GB). Estimates from the memory-bandwidth roofline on the methodology page. Seed-OSS-36B-Instruct VRAM calculator →

Get Seed-OSS-36B-Instruct running

The cheapest catalogued GPU that runs Seed-OSS-36B-Instruct is the AMD Radeon RX 7900 XTX (24 GB).

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
AMD Radeon RX 7900 XTX 24GB
24 GB VRAM · 355 W board power
2026 prices are volatile — check the current listing.

How to run Seed-OSS-36B-Instruct

No first-party Ollama, GGUF or MLX artifact. llama.cpp supports the architecture and community GGUF files exist (LM Studio's own community account publishes one); the official route is vLLM 0.10 or newer.

Specifications

Verified — Checked against the primary source — the model card or the vendor spec page — and corroborated by a second independent source.

Parameters
36.2 Billion
Context window
524,288
Architecture
Dense transformer (GQA, 64 layers)
Provider
ByteDance Seed
Licence
Apache-2.0
Specified at
Q4_K_M
System RAM
32 GB
Record updated
2026-10-06
LicenceApache-2.0Commercial use permitted

Commercial use permitted. No usage restrictions beyond attribution.

Limits and caveats

  • August 2025 release; newer reasoning models exist.
  • No first-party GGUF, Ollama or MLX artifact.
  • Very large KV cache at long context.
  • Card benchmark tables are vendor-run, with sampling temperature 1.1 / top_p 0.95 recommended.

See also Qwen3.8 27B

Quality and use cases

Scores as published by the model’s authors or an independent evaluator — quality, not throughput, and not measured by us.

Best forreasoninglong contexttool useagents

Seed-OSS-36B-Instruct — frequently asked questions

How much VRAM does Seed-OSS-36B-Instruct need?

About 23 GB at Q4_K_M — quantized weights plus framework overhead, before any KV cache. The cache grows with context length and is added on top; the table above folds it in. Apple Silicon counts unified memory toward the same figure.

Does Seed-OSS-36B-Instruct run on an RTX 4090 (24 GB)?

Yes. Seed-OSS-36B-Instruct needs about 23 GB at Q4_K_M, inside a 24 GB card, at an estimated 5 tokens/sec.

How do I run Seed-OSS-36B-Instruct locally?

No first-party Ollama, GGUF or MLX artifact. llama.cpp supports the architecture and community GGUF files exist (LM Studio's own community account publishes one); the official route is vLLM 0.10 or newer. Running the published tag would send your prompts to a hosted GPU rather than your own machine.