Seed-OSS — local AI model by ByteDance Seed
Written by Jakub Rusinowski · Last updated
Seed-OSS-36B is ByteDance Seed's dense 36B open-weight reasoning model with an Apache-2.0 licence, a controllable thinking budget, tool-call parser support and a native 512K context window. It dates from August 2025 and is kept in the library as a supported legacy option; the long context is a model capability, not a promise that a single consumer GPU can serve it.
Variants
The smallest Seed-OSS variant needs about 23 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache.
| Model | VRAM at Q4 | VRAM | Context | Run it |
|---|---|---|---|---|
| Seed-OSS-36B-Instruct → 36B | ~22.6 GB | 524,288 | cloud-hosted tag |
Memory is quantized weights plus overhead at Q4_K_M, from the same engine as the GPU & VRAM checker.
How to run Seed-OSS locally
Install Ollama, then pull the tag.
No first-party Ollama, GGUF or MLX artifact. llama.cpp supports the architecture and community GGUF files exist (LM Studio's own community account publishes one); the official route is vLLM 0.10 or newer.
Pick a size above for its own VRAM figure, speed estimate and install command.
Licence
Commercial use permitted. No usage restrictions beyond attribution.
Applies to: Seed-OSS-36B-InstructRecommended GPU
The cheapest catalogued GPU that runs Seed-OSS locally (min 23 GB VRAM) is the AMD Radeon RX 7900 XTX (24 GB).
Seed-OSS — frequently asked questions
How much VRAM does Seed-OSS need?
Seed-OSS needs about 23 GB VRAM at Q4_K_M quantization for its smallest variant. Variants: Seed-OSS-36B-Instruct (23 GB, Q4_K_M). On Apple Silicon, unified memory counts toward this requirement.
Can I run Seed-OSS on an RTX 4090 (24 GB)?
Yes — Seed-OSS runs on an RTX 4090 (24 GB) and other 24 GB cards such as the RTX 3090. Smaller variants also fit comfortably on 8–16 GB GPUs at Q4_K_M.
What quantization should I use for Seed-OSS?
Q4_K_M is the best balance of quality and VRAM for Seed-OSS in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.
How do I run Seed-OSS with Ollama?
Seed-OSS has no local Ollama tag — the published tag is cloud-hosted, so running it sends your prompts to a hosted GPU rather than your own machine.