Seed-OSS — local AI model by ByteDance Seed

Written by Jakub Rusinowski · Last updated

Seed-OSS-36B is ByteDance Seed's dense 36B open-weight reasoning model with an Apache-2.0 licence, a controllable thinking budget, tool-call parser support and a native 512K context window. It dates from August 2025 and is kept in the library as a supported legacy option; the long context is a model capability, not a promise that a single consumer GPU can serve it.

Variants

The smallest Seed-OSS variant needs about 23 GB of VRAM at Q4_K_M — quantized weights plus framework overhead, before any KV cache.

ModelVRAM
Seed-OSS-36B-Instruct →
36B
~22.6 GB

Memory is quantized weights plus overhead at Q4_K_M, from the same engine as the GPU & VRAM checker.

How to run Seed-OSS locally

Install Ollama, then pull the tag.

No first-party Ollama, GGUF or MLX artifact. llama.cpp supports the architecture and community GGUF files exist (LM Studio's own community account publishes one); the official route is vLLM 0.10 or newer.

Pick a size above for its own VRAM figure, speed estimate and install command.

Licence

Apache-2.0Commercial use permitted

Commercial use permitted. No usage restrictions beyond attribution.

Applies to: Seed-OSS-36B-Instruct

Recommended GPU

The cheapest catalogued GPU that runs Seed-OSS locally (min 23 GB VRAM) is the AMD Radeon RX 7900 XTX (24 GB).

Affiliate disclosure: Some links on this page are affiliate links — if you buy through them, LLM Configurator may earn a commission at no extra cost to you. As an Amazon Associate, LLM Configurator earns from qualifying purchases.
AMD Radeon RX 7900 XTX 24GB
24 GB VRAM · 355 W board power
2026 prices are volatile — check the current listing.

Seed-OSS — frequently asked questions

How much VRAM does Seed-OSS need?

Seed-OSS needs about 23 GB VRAM at Q4_K_M quantization for its smallest variant. Variants: Seed-OSS-36B-Instruct (23 GB, Q4_K_M). On Apple Silicon, unified memory counts toward this requirement.

Can I run Seed-OSS on an RTX 4090 (24 GB)?

Yes — Seed-OSS runs on an RTX 4090 (24 GB) and other 24 GB cards such as the RTX 3090. Smaller variants also fit comfortably on 8–16 GB GPUs at Q4_K_M.

What quantization should I use for Seed-OSS?

Q4_K_M is the best balance of quality and VRAM for Seed-OSS in most cases. Choose Q8_0 for near-lossless quality if you have spare VRAM, or smaller quants (Q3/Q2) only when memory is tight.

How do I run Seed-OSS with Ollama?

Seed-OSS has no local Ollama tag — the published tag is cloud-hosted, so running it sends your prompts to a hosted GPU rather than your own machine.