Team volume changes the equation
The economics that are marginal for one person become compelling in aggregate. Ten people making moderate use of an internal assistant, a RAG system over company docs, and some automation easily sum to 4M tokens/day. At GPT-4o prices that's roughly $570/month, forever, growing with headcount. One RTX 4090 workstation running Qwen 3 32B behind vLLM serves that same aggregate load — vLLM's continuous batching turns one 24 GB card into a genuine multi-user server — for about $153/month all-in during amortization and ~$45/month after. Break-even in about 5 months, and each additional user until saturation costs approximately nothing.
The catch is the sentence "behind vLLM." Someone on the team now owns a server: OS updates, the inference runtime, access control, a restart script, capacity judgment. None of it is hard — it's a Docker container and an SSO proxy, an afternoon for anyone who has deployed anything — but it must be owned. Teams where nobody wants that pager should read the cloud column without shame; an unowned server is worse than an invoice.
What small teams actually run
The pattern we see in workshops: three workloads dominate small-team AI value, and all three fit comfortably in the open-model class.
1. Internal assistant (chat over an OpenAI-compatible endpoint) — Qwen 3 32B-class quality is fully sufficient for drafting, summarizing, and internal Q&A. 2. RAG over company knowledge — the retrieval layer matters more than the model tier; a well-tuned local pipeline over your wiki and docs (guide) beats a frontier model with no context. Keeping embeddings + documents on your own network is also the quiet governance win: nothing about your knowledge base leaves the building. 3. Automation (ticket triage, report drafts, code review comments) — high-volume, latency-tolerant, quality-tolerant: the textbook local workload.
What does not fit: the occasional hard task — a thorny contract analysis, a nasty production bug. Keep a metered frontier key for those. A sensible policy is "local endpoint is the default; the cloud key requires a reason" — which also gives your data-governance story teeth, since the sanctioned default keeps prompts on-network (the full argument lives in our GDPR comparison).
Sizing the box
For a 5–25 person team: a single RTX 4090 (24 GB) covers 30B-class models with batching headroom — the $2,590 workstation build is the reference config. Heavier or 70B-ambitious teams should look at the dual RTX 3090 build ($2,870, 48 GB) which runs Llama 3.3 70B at Q4, or a 128 GB unified-memory box (GMKtec EVO-X2 class, $2,349) if silence and MoE models matter more than raw throughput. Concurrency rule of thumb: one modern 24 GB card with vLLM sustains a 10–25 person team's interactive load; batch jobs should run off-hours. If two boxes are entering the conversation, read the enterprise comparison — the calculus shifts again at rack scale.
The negotiation you should have first
Before buying hardware, price the alternative properly: at team volume you can drop to GPT-4o Mini/Gemini Flash-class models for the easy 80% of traffic and the cloud bill collapses (~$34/month at our preset volume — cheaper than the server's electricity). If your team's workload is genuinely Mini-class, cloud wins the pure cost fight, and the local case rests on governance, latency, and unlimited usage instead. Run both numbers in the calculator; let the spreadsheet, not the vibe, pick your side.