Cost matters less than everyone expects
Run the naive math and on-prem wins: 20M tokens/day is ~$2,850/month at GPT-4o list prices, while a fleet of ~11 dual-RTX-3090 nodes (48 GB each, ~$32k total at current street prices) running Llama 3.3 70B costs ~$2,200/month amortized over 24 months including power — and drops toward $900/month once amortized. A ~25% saving that compounds with volume.
Then reality complicates it in both directions. Cloud-side: enterprise agreements discount list prices materially, and routing the easy 80% of traffic to Mini-class models collapses the bill — Gemini Flash costs 4% of GPT-4o. On-prem-side: consumer-card fleets are the scrappy version; proper datacenter GPUs (or rented H100s for burst) cost more per unit but less per token at scale, while colo space, redundancy, and the platform engineer's salary all belong in the TCO. Model both honestly and the pure-cost gap usually lands within ±40% either way. At enterprise scale, cost is a tiebreaker, not the decision.
What actually decides it: control
Three control properties push large organizations on-prem, and none appear on an invoice:
- Data sovereignty. For regulated workloads — patient data, trading desks, defense, anything under GDPR's stricter categories — "inference happens inside our perimeter" converts an ongoing legal workstream into an architecture fact. Security reviews shorten; DPIAs simplify; the works council relaxes.
- Vendor independence. API models get deprecated, repriced, and re-policied on the vendor's schedule; a fleet running open weights (Llama 3.3, Qwen 3, DeepSeek-class MoEs) changes only when you choose. For systems embedded in core processes with multi-year lifetimes, that stability is worth real money.
- Auditability. Your logs, your retention, your explainability story — assembled from systems you control, which is what regulators and internal audit actually ask for.
The counterweights are real too: frontier quality (no open model matches the frontier on the hardest reasoning), organizational capability (an on-prem platform needs MLOps ownership — if that team doesn't exist, the project fails independent of the hardware), and utilization risk (a fleet sized for peak sits idle at trough; the cloud never does).
The hybrid architecture that wins in practice
Mature enterprise deployments in 2026 converge on the same shape:
1. An internal gateway (OpenAI-compatible) that all applications call — one URL, centralized auth, logging, and routing policy. 2. Owned/colocated open-model capacity behind it for the bulk tier: internal assistants, RAG over corporate knowledge, document pipelines, classification. This is 80–95% of token volume and the part where on-prem economics and governance both bite. 3. Frontier API contracts for the quality-critical tier, reached through the same gateway with data-classification rules enforced at routing time — sensitive classes never route out. 4. Rented GPU burst (RunPod, Lambda-class) for training runs and load spikes, so the owned fleet is sized for baseline, not peak.
This shape keeps the compliance surface small (one egress point), captures bulk-tier savings, and preserves frontier access where it earns its premium. It also de-risks the buy-in: start with the gateway plus rented capacity, prove utilization, then convert the baseline to owned hardware with real usage data instead of forecasts.
Sizing note for the calculator
The calculator below models the scrappy end (consumer dual-GPU nodes) and scales machine count with your volume automatically — useful for directional TCO, not a substitute for capacity planning. For a real deployment, benchmark your actual workload mix on one node first (our benchmark data gives starting points), then multiply. And involve procurement early: GPU lead times and colo contracts, not software, set enterprise timelines.