The fair fight is same-model, and it's nearly a tie
Most chatbot cost comparisons cheat: they price a local Llama 8B against GPT-4o. Price it against hosted Llama 8B — the same weights, served by Together.ai at $0.18 per million tokens or cheaper on DeepInfra — and a busy 4M-token/day support bot costs about $22/month in the cloud. Our $985 budget build spends ~$18/month on electricity to do the same job: the hardware never pays for itself on cost alone. We built the calculator below to show that, not hide it.
So why does anyone self-host a customer bot? Four reasons, all real, none of them "it's cheaper than the same model hosted."
Reason one: the transcripts
A support chatbot's traffic is customer PII by definition — names, order numbers, complaints, occasionally payment details customers paste despite your warnings. Routing that through any third party creates a processor relationship, a DPA, retention questions, and an entry in your privacy policy (the GDPR page walks through it). Self-hosting makes the answer to "who can see customer conversations?" one word long. For healthcare, finance, and EU-consumer-facing businesses, this alone decides it.
Reason two: the meter flips at frontier quality
If your flows genuinely need frontier-class handling — nuanced troubleshooting, multilingual empathy, complex policy application — the comparison becomes local-something vs GPT-4o at ~$570/month for our preset volume, and suddenly a 24 GB build running Qwen 3 32B with tight RAG breaks even in months. The trick most production bots use: RAG does the heavy lifting. A 14B model retrieving from your actual help docs answers support questions more accurately than a frontier model winging it from general knowledge — see the RAG comparison. Quality-per-dollar for support flows is mostly a retrieval problem, and retrieval is local-friendly.
Reason three: limits and policy
Rate limits bite bots at the worst time — the product-launch spike, the outage that floods support. Self-hosted concurrency is bounded by your hardware, not your API tier, and batching servers like vLLM sustain surprising concurrent load on one card. Policy control matters more subtly: your users' weird inputs are judged by your system prompt, not a provider's evolving content policy; nobody else's moderation layer can misfire on a legitimate customer question about, say, medication dosages or self-harm resources in a healthcare bot.
Reason four: the bill that doesn't scale with success
A hosted bot's cost scales linearly with adoption forever. A self-hosted bot's cost is a step function: the same box serves 10× the traffic until it saturates (then you buy one more box — the calculator models machine count honestly). Startups betting on volume growth like owning the flat part of that curve.
What to actually deploy
The boring reference stack: vLLM or Ollama serving Qwen 3 14B (16 GB card) or 32B (24 GB card), RAG over your help center (the quality multiplier), a thin orchestration layer with escalation-to-human, and a frontier API key wired in as a fallback route for conversations the local model flags as beyond it — the same router logic that wins for agents. Uptime is the tax: monitoring, restarts, and a failover story are yours now. If that sentence made someone on your team wince, the hosted-8B tier at $22/month is a perfectly honorable answer.