The compliance problem, stated precisely
GDPR doesn't prohibit cloud AI. What it does is attach obligations to every hop your personal data takes: you need a lawful basis for the processing, an Article 28 data-processing agreement with anyone who processes it for you, a Chapter V transfer mechanism if that processing happens outside the EU/EEA, and the ability to honor access and erasure requests wherever the data went. Every one of those obligations is satisfiable with a cloud provider — and every one of them is ongoing work that must survive vendor terms changing, subprocessor lists growing, and adequacy decisions being challenged (ask anyone who built on Privacy Shield before Schrems II invalidated it in 2020).
A local deployment collapses that stack. When inference runs on a machine in your office or EU datacenter rack, the model vendor never becomes a processor — there is no vendor in the data path. Open-weight models like Qwen 3, Llama 3.3, and Mistral Small are downloaded once (model weights contain no personal data) and run air-gapped if you want. The GDPR questions that remain — lawful basis, internal access control, retention of your own logs — are questions you already answer for every internal system.
What the cloud providers actually offer
Fairness requires the other side. OpenAI, Anthropic, Google, and Microsoft (Azure OpenAI) all offer enterprise API terms with data-processing agreements, no-training commitments, and configurable retention; Azure and Google offer EU-region processing, and Mistral's La Plateforme is EU-hosted by default. A careful legal team can build a defensible GDPR position on any of them, and thousands of EU companies have. The costs are the legal review itself, the monitoring of terms and subprocessors over time, the DPIA that now spans another organization, and residual transfer risk where US-headquartered providers are involved (the CLOUD Act discussion doesn't disappear because a DPA is signed).
The practical question for a mid-size company is rarely "is cloud AI legal?" — it's "which position do I want to defend in an audit, a customer security review, or a works-council negotiation?" "The data never leaves our infrastructure" is a one-sentence answer. The cloud equivalent is a binder.
Sensitive categories raise the bar
For Article 9 data — health, biometrics, beliefs — and for regulated verticals (legal privilege, banking secrecy, German §203 professional secrecy), the calculus tilts hard toward local. Many professional bodies and DPAs have issued cautious-to-negative guidance on feeding such data to third-party AI services. A local model isn't automatically compliant (access control and logging still matter), but it removes the disclosure-to-a-third-party event that triggers the strictest analysis. This is why hospitals, law firms, and public-sector bodies are disproportionately represented among on-premise LLM adopters.
The cost side-effect nobody expects
Here's the pleasant surprise: at department scale, the compliant option is not the expensive one. An RTX 4090 workstation (~$2,590, see /build) running Qwen 3 32B handles ~1M tokens/day of internal assistant load for about $119/month amortized (hardware over 24 months + electricity). The same volume through GPT-4o is ~$143/month — before you count the legal review hours, which bill higher than the hardware. Break-even against frontier API pricing lands around month 20; against the compliance-overhead delta it's arguably immediate. Run your own volumes in the calculator below.
Deployment checklist for on-premise LLMs
- Hardware: 24 GB VRAM minimum for 30B-class quality (GPU guide); 48 GB+ (dual-GPU or Mac Studio class) for 70B. Size with /analyzer.
- Serving: Ollama or vLLM behind your SSO; both expose OpenAI-compatible APIs so internal tools port over unchanged.
- Governance: treat prompts/outputs as personal-data processing; set log retention deliberately; document the system in your records of processing (Art. 30) — it's a short entry when there's no external processor.
- People: one engineer-day to stand up, ongoing effort near zero for chat/RAG workloads. Our local RAG guide covers document search over internal files.