EU · GDPR
GDPR-Compliant AI: Personal Data That Never Leaves Your Jurisdiction
Your DPO has been asked to approve a US AI vendor for a system full of EU personal data. Again.
Written by Jakub Rusinowski · Last updated 2026-08-01 · Hardware figures computed by our VRAM engine
This page is practitioner guidance on deployment mechanics, not legal advice — involve your counsel or DPO for decisions about your specific obligations.
GDPR-compliant AI is a question of who processes personal data, where, and on what legal footing — not of which model you pick. Self-hosting an open-weight model collapses three of the hardest problems at once: there is no processor to bind under Article 28, no third-country transfer to justify under Chapter V, and no vendor retention term to reconcile with your own. What remains is ordinary controller work you already do: lawful basis, transparency, records, and a DPIA where the processing warrants one.
Why this is hard
- Every hosted model adds a processor, and usually a sub-processor chain you cannot fully enumerate on request.
- Article 17 erasure is trivial for a database row and genuinely hard for a model that has already trained on the row.
- Post-Schrems II transfer assessments age badly, and a supervisory authority will ask what you did about it.
What the deployment gives you
No Chapter V transfer, because nothing transfers
Inference on hardware in your own facility or an EU datacentre means personal data never crosses a border. There is no adequacy decision to rely on, no standard contractual clauses to paper, and no transfer impact assessment to refresh every time the case law moves.
Article 30 records that survive an audit
A self-hosted deployment has one processing location, one retention policy, and one set of recipients — yours. Describing it in your record of processing activities takes a paragraph instead of a diagram with a dotted line leaving the page.
Erasure you can actually perform
Retrieval over an index you control makes Article 17 a delete operation with a verifiable result. This is the architectural reason to prefer RAG over fine-tuning on personal data: you can delete from an index, and you cannot meaningfully delete from frozen weights.
Data minimisation enforced before the prompt
A gateway that strips direct identifiers and scopes retrieval to the purpose turns Article 5(1)(c) into a pipeline stage. Minimisation asserted in a policy document is an intention; minimisation in code is a control.
Where the personal data goes — and where it stops
The open-weight model itself is the one component that comes from outside. It arrives once, as a file, and it contains no personal data of your data subjects; from then on every byte of personal data stays on your side of the boundary. That asymmetry is what makes the compliance story short.
- A DPIA where Article 35 is triggered — large-scale processing of special-category data or systematic evaluation still needs one, self-hosted or not.
- Retention on prompt and completion logs set explicitly. They routinely contain personal data and they are the control most often forgotten.
- Transparency that names AI processing in your privacy notice; self-hosting changes the recipients, not your Article 13 duty to describe the processing.
- Human review for anything with legal or similarly significant effects, so Article 22 does not apply to an automated decision nobody checked.
- Purpose limitation on the index: documents gathered for one purpose should not silently become the retrieval corpus for another.
What you have to satisfy
- Art. 28 GDPR — Processor obligations: Every hosted model is a processor requiring a compliant contract and documented sub-processors. Self-hosting removes the processor from the inference step entirely.
- Art. 17 GDPR — Right to erasure: Deletable from an index; effectively not deletable from trained weights. This single asymmetry should drive your RAG-versus-fine-tuning decision.
- Chapter V — Third-country transfers: Adequacy, SCCs, and transfer impact assessments apply the moment data leaves the EEA. Processing locally means the chapter never engages.
- Art. 35 GDPR — Data protection impact assessment: Required for high-risk processing regardless of deployment model. Self-hosting shortens the DPIA; it does not remove the obligation to do one.
Models that fit this deployment
| 模型 | VRAM (Q4) | 可运行于 | 上下文 | 许可 |
|---|---|---|---|---|
| Mistral Small 3.1 24B EU-published, Apache 2.0 — 16 GB card — A capable European open-weight model where provenance itself is part of the procurement argument. ollama pull mistral-small3.1 | 15 GB | 24 GB GPU (RTX 3090/4090) Mac: 24 GB 统一内存 | 125K | Apache 2.0 |
| Qwen 3 32B Department tier — one 24 GB card — Permissive licence, strong multilingual behaviour across EU languages, single-card deployment. ollama pull qwen3:32b | 20.6 GB | 24 GB GPU (RTX 3090/4090) Mac: 32 GB 统一内存 | 125K | Apache 2.0 |
| Llama 3.3 70B Instruct Organisation tier — 48 GB class — The reference open 70B for knowledge work; note the community licence when your legal team reviews terms. ollama pull llama3.3:70b | 43.1 GB | 2×24 GB GPUs or 48 GB card Mac: 64 GB 统一内存 | 128K | Llama Community |
Need AI your DPO will actually sign off on?
We help European organisations deploy AI that keeps personal data inside their own jurisdiction and their own control — architecture review, model selection, retrieval design, and the documentation your data protection officer needs to see. Vendor-neutral, practitioner-level, no platform to sell you.
常见问题
Go deeper
- The EU AI Act and on-premise AI — roles, timeline, obligations
- Open model licences — the five questions legal will ask
- What sovereign AI means, layer by layer
- Where EU-jurisdiction compute actually exists
- Local AI and GDPR — the full comparison
- Local vs cloud AI at enterprise scale
- Cost calculator — your volumes, current API prices
- Sovereign AI
- Legal AI
- Financial services AI
- Enterprise & Sovereign AI — the full hub