Local AI Compliance: What Self-Hosting Changes
Written by Jakub Rusinowski · Last updated 2026-09-13 · Assessment logic is deterministic and runs in your browser
This section is practitioner guidance on deployment and governance mechanics, not legal advice. Regulatory classification turns on facts about intended purpose and real-world use that a questionnaire cannot establish — involve your counsel or DPO for decisions about your specific obligations.
Self-hosting does not change which obligations apply to your AI system. The EU AI Act classifies by intended purpose, and an HR screening system carries the same duties whether it answers from a workstation under a desk or a vendor API in another jurisdiction. What self-hosting changes is your ability to demonstrate things: you own the logs, you pin the model version your documentation describes, and your data path has fewer parties in it. That is a real advantage and it is not an exemption.
The claim to stop believing first
"On-premise equals compliant" is the most common false statement in this space, and it is a category error rather than an exaggeration. The Act regulates the workload, not the rack. Nothing in Article 5, Article 6, Annex III or Article 50 turns on where the compute sits.
Two specific versions of the error are worth naming because vendors sell both:
"Local means no personal data processing." It means no personal data leaves your infrastructure. Processing is still processing, the GDPR still applies, and — in the one that catches people — inference logs on your own disk are a personal-data store you now own, with a retention period you probably have not set. A system designed not to store personal data frequently stores it anyway, in the log, forever.
"Open weights means exempt." The free and open-source carve-outs in the Act sit in the general-purpose AI model chapter and concern the duties of whoever trains and releases the weights. They do not reach the deployer of a system built on those weights. If your deployment lands in Annex III, the licence on the model file is irrelevant to that.
What self-hosting genuinely changes
Four things, and they are worth having:
1. The logs are yours. Article 12 requires high-risk systems to technically allow automatic event recording; Article 26(6) requires deployers to keep those logs for an appropriate period. On your own gateway the format, the retention and the completeness are your decisions, already flowing into whatever you use for everything else. With a managed API you get the vendor's export feature, with the vendor's retention window, which may be shorter than your obligation.
2. The version is pinned. Your risk assessment and your technical documentation describe a specific model revision. You change it when you decide to. An API model can change under you — and a risk assessment describing a moving target describes nothing. This is the most underrated practical advantage of self-hosting and the easiest to squander by pulling latest.
3. The accountability chain is shorter. Vendor role changes, terms updates and subprocessor churn each generate re-review work under a combined AI Act and GDPR posture. A pinned open-weight model converts that recurring workstream into a version decision you make on your own schedule.
4. The data path has fewer parties. This matters more for the data-protection half of the review than for the AI Act half — the GDPR page covers that side — but it removes a transfer analysis from the combined review, which is real work saved.
What self-hosting adds to your obligations
An honest accounting counts this column too. You are now operating an inference service, and it is yours to secure:
The endpoint. Model servers commonly default to binding broadly, because that is what the quickstart command does. An internal LLM reachable from more of the network than you intended is a data-exfiltration path with a chat interface attached. This is the most frequently found real misconfiguration in self-hosted deployments — more common than any model-level issue.
The retrieval corpus. If you built RAG over internal documents, the index is now an access-control boundary. Corpora routinely end up containing material with broader access than the index itself, turning the assistant into a query interface over documents its users could not otherwise open. Access control belongs on the retrieval layer, not only on the endpoint.
The model artefact. You downloaded weights from somewhere. Is the file you are running the file you intended? Checksums exist for this and are rarely checked.
The competence. The AI-literacy duty under Article 4 applies to you as deployer, and it applied from February 2025 — earlier than most of the Act. Running your own infrastructure means the people who need to understand the system's limitations include the ones operating the serving stack.
The on-premise deployment guide covers the architecture that addresses most of this; air-gapped deployment covers the case where the network boundary is the control.
Hardware choices that are also compliance choices
A few deployment decisions this site already helps with have compliance consequences that are not obvious:
Quantisation is a documentation fact, not just a performance one. Q4_K_M and F16 of the same model are different artefacts with different output characteristics. Your technical documentation should say which one is deployed, and a change of quantisation is a change of system — the same way a model version bump is. Use the VRAM calculator to decide it deliberately rather than discovering it from whatever fit.
Sizing decides whether oversight is affordable. A system that answers in two seconds supports a human reviewing every consequential output. One that takes forty does not, and the oversight procedure you wrote will quietly stop being followed. If human oversight is an obligation for your use case, throughput is a compliance parameter — check what your hardware actually delivers in the GPU checker before you design the procedure around it.
Offload to system RAM changes the failure mode. A model that partially fits VRAM and spills to host memory is slower and more variable. If your monitoring baselines were set on a machine where it fit, they will not transfer.
Assess your local deployment
The assessment asks about hosting mode because it changes the control set — not because it changes the classification. Both answers come out of the same profile.
Automated assessment based on the information you provide. Not legal advice, certification or an audit.
Frequently asked questions
Keep going
- The EU AI Act and self-hosted AI — timeline and deployer duties
- Local AI and GDPR
- GPU & VRAM checker
- On-premise LLM deployment architecture
- EU AI Act Compliance Assessment for AI Deployments
- Private AI Deployment: Governance for Controlled Environments
- ISO/IEC 27001 Controls for AI Deployments
- AI Compliance & Deployment Readiness — the full hub