Enterprise & Sovereign AI

Enterprise & Sovereign AI: Run AI on Infrastructure You Control

Written by Jakub RusinowskiLast updated July 12, 2026Hardware figures computed by our VRAM engine

Running AI on your own infrastructure means open-weight models (Qwen, Llama, Mistral, GPT-oss) served from hardware you control — a workstation for a team, a GPU server for a company. Businesses do it for three reasons: data that must not leave the building, cost at sustained volume, and independence from any single vendor or jurisdiction. This hub covers the why (sovereignty, compliance, economics) and the how (hardware, models, deployment, vendor selection) — with every hardware figure computed by the same engine as our GPU checker.

Two shifts made this page necessary. First, production AI inference is moving out of the public cloud: in Broadcom's Private Cloud Outlook 2026 — a survey of 1,800 senior IT decision-makers at enterprises across eight countries — the share using public cloud as the primary environment for production AI inference fell from 56% to 41% in a single year, while 56% now run or plan production inference on private cloud. Second, AI has become a sovereignty question: 77% of organizations now factor an AI vendor's country of origin into selection decisions (Deloitte, State of AI in the Enterprise — survey of 3,235 business and IT leaders across 24 countries, late 2025).

Neither shift is ideological. Inference at business volume is a predictable, steady-state workload — exactly the kind that owned hardware prices well against metered APIs — and the data flowing through internal AI assistants is exactly the data legal teams least want crossing borders or third-party processors. The result is that "can we run this ourselves?" has become a normal procurement question, asked by IT managers and compliance leads rather than enthusiasts.

This hub is written for that question. It is vendor-neutral: we don't sell hardware, platforms, or models — the recommendations come from the same open compute engine that powers this site's GPU compatibility checker and cost calculator, and where a page touches regulation it describes mechanics, not legal advice. Start with the guide that matches where you are: understanding the landscape, pricing a deployment, or choosing between vendors.

Built for your industry and its regulator

The deployment pattern is the same everywhere — open-weight models on infrastructure you control — but the thing you have to prove differs by sector. A hospital proves PHI never left the covered entity. A defence integrator proves the enclave has no route to the internet. A bank proves every prompt and completion is reconstructable years later. These pages start from the obligation and work back to the architecture.

Healthcare · HIPAAHIPAA-compliant AI

PHI never leaves the covered entity — no BAA to negotiate, because there is no third-party processor in the path.

See the deployment →
EU · GDPRGDPR-compliant AI

Article 17 erasure against frozen model weights, Article 28 without a processor chain, and Chapter V transfers that never happen.

See the deployment →
Sovereignty · JurisdictionSovereign AI

Weights, compute, data, and operations under one jurisdiction — the four layers, and which ones a hosted API can never give you.

See the deployment →
Isolated networksAir-gapped AI

A model server with no route to the internet: media-transfer update path, offline licence posture, and what breaks.

See the deployment →
Defence · ClassifiedDefence & military AI

IL5/IL6 enclave topology, ITAR-conscious model sourcing, and inference at the tactical edge where the link drops.

See the deployment →
Healthcare · ClinicalHealthcare AI

Clinical documentation, coding support, and RAG over the chart — with the de-identification and audit trail the regulator expects.

See the deployment →
Legal · PrivilegeLegal AI

Privilege-preserving review and drafting: no third-party disclosure, citation-grounded output, matter-scoped retrieval.

See the deployment →
Financial servicesFinancial services AI

Reconstructable prompt-and-completion records, model risk management under SR 11-7, and no customer data in a vendor log.

See the deployment →
Government · Public sectorGovernment & public sector AI

FedRAMP authorization boundaries, GovCloud and on-prem patterns, and the records obligations an AI assistant inherits.

See the deployment →
Manufacturing · OTManufacturing AI

Process knowledge and defect data that never leave the plant — OT-segmented inference under IEC 62443 zones.

See the deployment →

Four ways to deploy

Every vertical above runs on one of these. The choice is driven by your obligations and your appetite for operating hardware — not by model capability, which is the same in all four.

On-premise

GPUs in your own racks or server room. Nothing leaves the building; you own the hardware outright.

Best for: fixed sites, sustained volume, strict residency
Private cloud / VPC

Your cloud tenancy, your keys, your network boundary — rented GPUs instead of bought ones.

Best for: elastic demand, no datacentre of your own
Air-gapped enclave

No physical route to the internet. Models and updates arrive as reviewed media on a one-way path.

Best for: classified, ITAR-controlled, safety-critical
Sovereign hosted

EU- or country-hosted infrastructure under local operator control and local jurisdiction only.

Best for: residency mandates without owning hardware

What business hardware actually runs

One model per deployment tier — VRAM figures come from the same engine as the compatibility checker, so every row is independently verifiable.

ModelVRAM (Q4_K_M)Runs onContextLicense
Qwen 3 32B
Department tier — one 24–48 GB workstation
The proven internal-assistant class: summarization, drafting, RAG over internal documents. Apache 2.0.
ollama pull qwen3:32b
20.6 GB24 GB GPU (RTX 3090/4090)
Mac: 32 GB unified
125KApache 2.0
Llama 3.3 70B Instruct
Company tier — one 48–80 GB card
The reference open 70B for knowledge work; runs fully on a single RTX 6000 Ada / A100-class card.
ollama pull llama3.3:70b
43.1 GB2×24 GB GPUs or 48 GB card
Mac: 64 GB unified
128KLlama Community
GPT-OSS 120B
Server tier — one 80–96 GB card
GPT-4-class benchmark scores from an open-weight model on a single H100 or RTX PRO 6000.
ollama pull gpt-oss:120b
71.3 GB2×48 GB GPUs / big unified memory
Mac: 96 GB unified
128KApache-2.0
Qwen 3 235B-A22B (MoE)
Multi-GPU tier — 2–4 server cards
Frontier-adjacent MoE quality for organizations that need the ceiling in-house.
ollama pull qwen3:235b-a22b
142.7 GBMulti-GPU server — or use an API
Mac: 192 GB unified
125KApache 2.0

The guides

What Is Sovereign AI — and Why Country of Origin Now Matters

The four layers of AI sovereignty, why 77% of firms now vet vendor origin (Deloitte), and the practical control spectrum.

Read the guide →

On-Prem & Private AI: The 2026 Data

Every credible 2026 number on the inference shift, sovereignty screening, and cost — sourced, methodology included, one myth debunked.

Read the guide →

On-Prem LLM Total Cost of Ownership vs Cloud APIs at Company Scale

The four cost lines, two worked examples (1M and 20M tokens/day), and the honest column where cloud stays cheaper.

Read the guide →

On-Prem AI Hardware for Business: 10, 50, and 200-Seat Tiers

Sizing by headcount: one workstation for 10 seats, one pro card for 50, a GPU server for 200+ — with verifiable VRAM math.

Read the guide →

Choosing an On-Prem AI Vendor: The 30-Point Evaluation Checklist

30 RFP-ready questions across six dimensions — data path, model rights, sizing honesty, exit terms — with the disqualifying answers.

Read the guide →

On-Prem LLM Deployment Architecture: vLLM, SSO, and Logging

The production architecture in six steps: vLLM or Ollama behind an SSO gateway with audit logging — an engineer-week, not a program.

Read the guide →

Air-Gapped LLM Deployment for Healthcare and Legal

Who actually needs the air gap (healthcare and legal, specifically), why LLMs air-gap unusually well, and the six-step enclave build.

Read the guide →

Open-Model Licenses for Commercial Use: What Legal Will Ask

Apache/MIT vs Llama/Gemma community licenses — the five questions legal will ask, answered before they ask them.

Read the guide →

The EU AI Act and Self-Hosted AI: What Changes in August 2026

The August 2026 timeline, deployer vs provider roles, and what self-hosting changes under the Act (evidence) vs what it doesn't (obligations).

Read the guide →

What to demand from any AI vendor

We do not sell infrastructure, so these are the criteria we would apply on your behalf — the short version of our 30-question evaluation checklist.

SOC 2 Type II

Ask for the current report and the bridge letter — a Type I, or a report older than twelve months, is not the same control evidence.

99.9% uptime SLA

Get the remedy in writing, not just the number: what is measured, over what window, and what the credit actually pays.

Deployment model & data path

On-prem, VPC, or air-gapped — and a written statement of every hop your prompts take, including subprocessors.

Exit terms & model rights

Confirm you keep the weights, the adapters you trained, and your data in a portable format when the contract ends.

Rolling this out in your organization?

Jakub Rusinowski, the founder of LLM Configurator, runs corporate workshops and lectures on deploying local LLMs — hardware sizing, model selection, compliance-friendly architectures, and hands-on setup for your team. Direct, vendor-neutral, practitioner-level.

Ask about a workshop →

Frequently asked questions

What does "sovereign AI" actually mean for a business?

Keeping the four layers of an AI system — the data, the models, the compute, and day-to-day operational control — inside infrastructure and jurisdictions you choose. For a company that usually means open-weight models on owned or EU-hosted hardware instead of a foreign cloud API. It became a mainstream procurement criterion in 2025–2026: 77% of organizations now factor an AI vendor's country of origin into selection decisions, per Deloitte's State of AI in the Enterprise survey of 3,235 business and IT leaders.

Is on-premise AI actually cheaper than cloud APIs?

At sustained volume, usually yes. Inference at business scale is a steady-state workload: a one-time hardware cost amortized over 2–3 years plus electricity, versus a metered per-token bill forever. At roughly 1M tokens/day a single-workstation deployment already undercuts frontier API pricing; at company scale the gap widens. Below a few hundred thousand tokens a day, cloud APIs stay cheaper — run your own volumes in our break-even calculator before deciding.

What hardware does a company need to run AI on-premise?

By team size: a 24–48 GB GPU workstation (RTX 4090 / RTX 6000 Ada class) serves a department of ~10 with a 32B-class assistant; a single 48–96 GB pro card (RTX PRO 6000, A100/H100) serves ~50 with a 70B-class model; a 2–8 GPU server serves hundreds with frontier-adjacent MoE models. Every figure on this hub comes from the same VRAM engine as our free compatibility checker, so you can verify any configuration yourself.

Are open-weight models good enough for business use in 2026?

For the workloads businesses actually deploy — internal assistants, document summarization, RAG over company knowledge, drafting, coding help — yes. Qwen 3 32B and Llama 3.3 70B handle these well on single-card hardware, and GPT-oss 120B matches GPT-4o on standard benchmarks from one server GPU. Frontier cloud models keep an edge on the hardest reasoning tasks; most organizations pilot with their real documents before committing either way.

Does running AI on-premise automatically solve GDPR and compliance?

No — but it removes the hardest parts. With inference on your own hardware there is no third-party model processor, no cross-border transfer analysis for the inference itself, and no dependence on a vendor's retention terms. You still need lawful basis, access controls, and sensible logging — internal controls you already operate for other systems. Our GDPR comparison page walks through the mechanics; involve your DPO for specifics.

Keep going