ISOLATED NETWORKS
Air-Gapped AI: Local LLMs on Networks With No Route to the Internet
The environment where "we reviewed the vendor's security posture" is not an acceptable answer, because nothing connects outward at all.
Written by Jakub Rusinowski · Last updated 2026-08-01 · Hardware figures computed by our VRAM engine
An air-gapped AI deployment runs an open-weight model on hardware inside an isolated network that has no route to the internet, with models, dependencies, and updates crossing the boundary as reviewed media rather than over a link. Large language models suit this unusually well: once the weights are on disk, inference is a pure local computation with no licence check, no telemetry, and no callback — so the capability degrades gracefully into isolation in a way most enterprise software does not.
Why this is hard
- Most enterprise AI tooling assumes an outbound connection for licensing, telemetry, or model updates — and fails closed without one.
- A model that phones home once at startup is not air-gapped, and you often only discover it during accreditation.
- Package installs quietly reach the internet; an environment that has never been built offline usually cannot be.
What the deployment gives you
No outbound route, by construction
The inference host sits on the isolated segment with no default gateway to any external network. Weights load from local storage and inference is a self-contained computation — there is no licence server to reach and nothing to block.
A build that assumes offline from the start
Container images, model files, Python wheels, and CUDA components staged into an internal registry and mirror. The rule that saves projects: if it was not built and tested offline, it does not work offline.
A reviewable update path
Model and dependency updates arrive as signed, checksummed media through your existing transfer procedure. Every crossing is an auditable event with a named approver, which is what accreditation actually asks for.
Verified provenance before the gap
Weights are hashed, scanned, and recorded on the connected side, then verified again inside. Prefer safetensors over pickle-based formats — deserialisation of an untrusted checkpoint is code execution, and inside an enclave that is the worst place to find out.
How the model gets in — and nothing gets out
The whole design reduces to one question: what crosses the boundary, in which direction, under whose approval? Everything on the isolated side is ordinary infrastructure. The interesting engineering is entirely in the transfer step and in having built the enclave to run without a network in the first place.
- Verify checksums on both sides of the gap. A hash computed only where the file was downloaded proves nothing about what arrived.
- Pin exact versions of every dependency, including CUDA and driver builds — "latest" is meaningless in an environment that cannot fetch it.
- Prefer safetensors weights; loading a pickle checkpoint executes code, and an enclave is the last place you want to discover that.
- Keep a rebuild runbook that assumes zero connectivity and test it. The failure mode is discovering the gap during an incident, not during setup.
- Log the transfer events, not just the system events. During accreditation the boundary crossings are what gets examined.
What you have to satisfy
- NIST SP 800-53 — SC-7 boundary protection: The control family that isolation is assessed against; an enclave with a documented, monitored transfer procedure is the pattern assessors expect to see.
- IEC 62443-3-3 — Zones and conduits: The industrial framing of the same idea — inference belongs in a defined zone with an explicit, minimal conduit rather than a general network path.
- NIST SP 800-171 — Controlled unclassified information: Where CUI is in scope, isolating the processing environment materially reduces the assessment surface for the AI system itself.
Models that fit this deployment
| 模型 | VRAM (Q4) | 可运行于 | 上下文 | 许可 |
|---|---|---|---|---|
| Qwen 3 14B Compact enclave tier — 12 GB card — Where the enclave has limited power and cooling; still capable enough for retrieval and drafting. ollama pull qwen3:14b | 9.7 GB | 12 GB GPU (RTX 3060 12GB / 4070) Mac: 16 GB 统一内存 | 125K | Apache 2.0 |
| Qwen 3 32B Standard enclave tier — 24 GB card — The default choice: one card, permissive licence, no runtime dependency on anything external. ollama pull qwen3:32b | 20.6 GB | 24 GB GPU (RTX 3090/4090) Mac: 32 GB 统一内存 | 125K | Apache 2.0 |
| GPT-OSS 120B High-capability enclave — 2×48 GB — When the isolated environment has to match what colleagues get from a frontier API outside. ollama pull gpt-oss:120b | 71.3 GB | 2×48 GB GPUs / big unified memory Mac: 96 GB 统一内存 | 128K | Apache-2.0 |
Standing up AI inside an isolated environment?
We help teams design and build air-gapped AI enclaves — staging and transfer procedure, offline dependency mirroring, model selection for constrained hardware, and the runbook your accreditation process will ask to see. Vendor-neutral and practitioner-level.
常见问题
Go deeper
- Air-gapped LLM deployment — the six-step build guide
- Production deployment architecture, step by step
- Hardware sizing by team size, with verifiable VRAM math
- Sovereign cloud providers in the EU
- GPU compatibility checker — size the enclave hardware
- Build planner — spec a machine for an isolated site
- Local vs cloud AI at enterprise scale
- Defence & military AI
- Government & public sector AI
- Manufacturing AI
- Enterprise & Sovereign AI — the full hub