DEFENCE · CLASSIFIED
Military & Defence AI: Inference Inside the Enclave, Including at the Edge
Capability that stops working when the satellite link drops is not capability — it is a demonstration.
Written by Jakub Rusinowski · Last updated 2026-08-01 · Hardware figures computed by our VRAM engine
This page is practitioner guidance on deployment mechanics, not legal advice — involve your counsel or DPO for decisions about your specific obligations.
Defence AI deployment means running models inside the accredited enclave that already holds the data, rather than sending that data to a commercial API. For most defence organisations this resolves to open-weight models on accredited on-premise or government-cloud infrastructure at the appropriate impact level, plus a genuinely disconnected edge tier for platforms and forward elements that cannot assume a link. The technical work is unremarkable; the difficulty is accreditation, provenance, and export-control awareness in model sourcing.
Why this is hard
- Mission data cannot traverse a commercial API, so the tools everyone else uses are simply unavailable.
- Connectivity at the edge is contested, intermittent, and sometimes deliberately absent — cloud inference assumes none of that.
- Model provenance is a supply-chain question here, and "we downloaded it from a public hub" is not a provenance answer.
What the deployment gives you
Inference inside the accredited boundary
The model runs at the impact level where the data already lives, so no new cross-domain flow is created. The AI system inherits the enclave's existing accreditation posture instead of opening a fresh boundary discussion.
Edge inference that assumes no link
A compact model on a ruggedised workstation or embedded GPU keeps working when the connection is degraded, denied, or intentionally silent. Sizing follows the platform power and thermal envelope, not the datacentre catalogue.
Provenance recorded, not assumed
Weights hashed, scanned, and version-pinned on the connected side, with the record following the artefact into the enclave. Publisher jurisdiction, licence terms, and export-control posture reviewed before anything is staged.
Defence-in-depth alignment
Identity from the enclave directory, least privilege on the retrieval corpus, full audit of prompts and completions, and no egress from the inference host — controls your assessors already recognise, applied to one more system.
Two tiers: the enclave and the edge
Defence deployments almost always need both, and they are sized very differently. The enclave tier is a normal GPU deployment inside an accredited boundary. The edge tier is a much smaller model on constrained hardware whose defining requirement is that it never expects a network.
- Size the edge tier to the platform, not the model you wish you could run — power, thermals, and shock envelope decide this before capability does.
- Treat model artefacts as a supply-chain item with a recorded publisher, licence, hash, and review date, the same as any other dependency.
- Keep the human decision-maker in the loop by design. These systems support analysis and staff work; they are not the place for autonomous consequential action.
- Plan the update cadence around the transfer procedure that already exists, rather than proposing a new pathway across the boundary.
- Review licence field-of-use terms specifically — some open model licences carry acceptable-use restrictions that a defence programme has to read closely.
What you have to satisfy
- DoD CC SRG IL5 / IL6 — Impact levels: IL5 covers controlled unclassified national-security information; IL6 covers classified up to SECRET. The level sets where inference may run and who may operate it.
- ITAR / EAR — Export control: Technical data and some software are controlled. Model sourcing, foreign-national access to the enclave, and cross-border support all need review with your export-control officer.
- NIST SP 800-171 — CUI in nonfederal systems: The baseline for defence contractors holding controlled unclassified information; isolating inference keeps the AI system inside an already-assessed boundary.
- NIST SP 800-53 SC-7 — Boundary protection: The control assessors examine for enclave isolation — a documented, monitored transfer path rather than an ad-hoc one.
Models that fit this deployment
| 模型 | VRAM (Q4) | 可运行于 | 上下文 | 许可 |
|---|---|---|---|---|
| Qwen 3 14B Edge tier — 12 GB, ruggedised — Fits an embedded or mobile GPU envelope while still handling summarisation and structured extraction. ollama pull qwen3:14b | 9.7 GB | 12 GB GPU (RTX 3060 12GB / 4070) Mac: 16 GB 统一内存 | 125K | Apache 2.0 |
| Qwen 3 32B Enclave workstation tier — 24 GB — Analyst-facing assistant on a single card; permissive licence simplifies the legal review. ollama pull qwen3:32b | 20.6 GB | 24 GB GPU (RTX 3090/4090) Mac: 32 GB 统一内存 | 125K | Apache 2.0 |
| GPT-OSS 120B Enclave server tier — 2×48 GB — Where the enclave has to match commercial frontier capability for staff work. ollama pull gpt-oss:120b | 71.3 GB | 2×48 GB GPUs / big unified memory Mac: 96 GB 统一内存 | 128K | Apache-2.0 |
Deploying AI inside an accredited or disconnected environment?
We help defence organisations and their suppliers architect AI that runs where the data already is — enclave sizing, edge hardware selection within real power and thermal envelopes, model provenance review, and the documentation your accreditation process expects. Vendor-neutral and practitioner-level.
常见问题
Go deeper
- Air-gapped LLM deployment — the enclave build guide
- Open model licences — field-of-use and acceptable-use terms
- Hardware sizing with verifiable VRAM math
- Build planner — spec ruggedised and enclave hardware
- GPU compatibility checker — verify the edge envelope
- Local vs cloud AI at enterprise scale
- Air-gapped AI
- Government & public sector AI
- Sovereign AI
- Enterprise & Sovereign AI — the full hub