Enterprise AI statistics have a laundering problem: a number appears in a vendor deck, gets repeated without its methodology, mutates in transit, and two blog-generations later is "well known." This page is the antidote for the on-prem/sovereign corner of the field. Each statistic below states who measured it, how, and when. Where a figure is vendor-commissioned, that's labeled — commissioned research isn't worthless, but you should know who paid. And where a widely repeated number turns out to have no traceable source, we say so instead of repeating it.
Figures are current as of July 2026; this page is revised as new survey waves publish (see the last-updated date above).
Where inference runs: the shift out of public cloud
Sovereignty: vendor origin as a selection criterion
Cost: what the credible comparisons show
The number you should NOT cite
A claim circulates that "55% of enterprise AI inference is now performed on-premises or at the edge, up from 12% in 2023." We attempted to verify it for this page and could not find any primary source — no survey, no analyst report, no methodology. The credible measurements above tell a different (and more modest) story, and IDC-cited figures still put roughly two-thirds of enterprise AI compute in the cloud. The likely origin is a garbled version of a Gartner prediction about edge analytics — "by 2025, 55% of all data analysis by deep neural networks will occur at the point of capture in an edge system, up from less than 10% in 2021" — which measures something else entirely (data analysis at point of capture, not enterprise inference share).
If you've used the 55% figure in a deck: the defensible replacement is the Broadcom 56→41% shift, which is sourced, recent, and directionally similar. If you've seen other uncited on-prem AI numbers you'd like run down, tell us — corrections and verifications are added to this page.
Citing this page
Every statistic above should be cited to its original publisher (Broadcom/Radius Tech, Deloitte, Cloudian, Gartner), not to us — this page exists to make those chains easy to follow. If you link this roundup as a secondary source, it lives at a stable URL and each revision updates the last-updated date above. For dated, citable snapshots of the local-AI hardware and model landscape more broadly, see our biweekly reports.
Frequently asked questions
What percentage of enterprise AI inference runs on-premises?
No credible survey measures "share of inference" directly. The best-sourced adjacent figures: public cloud as the primary environment for production AI inference fell from 56% to 41% of enterprises year-over-year, and 56% now run or plan production inference on private cloud (Broadcom Private Cloud Outlook 2026, n=1,800 senior IT decision-makers). The widely repeated "55% on-prem, up from 12% in 2023" claim has no traceable primary source and should not be cited.
Is the "77% factor country of origin" statistic reliable?
It is among the best-sourced statistics in this space: Deloitte's State of AI in the Enterprise, surveying 3,235 business and IT leaders across 24 countries and six industries in August–September 2025. Deloitte does not sell AI infrastructure, the sample is large and multi-country, and the fieldwork window is disclosed. Cite it with attribution: "per Deloitte's State of AI in the Enterprise survey."
Are enterprises really moving AI workloads out of the public cloud?
Two independent 2026 surveys point the same direction with different strengths: Broadcom/Radius Tech (n=1,800, disclosed methodology) measured public-cloud-primary inference falling 56%→41% year-over-year; Cloudian's vendor-commissioned survey found 93% repatriating, in progress, or evaluating. The honest summary: a significant, measurable shift toward private infrastructure for inference specifically — not a wholesale cloud exodus.
How much cheaper is on-premise AI than cloud APIs?
Above a volume threshold, materially: Deloitte's TMT analysis estimates 50%+ savings over three years, and our own calculator (verified per-token prices, street hardware prices) shows a single-GPU workstation undercutting GPT-4o API spend within two years at ~1M tokens/day. Below a few hundred thousand tokens/day, APIs stay cheaper. The threshold depends on your volumes — compute it, don't assume it.