Broadcom surveyed 1,800 senior IT decision-makers and found production inferencing on private cloud at 56% while public cloud fell to 41%. The repatriation story is real, but it is redistribution rather than retreat.
Key Takeaways
For most of the last three years the assumption underneath enterprise AI planning was that inference belonged in the public cloud. The models were there, the accelerators were there, and nobody wanted to buy GPUs on a depreciation schedule for a capability whose requirements were changing every quarter. Broadcom's Private Cloud Outlook 2026, a survey of 1,800 senior IT decision-makers across the Americas, Europe and Asia-Pacific, says that assumption has inverted.
Production AI inferencing now runs, or is planned to run, on private cloud at 56% of the organisations surveyed. Public cloud's share fell to 41%, down 15 percentage points from 56% the year before. A 15-point move in a single year on a workload this strategically loaded is not a preference drifting. It is a decision being made repeatedly, by different organisations, on grounds solid enough to override the convenience of staying where the models already are.
The workloads that made the public cloud business case were bursty, unpredictable, and cheap to leave idle. Inference is close to the opposite. Production inference is steady-state, latency-sensitive, and runs continuously against accelerators that are expensive precisely because they are always busy. Pay-as-you-go pricing is a poor fit for a workload that never stops, and the more successful an AI feature becomes the worse the arithmetic gets. That is a rare shape in enterprise computing: a workload whose unit economics improve with owned infrastructure.
Data gravity compounds it. Inference wants to sit next to the corpus it reasons over, and for most enterprises that corpus is the accumulated content of twenty years of internal systems, much of it subject to residency rules that were written before anyone was sending documents to a model. Moving the model to the data is often cheaper and always simpler than moving the data to the model.
“Enterprise AI has found its infrastructure home. And it is private cloud.”
Prashanth Shenoy, Vice President of Marketing, VMware Cloud Foundation Division at Broadcom
It would be easy to read this as the cloud reversal that on-premises vendors have been predicting since roughly 2012, and that reading would be wrong. Broadcom found 83% of enterprises considering or already executing repatriation, with 50% having moved something back. But IDC puts full repatriation at only 8% to 9% of enterprises. Almost nobody is leaving. What is happening is that workload placement has become a per-workload decision again, made on cost and latency and sovereignty rather than on a corporate cloud-first mandate.
It is worth being precise about which number means what, because the headline figures in this debate get conflated constantly. The widely circulated claim that 86% of CIOs are moving workloads back originates in a Barclays CIO survey and refers to moving some workloads, on a spectrum that includes a single database. IDC's 8% to 9% counts only comprehensive exits. Both are true and they describe completely different behaviours, which is how the same underlying trend gets reported as either a mass migration or a statistical illusion.
That is a more demanding operating model than either extreme. A cloud-first policy is easy to administer because it removes the decision. A genuinely hybrid estate requires someone to own placement criteria, re-evaluate them as pricing and hardware change, and maintain security controls that behave identically in both environments. This publication's earlier analysis of total cost of ownership found the gap between on-premises and cloud far narrower than most models assume, and a narrow gap is exactly the condition under which per-workload judgement beats policy.
The security model is where repatriated AI tends to disappoint. In public cloud, a great deal of workload isolation, key management and audit logging arrives as a service default. Rebuilding that around a private inference cluster is real engineering, and it is usually scoped after the GPUs are ordered rather than before. An inference endpoint with access to the enterprise document store is one of the most sensitive things in the estate, and it will not inherit the controls the rest of the platform has.
There is an operating-cost line that rarely makes the business case. A hybrid estate needs people who can run both halves competently, and the skills for private AI infrastructure, GPU scheduling, accelerator utilisation, and inference-layer capacity planning, are scarce and currently expensive. Organisations that repatriate to save on compute sometimes discover they have converted a predictable cloud bill into an unpredictable hiring problem.
Enterprise AI spent three years being an argument about models. It is now an argument about where the electricity goes, which is a more familiar problem and, for once, one that infrastructure teams are better equipped to answer than anyone else.
Whitepaper
Repatriated inference does not inherit the cloud defaults it used to run behind, which is the control gap this addresses directly.
Download
Whitepaper
Inference moves to sit beside the corpus, so knowing where regulated content actually lives becomes a placement decision rather than a filing one.
Download
Whitepaper
The other half of a hybrid estate still runs in public cloud, and the risk case for what stays there deserves the same scrutiny as what moves.
Download
Analysis of 300 enterprise workloads finds the true cost gap is far narrower than most TCO models suggest.
Edge deployments have moved from pilot to production across manufacturing, retail and healthcare.