After two decades of moving to the cloud, enterprises are actively rebuilding the internal service provider model. The difference this time: it has to span cloud, on-prem, and edge, with AI workloads that are unrecognizable from the ones pre-dating the move to cloud.
For almost twenty years, the story of enterprise IT has been a one-way trip out of the datacenter. First virtualization, then the public cloud, then “cloud-first” as policy: if a workload could run somewhere else, it eventually did. Along the way, IT organizations largely got out of the business of being infrastructure providers to their own company. That job moved to AWS, Azure, and Google.
That trip is no longer one-way. Rising GPU and inference workloads, an edge computing footprint that’s growing faster than almost anything else in IT, and a decade of unpredictable cloud bills are pulling infrastructure decisions back inside the perimeter, not as a retreat from cloud, but as a much more deliberate, workload-by-workload placement strategy. And that shift is quietly resurrecting a job IT hasn’t had to do at scale in years: acting as its own infrastructure service provider, with the self-service, governance, and reliability a cloud provider is expected to deliver.
Three forces are pulling infrastructure back inside the perimeter
None of these forces is “the cloud is dying.” Public cloud spend keeps growing: Gartner still projects worldwide public cloud spending to grow 21.5% in 2025, and IaaS specifically to grow even faster, at 24.8% (Digital Chiefs). What’s changed is that “cloud-first, by default” has given way to “wherever it makes sense,” and increasingly, that’s on-prem or at the edge.
- Cost and control are driving a deliberate rebalancing. The headline stat (86% of CIOs planning to move some workload back to on-prem or private cloud, per the Barclays CIO Survey, Q4 2024) is real, but it’s frequently overstated in how it’s used: it counts a company moving one database the same as one moving 40% of its estate. Only an estimated 8–9% of enterprises plan full repatriation (IDC, via Digital Chiefs). The more accurate picture, from a Broadcom survey of 1,800 IT leaders, is that the vast majority of enterprises now run a deliberate mix:

Only 15% of enterprises run public cloud only and 10% run private cloud only. The rest deliberately mix both. Source: Broadcom global IT leader survey (1,800 respondents), cited via InfoWorld.
93% of respondents said they intentionally balance private and public cloud resources, and 84% run both traditional and cloud-native applications inside their private cloud (InfoWorld). Cost is the entry point: 84% of organizations cite managing cloud spend as their top challenge, and Flexera estimates 27% of cloud infrastructure spend is wasted on underused resources (Flexera State of the Cloud 2025). But security, compliance, and data sovereignty now outrank cost as the top reason to repatriate a given workload, particularly with the EU AI Act, GDPR, and similar rules in play.
Edge computing is one of the fastest-growing categories in IT. Real-time processing, industrial IoT, and telco 5G buildouts are pushing serious compute out of centralized datacenters entirely:

The edge computing market is projected to grow more than 10x between 2025 and 2035. Source: GM Insights, Edge Computing Market Report.
That’s a 28% CAGR from 2026 to 2035 (GM Insights), against a backdrop of roughly 29 billion connected devices expected globally by 2030 and IDC’s estimate that more than 60% of organizations will be running edge analytics by 2027. This isn’t a side project. For many industrial, retail, and telco organizations, it’s becoming a primary infrastructure surface IT has to run and govern like any other environment.
- AI and GPU demand are forcing the placement question for every workload. The scale of this one is hard to overstate: IDC forecasts worldwide AI infrastructure spending will reach $758 billion by 2029, with GPU-equipped (“accelerated”) servers growing at a 42% five-year CAGR and eventually accounting for more than 94% of that spend (IDC). Training may stay in the cloud, but production inference, the part that actually touches customers, is moving somewhere else:

Public cloud’s share of production AI inference fell 15 points in a single year. Source: Broadcom Private Cloud Outlook 2026 (1,800 IT leaders), cited via Campus Technology.
That’s a genuine reversal in a single year. Public cloud’s share of production AI inference fell from 56% to 41%, while 56% of enterprises now say they are running, or planning to run, that same workload on private cloud instead (Broadcom, Private Cloud Outlook 2026, a survey of 1,800 IT leaders). Of the enterprises repatriating any workload at all, 43% say what they’re specifically moving is AI training, large language models, or inference, not a general-purpose application (Campus Technology). Cost is again a proximate driver, GPU capacity is scarce and expensive, and Nvidia has estimated on-prem AI/ML infrastructure can run roughly 30% cheaper than equivalent public cloud capacity. But latency (inference running next to the data it scores) and data residency requirements matter just as much: Gartner forecasts 60% of financial firms outside the US will adopt sovereign or on-premises AI deployments by 2028.
The response: IT is re-becoming an internal service provider
Put those three forces together and the implication is unavoidable: IT organizations now have to run infrastructure across public cloud, private cloud, and an expanding edge, simultaneously, for the same set of internal customers who no longer accept ticket queues and week-long lead times, because they’ve spent a decade getting used to what a cloud provider’s self-service console feels like.
The data shows this shift is already well underway. Gartner’s own framing for it is almost a direct description of the trend:
The share of large engineering organizations with a platform team is projected to nearly double, 2022–2026. Source: Gartner, cited via Roadie.io.
By 2026, Gartner projects 80% of large software engineering organizations will have established a platform team functioning as an internal service provider, up from 45% in 2022 (Gartner, cited via Roadie.io). Where it’s worked, the payoff is measurable, Spotify cut time-to-first-pull-request for new engineers by 55% after building an internal developer platform on Backstage.
But there’s a catch, and it’s the hard part of this whole story. The internal service providers being rebuilt today have a much harder job than the ones that existed before cloud took over. A pre-cloud internal IT shop mostly ran one kind of environment, on infrastructure it fully controlled. Today’s version has to deliver that same self-service, governed experience across public cloud, private cloud, and edge locations that can each have different latency, connectivity, compliance, and hardware constraints, without the dedicated army of engineers a hyperscaler has for the job.
Why “be your own cloud provider” is harder than it sounds
| Challenge | What it looks like in practice |
| The queue-or-chaos choice | Every request routed through a human ticket is a request that could have been instant, but the alternative most teams reach for is direct, ungoverned access, which trades one problem for a worse one |
| Tool fragmentation across environments | Provisioning, cost tracking, observability, and security policy typically live in separate systems for cloud, on-prem, and edge, with no single place enforcing them consistently |
| Governance that doesn’t scale by hand | Manually enforced spend limits, access rules, and configuration standards break down the moment they have to apply across three infrastructure types instead of one |
| Distributed edge operations | Edge locations multiply the number of places infrastructure has to be provisioned, monitored, and kept in policy, often with intermittent connectivity and no local ops staff |
| Cost attribution across a mixed estate | When spend spans cloud invoices, private cloud capacity, edge hardware, and GPU capacity, tracing a bill back to an owner, team, or project gets exponentially harder |
This is precisely the gap between “IT wants to act like a cloud provider again” and “IT actually can.” A cloud provider’s self-service experience is backed by a purpose-built control plane. Recreating that internally, across environments a hyperscaler never has to unify, is a materially different, and harder, engineering problem.
How Torque closes that gap
This is the exact use case Torque is built for: giving IT organizations the same self-service, governed experience a cloud provider offers, across public cloud, private cloud, edge, and bare metal, without building that control plane from scratch. Torque’s own positioning makes the scope explicit: the platform unifies “cloud, private cloud, data center, and edge under a single operational model” instead of leaving teams to run separate tools and separate policies for each one (Quali, Hybrid Infrastructure).
Meet requesters where they already work. Torque exposes infrastructure through a self-service catalog, IDE and internal developer portal integrations, a CLI, and a REST API, so “self-service” doesn’t mean forcing every team onto one interface. It means infrastructure shows up wherever the request is already happening.
One policy layer, enforced everywhere. Every request, regardless of which environment it targets, is evaluated against the same governance before anything is provisioned, whether the target is a public cloud account, a VMware cluster, a bare-metal server, or a GPU pool. That’s the direct fix for tool fragmentation: instead of separate, inconsistent rules per environment, there’s one enforcement point across all of them.
Shadow infrastructure gets discovered, not just blocked. Rather than only policing new requests, Torque continuously scans for resources that were deployed outside it, in any environment, and converts them into governed Infrastructure as Code instead of leaving them as invisible risk. That’s a direct answer to the queue-or-chaos choice: infrastructure that teams stood up on their own in the past doesn’t have to be hunted down and re-provisioned by hand, it gets brought under governance automatically (Quali, Torque AI infrastructure expansion announcement).
Cost and context built in before launch, not audited after. Spend limits, approved instance types, and budget thresholds are active from the moment a resource is requested, and every environment is automatically tagged by owner, team, purpose, and business priority, so cost attribution across a mixed cloud/on-prem/edge estate stops being a forensic exercise.
Day-2 automation instead of manual upkeep. Once an environment is live, Torque continuously monitors for configuration drift, manages its lifecycle, and optimizes it against utilization, cost, and business priority: the ongoing work that a distributed edge footprint makes impossible to do by hand at scale.
Proven at the edge, not just architected for it. Torque already orchestrates NVIDIA DGX Spark deployments, letting autonomous agents run inference workflows locally, without depending on cloud APIs, for real-time decision-making in disconnected or regulated environments across industries including retail, defense, and healthcare (Quali, Torque and NVIDIA DGX Spark announcement). As Quali CEO Lior Koriat put it, “Whether you’re fine-tuning LLMs or running secure, autonomous agents at the edge, Torque ensures that teams can scale fast, stay compliant, and control cost.”
Built for sovereign and regulated AI. For workloads that cannot leave a jurisdiction, Torque’s control plane can run entirely self-hosted, inside the customer’s own infrastructure, whether on-premises, private cloud, or a sovereign cloud region, using OPA-based policy-as-code to govern what models run where and under what rules, with every action logged for audit (Quali, Torque Sovereign AI announcement). That capability matters commercially as much as technically: McKinsey estimates data and AI sovereignty requirements will shape 30 to 40 percent of global AI spending by 2030, and the EU AI Act is moving toward full enforcement in August 2026 (McKinsey, cited via the same Quali announcement).
It plugs into the operational stack you already run. Observability (Datadog, Dynatrace, Prometheus, Splunk), incident management (PagerDuty, Opsgenie), and infrastructure-as-code (Terraform, Ansible) integrate directly, so becoming an internal service provider doesn’t mean ripping out the tools your teams already trust.
The pendulum swinging back toward on-prem and edge isn’t a rejection of the cloud era. It’s enterprises applying the lesson the cloud era taught them: infrastructure should be self-service, governed, and cost-transparent, no matter where it physically runs. The organizations pulling this off are rebuilding IT as an internal service provider with cloud-provider sophistication, across every environment they operate: cloud, private cloud, and edge alike. Torque is the platform that makes that possible without a multi-year build.
See it in action: explore how Torque delivers governed self-service infrastructure at quali.com/internal-infrastructure-service.
Sources
- Barclays CIO Survey, Q4 2024, cited via Tasrie IT
- Digital Chiefs: “Cloud Repatriation 2026 Is a Statistical Illusion”
- InfoWorld: “The Private Cloud Comeback” (Broadcom survey, 1,800 IT leaders)
- Puppet: “The Great Repatriation: Why Enterprises Are Moving Workloads Back On-Premises” (Flexera State of the Cloud 2025)
- GM Insights: Edge Computing Market Size & Share Report
- IDC: Artificial Intelligence Infrastructure Spending to Reach $758Bn by 2029
- Broadcom: Private Cloud Outlook 2026 press release
- Campus Technology: “Research: Enterprise AI Workloads Are Tipping Toward Private Cloud”
- io: “Platform Engineering in 2026: Why DIY Is Dead” (Gartner, Spotify/Backstage data)
- Quali: Internal Infrastructure Service
- Quali: Hybrid Infrastructure
- Quali: Torque and NVIDIA DGX Spark announcement, via PR Newswire
- Quali: Torque AI infrastructure expansion announcement, via PR Newswire
- Quali: Torque Sovereign AI announcement (McKinsey sovereignty-spend estimate), via PR Newswire






