Tuesday, August 11, 2026 · Week 33 DE · EN · FR · ES Dark
Guides

When GPUs Devour the SaaS Budget

Servers and data centers are growing faster than software: Why IT budgets are shifting to GPUs and what FinOps levers cloud teams need now.

By Alec Chizhik July 15, 2026 6 min read
When GPUs Devour the SaaS Budget

Servers and data centres are projected to grow by more than 30 per cent in 2026 according to industry forecasts, while software lags behind. Within the same IT budget, this often translates to: GPU hours and cluster capacity increase, while SaaS and licence positions are trimmed. For cloud teams, this isn’t a market narrative-it’s a budget and architecture question.

Key Takeaways

  • Two speeds in the budget. AI infrastructure and data centres are pulling the growth engine, while software continues to expand-yet noticeably slower than the hardware layer.
  • Rearrangement, not universal growth. Teams sharing the same pot feel the trade-off: inference and training often displace seats, tools, and consultant budgets.
  • FinOps needs a third axis. Traditional cloud tagging isn’t enough. Teams require workload placement, reserved vs. on-demand rules, and hard kill criteria for idle GPUs.

Related:Model loading devours the expensive TPU hour  /  Sovereign cloud fails at the connection, not the code

What the numbers really mean

Gartner expects global IT spending to keep expanding in 2026, driven chiefly by AI infrastructure and data-centre systems. In its latest forecast, data-centre systems surge by more than 50 per cent and servers by over 30 per cent. Software is still growing, but in the low-to-mid single digits and, in parts of the forecast, has been revised downward. The message isn’t “software is dying.” The message is: the next euro in the IT budget is first spent on compute power.

At the same time, pure AI outlays are ballooning globally to levels that dwarf traditional software budgets, anchored by AI-optimised servers and infrastructure. Those who read this as a hyperscaler story miss the mid-market effect: many DACH organisations aren’t enlarging their overall IT pot at the same pace. They’re reshuffling. And reshuffling means seat licences, shadow SaaS, and “nice-to-have” platforms come under pressure the moment the GPU invoice or colocation megawatt commitment lands.

On X and in vendor communications, this condenses into an industry rule: a tech giant warns on revenue, the market reads it as a shift from software to AI hardware. Regardless of individual quarterly figures, the mechanism is tangible for practitioners. Fixed IT budgets behave like communicating vessels.

// Metric
+50 %
Magnitude of the growth Gartner forecasts for data-centre systems in 2026-well above the expansion rate of the software pillar in the same forecast universe.
// Source: Gartner Worldwide IT Spending Forecast 2026 (growth rates, without absolute prices)

Why GPU Costs Break SaaS Logic

Traditional cloud FinOps is optimized for elastic, short-lived workloads: instances you shut down, storage you tier, licenses you count per seat. Persistent GPU demand shatters these assumptions. Training runs that span days or weeks, inference clusters running at high utilization, and expensive idle times create cost profiles that neither “pay-as-you-go by default” nor “one SaaS dashboard per team” can capture.

Three structural differences:

  1. Granularity. A SaaS license is predictable per month. A GPU hour fluctuates with queue depth, spot market, region, and model size.
  2. CapEx pressure. Securing capacity (reserved instances, committed-use discounts, on-prem racks) locks budget before any benefit is realized. Staying flexible means paying premium rates and risking underutilized gaps.
  3. Ownership. SaaS is often bought by the line of business. GPU capacity lands in the platform team, cloud center of excellence, or data-science group-and collides with the traditional software budget of the business IT function.

That’s why the shift inside companies often feels unfair: the AI team scales, the business unit loses a tool. FinOps must surface this conflict line instead of burying it inside a single cloud invoice.

What Usually Gives in the Budget

No drama, just patterns from the field:

  • Duplicate productivity suites and unused seats after M&A or tool sprawl
  • Observability and dev tooling with overlap (three APM licenses, two CI systems)
  • Consulting and implementation in favor of “we build internally with agents”
  • Feature modules in ERP/CRM that nobody adopted
  • Over-provisioned non-prod cloud that suddenly competes with the GPU pool for the same hyperscaler commitment

This isn’t a call to slash software across the board. It’s the observation that CFOs and CIOs can only spend each euro once. If the GPU roadmap isn’t reconciled with the license landscape, the reallocation lands as a surprise in the Q3 forecast.

// contributes
  • Clear workload classes (training, batch inference, online inference, experiments)
  • Commit strategies per class instead of one GPU blanket fee
  • License rightsizing alongside capacity planning
// open issues
  • Idle GPUs without owners and without auto-stop
  • “Cloud first” for steady workloads without a cost model
  • SaaS cuts based on gut feel instead of usage data

FinOps Checklist for the Reallocation Phase

Pragmatic steps that can be completed in a single planning session:

  1. A single currency. Convert GPU hours, tokens and SaaS seats into euros per month per team. Without a common unit, the debate remains political.
  2. Workload placement. What must run on-demand in the public cloud, what warrants reserved or committed capacity, and what belongs in colocation or your own racks once utilisation stabilises?
  3. Kill-switch for idle. Terminate unused notebook instances and experimental clusters after X hours. This is the fastest lever before any licence kill.
  4. Seat audit with data. Review 90 days of usage before cancellation. Otherwise you cancel licences only to repurchase them at a higher price next quarter.
  5. Separate commit vs. flex. Lock in baseline load, keep peaks flexible. A 100 % on-demand GPU pool is rarely FinOps; it’s usually convenience.
  6. Factor in power and location. Capacity without grid access and permits is dead budget-see site risks in data-centre projects.
  7. Document the exit. When the model changes or the cluster shrinks: which licences and commits remain, and who pays?

// Definition

What is budget reallocation in the AI context? The shift of IT spending within a fixed budget from traditional software and service line items toward compute power for training and inference (GPUs, AI servers, specialised cloud instances, colocation). It’s an allocation problem, not proof that software is irrelevant.

What Platform Teams Must Change Now

The operational answer isn’t “more dashboards.” It’s governance with teeth:

Showback before chargeback. First make costs visible per team and workload class, then implement internal billing. Chargeback without showback only spawns shadow accounts.

Quotas instead of hope. Every experiment team receives a GPU-hour quota per sprint. Exceeding it requires approval, not a retrospective email to FinOps.

Review software and compute together. Quarterly joint review: licence owners and cluster owners in one room. Reviewing only SaaS reveals half the truth. Reviewing only GPUs silently dismantles productivity tools.

The market can keep talking about global record forecasts. In the engine room, the question is whether you plan the reallocation-or endure it.

Frequently Asked Questions

What is budget reallocation in the context of AI?

The shift of IT funds from software licences and traditional services towards GPU and AI infrastructure within the same or only slightly growing budgets. The cause is the stronger growth of the infrastructure pillar compared to many software positions.

Does that mean SaaS is becoming unimportant?

No. SaaS continues to grow, but often more slowly than AI hardware and data-centre capacity. Within a limited budget, unused seats and duplicate tools are the first to go. Mission-critical systems remain. The lever lies in rightsizing, not blanket elimination.

Why don’t classic cloud FinOps tags suffice?

Because GPU workloads have different cost profiles: long runtimes, expensive idle time, committed-use discounts and placement decisions. Tags help with attribution but do not replace workload classes, quotas and kill switches.

Cloud, colocation or on-prem GPUs – which is cheaper?

It depends on utilisation, runtime and operational capability. Short-lived experiments and peaks often favour public cloud. Stable, high utilisation over months may justify reserved instances, colocation or on-prem capacity. Without utilisation data, “on-prem is cheaper” is an assertion, not a calculation.

What should the next FinOps review include?

GPU hours and idle rate per team, SaaS seats with 90-day utilisation, open committed-use discounts, placement decisions per workload class and an explicit list of which licence positions are being sacrificed for capacity – including owner and deadline.

Editor’s Reading List

mybusinessfuture

When a German AI model actually pays off

digital-chiefs

The bill for a decade of siloed solutions

securitytoday

What is KRITIS? Operators, obligations and thresholds

Image source: AI-generated (July 2026)

Also available in

FrançaisEspañolDeutsch
MBF Media Newsletter

The monthly briefing for decision-makers

Once a month, the MBF Media Newsletter gathers what matters from cloudmagazin, MyBusinessFuture, Digital Chiefs and SecurityToday, curated by the editorial team.

25,000 IT and business decision-makers read this newsletter. Read along.

Subscribe for free
MBF Media Newsletter, aktuelle Ausgabe auf dem iPhone
A magazine by Evernine Media GmbH