Posted on Mar 11, 2026 · Updated Mar 11, 2026 · 9 min read

Your Cloud Bill Is 27% Waste — Here's Where Engineering Teams Lose Control

Here's a number that should make every engineering lead uncomfortable: $44.5 billion in cloud infrastructure waste, projected for 2025 alone. That's according to a Harness survey of 700 engineering leaders who pointed to one root cause: the disconnect between FinOps teams setting budgets and developers actually provisioning resources.

The problem isn't laziness or incompetence. It's that 55% of developers base cloud purchasing decisions on guesswork, and only 33% can even see which workloads are overprovisioned. You can't fix what you can't see. This article breaks down where cloud waste actually comes from, why it's getting worse despite the rise of FinOps, and what engineering leads can concretely do about it.

TL;DR

Organizations waste 27% of cloud spend on idle and oversized resources (Flexera 2025). Kubernetes clusters average just 10% CPU utilization. The fix isn't another dashboard — it's embedding cost visibility into engineering workflows, automating rightsizing, and making waste everyone's problem.

Data center server room corridor with rows of servers and green indicator lights

How much cloud spend is actually wasted?

Twenty-seven percent. That's the share of cloud IaaS and PaaS spend that goes to waste, according to Flexera's 2025 State of the Cloud Report. Apply that to Gartner's $723 billion global cloud spend forecast, and you're looking at roughly $195 billion burned on resources nobody needs. For a team spending $500K annually, that's $135K going nowhere.

Where does it actually go? The waste breaks down into a few predictable buckets. Idle resources that nobody shut down. Instances sized for peak load that run 24/7 at 10% utilization. Savings Plans and Reserved Instances that nobody bought because nobody owns the decision. And the storage volumes, snapshots, and load balancers that outlive the projects they were created for.

Developer Cost Optimization Gaps% of developers NOT using each optimization practice0%25%50%75%Spot orchestration71%Instance rightsizing61%Reserved/Savings Plans58%Idle resource cleanup48%Source: Harness FinOps in Focus 2025 (700 engineering leaders surveyed)
The majority of developers skip basic cost optimization practices entirely — not from negligence, but from lack of tooling and visibility.

The Harness data paints a stark picture: 71% of developers don't use spot instances, 61% don't rightsize, and 58% haven't adopted reserved pricing. These aren't exotic techniques. They're table stakes for cost efficiency, and most engineering teams simply aren't doing them. For a detailed look at industry-wide waste trends, see the state of cloud waste in 2026.

What makes this worse is the detection lag. It takes an average of 31 days to identify cloud waste and 25 days to rightsize overprovisioned resources (Harness 2025). That's nearly two months of paying for resources before anyone even notices they're unnecessary.

Why is overprovisioning getting worse, not better?

Average CPU utilization across Kubernetes clusters has dropped to just 10% — down from 13% the prior year. Memory utilization sits at 23%. This comes from CAST AI's 2025 Kubernetes Cost Benchmark analyzing 2,100+ organizations. Let that sink in: you're paying for ten CPUs and using one.

Kubernetes Resource Utilization vs. ProvisionedAverage across 2,100+ organizations (CAST AI, 2025)10%utilizedCPU90% wasted23%utilizedMemory77% wastedSource: CAST AI 2025 Kubernetes Cost Benchmark Report
Kubernetes clusters are becoming less efficient year-over-year — CPU utilization dropped from 13% to 10% as infrastructure complexity grows.

Why is this trending the wrong direction? Three reasons.

First, safety margins compound. A developer requests 2x the memory they need "just in case." The platform team adds another 1.5x buffer for node-level headroom. The autoscaler has its own thresholds. You end up provisioning 3-4x what the application actually consumes — and nobody reviews the math after deployment.

Second, microservices multiply the problem. A monolith might waste resources in one place. Break it into 40 services, each with its own resource requests and limits, and you've created 40 opportunities for overprovisioning. Each service gets its own "just in case" buffer, and the aggregate waste grows with every new service added.

Third, nobody owns the review cycle. Rightsizing requires someone to look at utilization data, propose changes, test them, and deploy — repeatedly, across hundreds of workloads. Without automation, it doesn't happen. And 61% of developers aren't rightsizing at all.

Performance analytics dashboard displaying colorful graphs and charts on a laptop screen

The FinOps-engineering disconnect

Fifty-two percent of engineering leaders acknowledge that the disconnect between FinOps and development teams drives cloud waste (IT Pro / Harness 2025). The irony? Everyone agrees FinOps matters. Ninety-six percent of tech executives say it's important to their cloud strategy. Yet only 14.2% of organizations operate at an advanced "Run" maturity level (FinOps Foundation / nOps).

What does the disconnect actually look like? FinOps teams produce monthly cost reports. Developers never read them. Finance sets budgets based on last quarter plus 10%. Engineers provision based on what worked for the last project. Nobody talks to each other until the bill spikes.

The visibility gap is the core issue. Only 43% of developers have real-time data on idle resources. Thirty-nine percent can't see unused or orphaned resources. And 33% lack any visibility into over- or under-provisioned workloads. You're asking developers to optimize costs they can't measure in systems they can't observe.

Meanwhile, FinOps teams are expanding — 59% of organizations grew their FinOps headcount in 2025, up from 51% in 2024 (Flexera 2025). And 87% now use cost efficiency as their primary metric for cloud goals, a 22-point jump from the prior year. The intent is there. The execution isn't.

So why doesn't more headcount solve it? Because FinOps teams often operate as a reporting function, not an embedded engineering practice. They can tell you what you spent last month. They can't tell a developer in a pull request that the Deployment manifest they're about to merge requests 4 GB of memory for a service that peaks at 800 MB.

AI infrastructure is repeating cloud's mistakes

Ninety-eight percent of organizations now manage AI spend — a staggering jump from 63% in 2025 and just 31% in 2024, according to the FinOps Foundation's 2026 State of FinOps report. FinOps for AI is now the #1 forward-looking priority. But awareness doesn't equal readiness.

AI Spend Management Adoption% of organizations actively managing AI infrastructure costs0%25%50%75%100%20242025202631%63%98%Source: FinOps Foundation State of FinOps 2024-2026
AI spend management adoption tripled in two years — but maturity hasn't kept pace with awareness.

A ClearML survey of Fortune 1000 companies found that 44% of enterprises manually assign workloads to GPUs or have no GPU utilization strategy at all. And 53% say cost control is their primary workload management challenge. Sound familiar? It's the same pattern cloud compute went through a decade ago — expensive hardware, manual allocation, zero visibility.

The difference is the price tag. A single H100 GPU instance runs $25-35/hour. Leave a training cluster idle over a weekend and you've burned $2,000+ before Monday morning. The margin for error that exists with $0.10/hour CPU instances doesn't exist with GPU compute. Every hour of waste is amplified by an order of magnitude.

The playbook for fixing AI waste isn't new — it's the same playbook cloud infrastructure teams should have followed from the start. Automate scheduling. Track utilization in real time. Use spot instances where possible (CAST AI reports 59-77% savings on compute with spot adoption). Don't wait until the GPU bill forces the conversation.

What can engineering leads actually do about it?

Eighty-four percent of organizations cite cloud spend as their top cloud challenge (Flexera 2025). If you're an engineering lead, you have more leverage over this than you think. Here's a concrete framework.

1. Make waste visible at the point of decision

The biggest lever is shifting cost visibility from monthly reports to real-time feedback. Show developers what their services cost in the tools they already use — CI/CD pipelines, pull request comments, deployment dashboards. When a developer can see that their resource request will cost $1,200/month before they merge, behavior changes. Our cloud cost optimization checklist has a step-by-step rundown of tooling and process changes that work for engineering teams.

2. Automate the boring stuff

Rightsizing, idle resource cleanup, and spot instance orchestration shouldn't be manual. CAST AI's data shows that even partial spot adoption delivers 59% average compute savings. Exclusive spot usage reaches 77%. And relocating GPU workloads to cost-effective regions delivers 2x-7x savings beyond spot pricing alone. Set up automation for the optimizations your team will never get to manually. Teams that have done this systematically have documented how to cut cloud bills by 40% or more.

3. Tag everything, enforce it in CI

You can't allocate costs you can't attribute. Every resource needs an owner, a team, an environment tag, and an expiry date for non-production resources. Block deployments that don't meet your tagging policy. It sounds strict, but untagged resources are how $135K per $500K in spend disappears without anyone noticing.

4. Set resource request defaults that make sense

Most Kubernetes resource waste comes from copy-pasted Deployment manifests with generous defaults. Create right-sized templates for common workload profiles — a web API doesn't need the same resources as a batch processing job. Review requests and limits quarterly against actual utilization data.

5. Make cost a team metric

Add cost-per-request or cost-per-transaction to your team's service dashboards alongside latency and error rates. When waste is visible and owned, it gets fixed. When it lives in a monthly PDF from finance, it doesn't.

Engineering team collaborating around screens showing performance metrics and cloud data

Frequently asked questions

What percentage of cloud spend is wasted?

Organizations waste approximately 27% of their cloud IaaS and PaaS spend on idle, overprovisioned, or forgotten resources (Flexera 2025). Harness estimates this translates to $44.5 billion in infrastructure waste industry-wide for 2025. The percentage rises for organizations without active FinOps practices.

What is cloud overprovisioning?

Overprovisioning means allocating more compute, memory, or storage than a workload actually needs. CAST AI's 2025 benchmark found average Kubernetes CPU utilization at just 10% — meaning 90% of provisioned CPU goes unused. The root cause is typically "just in case" sizing combined with no review cycle after deployment.

How long does it take to find cloud waste?

Without automated tooling, it takes an average of 31 days to identify cloud waste and 25 days to rightsize overprovisioned resources (Harness 2025). Automated platforms can surface waste within hours and apply rightsizing recommendations continuously.

Does FinOps actually reduce cloud costs?

Yes, but maturity matters. Ninety-six percent of executives call FinOps important, yet only 14.2% operate at an advanced maturity level. Organizations with mature FinOps practices report significantly lower waste rates. The key is moving from reporting (telling teams what they spent) to automation (preventing waste before it occurs).

How much can spot instances save?

Partial spot instance adoption yields an average 59% compute cost reduction, while exclusive spot usage delivers 77% savings (CAST AI 2025). GPU workloads relocated to cost-optimized regions can achieve 2x-7x additional savings. Spot works best for stateless, fault-tolerant workloads with proper orchestration.

Estimate your cloud costs — for free

Compare AWS, Azure, and GCP pricing side by side with our free calculator, and dig into the guides to learn how to cut cloud waste. No sign-up required.