Posted on Mar 13, 2026 · Updated Mar 13, 2026 · 13 min read

The Art of Kubernetes Request Sizing: Stop Wasting 70% of Your Cluster

For three consecutive years, roughly 70% of requested Kubernetes CPU and memory resources have gone completely unused (CAST AI 2025 Kubernetes Cost Benchmark). Average CPU utilization sits at 10%. Memory at 23%. The numbers haven't improved — they've gotten worse. And it's not because teams are lazy. It's because Kubernetes request sizing is genuinely hard to get right.

This guide covers the mechanics of CPU and memory requests and limits, why teams consistently over-provision, and a practical framework for right-sizing that doesn't sacrifice reliability. If you've ever set resources.requests.memory: 2Gi because "it seemed about right," this article is for you. For broader context on where Kubernetes costs come from, start with our guide to understanding Kubernetes costs and our article on the hidden costs of Kubernetes.

TL;DR

68% of Kubernetes pods request 3–8x more memory than they use (Wozz, 2026). Right-sizing saves 30–50% on overprovisioned resources. Kubernetes 1.35 now supports in-place pod resizing without restarts. Set requests to P95 usage, remove CPU limits, and use VPA in recommendation mode before automating.

Cloud computing technology network illustration showing interconnected nodes and data flow

How bad is the Kubernetes resource waste problem?

Average CPU utilization across Kubernetes clusters is 10% — down from 13% the prior year. Memory utilization is 23%. That data comes from CAST AI's 2025 benchmark analyzing 2,100+ organizations across AWS, Azure, and GCP. In plain terms: for every ten CPU cores you're paying for, nine sit idle.

The waste compounds at multiple layers. There's a 40% gap between what clusters provision and what pods request. Then another gap between what pods request and what applications actually use. Stack those together and you get single-digit utilization as the industry norm, not the exception.

The Three-Layer Waste StackHow 100% of provisioned capacity becomes 10% utilizationProvisioned Capacity: 100%Nodes allocated by the cluster autoscaler40% gap (CPU) / 57% gap (memory)Pod Requests: ~60% of provisionedWhat developers ask for in resource specs70% of requests go unused (CAST AI, 3 years running)10%CPU23%MemoryActual UsageSource: CAST AI 2025 Kubernetes Cost Benchmark (2,100+ organizations)
Waste compounds at every layer: cluster autoscaler provisions more than requested, developers request more than needed, and applications use a fraction of what's requested.

According to CAST AI's 2025 Kubernetes Cost Benchmark analyzing 2,100+ organizations, approximately 70% of requested CPU and memory resources are never utilized — a figure that has remained consistent for three consecutive years (CAST AI, 2025 ). This persistent waste pattern suggests the problem is systemic, not behavioral.

The financial impact scales with cluster size. Sysdig estimates organizations running ~150 nodes overspend by roughly $980,000 per year. At 1,000+ nodes, the waste exceeds $10 million annually (Sysdig). And 59% of containers have no CPU limits defined at all — meaning there's not even a ceiling on the waste.

What do requests and limits actually do?

Kubernetes resource requests and limits are the two most consequential numbers in any pod spec, and most teams confuse them. Requests determine scheduling. Limits determine termination. Getting the distinction wrong causes either waste (too high) or instability (too low).

Requests tell the scheduler "this pod needs at least this much CPU and memory to run." The scheduler won't place the pod on a node unless that much capacity is available — even if the pod never actually uses it. This is the number that drives cluster sizing and cost. If you request 2 GiB of memory but only use 200 MiB, you've reserved 1.8 GiB that nothing else can use.

Digital network illustration with cloud computing data connections across a blue circuit board

Limits tell the kubelet "kill this pod if it tries to use more than this." For memory, exceeding the limit triggers an OOMKill. For CPU, exceeding the limit triggers throttling — the container gets fewer CPU cycles. The consequences are different, and the strategy should be different too.

Here's the critical insight most teams miss: CPU requests and limits work fundamentally differently from memory requests and limits. CPU is compressible. If a container exceeds its CPU request, it gets throttled but keeps running. Memory is incompressible. If a container exceeds its memory limit, it gets killed. This asymmetry means the optimal strategy is different for each.

An increasing number of production teams are removing CPU limits entirely and keeping only CPU requests. Why? CPU limits cause throttling even when the node has spare capacity. Your container gets artificially constrained to its limit, even if nobody else needs those cycles. This manifests as latency spikes that are incredibly hard to debug — the application looks healthy in every metric except response time.

Memory limits, on the other hand, are essential. Without them, a memory leak in one container can destabilize an entire node by consuming all available memory and triggering the kernel OOM killer against random pods.

Why do teams consistently over-provision?

A Wozz study of 3,042 production clusters in January 2026 found that 68% of pods request 3–8x more memory than they actually use. And 64% of engineering teams add "just to be safe" headroom of 2–4x after a single OOM incident. Only 12% of teams can state their P95 memory usage without looking it up.

How Common Are Sizing Mistakes?Frequency of resource configuration anti-patterns in production0%33%66%100%Can't state P95 memory88%HPA scales on CPU only80%Pods request 3-8x needed68%"Just in case" 2-4x buffer64%No CPU limits defined59%Unscaled + daily CPU spikes46%Monitoring gapsBehavioral patternsConfig gapsScaling gapsSources: Wozz (2026), Datadog (2025), Sysdig (2023)
88% of teams can't state their P95 memory usage — meaning almost nobody sizes requests based on actual data.

The root cause isn't ignorance. It's asymmetric risk. Under-provisioning causes immediate, visible pain: OOMKills, restarts, pager alerts at 3 AM. Over- provisioning causes diffuse, invisible pain: higher cloud bills that nobody ties to a specific decision. Given the choice, every rational engineer picks "waste money" over "get paged."

The copy-paste problem compounds it. A developer sets generous requests for a new service based on a similar existing service. That service's requests were themselves copied from an older service. Nobody goes back to validate. You end up with entire clusters where every pod inherits the same inflated defaults, regardless of actual workload characteristics.

A January 2026 study by Wozz analyzing 847,293 pods across 3,042 production clusters found that 68% of pods request 3–8x more memory than they actually use, while 64% of engineering teams add 2–4x "safety headroom" after a single OOM incident (Wozz, 2026 ). Only 12% of teams can state their P95 memory usage without checking dashboards.

How to right-size: a practical framework

Right-sizing doesn't mean making requests as small as possible. It means making them accurate. Organizations that implement right-sizing typically achieve 30–50% cost reduction on overprovisioned resources while maintaining application performance (OneUptime, 2026 ). Here's a step-by-step framework that works.

Step 1: Measure actual usage for 7–14 days

Don't size based on load tests or educated guesses. Collect real production metrics. Use Prometheus with container_cpu_usage_seconds_total and container_memory_working_set_bytes over at least 7 days, covering both weekday and weekend traffic patterns. You need the P50, P95, and P99 for both CPU and memory.

Step 2: Set CPU requests to P95 usage

Take the P95 CPU usage value and round up to the nearest 50m increment. This ensures the pod gets scheduled on a node with enough capacity for normal operation. The P99 spike will still be handled — Kubernetes allows bursting above requests as long as the node has spare capacity.

Step 3: Remove CPU limits (in most cases)

CPU is compressible. A container that exceeds its CPU request gets throttled, but it doesn't get killed. CPU limits prevent your container from using available node capacity, causing artificial latency spikes. Unless you have strict multi-tenant isolation requirements, remove CPU limits and let the scheduler manage contention through requests.

Step 4: Set memory requests to P99 usage + 10% buffer

Memory is unforgiving. Exceed the limit and the container dies. Use the P99 value (not P95) and add a 10% buffer. This accounts for garbage collection spikes in JVM languages, memory fragmentation, and the occasional request pattern that exceeds normal usage. The 10% buffer is your OOM insurance — cheap compared to the alternative.

Step 5: Set memory limits to 1.5–2x the request

Memory limits are your safety net against leaks. A limit of 1.5–2x the request gives enough headroom for legitimate spikes while catching runaway memory growth before it crashes the node. For JVM applications with fixed heap sizes, the limit can be tighter. For Node.js or Python with less predictable memory patterns, lean toward 2x.

Analytics dashboard showing performance monitoring charts and resource usage metrics on a laptop

Step 6: Review and adjust quarterly

Traffic patterns change. Code changes. Dependencies update. A resource configuration that was right three months ago may be 2x too generous or dangerously tight today. Build a quarterly review into your sprint cycle. It takes 30 minutes per service and can save thousands per month. Once requests are accurate, the logical next step is Kubernetes cost allocation — attributing those accurate costs to the right teams.

VPA, Goldilocks, and the right-sizing tool landscape

You don't have to do this manually. The Kubernetes ecosystem has mature tools for automated right-sizing. Eighty percent of HPA deployments scale on CPU and memory utilization only — just 20% use custom metrics (Datadog 2025). And 46% of unscaled workloads experience multiple significant CPU spikes per day, meaning they'd benefit from autoscaling but don't have it.

Annual Savings from Right-Sizing by Cluster SizeEstimated 30-50% savings on overprovisioned resourcesBefore right-sizingAfter right-sizing$0$2.5M$5M$10M~50 nodes$327K-$130K~150 nodes$980K-$390K~500 nodes$3.3M-$1.3M1,000+ nodes$10M+-$4MWozz case study: $47.2K/mo → $11.1K/mo (76% savings) from memory right-sizing aloneSources: Sysdig, CAST AI, Wozz (2023-2026)
At 150 nodes, right-sizing saves roughly $390K per year. At 1,000+ nodes, it's a $4M+ annual opportunity.

Vertical Pod Autoscaler (VPA)

VPA observes actual resource usage and recommends (or automatically sets) optimal requests and limits. Start in Off mode (recommendation only) to see what it suggests without any changes. Then graduate to Initial mode (sets resources only at pod creation) and eventually Auto mode (updates running pods). As of Kubernetes 1.35, VPA supports InPlaceOrRecreate mode, which resizes pods without eviction when possible.

Goldilocks

Goldilocks, from Fairwinds, runs VPA in recommendation mode for every deployment in a namespace and presents the results on a dashboard. It's the fastest way to see "here's what every pod should be requesting, based on real usage." No automatic changes. Just data. Install it, wait a week, and review the dashboard.

KEDA (event-driven scaling)

For workloads driven by queues, events, or custom metrics, KEDA scales pod replicas based on external signals rather than CPU and memory. This is critical for the 80% of HPA deployments that only scale on CPU — many workloads have usage patterns that CPU metrics can't capture, like queue depth or request rate.

Karpenter (node-level right-sizing)

Karpenter (originally from AWS, now multi-cloud) replaces the cluster autoscaler with a node provisioner that matches node types to actual workload requirements. Instead of pre-defining node groups, Karpenter selects instance types dynamically based on pending pod resource requests. It's right-sizing at the infrastructure layer.

Kubernetes 1.35: in-place pod resize changes everything

In December 2025, Kubernetes 1.35 graduated In-Place Pod Resize to GA (kubernetes.io). This is the most significant change to resource management since VPA was introduced. Before 1.35, changing a pod's resource requests required restarting it. Now you can resize a running pod without eviction.

Why does this matter? The biggest objection to automated right-sizing has always been "I can't restart production pods during business hours." VPA in Auto mode would evict and recreate pods to apply new resource values. For stateful workloads, long-running connections, or latency-sensitive services, that's unacceptable. In-place resize eliminates this barrier.

Cloud computing illustration showing data storage and network connectivity in the cloud
Right-Sizing Tool ComparisonWhich tool fits which problemToolWhat It SizesAutomationNo Restart?StatusVPACPU/memory requestsRecommend → AutoWith 1.35StableGoldilocksCPU/memory (via VPA)Recommend onlyN/A (read-only)StableKEDAPod replicas (events)Fully automaticAdds/removes podsCNCF GradKarpenterNode instance typesFully automaticNode-level onlyStableIn-Place ResizeRunning pod resourcesManual / VPA-drivenYes (GA)K8s 1.35Recommended stack: Goldilocks (discover) → VPA + In-Place Resize (automate) → Karpenter (nodes)Add KEDA for event-driven workloadsSources: kubernetes.io, CNCF, Fairwinds (2025-2026)
Start with Goldilocks for visibility, then graduate to VPA with in-place resize for automated right-sizing.

In-Place Pod Resize graduated to GA in Kubernetes 1.35 (December 2025), enabling VPA to resize running pods without eviction for the first time (kubernetes.io). This removes the primary barrier to automated right-sizing in production: the requirement to restart pods when adjusting resource requests and limits.

The practical implication: you can now run VPA in InPlaceOrRecreate mode. It'll try to resize the pod in-place first. If the node can't accommodate the new size, it falls back to eviction and recreation. This hybrid approach gives you automated sizing with minimal disruption — exactly what production teams have been waiting for.

Frequently asked questions

What should Kubernetes CPU requests be set to?

Set CPU requests to the P95 actual usage value over 7–14 days of production traffic. Round up to the nearest 50m increment. This ensures scheduling accuracy while allowing bursting for short spikes. CAST AI's 2025 benchmark shows average CPU utilization at just 10% of allocated resources, indicating most teams over-request by 5–10x.

Should I set CPU limits on Kubernetes pods?

In most cases, no. CPU limits cause throttling even when the node has spare capacity, resulting in unexpected latency spikes. Sysdig found that 59% of containers have no CPU limits defined. Removing CPU limits and relying solely on requests for scheduling is an increasingly common practice for non-multi-tenant clusters.

How much memory buffer should I add to Kubernetes requests?

Add 10% buffer above your P99 memory usage. Memory is incompressible — exceeding the limit kills the pod. The Wozz 2026 study found that 68% of pods request 3–8x more memory than needed, so even with a 10% buffer, you'll likely reduce requests significantly. Set limits to 1.5–2x the request as a safety net.

What is Kubernetes in-place pod resize?

In-place pod resize, GA in Kubernetes 1.35 (December 2025), allows changing a pod's CPU and memory resources without restarting it. Previously, changing resource requests required pod eviction and recreation. This enables VPA's InPlaceOrRecreate mode for automated right-sizing with minimal production disruption.

How much can Kubernetes right-sizing save?

Right-sizing typically saves 30–50% on overprovisioned resources. At 150 nodes, that's roughly $294K–$490K per year. At 1,000+ nodes, savings exceed $3–5 million annually (Sysdig, CAST AI). A Wozz case study showed a 76% reduction — from $47,200/month to $11,100/month — through memory right-sizing alone.

Estimate your cloud costs — for free

Compare AWS, Azure, and GCP pricing side by side with our free calculator, and dig into the guides to learn how to cut cloud waste. No sign-up required.