Saturday, October 3, 2026 · Week 40 DE · EN · FR · ES Dark
Tech & Gadgets

High Kubernetes Requests Pay for Idle Nodes

Placement and scaling run on reserved CPU and reserved memory, not on measured consumption.

By Alec Chizhik September 24, 2026 8 min read
High Kubernetes Requests Pay for Idle Nodes

The cloud bill of a Kubernetes cluster follows the reserved resources, because schedulers and node autoscalers operate on requests. Limits, namespace budgets and autoscaling only take effect when they are tied to this reservation. Anyone who keeps requests permanently high pays for provisioned capacity.

Key takeaways

  • Requests determine the node count. Schedulers and node autoscalers operate on reservations, and the cloud bill follows the running instances.
  • Limits protect neighbors at runtime. They throttle CPU or trigger OOMKills, but on their own they never lower the reservation, nor the fleet that depends on it.
  • Quotas must cap requests. ResourceQuotas and LimitRanges only have a cost impact when they limit the reservation and the defaults.
  • Three metrics on a weekly cadence. Request utilization, unused request capacity and limit events reveal waste and collateral damage.

Related:Observability Quietly Eats the Cloud Budget  /  EKS 1.36 Gets Expensive When FinOps Discipline Is Missing

What are Kubernetes requests? Kubernetes treats CPU and memory requests as a reservation for placement. The scheduler only assigns the pod to a node that still has this amount free. The Cluster Autoscaler counts the same figure when deciding whether a new node is needed. Anyone who keeps requests permanently high pays for provisioned capacity. Limits cap consumption at runtime.

Why Requests Determine the Node Count

Kubernetes treats CPU and memory requests as a reservation for placement. The scheduler adds up the requests of the pods already running on a node and only accepts a new pod if the remaining request capacity is sufficient. Measured usage plays no role in this step. A service with a light load and a large request blocks the same capacity as a service that actually fills its reservation. System services and DaemonSets are subtracted from the same node total before application pods find room.

Node autoscalers watch pods stuck in Pending because not enough request capacity is free anywhere. An extra node then appears, and with it another instance on the cloud bill. When requests grow across many namespaces and replicas, the fleet grows even with flat traffic. Billing at the major providers is tied to running instances, to Persistent Volumes and to the network. After a load test, such nodes stick around unless someone pulls back requests or the replica count.

Requests thus determine how much capacity the platform actually provisions. Set them permanently high and you pay for provisioned capacity. Set them too tight and you get Pending, eviction under pressure and scale-ups that can linger after the peak. A platform that only comments on bills without touching specs is not in control. Fragmentation amplifies this, because a few large pods leave leftover capacity too small for other requests.

Many internal chargeback models allocate node costs according to the requests. That is consistent, because these are exactly the fields that drive packing and often the node count. Without reliable requests, every budget conversation becomes guesswork. The team then argues over shares while the fleet keeps growing unnoticed. This yields the first checkpoint: does the spec state what really binds the nodes?

Limits Cap the Damage at Runtime

Limits cap CPU and memory as soon as a container crosses the boundary. On CPU this produces throttling, which shows up in the tail latency of incoming requests. On memory, the outlier typically ends in an OOMKill. Both protect neighboring workloads on the same node and contain the local blast radius. Neither changes the bill as long as the requests stay untouched and the nodes keep running.

The quality-of-service class follows from request and limit. If both are set and equal, the pod is Guaranteed. If the request is smaller than the limit, it is Burstable. If both are missing, the pod lands in BestEffort and is evicted first when resources run short. Teams often set request equal to limit to force Guaranteed, which raises the request along with the limit and cuts packing density.

A limit without a sound request reverses the problem. The scheduler packs tightly, while at runtime the container may run far beyond its reservation. Under load, CPU steal and memory pressure hit the neighbors. Costs and stability then both deteriorate, because the node ends up fuller than the packing logic assumed. Limits are a brake on operations.

Rightsizing Needs Usage Plus Headroom

Rightsizing sets requests to the observed load plus a deliberately chosen headroom. That takes usage time series spanning days and weeks, because a single peak from a deploy window is no yardstick. Measure only midday on weekdays and you underestimate overnight batch jobs. Measure only the deploy window and you overestimate steady-state operation. CPU and memory follow different failure patterns, because CPU can be throttled and memory cannot.

Vertical autoscalers can suggest or set requests. That shortens the manual loop and shifts the responsibility to observation windows, lower bounds and stability. Windows that are too short create oscillation between size classes. Every change to a request changes the packing and can remove or add nodes. Rightsizing without an eye on the node autoscaler stays incomplete, because the savings only materialize once instances actually disappear.

Headroom is part of the design. Batch jobs, heaps, caches and startup spikes need slack, and that slack has to be written into the requests. The rule should be stated explicitly, for example relative to a high usage percentile, and reviewed periodically. Otherwise the reserve ossifies into unjustified overprovisioning. A reserve nobody can explain anymore is, in practice, a second hidden budget.

No single formula fits every service class. Stateless frontends, queue workers and memory-heavy runtimes have different profiles. The review should address the class and its load shape. A global cut across all namespaces systematically produces throttling in one place and idle nodes in another.

Namespace Budgets Stop Silent Capacity Growth

Without a cap, the request total grows with every deployment, every sidecar and every replica. ResourceQuotas on CPU and memory requests limit what a namespace may reserve in total. LimitRanges set defaults, minima and maxima per container. That leaves fewer pods without a request. Missing defaults are a common source of silent reservation, because a supposedly small service gets retrofitted later.

A budget that caps only limits lets the reservation continue. A budget only on the pod count ignores bloated specs. The quota should target what binds nodes: CPU requests, memory requests, PersistentVolumeClaims and, where needed, the number of workloads. How a provider later mirrors this in the bill varies.

Quotas without a process get bypassed. New namespaces, a second cluster and requests in shared system namespaces undercut the cap. Increases belong in the same review as a capacity expansion. A team asking for more reservation has to show usage and limit events. Otherwise the platform is funding unsubstantiated hope.

Separate budgets for production and non-production keep load tests from eating up the production reservation. A shared cluster without separate quotas turns exactly that into the routine surprise. The platform should make visible how much of a quota is already bound by requests. Without that visibility, a rejected deploy comes as a surprise and the team starts hunting for a way around the cap.

Autoscaling Amplifies Every Error in the Spec

The Horizontal Pod Autoscaler changes the replica count based on CPU, memory or custom metrics. Relative CPU utilization depends on the request and therefore on the chosen reservation. Every replica adds its requests in full to the packing. A request set too high produces expensive copies and earlier demand for nodes. A request set too low makes utilization look high, so the autoscaler scales early, while the scheduler still packs those undersized pods poorly.

Vertical and horizontal scaling without clear ownership creates oscillation. The vertical path changes the size, the horizontal one the count. The node autoscaler responds by adding instances. During the transition, old and new sizes often run in parallel. That is exactly where costs and Pending queues climb together, often along with throttling, because the new size is not yet stable.

Scaling bounds must match the namespace quotas. A high maximum with a tight quota creates Pending instead of capacity. A wide quota without a maximum creates node growth. The chain stays fixed: first request size, then replica count, then the namespace cap, and last the node count. Skip a step and the effect only shows up in the billing.

Stabilization windows belong in the same review. Windows that are too short create flapping and constant repacking. Windows that are too long keep nodes after the peak that nobody needs anymore. Autoscaling multiplies the spec already written in the file.

Which Three Metrics Count Every Week

For a weekly cross-section of the cost picture, a lean set that ties operations to the bill is enough. First, request utilization: measured CPU and memory usage relative to the request, split by workload and namespace, aggregated over the week. If it stays permanently very low, the reservation is bloated. If it sits permanently near the full reservation, the reserve is missing or the request is too tight. CPU and memory have to be read separately, because they call for different corrections.

Second, unused request capacity in the cluster. This is the delta between reserved capacity, packable capacity and measured load. It shows up as underutilized nodes, as a low packing ratio or as namespaces with a high reservation and low usage. This figure explains why the bill does not fall with traffic. It gives the app team and the platform a shared number.

Third, limit events: CPU throttling time and OOMKills per workload. If they rise after a rightsizing pass, the reserve was too aggressive. If they stay high while request utilization is low, request and limit are set incorrectly relative to each other. Cost control that only cuts the bill and ignores these events moves the damage into the incident. Throttling shows up as latency in the user path. Three metrics, one fixed weekly cadence.

Whoever May Change the Spec Controls the Bill

The mechanics only work when it is clear who may change requests, limits and quotas. App teams need leeway in their namespace. The platform needs a veto as soon as the request total forces new nodes. A self-service portal without reference to a quota turns every scaling action into an order charged to the platform.

Changes to requests belong in the same visible path as other production-affecting specs. A change that doubles CPU requests is a capacity decision. It should include usage series and the three weekly metrics. Otherwise a default copied from a tutorial stays in production for months, holding capacity nobody can justify anymore.

The bill keeps following the instances. Teams that check requests, limits and namespace caps weekly against the same three metrics keep node count and blast radius in one shared review. Those who push this to a ticket after the billing month arrive too late.

Frequently Asked Questions

Do Limits Lower Our Cloud Bill Directly?

Limits throttle CPU or kill memory hogs and protect neighbors on the node. The bill follows the instances that the scheduler and node autoscalers derive from the requests. A limit without a matching request can even worsen packing, because more is drawn at runtime than was ever reserved.

Which Three Metrics Should We Read Every Week?

Request utilization relative to the reservation, unused request capacity in the cluster and limit events such as throttling time and OOMKills. Thresholds for good or bad are not universal and depend on the service class and the accepted risk. The value lies in the fixed rhythm and in tying deviations back to spec, quota and autoscaler.

Is a ResourceQuota on the Namespace Enough as a Cost Cap?

Only if it caps CPU and memory requests and LimitRanges set the defaults. A quota on limits alone or on the pod count alone lets the reservation keep growing.

Image source: AI-generated (September 2026)

Translated from the German original using artificial intelligence. The German version is authoritative.

Also available in

FrançaisEspañolDeutsch
MBF Media Newsletter

The monthly briefing for decision-makers

Once a month, the MBF Media Newsletter gathers what matters from cloudmagazin, MyBusinessFuture, Digital Chiefs and SecurityToday, curated by the editorial team.

Around 23,500 IT and business decision-makers read this newsletter. Read along.

Subscribe for free
MBF Media Newsletter, aktuelle Ausgabe auf dem iPhone
A magazine by Evernine Media GmbH