Kubernetes resource limits: balancing predictability and efficiency
A 2023 analysis argues that setting CPU and memory limits in Kubernetes improves performance predictability, even if it reduces raw cluster efficiency.
Dieser Artikel ist nur auf Englisch verfügbar.
In a post on the Kubernetes Blog in November 2023, Milan Plžík from Grafana Labs presented a detailed argument for using resource limits in container orchestration. While many engineers advocate for removing CPU limits to boost speed, this perspective highlights how limits provide essential stability and predictability for production systems.
What happened
The technical community has seen a surge in advice suggesting that Kubernetes users should stop setting CPU limits on their pods. Proponents of this view argue that limits artificially throttle performance and waste paid compute power that could otherwise be utilized during idle periods. Articles such as "For the Love of God, Stop Using CPU Limits on Kubernetes" have popularized the idea that removing these constraints allows services to run faster by borrowing unused cycles from neighboring workloads.
Plžík, a Site Reliability Engineer at Grafana Labs, challenged this prevailing wisdom by focusing on the operational risks of unlimited resource consumption. He noted that while removing limits might improve immediate performance metrics, it introduces significant unpredictability. When pods compete for shared node resources without defined boundaries, their behavior becomes highly dependent on the specific mix of other applications running on the same machine. This variability makes it difficult to guarantee consistent service levels, especially during traffic spikes or infrastructure changes.
The core of the argument is that the hidden cost of extra resources is a lack of observability and control. Without limits, it becomes nearly impossible to determine exactly how much capacity a pod had available at any given moment. This uncertainty complicates capacity planning for high-traffic events like Black Friday, where historical data may not reflect future constraints if the underlying bin-packing of pods changes. Consequently, what appears to be efficient resource usage can quickly turn into cascading failures when spare capacity disappears.
How it works
Kubernetes manages resources through requests and limits. A request specifies the minimum amount of CPU or memory a pod needs to run, while a limit defines the maximum amount it can use. When limits are removed or set very high, pods operate in a best-effort mode regarding upper bounds. They can consume any free resources on the node, but this creates a "special snowflake" scenario where each pod’s performance depends entirely on its neighbors’ current load.
If a pod exceeds the physical resources of its node, it faces throttling or out-of-memory (OOM) kills, similar to hitting a configured limit. However, without explicit limits, these events are harder to predict and debug. The system lacks clear signals about when a workload is approaching its breaking point because it is constantly absorbing variable amounts of spare capacity. This makes profiling and performance tuning difficult, as data samples may not capture rare but critical spikes in resource usage.
To restore predictability, Plžík suggests two primary configuration strategies. The first is "fixed-fraction headroom," where limits are set to a small percentage above the requests. This allows some burstability while bounding the overcommit ratio on each node. The second strategy is setting requests equal to limits. This places the pod in the Guaranteed Quality of Service (QoS) class, ensuring it receives dedicated resources and is only evicted after lower-priority pods. Both approaches sacrifice some potential efficiency to gain stable, reproducible performance characteristics.
Key details
- Pods without limits consume extra node resources, making their performance dependent on the unpredictable load of neighboring pods.
- Observability suffers because it is difficult to track exactly how much spare capacity a pod used at any specific moment without extensive data mining.
- Historical performance data from unlimited pods may be misleading for capacity planning if cluster bin-packing changes during high-traffic events.
- Setting requests equal to limits assigns the Guaranteed QoS class, which protects pods from eviction until BestEffort and Burstable pods are removed.
- Fixed-fraction headroom allows limited bursting while keeping per-node overcommit within a known bound, reducing performance variance.
- Removing limits eliminates incentives for product teams to optimize their code, as they rely on free spare resources rather than efficient design.
Why it matters
For software engineers and platform teams, the decision to use limits is a trade-off between raw efficiency and operational reliability. While removing limits can squeeze more value out of hardware by utilizing idle cycles, it transfers the risk of resource contention to the application layer. This can lead to sudden latency spikes or OOM kills when the cluster is under pressure, requiring urgent and reactive capacity increases. In contrast, setting limits provides a safety net that forces workloads to behave within known parameters, making incidents easier to diagnose and prevent.
This approach also impacts the economic model of cloud infrastructure. When teams have access to unlimited spare resources, they may neglect optimization efforts, leading to bloated services that perform poorly under constrained conditions. By enforcing limits, organizations encourage developers to right-size their applications and handle resource scarcity gracefully. This discipline helps maintain service level agreements (SLAs) and ensures that performance remains consistent regardless of the surrounding cluster state.
Furthermore, predictable resource usage simplifies the work of site reliability engineers. Instead of debugging complex interactions between competing pods, they can rely on defined boundaries to isolate issues. This clarity is crucial during major upgrades or scaling events, where unexpected resource contention can cause widespread disruptions. Ultimately, the goal is not just to save money on compute costs, but to build a resilient system that behaves consistently under stress.
What you can do
- Audit your current Kubernetes workloads to identify pods running without CPU or memory limits.
- Implement fixed-fraction headroom by setting limits slightly higher than requests for services that need occasional burst capacity.
- Configure critical, latency-sensitive services with requests equal to limits to ensure Guaranteed QoS and priority during resource pressure.
- Review historical monitoring data to detect if performance improvements are due to genuine optimization or simply access to spare node resources.
- Establish guidelines for product teams to right-size resource requests based on actual usage patterns rather than worst-case assumptions.
- Use automated tools like Karpenter or node auto-provisioning to improve bin-packing efficiency when using strict request-limit configurations.


