Tuning Linux swap parameters for stable Kubernetes nodes
A deep dive into kernel parameters like swappiness and watermarks reveals how to safely enable swap in Kubernetes v1.34 without triggering OOM kills.
In a post on the Kubernetes Blog in August 2025, Ajay Sundar Karuppasamy from Google detailed how to tune Linux swap for Kubernetes clusters. The article addresses the upcoming stabilization of the NodeSwap feature in Kubernetes v1.34, which allows nodes to use disk space as virtual memory. This shift requires careful configuration of kernel parameters to balance resource utilization with system stability.
What happened
For years, the standard advice for Kubernetes operators was to disable swap entirely to ensure predictable performance. However, the introduction of the NodeSwap feature changes this paradigm by allowing Linux nodes to offload less frequently used memory pages to secondary storage. This capability aims to reduce out-of-memory (OOM) kills and improve overall resource efficiency. Despite these benefits, enabling swap is not a simple toggle. Without precise tuning, it can lead to severe performance degradation and interfere with the Kubelet’s ability to manage pod evictions gracefully.
The article highlights that misconfigured swap settings can cause the kernel to behave unpredictably under memory pressure. In tests using default settings, nodes experienced unexpected restarts and premature OOM kills when subjected to high memory allocation rates. The core issue lies in the interaction between the Linux kernel’s page replacement algorithms and Kubernetes’ eviction logic. If the kernel does not reclaim memory quickly enough, or if it reclaims the wrong type of memory, the node can become unstable before the Kubelet has a chance to act.
To address these challenges, the author conducted a series of stress tests on Google Kubernetes Engine (GKE) nodes running Kubernetes v1.33.2. The tests involved custom Go applications designed to simulate various memory access patterns and pressures. By adjusting specific kernel parameters, the author demonstrated how to create a safer operational window for memory management, preventing critical failures during sudden spikes in demand.
How it works
Linux manages memory in pages, typically 4KiB each. When physical RAM is full, the kernel must decide which pages to move to swap space. It distinguishes between anonymous memory, such as heap and stack data, and file-backed memory, like executable code and caches. Anonymous pages must be written to a swap device to be reclaimed, while clean file-backed pages can simply be discarded. The kernel uses several parameters to guide these decisions, primarily vm.swappiness, which controls the preference for swapping anonymous pages versus dropping file cache.

Two other critical parameters are vm.min_free_kbytes and vm.watermark_scale_factor. The min_free_kbytes setting defines a safety buffer of free memory. When available memory drops below this threshold, the kernel aggressively reclaims pages. The watermark_scale_factor determines the gap between the low, min, and high memory watermarks. A larger gap gives the background reclaim process, known as kswapd, more time to move pages to swap gradually. This prevents the system from hitting the critical min watermark too quickly, which would otherwise block process allocations and trigger OOM kills.
Key details
- The NodeSwap feature is expected to graduate to stable status in Kubernetes v1.34.
- Tests were conducted on GKE nodes with 8GiB RAM and 50GB swap on pd-balanced disks.
- Default settings (
swappiness=60,min_free_kbytes=68MB) led to OOM kills under high load. - Increasing
min_free_kbytesto 512MiB forces earlier memory reclamation, providing a larger safety buffer. - Setting
watermark_scale_factorto 2000 widens the swapping window, allowingkswapdto work more effectively. - Critical system components like kubelet and container runtime should have swap disabled via cgroups.
Why it matters
For software engineers and platform teams, enabling swap introduces a complex trade-off between capacity and latency. Swapping is significantly slower than accessing RAM, so if an application’s active working set is moved to disk, performance will suffer due to increased I/O wait times. However, properly tuned swap can prevent abrupt application crashes caused by OOM kills, which are often more disruptive than temporary latency spikes. Understanding these kernel mechanics is essential for building resilient systems that can handle memory overcommitment safely.

Furthermore, improper tuning can mask underlying issues such as memory leaks. Instead of failing fast with an OOM kill, a leaky application might slowly degrade node performance by consuming swap space, making diagnosis difficult. It also risks bypassing Kubernetes’ graceful eviction mechanisms. If the kernel triggers an OOM kill before the Kubelet can evict pods based on policy, higher-priority workloads may be terminated unexpectedly. Therefore, aligning kernel behavior with Kubernetes’ expectations is crucial for maintaining cluster health.
What you can do
- Start with
vm.swappiness=60for general-purpose workloads, but adjust based on whether your apps are I/O-sensitive or cache-heavy. - Set
vm.min_free_kbytesto approximately 2-3% of total node memory (e.g., 500MB for an 8GiB node) to create a sufficient safety buffer. - Increase
vm.watermark_scale_factorto 2000 to expand the window for background memory reclamation and prevent sudden OOM events. - Protect critical system daemons by setting
memory.swap.max=0in their cgroups to ensure they remain responsive under pressure. - Benchmark these settings in a test environment with your specific workloads, as performance varies with disk type and CPU load.
- Monitor I/O wait times and swap usage closely to detect thrashing or hidden memory leaks early.


