Kubernetes v1.34 enables node swap for higher AI workload density
Kubernetes v1.34 reaches General Availability for node swap, allowing clusters to use fast NVMe SSDs to page out idle memory and significantly increase pod density for AI workloads.
Kubernetes support for running nodes with swap enabled has reached General Availability in version 1.34. This update allows cluster operators to use fast local storage to handle memory overflow, addressing a critical bottleneck for modern agentic AI workloads that often sit idle while consuming large amounts of RAM.
What happened
Memory capacity is frequently the first hard limit encountered in Kubernetes clusters. Nodes often exhaust their physical RAM long before CPU resources are fully utilized. This constraint has become more acute with the rise of agentic AI workloads, which require substantial memory footprints to initialize and execute untrusted code. After this initial burst of activity, these agents often enter long periods of idleness while waiting for user prompts. Keeping this dormant state resident in expensive physical RAM limits the number of pods a single node can host, driving up infrastructure costs.
The release of Kubernetes v1.34 changes this dynamic by officially supporting node swap. By enabling the Linux kernel to page out anonymous memory to disk, swap acts as a buffer during traffic spikes or periods of heavy memory oversubscription. When backed by high-speed NVMe solid state drives (SSDs), this approach allows nodes to offload dormant memory pages and pack significantly more pods onto each machine. Benchmarks indicate that this method can yield density gains of up to three times, often with minimal impact on latency.
Historically, swap was discouraged in Kubernetes environments for two primary reasons. First, memory accounting under cgroup v1 treated memory and swap as a combined limit, making it difficult to isolate and predict a container's real memory usage. Second, paging to traditional spinning disks introduced severe latency penalties. The new support relies on cgroup v2, which provides separate swap accounting, and pairs with fast NVMe storage to mitigate latency issues, making the feature practical for production use.
How it works
Node swap functions by allowing the operating system to move inactive memory pages from RAM to a designated swap space on disk. In Kubernetes v1.34, this is managed through the kubelet configuration. Operators set failSwapOn to false and define the swapBehavior as LimitedSwap. This configuration tells the node to use swap only when necessary, rather than relying on it as primary memory.
The effectiveness of this mechanism depends heavily on the speed of the underlying storage. By routing swap to Local SSDs, the I/O wait times associated with paging are drastically reduced. This setup works best with Burstable Quality of Service (QoS) classes, where container memory limits are set higher than requests. The node then automatically ration swap space based on idle application memory usage, keeping active processes in fast physical RAM while moving dormant data to disk.
Key details
- Kubernetes v1.34 marks the General Availability of node swap support.
- The feature requires cgroup v2 for independent swap accounting and isolation.
- Benchmarks show a 50% reduction in RAM footprint for Linux kernel builds, dropping from 600 MB to 300 MB.
- Headless Chrome pods using gVisor saw a 100% density increase, rising from 80 to 160 concurrent pods.
- Isolated Python sandboxes achieved a 200% density improvement, scaling from 80 to 240 concurrent sessions.
- Latency increases at peak density are primarily driven by CPU competition rather than swap I/O.
Why it matters
For engineers building platforms for AI agents or CI/CD pipelines, this development offers a direct path to reducing infrastructure costs. Agentic workloads are inherently bursty; they spike during initialization and code execution, then remain idle. Without swap, operators must provision enough RAM to handle the peak, leaving most of that memory wasted during idle periods. Node swap allows teams to right-size memory requests for active usage while using disk space as an insurance policy for bursts.
This also impacts security architectures. Secure execution environments like gVisor or Kata Containers add memory overhead due to isolation requirements. Previously, this overhead strictly limited pod density. With node swap, the additional memory cost of these secure runtimes can be paged out when not actively used. This means teams can maintain strict security boundaries for untrusted code without sacrificing the economic benefits of high-density scheduling.
What you can do
- Upgrade your Kubernetes clusters to version 1.34 or later to access GA node swap features.
- Configure kubelet with
failSwapOn: falseandmemorySwap.swapBehavior: LimitedSwap. - Ensure your nodes are equipped with fast NVMe Local SSDs to minimize swap latency.
- Set container memory limits higher than requests to enable Burstable QoS behavior.
- Benchmark your specific workloads to determine the optimal ratio of RAM to swap space.
- Monitor CPU contention at high densities, as this becomes the primary bottleneck before swap I/O.



