Kubernetes 1.32 introduces API streaming to fix list request memory spikes
Kubernetes 1.32 graduates watch lists to beta, allowing clients to stream large resource collections and prevent API server out-of-memory crashes.
In a post on the Kubernetes Blog in December 2024, engineers from Upbound, Google, and Red Hat detailed a critical memory efficiency improvement for the Kubernetes API server. The update addresses how large clusters handle bulk data retrieval, introducing a streaming mechanism that prevents sudden memory exhaustion during heavy load.
What happened
Managing large Kubernetes clusters often leads to significant memory overhead when components issue list requests. In the traditional implementation, the kube-apiserver must assemble the entire response in memory before sending any data to the client. If the response body is hundreds of megabytes, or if multiple requests arrive simultaneously after a network outage, this process can quickly consume all available RAM. While existing API Priority and Fairness mechanisms protect against CPU overload, they offer limited protection against these unbounded memory spikes.
The investigation revealed that this memory allocation happens because the server fetches data from the database, deserializes it, and then constructs the final response format all at once. This sequence creates a massive temporary memory footprint that neither Go’s garbage collection nor configured memory limits can effectively manage during sudden spikes. In high-availability setups, this can lead to a cascading failure where one API server crashes due to an out-of-memory (OOM) condition, shifting the same heavy requests to other nodes and causing them to fail as well.
To solve this, the Kubernetes team graduated the watch list feature to beta in version 1.32. This allows clients to opt-in to streaming lists by switching from standard list requests to a specialized form of watch requests. By serving these requests from the watch cache, the server streams each item individually rather than buffering the whole collection. This change ensures that memory overhead remains constant, bounded only by the maximum size of a single object plus minor allocations, drastically improving stability for clusters with many large objects.
How it works
The core mechanism relies on shifting from batch processing to streaming. Traditional list requests require the server to hold the entire dataset in RAM during serialization. In contrast, the new watch list approach leverages the existing watch cache, which is an in-memory cache designed to scale read operations. When a client uses this method, the API server sends objects one by one as they are retrieved from the cache, avoiding the need to allocate memory for the full response body at once.

This architectural shift decouples memory usage from the total number of objects in a collection. Instead of memory consumption growing linearly with the size of the list, it stays flat regardless of how many items are being returned. This makes the API server resilient even when handling thousands of large resources, such as Secrets with substantial payloads, without risking OOM kills.
Key details
- The watch list feature reached beta status in Kubernetes 1.32.
- Clients must explicitly enable the
WatchListClientfeature gate in client-go to use streaming lists. - Synthetic tests showed memory usage stabilized at 2 GB with streaming enabled, compared to 20 GB with it disabled.
- The feature requires etcd versions 3.4.31+ or 3.5.13+.
- In Kubernetes 1.33, new feature gates
StreamingCollectionEncodingToJSONandStreamingCollectionEncodingToProtobufwere introduced for server-side streaming without client changes. - The
WatchListfeature gate is disabled by default in Kubernetes 1.33, though it was enabled by default for kube-controller-manager in 1.32.
Why it matters
For platform engineers and SREs managing large-scale clusters, this update addresses a fragile point in control plane stability. Memory exhaustion in the API server is difficult to diagnose and recover from, often resulting in complete control plane outages. By adopting streaming lists, teams can run larger clusters with more complex resources without fearing that a routine reconciliation loop or a post-outage sync will crash their management layer.
This change also has implications for how API costs are calculated in future versions. Currently, API Priority and Fairness assigns a low cost to list requests to maintain parallelism for typical use cases. As the ecosystem shifts toward watch lists, the system can safely increase the cost estimation for traditional list requests. This will provide better protection against legacy clients or misconfigured tools that continue to issue expensive bulk requests, ensuring fairer resource distribution across the cluster.
What you can do
- Upgrade your cluster to Kubernetes 1.32 or later to access the beta watch list feature.
- Verify that your etcd version is at least 3.4.31 or 3.5.13 to ensure compatibility.
- Update Golang-based clients to enable the
WatchListClientfeature gate in client-go. - Monitor memory usage on your API servers during peak load to identify if large list requests are still causing spikes.
- Encourage third-party controllers and operators in your environment to adopt the streaming API during the beta phase.
- Plan for future upgrades to Kubernetes 1.33 to leverage server-side streaming encoding features that require no client code changes.


