Kubernetes v1.33 introduces streaming list responses to cut API server memory use
Kubernetes v1.33 adds streaming encoding for List responses, reducing kube-apiserver memory usage by up to 20x during large dataset fetches and improving cluster stability.
In a post on the Kubernetes Blog in May 2025, maintainers Marek Siarkowicz and Wei Fu described a major architectural change in Kubernetes v1.33. The update introduces streaming encoding for List responses, a feature designed to drastically reduce memory consumption in the API server when handling large datasets. This improvement addresses long-standing stability issues in large-scale clusters where fetching extensive resource lists previously risked causing out-of-memory errors.
What happened
Operating large Kubernetes clusters has always involved managing the trade-off between data accessibility and resource consumption. One of the most persistent challenges has been handling List requests, which fetch substantial datasets from the cluster state. In previous versions, the API server would serialize an entire response into a single contiguous block of memory before sending it to the client. Even though HTTP/2 can split responses into smaller frames for transmission, the underlying server held the complete data buffer in memory until the entire transfer finished. This meant that if network congestion slowed down the transfer, hundreds of megabytes of memory remained locked for seconds or minutes.
This inefficiency became critical at scale. When multiple large List requests occurred simultaneously, the cumulative memory usage could spike rapidly, leading to Out-of-Memory (OOM) situations that compromised cluster stability. The issue was compounded by how Go’s encoding/json package manages memory. It uses sync.Pool to reuse buffers, which is efficient for steady workloads but problematic for sporadic large responses. Once a large response expanded the memory pool, those oversized buffers remained reserved for subsequent small requests, preventing garbage collection and keeping memory usage artificially high long after the heavy load had passed.
How it works
The new streaming encoder changes how the API server processes List responses by focusing on the Items field, which contains the bulk of the data in collection structures. Instead of encoding the entire array as one monolithic block, the server now processes and transmits each item individually. As each chunk is sent to the client, the memory associated with it is freed immediately. This incremental approach ensures that the API server’s memory footprint remains predictable and manageable, regardless of the total number of objects in the list.

Because Kubernetes objects are typically limited to 1.5 MiB in etcd, streaming keeps individual memory allocations small. The system maintains strict backward compatibility by validating Go struct tags before activation, ensuring the output is byte-for-byte identical to the original encoder. Standard encoding still handles all fields except Items, so clients do not need to change their code or even be aware that the underlying mechanism has shifted. This seamless integration supports all Kubernetes List types, including built-in lists and Custom Resource UnstructuredLists.
Key details
- Version: The feature was introduced in Kubernetes v1.33, announced in May 2025.
- Memory reduction: Benchmarks showed a 20x improvement in memory usage for large list operations, dropping from 70-80GB to just 3GB.
- Mechanism: The encoder streams individual items within the Items field rather than serializing the whole array at once.
- Compatibility: No client-side changes are required; the output remains byte-for-byte consistent with previous versions.
- Trigger: The streaming encoder activates only after rigorous validation of struct tags to ensure safety.
- Scope: It applies to all Kubernetes List types, including standard resources and custom resources.
Why it matters
For engineers building and maintaining large-scale Kubernetes infrastructure, this update directly impacts reliability and cost efficiency. High memory usage in the kube-apiserver often forces teams to over-provision hardware to handle peak loads, increasing operational costs. By reducing the memory footprint of large List requests by such a significant margin, organizations can run leaner control planes without sacrificing performance. It also reduces the risk of unexpected OOM kills, which can cause service disruptions and complicate debugging efforts during incident response.
Furthermore, this change improves the predictability of cluster behavior under load. In environments where monitoring tools, controllers, or CI/CD pipelines frequently list large numbers of resources, the previous memory spikes could create cascading failures. With streaming encoding, these operations no longer hold onto excessive memory during transmission delays. This allows the API server to handle more concurrent requests and larger datasets smoothly, making it easier to scale clusters to support thousands of nodes and tens of thousands of pods.
What you can do
- Upgrade your control plane to Kubernetes v1.33 or later to benefit from the streaming encoder.
- Monitor kube-apiserver memory usage during large List operations to observe the reduction in peak consumption.
- Review any custom controllers or operators that perform large List calls to ensure they handle streaming responses gracefully, though no code changes are strictly required.
- Check your cluster’s horizontal pod autoscaler settings, as reduced API server load may allow for tighter scaling thresholds.
- Validate that your monitoring stack does not rely on fixed buffer sizes for API response parsing, although the byte-for-byte compatibility should prevent issues.
- Consider adjusting resource requests and limits for the kube-apiserver if you previously over-provisioned memory to mitigate OOM risks.


