How controller-runtime cache prevents API server overload
A deep dive into the list-watch pattern and local caching in Kubernetes controllers, explaining why reads are cheap but consistency is eventual.
In a post on the Kubernetes Blog in July 2026, engineers detailed the internal mechanics of the controller-runtime cache. The article clarifies how Go-based controllers interact with the Kubernetes API server, correcting common misconceptions about data consistency and memory usage in production environments.
What happened
The publication addresses a widespread misunderstanding among developers building Kubernetes operators: that every read operation inside a reconciler triggers a direct HTTP request to the API server. In reality, controller-runtime relies on a local in-memory cache populated via a list plus watch pattern. This architectural choice ensures that controllers can handle hundreds of reconciliations per second without overwhelming the control plane or etcd.
However, this efficiency comes with specific trade-offs. Because controllers read from a local copy rather than the live source of truth, they may encounter stale data immediately after a write operation. The article highlights that while writes go directly to the API server, reads are served from memory, meaning the system prioritizes availability and low latency over strong immediate consistency. Developers who ignore this model risk creating controllers that consume excessive memory or perform inefficient linear scans over large datasets.
The explanation serves as a corrective guide for engineers who have observed unexpected behavior in high-load scenarios. By mapping out the flow from the API server to the local informer store, the authors provide a coherent mental model for debugging issues related to memory pressure, network traffic, and reconciler logic errors.
How it works
The core mechanism relies on the Reflector, a component within client-go that maintains a continuous watch on specific resource types. At startup, the Reflector fetches an initial snapshot of the objects it cares about. Modern implementations use a streaming list approach, where the API server sends synthetic ADDED events for existing objects before switching to live change events. This eliminates the need for a separate bulk list call, reducing connection overhead.
Once the initial state is captured, the Reflector keeps a watch open, using resourceVersion to ensure no events are missed. If the connection drops, the Reflector reconnects using the last known version. If that version is too old, the API server returns a 410 Gone error, forcing the Reflector to fetch a fresh snapshot and restart the process. This ensures the local cache remains eventually consistent with the cluster state.
Incoming changes are processed through a delta queue. Recent versions of client-go (1.36+) use RealFIFO, a strictly ordered queue that preserves the global sequence of events. Unlike older implementations that deduplicated events per object, RealFIFO passes every notification through in order. This means if an object is updated three times rapidly, the controller’s event handler will receive three distinct update notifications, ensuring no intermediate state is silently dropped before reaching the application logic.
Key details
r.Get()andr.List()inside a reconciler read from a local in-memory cache, not the API server.- Writes bypass the cache and go directly to the API server, meaning subsequent reads may return stale data.
- The cache uses a list + watch pattern, with modern implementations using streaming lists for initial synchronization.
RealFIFOreplacedDeltaFIFOin recentclient-goversions, removing automatic deduplication in the informer layer.- Memory consumption is driven by the size of the local cache and the number of indexed fields, not just the number of objects.
APIReaderprovides direct access to the API server but should be used sparingly due to higher latency and load.
Why it matters
For software engineers building operators, understanding this cache model is critical for performance tuning. Since reads are local, adding more indexes or watching unnecessary resource types can lead to gigabytes of memory usage without any visible increase in API server load. This hidden cost often manifests only under production scale, making it difficult to diagnose during development. Recognizing that List() operations can trigger linear scans over the entire local store helps developers avoid writing inefficient filtering logic that degrades controller responsiveness.
Furthermore, the eventual consistency model impacts how reconcilers handle state transitions. Assuming that a Get() call immediately reflects a previous Update() can lead to race conditions and logical errors. Engineers must design reconcilers to be idempotent and resilient to stale reads, treating the local cache as a hint rather than a definitive source of truth. This shift in mindset prevents fragile controllers that fail when the cluster is under heavy load or when network partitions cause temporary desynchronization.
What you can do
- Audit your controller’s
Watchesto ensure you are only caching resource types strictly necessary for reconciliation logic. - Use predicates to filter events at the informer level, preventing unnecessary reconcile requests from entering the workqueue.
- Avoid assuming immediate consistency after writes; design reconcilers to handle cases where local cache lags behind the API server.
- Monitor memory usage of your controller pods, as large caches with many indexes can cause out-of-memory crashes.
- Use
APIReaderonly for specific cases where strong consistency is required, such as validating final state before completing a workflow. - Test controller behavior under high churn rates to identify issues with stale reads or queue backlogs.



