Kubernetes v1.35 enforces cgroup v2 migration for Linux nodes
Kubernetes v1.35 defaults failCgroupV1 to true, blocking kubelet startup on legacy cgroup v1 nodes and requiring immediate migration or configuration overrides.
The Kubernetes project has moved cgroup v1 support into maintenance mode, signaling a definitive shift toward the unified resource management interface of cgroup v2. Starting with Kubernetes v1.35, the kubelet will refuse to start on nodes still using cgroup v1 unless administrators explicitly override this behavior. This change affects all Linux-based clusters, requiring operators to verify their kernel versions, container runtimes, and node configurations before upgrading.
What happened
Control groups, or cgroups, are a Linux kernel feature that allows the operating system to allocate resources like CPU and memory to specific processes. Kubernetes relies on this mechanism to ensure containers do not interfere with one another. While cgroup v2 has been stable in Kubernetes since version 1.25, support for the older v1 interface is now being phased out. In Kubernetes v1.35, the failCgroupV1 configuration parameter defaults to true. This means that if a node is running cgroup v1, the kubelet process will fail during startup, effectively taking the node offline.
Administrators who are not yet ready to migrate can temporarily set failCgroupV1: false in their kubelet configuration file. However, this is only a stopgap measure. The Kubernetes deprecation policy indicates that full removal of cgroup v1 support is forthcoming, tracked under KEP-5573. For clusters managed by kubeadm, the enforcement is even stricter. The SystemVerification preflight check, part of the k8s.io/system-validators tool, now returns an error during initialization, joining, or upgrading if it detects cgroup v1 on a node running kubelet v1.35 or later. Previously, this was merely a warning.
This transition is not just about compliance; it unlocks modern resource management capabilities. Cgroup v2 offers a single unified hierarchy, which simplifies the interface compared to the multiple hierarchies of v1. It provides stronger isolation and supports new features like Pressure Stall Information (PSI) and improved memory quality of service (QoS). Clusters remaining on v1 miss out on these optimizations and face increasing compatibility issues with newer container runtimes and monitoring tools.
How it works
Cgroup v2 replaces the complex, multi-hierarchy structure of v1 with a single tree where all controllers (CPU, memory, I/O) are attached. This unification allows for more consistent resource accounting and prevents scenarios where a process is limited by one controller but not another due to hierarchy mismatches. In Kubernetes, the kubelet interacts with these controllers to enforce limits defined in Pod specifications. With v2, the kubelet can use features like memory.high for throttling and memory.min for hard protection, which were not possible or reliable in v1.
The migration also changes how out-of-memory (OOM) events are handled. In cgroup v2, the kubelet sets memory.oom.group for each container cgroup. When an OOM event occurs, the kernel kills all processes within that container simultaneously, rather than picking them off one by one. This prevents partially functioning containers from lingering in a broken state. Additionally, cgroup v2 enables delegation, allowing rootless containers to manage their own cgroups safely via systemd, a capability that was risky or impossible in v1.
Key details
- Default Enforcement: In Kubernetes v1.35,
failCgroupV1defaults totrue, causing kubelet startup failure on cgroup v1 nodes. - Kernel Requirements: A Linux kernel version 5.8 or later is required for cgroup v2 support, with 5.9+ recommended for Memory QoS stability.
- Runtime Support: Containerd v1.4+ and CRI-O v1.20+ support cgroup v2; automatic driver discovery requires containerd v2.0+ or CRI-O v1.28+.
- Memory QoS: The alpha Memory QoS feature, which uses
memory.highandmemory.low, is exclusively available on cgroup v2 nodes. - Tooling Updates: Monitoring tools must be updated; cAdvisor v0.43.0+ is recommended, along with compatible versions of Java and Node.js libraries.
- CPU Weight Conversion: Newer OCI runtimes like crun v1.23 and runc v1.3.2 use a non-linear conversion from v1
cpu.sharesto v2cpu.weight, improving granularity for small CPU requests.
Why it matters
For software engineers and platform teams, this shift means that legacy infrastructure configurations are no longer viable. If you upgrade your control plane to v1.35 without migrating your worker nodes, your cluster will break. Nodes will fail to join, and existing nodes may fail to restart after a reboot. This creates a hard dependency on operating system updates, as many older Linux distributions default to cgroup v1. Teams must now coordinate kernel upgrades across their fleet, which can be a significant operational hurdle in large, heterogeneous environments.
Beyond the immediate migration pain, staying on v2 unlocks better resource efficiency. Features like PSI provide real-time visibility into resource contention, allowing for more accurate autoscaling decisions. The improved OOM handling ensures that application failures are cleaner and easier to debug. Furthermore, as in-place vertical scaling becomes stable, cgroup v2 is required for accurate aggregate enforcement. Ignoring this migration path will eventually leave clusters unable to adopt these performance and reliability improvements.
What you can do
- Check Current Status: Run
stat -fc %T /sys/fs/cgroup/on your nodes. If it returnscgroup2fs, you are already on v2. - Verify Kernel Version: Ensure all Linux nodes are running kernel 5.8 or later. Upgrade the OS if necessary before attempting the Kubernetes upgrade.
- Update Container Runtime: Confirm you are using containerd v1.4+ or CRI-O v1.20+. Ideally, move to containerd v2.0+ to leverage automatic cgroup driver discovery.
- Configure Kubelet: If you cannot migrate immediately, set
failCgroupV1: falsein your kubelet config, but plan to remove this override soon. Match the cgroup driver to your runtime, preferably usingsystemd. - Update Monitoring Stack: Upgrade cAdvisor to v0.43.0 or later and verify that your Prometheus scrapers and other monitoring tools support cgroup v2 metrics.
- Test Memory QoS: If you use advanced resource management, test the alpha Memory QoS features in a staging environment to understand how
memory.highthrottling affects your workloads.



