Stop giving AI agents root access to production systems
CNCF Ambassador Mauro Morales argues against granting AI agents root privileges, proposing immutable infrastructure and software factories as safer alternatives for autonomous system management.
Este artículo está disponible solo en inglés.
Giving autonomous AI agents root access to production machines is a dangerous shortcut that sacrifices reproducibility for convenience. In a recent CNCF article, Ambassador Mauro Morales outlines why developers should instead treat agents as workloads that propose changes through established software factories rather than executing them directly.
What happened
The cloud native community is currently grappling with how much autonomy to grant AI agents operating within Kubernetes clusters and other infrastructure. This question arises frequently among engineers building agent-based tools, security teams securing installations like OpenClaw, and platform leads managing cluster controls. While there is no universal answer, Morales provides an architectural perspective focused on the operating system and software delivery layers.
Morales acknowledges the temptation of granting superuser access, noting that it enables impressive feats such as migrating services between machines without downtime. However, he warns that experimentation differs vastly from production requirements. Relying on instruction files like AGENTS.md to prevent catastrophic errors in a nondeterministic system is insufficient. The core issue is not just failure prevention but reproducibility: teams must know exactly what changed and who or what initiated the change.
The argument extends beyond simple bug avoidance. Just as human operators are restricted from unrestricted root access in production environments, AI agents should face similar constraints. The goal is to maintain a clear audit trail and ensure that every state transition is intentional, reviewed, and traceable back to its source.
How it works
The proposed model relies on immutable infrastructure patterns where the base operating system is treated as an image updated as a unit rather than modified package by package at runtime. Technologies like bootc, Flatcar Container Linux, and Kairos support this approach by allowing selected filesystem parts to be read-only or by replacing the entire base system image during updates. This ensures the machine moves from one defined state to another, with runtime serving only to execute pre-built states rather than improvise new ones.
Agents operate as isolated workloads within this framework. Instead of having direct host access, an agent runs in a short-lived execution environment, such as a container pod or virtual machine, depending on the threat model. The agent inspects the current system state, diagnoses issues, and formulates a proposal for the next desired state. This proposal is then submitted as a versioned definition, triggering the software factory pipeline.
The software factory handles source control, review, continuous integration, testing, image building, signing, and deployment. Only after passing through these gates does the new system image become available for consumption. This process separates the agent’s observational capabilities from its executive power, ensuring that any change to the host follows the same rigorous path as human-authored code.
Key details
- Root access should never be granted to AI agents in production environments due to the nondeterministic nature of their outputs.
- Immutable systems prevent configuration drift by treating the base OS as an image that is replaced entirely rather than modified in place.
- Agents should run in isolated execution environments like pods or VMs, separated from the host kernel according to their specific threat model.
- The software factory pipeline must retain provenance data to trace any problematic state back to the specific commit and image that introduced it.
- Self-healing actions, such as restarting workloads, can be authorized autonomously, but self-improvement changes require full pipeline review.
- Agent identity must be restricted to proposing changes, without permissions to approve, merge, sign, or force-deploy new images.
Why it matters
For software engineers and DevOps leaders, this approach shifts the focus from trusting the AI model to trusting the boundaries around it. By removing the ability of agents to mutate the running system directly, teams eliminate a major class of unrecoverable errors. This separation of concerns allows organizations to leverage AI for diagnostics and optimization without compromising the stability and security of their production infrastructure.
Furthermore, this model integrates AI operations into existing DevOps best practices. It ensures that agent-driven changes are subject to the same testing, signing, and review processes as traditional code. This consistency simplifies compliance, auditing, and incident response, as every system state has a verifiable history. It transforms AI from a potential risk vector into a controlled contributor within the software delivery lifecycle.
What you can do
- Audit your current AI agent deployments to ensure none have unrestricted root or superuser access to production hosts.
- Implement immutable infrastructure patterns using tools like bootc, Flatcar, or Kairos to enforce state-based updates.
- Configure agent harnesses to run in isolated containers or VMs, limiting their access to host resources and credentials.
- Establish a software factory pipeline that requires review and testing for any system image changes proposed by automated agents.
- Define clear policies distinguishing between autonomous self-healing actions and proposed self-improvement changes that require human approval.
- Restrict agent identities to prevent them from approving their own pull requests or signing deployment artifacts.



