Cloudflare Containers rearchitected for on-demand AI agent sandboxes
Cloudflare updated its Containers infrastructure to support runtime image selection and filesystem snapshots, reducing startup times by over six times for AI agent workloads.
Cloudflare has fundamentally rearchitected its Containers platform to better support the dynamic needs of AI agents. Published on September 30, 2026, this update introduces a new scheduling policy that allows application code to select container images and instance types at runtime rather than at deployment. The changes aim to eliminate the latency penalties associated with traditional container orchestration, enabling sandboxes to start in under a second.
What happened
Previously, Cloudflare Containers operated similarly to traditional application deployments. Developers had to define the container image and compute resources during the build process, deploying each configuration as a separate application. This model required distinct Durable Object namespaces for every combination of image and instance type. If an agent needed both a small Node.js environment and a large Python build environment, developers had to manage two separate applications and route traffic between them manually. Every change to the environment required a new deployment cycle, making it difficult to adapt to the unpredictable resource needs of AI agents.
The new architecture shifts these decisions into the application logic itself. By introducing the durable_object scheduling policy, Cloudflare allows code to choose the specific image and instance size when a task arrives. This means a single Durable Object class can now spin up different types of sandboxes side by side. For example, an agent can request a lightweight environment for simple queries and a heavy-duty build environment for complex tasks without any pre-provisioning. This shift transforms infrastructure configuration from a static deployment artifact into dynamic code that executes at request time.
Alongside this flexibility, Cloudflare addressed the critical issue of startup latency. AI agents often create sandboxes on demand for individual tasks, meaning users wait for the environment to initialize before any work begins. The previous global control plane model introduced significant overhead as it resolved configurations and coordinated placement across the network. The new system localizes this process, starting containers on the same machine as the controlling Durable Object whenever possible. This reduces the median startup time from just over four seconds to 648 milliseconds, a more than six-fold improvement verified by independent benchmarks.
How it works
The core mechanism enabling this speed and flexibility is the tight integration between Containers and Durable Objects. Each container instance is attached to a Durable Object, which acts as a persistent, programmable controller. In the new model, the Durable Object does not just manage the lifecycle; it directly controls the container’s configuration via the native ctx.container API. When a request arrives, the code checks the task requirements and selects an image from a predefined list in the configuration file. The scheduler then looks for capacity on the local host first, favoring machines that already have the required image or snapshot in local storage to avoid download delays.
To further reduce startup times, the runtime no longer boots a virtual machine from scratch for every request. Instead, it restores a prepared virtual machine that is already initialized but unassigned. This approach reuses networking and filesystem setups, batching operations that were previously performed sequentially. Additionally, Cloudflare introduced a ready-to-use system image called cloudflare/debian-trixie. This base image includes Debian Trixie Slim and Node.js 24.20.0 LTS, distributed across hosts in advance. Agents can start this sandbox instantly and then use execution commands to clone repositories or install packages, bypassing the need to build and push custom Docker images for simple tasks.
Key details
- Runtime Configuration: The new
durable_objectscheduling policy allows code to select container images and instance types at runtime, removing the need for separate deployments for each environment type. - Startup Performance: Median startup time dropped from 4.049 seconds to 648 milliseconds, with the 95th percentile improving from 5.839 seconds to 910 milliseconds.
- Filesystem Snapshots: A public beta feature enables saving and restoring container filesystems, allowing agents to pause and resume long-running tasks without losing state or repeating setup steps.
- Prepared Base Image: The
cloudflare/debian-trixieimage is pre-distributed across hosts, allowing agents to start a Linux environment immediately without building custom images. - Burst Capacity: Preliminary tests showed the system could start 100,000 containers in 5.387 seconds across six locations, demonstrating high scalability for burst workloads.
- Rollout Control: Rollouts are now managed via code within the Durable Object, allowing strategies like canary releases or pinning active projects to specific images without platform-level configuration changes.
Why it matters
For engineers building AI agent platforms, this update removes a major bottleneck in user experience. Traditional container orchestration is designed for long-running services, not ephemeral tasks that must start instantly. By moving configuration to runtime, developers can build more efficient systems that only consume resources when needed. This is particularly important for coding agents, evaluation frameworks, and reinforcement learning systems that require thousands of isolated environments. The ability to start a sandbox in under a second means users spend less time waiting for environments to provision and more time interacting with the agent.
The introduction of filesystem snapshots also changes how state is managed in serverless environments. Previously, preserving an agent’s work required complex external storage solutions or keeping containers running indefinitely, which increased costs. Now, an agent can save its workspace, terminate the container, and restore it later exactly where it left off. This capability supports asynchronous workflows and long-running tasks that may span days, making it feasible to build more sophisticated agent applications on serverless infrastructure without managing persistent servers.
What you can do
- Update your
wrangler.jsoncto include thedurable_objectscheduling policy and declare the images your Durable Object can access. - Refactor existing container logic to select images and instance types based on task parameters within the Durable Object class.
- Test the
cloudflare/debian-trixiebase image for tasks that do not require custom Docker builds to leverage pre-distributed assets. - Implement filesystem snapshotting in your workflow to save agent state after significant milestones and restore it upon resumption.
- Use Durable Object storage to manage rollout strategies, such as pinning specific IDs to older images during migration periods.
- Monitor startup metrics using the new
ctx.containerAPI to ensure your agents are benefiting from the localized scheduling improvements.
