Cloudflare's multi-agent security harness reduces alert fatigue
Cloudflare replaced a single AI agent with a specialized multi-agent system to handle security alerts, reducing hallucinations and improving evidence grounding.
On October 7, 2026, Cloudflare detailed its new multi-AI-agent architecture for security operations within its Managed Defense service. The system addresses the "alert paradox," where high volumes of security notifications overwhelm human analysts by automating data gathering and initial triage. This approach uses specialized agents rather than a single general-purpose model to improve accuracy and reduce hallucinations.
What happened
Security teams often face a flood of alerts that require manual correlation to determine if they represent a genuine threat or noise. Cloudflare’s engineering team found that their initial prototype, which relied on a single general-purpose AI agent to manage entire investigations, struggled with reliability. The agent frequently hallucinated claims unsupported by evidence because telemetry, policy details, and threat intelligence were flattened into a single prompt. This caused the distinct roles of different data sources to merge, leading to three specific failure modes: context being mistaken for authority, scope drifting across accounts or time ranges, and ambiguous failure states where timeouts looked like negative results.
To solve these issues, Cloudflare restructured its workflow into a deterministic reconnaissance phase followed by specialized AI analysis. Before any large language model is invoked, application code executes fixed reconnaissance workflows using versioned API calls. This process collects customer identity, detection history, traffic baselines, and enforcement outcomes, storing each piece of data with its source and timestamp. By separating evidence collection from inference, the system ensures that the input data is consistent and reproducible, allowing specialists to interpret static snapshots rather than fetching live data that might change between runs.
The new harness filters out noise early using Clef, Cloudflare’s open-source decision model running on Workers AI. Alerts with a high likelihood of being false positives are classified as passive and kept out of the active queue. For alerts requiring deeper review, a coordinator agent dispatches four specialist agents in parallel: one for traffic analysis, one for customer context, one for global telemetry, and one for threat intelligence. A synthesis agent then combines these typed findings into a single advisory, restricted to an approved vocabulary and unable to fetch new evidence independently.
How it works
The core innovation lies in keeping evidence collection and scope enforcement in application code rather than relying on prompt engineering. The system creates a versioned evidence package that includes the subject, scope, time anchor, and admitted evidence. Specialist agents must cite items from this package, and application code validates that every citation exists and supports the claim. If a lookup fails or times out, the system explicitly records the gap as "not checked" rather than conflating it with "checked and not found." This distinction prevents the model from making assumptions based on missing data.
Global context is integrated without compromising customer privacy. The global telemetry specialist compares an alert against aggregate patterns seen across Cloudflare’s network, such as whether an IP is scanning thousands of sites or appearing for the first time. It never accesses individual records from other customers. The synthesis agent weighs this global reputation against the specific customer’s history, ensuring that widespread patterns do not automatically trigger incident responses for isolated benign activities. Finally, Clef scores the collected evidence to determine if it is sufficient for a decision, selecting from a reduced list of attack classifications.
Key details
- Cloudflare uses OpenAI Daybreak and Anthropic models, including GPT-5.6 Cyber and Mythos, for deep analysis after initial triage.
- The system distinguishes between three evidence states: not checked, checked with no match, and checked with evidence supporting absence.
- Specialist agents run in parallel under a coordinator, with a synthesis agent combining their findings into a final advisory.
- Deterministic code handles evidence admission and validation, preventing models from crossing tenant boundaries or acting autonomously.
- The architecture relies on Cloudflare Workers, Workflows, D1, R2, and Durable Objects to maintain state and coordinate stages.
- Human analysts remain responsible for final decisions, with the AI providing consolidated views of related alerts and recommended next steps.
Why it matters
For software engineers and security leads, this shift from single-agent to multi-agent architectures highlights a critical lesson in building reliable AI systems: models should not be trusted with data retrieval or scope definition. By moving these responsibilities to deterministic code, Cloudflare reduces the risk of hallucinations and scope drift, which are common pitfalls in agentic workflows. This pattern of "recon first, inference second" ensures that AI models operate on verified, static contexts, making their outputs more auditable and consistent.
The practical impact is a significant reduction in the time analysts spend correlating disparate alerts. Instead of manually checking traffic logs, historical dispositions, and global threat feeds, analysts receive a pre-packaged advisory with cited evidence. This allows teams to focus on remediation and strategic mitigation rather than repetitive data gathering. As AI tools become more prevalent in security operations, understanding how to ground them in verifiable evidence will be essential for maintaining trust and operational efficiency.
What you can do
- Separate data retrieval from model inference in your AI applications to prevent scope drift and ensure reproducible inputs.
- Implement strict validation layers that check if model citations exist in the provided evidence package before accepting outputs.
- Use specialized, narrow-scope agents for distinct tasks like traffic analysis or threat intelligence rather than relying on a single generalist model.
- Explicitly track and label missing data states to distinguish between "not found" and "not checked" in your system logs.
- Maintain human-in-the-loop controls for final decision-making, especially in high-stakes domains like security and compliance.
- Evaluate your current alert triage processes to identify repetitive tasks that could be automated with deterministic rules before involving AI.



