Security & privacy

Cloudflare uses AI agents to stress-test its own web application firewall

Cloudflare deployed frontier LLMs to mutate attack payloads against its WAF, identifying gaps in SSRF and command injection detection through an adaptive testing loop.

Cloudflare engineers recently subjected their own Web Application Firewall (WAF) to a rigorous stress test using frontier large language models. Published on September 29, 2026, the study details how an automated system acted as an adversarial hacker, iterating through thousands of payload variations to find blind spots in real-time defenses. The experiment revealed specific weaknesses in handling obfuscated server-side request forgery and command injection attempts, leading to immediate updates in Cloudflare’s managed rulesets.

What happened

The security team built a custom Python-based testing harness to evaluate whether modern AI models could bypass WAF protections more effectively than traditional static or dynamic analysis tools. Unlike standard penetration testing, this system used LLMs to dynamically mutate attack vectors based on live HTTP responses. The tester had no access to source code or internal WAF rules, relying solely on selected response data to guide its next move. This black-box approach mimicked how an external attacker might probe a protected application without prior knowledge of its infrastructure.

The test ran against an authorized customer staging environment configured with Cloudflare’s strictest settings, including WAF Attack Score blocking at 30 or below and OWASP Core Ruleset at Paranoia Level 3. Over the course of the experiment, the system executed 1,107 mutation attempts across 45 scenarios covering six major attack categories: cross-site scripting, SQL injection, command injection, server-side request forgery, path traversal, and Log4j exploits. While the majority of attacks were blocked, the process identified 49 specific findings that required human review and subsequent mitigation.

Most of the successful bypasses fell into two categories: command injection and server-side request forgery. For instance, in one SSRF scenario, the model discovered that representing a cloud metadata IP address with a trailing dot allowed it to evade detection where standard decimal or octal representations failed. These findings were not confirmed exploits but rather leads that indicated the WAF’s normalization logic might treat certain edge cases differently. The team used these insights to refine their detection engines, resulting in new rules for obfuscated hosts and restricted protocols.

How it works

The testing system operates on an adaptive loop driven by two distinct LLM calls per iteration. The first call, the proposal phase, receives the initial request context and a history of previous results to suggest a new variation. It might change the encoding, move the payload to a different part of the HTTP request, or alter the destination format. The second call, the review phase, analyzes the response status, headers, and body to determine if the attempt was successful or blocked. This feedback loop allows the model to learn which mutations are effective without ever seeing the underlying WAF rule expressions or attack scores.

Crucially, the code maintains strict control over the execution environment. The LLM does not send requests directly; instead, it generates suggestions that the Python harness validates and executes. Before each request, the system checks the target hostname against an allowlist, disables redirects, and enforces a hard limit on attempts. After each response, the system records structured evidence, treating any returned text as untrusted input. This design ensures that the testing process remains safe and reproducible, preventing the AI from causing unintended damage or leaking sensitive data during the evaluation.

Key details

  • The test generated 1,107 mutation attempts, with 558 explicitly blocked by the WAF before reaching the application.
  • Human triage reduced the raw output to 49 valid findings, 48 of which were related to command injection or SSRF.
  • The WAF configuration included WAF Attack Score blocking at 30 or below and OWASP Core Ruleset at Paranoia Level 3.
  • New detections for SSRF obfuscated hosts and restricted protocols were added to the Managed Ruleset following the test.
  • The system used two separate LLM calls per iteration: one for proposing mutations and one for reviewing responses.
  • Requests that bypassed the WAF were treated as leads for investigation, not confirmed exploits, requiring further validation.

Why it matters

For software engineers and security leads, this study highlights the limitations of static rule-based defenses against adaptive adversaries. Traditional penetration tests often follow predefined scripts, missing novel encoding techniques or logical flaws that an AI can discover through iteration. By demonstrating that LLMs can find gaps in even highly configured WAFs, Cloudflare underscores the need for continuous, automated security testing that evolves alongside threat landscapes. It suggests that relying solely on signature-based detection is insufficient when attackers can use AI to generate infinite variations of known exploits.

The findings also reinforce the importance of defense in depth. Even when a WAF fails to block a specific mutated request, the application itself must remain secure. In the SSRF example, the bypass did not result in data exfiltration because the application layer likely had its own protections or the request was malformed in a way that prevented actual exploitation. This reminds developers that patching software and maintaining up-to-date dependencies are critical complements to network-level security controls. A WAF is a shield, not a cure, and its effectiveness depends on the resilience of the entire stack.

What you can do

  • Enable all available managed rulesets and set WAF Attack Score thresholds to recommended levels for your risk profile.
  • Run existing application security tests against staging environments protected by the same WAF configurations as production.
  • Implement positive security controls that define expected request shapes, reducing the attack surface for unknown variants.
  • Review security events in log-only mode before enforcing new rules to identify false positives and legitimate traffic patterns.
  • Keep application dependencies and frameworks updated to mitigate vulnerabilities that might be exposed if WAF layers fail.
  • Consider using AI-driven testing tools to supplement manual penetration tests, focusing on adaptive mutation of payloads.

Tools from the Bytechap store

Keep reading

All stories