Cloudflare Uses an AI Harness to Probe and Harden its WAF

Cloudflare placed frontier AI models inside a controlled testing harness to probe its Web Application Firewall (WAF), according to an account by Matt Foster. The models used already-blocked attack payloads as starting points, proposing changes to encoding, placement, or delivery based on responses from earlier attempts, rather than replaying fixed test cases. The models had no access to Cloudflare's WAF rules, source code, or internal security signals, making the test black-box from their perspective.
A Python harness handled the work Cloudflare did not delegate to the models: constructing and replaying HTTP requests, maintaining scenario state, enforcing limits, and collecting responses. One model call proposed the next mutation while another reviewed the resulting response, so later attempts could adapt without the models directly controlling request execution.
Across 45 scenarios, the system generated 1,107 attempts. Of those, 607 produced the post-triage result set: 558 requests blocked by the WAF and 49 findings considered relevant for further remediation. Forty-eight of the 49 findings involved command injection or server-side request forgery (SSRF).
In one SSRF test, the tester repeatedly changed the representation and placement of a cloud metadata address, trying decimal, octal, and other forms. A request using a decimal representation was eventually blocked. On the next attempt, the model kept the same request shape but switched to a trailing-dot representation; the client then encountered a redirect rather than a WAF block. Cloudflare preserved the result for investigation instead of treating it as evidence the attack succeeded.
Human review remained the final validation step. Reviewers checked whether requests had reached the target, remained malicious, were clearly unblocked, fell within the WAF's responsibility, and could be safely reproduced. Surviving cases were replayed and evaluated as candidates for changes to rules, normalization, or other mitigations. The work contributed three changes to Cloudflare's Managed Ruleset: two new detections, SSRF - Obfuscated Host and SSRF - Restricted Protocol, plus an improvement to the existing SSRF - Cloud rule.
The article notes similar harness patterns elsewhere: Cloudflare's Vulnerability Discovery Harness separates discovery from independent validation; Google Mandiant's Agentic Vulnerability Discovery Harness chains specialised agents through source-code analysis, hypothesis generation, and verification; OpenAI's Codex Security builds a repository threat model, searches for vulnerabilities, and attempts reproduction in an isolated environment before proposing fixes for human review; and Google's PageBreak validates whether AI-generated vulnerability hypotheses are exploitable.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Cloudflare placed frontier AI models inside a controlled testing harness to probe its Web Application Firewall (WAF), using blocked attacks as starting points for models to generate and refine new variations. By Matt Foster