AivexaNewsSearch
AI news for builders and product teamsChecked every hour

We tested our own WAF with frontier AI models. Here’s what we found

Collected Oct 8, 2026

Cloudflare ran an LLM-driven testing tool against an authorized customer staging environment protected by its own Web Application Firewall, recording 1,107 mutation attempts across 45 scenarios. The company published the results, describing the exercise as a foundational building block of its WAF development lifecycle.

The tester uses frontier models in two roles: a proposal call that receives the starting request, context and a short history of earlier results and suggests the next variation, and a review call that sees the request context, response status, selected headers and response body. Neither call has access to rule expressions, rule IDs, WAF Attack Score details or the identity of the security layer that acted, and neither can deploy a rule or change enforcement. The system is written in Python and handles HTTP replay, scenario orchestration, state tracking and result collection; code, not the models, sends each request, checking the target hostname against an allowlist, disabling redirects and enforcing an attempt limit. Response text may reappear in later prompts, so the tester treats it as untrusted input.

Of the 45 scenarios, 44 covered six attack categories: cross-site scripting, SQL injection, command injection, server-side request forgery, path traversal or local file inclusion, and Log4j; the remaining scenario covered log injection and is reported separately. The test zone ran WAF Attack Score blocking scores of 30 or below, all Cloudflare Managed Ruleset enabled, and the OWASP Core Ruleset at Paranoia Level 3.

After triage, the post-triage result set was 607: 558 blocked requests plus 49 WAF-relevant findings. XSS, LFI, SQLi and Log4j had near full coverage, while 48 of the 49 findings belonged to command injection and SSRF. Requests that were not blocked were reviewed against five checks covering whether a valid request was sent, whether it was clearly not blocked, whether it remained malicious, whether the behavior belonged to the WAF, and whether engineers could reproduce it safely. In one SSRF session, the tester sent the same cloud metadata address as integer, octal and trailing-dot forms; the WAF blocked all but one, where the client encountered a redirect rather than a block, prompting investigation into whether the trailing dot changes how the destination is read.

The work contributed to three Managed Ruleset changes in the July 21 release: new detections for SSRF - Obfuscated Host and SSRF - Restricted Protocol, and improvement of the existing SSRF - Cloud rule. Candidate rules were tested against live traffic and assessed for false-positive risk before protecting customers. A future post will cover a white-box approach where the model knows both application vulnerabilities and the WAF rules.

Why it matters: Cloudflare customers get the new SSRF detections from the July 21 release. Cloudflare advises checking that Managed Rules and WAF Attack Score are configured correctly, running Managed Rules in log first and reviewing Security Events before switching to Block, or asking an account team to enable Attack Signature Detection. It notes that a bypassing payload still needs an exploitable application, so patching remains a strong defense.

Read at Cloudflare AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

We built a WAF tester that adapted each request based on what the WAF blocked or passed. This helped us explore variations that a fixed test might miss. We ran it across six attack categories on an authorized staging environment and discovered detection gaps worth fixing. Here’s how the loop worked, what got through, and what we did about it.