Cut your AI spend with AI Gateway's Auto Router

Cloudflare released Auto Router in public beta on Thursday, available through AI Gateway. Developers set their model to cloudflare/auto, and the router automatically sends each request to a model it judges capable enough for the task, without the end user choosing a model. Cloudflare reported early internal results through its OpenCode harness showing cost savings of up to 30% compared with using only frontier models such as OpenAI Sol and Anthropic Claude Opus.
On Cloudflare's internal general knowledge work benchmark, which uses simulated workspace tools across email, calendars, Slack, files, travel and finance, cloudflare/auto completed 252 of 291 trials for a success rate of 86.6% at a total cost of $2.10, or $0.0084 per success. Anthropic Claude Opus 5.5 succeeded on 281 of 291 trials (96.6%) for $5.91 total and $0.0210 per success, while OpenAI GPT-6 Sol succeeded on 245 of 291 (84.2%) for $2.64 and $0.0108 per success. The benchmark used 97 tasks with three samples per model per task.
When a request arrives, AI Gateway builds a pool of models that can serve it, filtering out those that do not support the request format or execution mode and accounting for credentials, billing configuration, access control policies and spend limits. Unhealthy upstream providers are excluded during downtime and returned automatically after an outage. A multi-head classification model running on Workers AI and deployed on GPUs across Cloudflare's edge network analyzes the most recent messages, assigning probabilities across 14 task categories such as coding, planning, research and data analysis, and rating the request from one to five on complexity, ambiguity, stakes and dependence on earlier context. A scoring matrix combines these signals with model benchmark results and token prices; the router selects the model with the highest utility, defined as expected quality minus an adaptive cost penalty.
The router also prices cache reads and writes. Within a turn it tends to keep the same model, while across turns it applies a switching penalty that grows with the tokens already in context, charging non-cached candidates the full cost of rewriting the context. The router returns a ranked list and AI Gateway falls back to another eligible model if the top provider cannot serve the request. Cloudflare said the same classifier will support other routing profiles, including a planned cloudflare/auto-best that selects highest expected quality without the cost tradeoff. Planned work includes expanding supported models, zero-data-retention filtering, provider capacity, reasoning-level selection, Responses API and WebSockets support, and structured decision models as a first-pass classifier.
Why it matters: Organizations adopting AI through AI Gateway can reduce token spend automatically instead of relying on individuals to pick cheaper models request by request, while users keep access to the most capable models. The Auto Router is free while in beta.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Cloudflare AI Gateway now features a model router that evaluates request complexity using an edge-deployed classifier to select the optimal model. By balancing expected output quality against token costs, organizations can dramatically cut AI spend while maintaining performance.