AI Search is now generally available

Cloudflare announced that AI Search, its fully managed index and retrieval pipeline that combines Workers AI, Vectorize, R2, and Browser Run, is now generally available. Cloudflare says it uses the product to power search on its own blog and developer documentation, and that developers have used it for internal documentation and website search since it launched over a year ago.
The release adds native image embeddings, optical character recognition for PDFs, and support for larger files. Text files such as Markdown, HTML, CSV, and JSON, plus PDFs, can now be up to 10 MiB, up from 4 MiB. For scanned PDFs with no extractable text, OCR can be enabled so AI Search reads text from each page before chunking and embedding. OCR is available to every account.
Native multimodal retrieval is available with the Qwen3-VL-Embedding model. At query time, AI Search checks whether an instance's embedding model supports images; if so, a query image is embedded directly into the same vector space as indexed images and text. With a text-only model, image queries are converted to text using ToMarkdown and searched as a caption. Cloudflare says the earlier approach detected objects, generated a caption, and embedded that text, while the new implementation keeps captions for textual understanding and embeds image pixels directly. To keep richer representations efficient, AI Search uses Matryoshka Representation Learning, which allows smaller embeddings to retain useful information. The query flow optionally rewrites the query, embeds it, runs vector and keyword search in parallel, fuses and optionally reranks results, then returns top chunks or passes them to a generation model.
Billing begins November 1, 2026, with a reminder email planned beforehand. A free monthly allotment remains on all Workers plans: 5 million ingestion tokens covering any supported file type, 10 GB stored data, 1,000 semantic queries, and 1,000 full-text queries. The free query allotment changed from a shared pool of 2,000 queries. Paid rates are $0.75 per 1 million tokens for base ingestion, an added $0.50 per 1 million tokens for image processing, $2.00 per GB-month for storage, $0.75 per 1,000 semantic queries, and $0.10 per 1,000 full-text queries. Embedding and reranking are free with select Workers AI models, with third-party models billed separately. Parsing, chunking, embedding, keyword indexing, and reranking are included, and there are no instance hours, capacity units, or monthly minimums.
Why it matters: Teams indexing on Cloudflare can now ingest larger PDFs and image-heavy collections and search them with image queries, but should plan for usage-based billing starting November 1, 2026, and for the split free query allotment.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
AI Search is now generally available. It embeds image pixels directly for visual search, runs optical character recognition on scanned PDFs, accepts files up to 10 MiB, and works with any chat model. Here's what's new and how pricing works.