Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash

Cloudflare released Clef-omni on Thursday, an open-weight decision model that accepts audio, video, image and text input in one API call. The launch follows last week's release of Clef and Clef-flash, Cloudflare's first open-weight decision models. Alongside the new model, Cloudflare cut the price of Clef-flash and sped up the hosted version of Clef on Workers AI. Clef-omni weighs in at $0.15 per million input tokens, Clef stays at $0.24, and Clef-flash falls from $0.09 to $0.038 per million input tokens.
The family is now three models with overlapping but distinct roles. Clef is positioned as the higher-quality general model with a 64k context window; Clef-flash is the cheap, high-throughput option; Clef-omni is the multimodal entry point. Cloudflare says Clef is fully Jev-API compatible and available through AI Gateway, so existing callers can switch by changing the model ID. Weights for Clef-omni are published on Hugging Face.
The headline capability is input surface area. Clef-omni takes audio in wav or mp3, video in mp4 or webm, images, and text. The example request in Cloudflare's announcement posts images, audio and video as base64 data URLs, plus a set of named questions with a type field. That lets a caller ask, in one shot, whether a label is visible in a photo, whether a machine sounds normal in a recording, and whether a fan is spinning in a clip. The practical pitch is pipeline collapse: instead of chaining a speech-to-text model, a vision model and a classifier, one call returns constrained scores.
Architecturally, Clef-omni is built on a Qwen3-Omni-30B-A3B-Instruct mixture-of-experts foundation. Cloudflare keeps the comprehension backbone and discards the text-to-speech output components. Since Clef models are not LLMs and do not generate output tokens, the system skips token generation entirely: a prefill pass scores all modalities and all valid option values at once, avoiding any transcribe-or-caption step. Media maps into a unified sequence, with audio synced against video frames for joint processing. Candidate values are pulled from internal embeddings using a two-stage attention routing scheme, in which each option collects evidence across the input before field vectors cross-attend over the full context to produce confidence scores. A lexical grammar constrains outputs so option semantics survive.
Training follows the Clef recipe: the Qwen3 backbone is frozen, low-rank adapters (LoRA) are trained, and post-training mixes label-smoothed cross-entropy loss with Brier score calibration. LoRA is a parameter-efficient technique that trains a small number of additional weights rather than the full network; Brier scoring is a proper scoring rule that rewards calibrated probabilities. The stated goal is robustness to schema variations, field ordering and prompt structure.
Latency is a selling point. Cloudflare reports text-only decisions at roughly 130 ms median, image inputs near 150 ms, audio clips in a few hundred milliseconds, and a full 21-second video with sound scored in about 1.5 seconds, all in a single call.
On benchmarks, Clef-omni lands near its siblings rather than dominating them. BFCL case exact: 98.2 (Clef 98.47, Clef-flash 98.76, Jev 95.75). API-Bank accuracy: 92.7 (Clef 91.93, Clef-flash 93.11, Jev 88.19). BANKING77 macro-F1: 94.8 versus Jev's 79.74. CLINC150+OOS macro-F1: 97.7 versus Clef-flash's 66.77 and Jev's 89.27. Areas where it trails: ToolRet nDCG@10 66.6 against Clef's 69.19; Home appliances case exact 69.3 against Clef-flash's 97.73; When2Call accuracy 63.3 against Jev's 80.97; BRIGHT nDCG@10 42.0 against Jev's 47.52; PhishNChips accuracy 73.2 against Clef's 79.60. On TypeSafe workflow evals it reports invoice processing exact actions 60.2, primary action 82.0, customer service exact actions 71.6, security incidents exact actions 61.7 and agent trace observability primary action 65.8.
Clef-flash's price cut comes with a trade-off. The hosted context window shrinks from 64k to 24k. Cloudflare says its usage data shows only 0.24% of requests exceed 24k input tokens, and points users with larger contexts to Clef. Hugging Face weights are unchanged and were trained for a 256k context window for self-hosting. That is a fairly wide gap between hosted and self-hosted limits, and it means anyone relying on the advertised 64k needs to either move up to Clef or run the weights themselves.
Clef itself gets faster without new weights, because the changes sit in the serving layer. Cloudflare moved serving to SGLang and contributed the model upstream (PR #42721), landing in SGLang 0.5.22. Reported median/p95 pairs improve from 262/438 ms to 152/351 ms at roughly 800 tokens (1.7x median speedup), 616/777 ms to 305/531 ms at about 3,400 tokens (2.0x), and 2,721/3,250 ms to 1,635/1,805 ms at about 16,000 tokens (1.7x). Updated SGLang launch commands are on the Hugging Face repo and in the SGLang cookbooks.
Cloudflare also sketched internal usage: a public GitHub docs repo using Clef to detect and close spam issues, the EmDash content management system moderating plugin libraries for phishing, a data loss prevention team scanning for personally identifiable information such as government IDs, and a threat intelligence team detecting malicious domains. The framing is that classification now lives at the model layer rather than requiring a specialized team and a labeled corpus.
Why it matters: teams building agentic or intake-heavy workflows that currently glue together speech-to-text, vision and classifier models can collapse that stack into one schema-constrained call, and multimodal support arrives at a lower per-token price than Clef. The constraint to plan around is Clef-flash's 24k hosted context, which changes the calculus for anyone running long documents or large payloads. Self-hosting remains the escape hatch, though it shifts cost and operations back onto your own infrastructure.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
We are expanding the Clef decision model family with Clef-omni, natively processing audio, video, images, and text in a single pipeline. We’ve also lowered Clef-flash pricing and boosted Clef inference speeds by up to 2.0x.