Cloudflare Expands Clef Decision Model Family with Multimodal Capabilities
Advancing Decision Models Beyond Text
Cloudflare has announced a significant expansion of its Clef family of open-weight decision models, introducing Clef-omni, a new iteration capable of processing audio, video, image, and text inputs. This development marks a shift in how decision-based AI models interact with complex data, moving away from traditional text-only paradigms. By integrating multiple modalities into a single API call, developers can now streamline workflows that previously required complex, multi-stage pipelines for speech-to-text transcription or frame-by-frame video analysis.
Clef-omni is built upon a Qwen3-Omni-30B-A3B-Instruct mixture-of-experts (MoE) foundation. The architecture is specifically optimized for decision-making rather than generative language tasks, allowing it to bypass the computational overhead associated with token generation. Instead, the model performs a prefill pass across the entire payload, scoring all modalities simultaneously. This approach enables rapid processing, with the platform reporting that a 21-second video clip with audio can be scored in approximately 1.5 seconds.
Infrastructure Optimizations and Pricing Adjustments
Alongside the introduction of Clef-omni, Cloudflare has implemented infrastructure improvements to the existing Clef model, resulting in significant latency reductions. By migrating to the SGLang serving framework, the company has achieved median speed improvements of up to 2.0x for specific input sizes. These optimizations were achieved at the serving layer, meaning the underlying model weights remain consistent while performance increases for users of the hosted Workers AI service.
The company also adjusted its pricing structure for the Clef-flash model, reducing costs to $0.038 per million input tokens, positioning it as a more affordable option for high-volume agentic workflows. To facilitate this price reduction, the context window for the hosted version of Clef-flash has been adjusted to 24k tokens. Cloudflare noted that internal usage data indicated only 0.24% of requests exceeded this threshold. For applications requiring larger context, the standard Clef model remains available with a 64k context window.
Practical Applications in Security and Operations
The Clef model family is increasingly being utilized for automated classification tasks that previously required specialized machine learning teams or extensive custom training. Within Cloudflare’s own internal operations, the models are currently deployed for diverse security and management use cases. These include detecting and closing spam on public repositories, moderating plugin libraries for phishing attempts, and scanning data for personally identifiable information (PII) to support data loss prevention (DLP) efforts. Additionally, the threat intelligence team employs the technology to assist in identifying malicious domain activity.
By offering these models with open weights, Cloudflare aims to lower the barrier for developers looking to integrate decision-making capabilities into their own applications. The models are designed to be resilient to variations in schema and prompt structures, providing a standardized way to parse input data with precision across different domains.