Cloudflare has launched a novel open-weight decision-making model named Clef-omni. Building on its pre-existing proficiency in processing text and images, this model now extends its support to encompass audio and complete video formats. This enhancement allows for the concurrent processing of multimodal content via a single API call. Clef-omni inherently supports a variety of audio and video formats, thereby obviating the necessity for developers to implement supplementary speech-to-text transcription and audio-video separation procedures.
It is constructed on the foundation of Qwen3-Omni-30B-A3B-Instruct, preserving its fundamental comprehension abilities. Its primary focus lies in structured decision-making tasks, eschewing the generation of conventional text outputs, and delivering responses with remarkable speed.
Moreover, Cloudflare has slashed the usage fee for Clef-flash, reducing the cost per million input tokens from $0.09 to $0.038. Concurrently, it has modified the context window for the hosted version from 64k to 24k. This suite of models employs a tiered pricing strategy to cater to the diverse cost and performance requirements of different developers.
