Skip to content

Cloudflare's Clef-omni adds audio, video decisions

Short answer

Cloudflare released Clef-omni, a decision model that scores audio, video, image and text in one API call instead of separate transcription pipelines. It also cut Clef-flash input pricing from $0.09 to $0.038 per million tokens and made Clef inference up to 2x faster, while trimming Clef-flash's context window from 64k to 24k tokens.

What this means for operators

For a support or ops team using classification models to triage tickets, screen uploads for PII, or flag suspicious content, Clef-omni removes the need to stitch together separate speech-to-text and video-frame extraction steps before scoring — a customer complaint with a voice memo or video attachment can now go through one call. The Clef-flash price cut (now below Jev's pricing) and the 2x speed gain on Clef make it cheaper and faster to run these decision checks at the volume a 10-200 person company's ticket or document queue actually generates, though teams with inputs over 24k tokens per request will need to use the pricier Clef model instead of Clef-flash.

Cloudflare expanded its open-weight Clef family of decision models — lightweight models built for classification and schema-constrained scoring rather than text generation — with three changes: a new multimodal model, a price cut, and a speed upgrade.

Clef-omni is the first model in the family to natively process audio (wav/mp3) and video (mp4/webm) alongside text and images in a single pipeline, built on a Qwen3-Omni-30B-A3B-Instruct foundation with the text-to-speech components stripped out. Cloudflare reports median response times of about 130ms for text-only input, 150ms for images, and roughly 1.5 seconds for a full 21-second video clip with sound. It launches at $0.15 per million input tokens, with weights published on Hugging Face.

Clef-flash pricing drops from $0.09 to $0.038 per million input tokens, which Cloudflare says makes it cheaper than TypeSafe's Jev model. The trade-off is a smaller hosted context window — cut from 64k to 24k tokens — though the underlying weights still support 256k tokens for anyone self-hosting. Cloudflare says only 0.24% of its observed requests exceeded 24k input tokens, which is why it made the cut rather than leave the window unchanged.

Clef itself keeps its $0.24 per million token price and 64k context window but gets serving-infrastructure optimizations — including a move to SGLang — that cut median latency by up to 2.0x depending on input size, with no changes to the model weights.

Cloudflare lists internal use cases including closing spam issues on its GitHub docs repo, moderating plugin libraries for phishing in its EmDash content management system, scanning for personally identifiable information, and detecting malicious domains. Clef remains API-compatible with Jev and is available through Cloudflare's AI Gateway.

Source: Cloudflare Blog

Next step

Visibility Analyzer

An AI news item will not tell you how assistants see your own site. The Visibility Analyzer checks that on your live site: which answers cite you, which pages an engine cannot retrieve, and what to fix first. Free to run.

Run a free visibility audit

Free to run. No card.

Fee
Free
Length
One run, minutes

Free tier: two analyses a day, no card required.