AI Briefing
KOSign in

Cloudflare releases Clef-omni with audio and video support, cuts Clef-flash pricing, and speeds up Clef

·2026.10.10 03:27

Key point

Clef-omni handles audio and video inputs natively, while Clef-flash drops to $0.038 per million input tokens but reduces its hosted context window to 24k.

Details

Cloudflare has expanded its Clef family of open-weight decision models with the release of Clef-omni, which natively accepts audio, video, image, and text inputs. This marks a shift from text-only decision models, allowing developers to bypass cascading transcription pipelines by processing multimodal data in a single API call. The model is built on a Qwen3-Omni-30B-A3B-Instruct mixture-of-experts foundation, utilizing low-rank adapters (LoRA) and a two-stage attention routing mechanism to score parameters across all modalities simultaneously.

Performance and Pricing Updates

Clef-omni delivers calibrated decisions with median latencies of approximately 130 ms for text, 150 ms for images, and 1.5 seconds for a 21-second video clip. It is priced at $0.15 per million input tokens.

In parallel, Cloudflare reduced the price of Clef-flash to $0.038 per million input tokens, making it cheaper than the competitor model Jev. This price cut comes with a trade-off: the hosted version of Clef-flash now has a context window of 24k tokens, down from the previously advertised 64k. Cloudflare notes that only 0.24% of requests exceed this limit, though the open weights on Hugging Face still support a 256k context window for self-hosting. Clef remains at $0.24 per million input tokens with a 64k context window.

Infrastructure Optimizations

The hosted Clef model has seen significant speed improvements due to infrastructure optimizations, including a migration to SGLang for serving. Median response times have improved by 1.7× to 2.0× across various input sizes:

  • ~800 tokens: Median dropped from 262 ms to 152 ms.
  • ~3,400 tokens: Median dropped from 616 ms to 305 ms.
  • ~16,000 tokens: Median dropped from 2,721 ms to 1,635 ms.

These updates are available via AI Gateway and are fully compatible with the Jev API. Cloudflare highlights internal use cases such as spam detection, phishing moderation, and PII scanning, emphasizing that decision capabilities are now accessible without specialized machine learning teams.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.