Cloudflare launches Clef, open multimodal decision models on Workers AI
Cloudflare announced Clef on 1 October 2026: two decision models, Clef (27B, from Qwen3.8-27B) and Clef-flash (9B, from Qwen3.5-9B). Weights went up on Hugging Face the evening before under Apache-2.0, and both run on Workers AI as @cf/cloudflare/clef ($0.24 / M input tokens) and @cf/cloudflare/clef-flash ($0.09 / M), with a 65,536-token context, 1–64 questions per call and up to four images.
How it works. A frozen Qwen backbone with a rank-256 LoRA (shipped merged) and a separate joint schema head that scores every option of every question in one prefill pass. Training used label-smoothed cross-entropy plus Brier loss, with RLCD as a secondary objective, on internal synthetic data. Cloudflare says the API is “fully compatible with Jev and SystemOne”. Fine-tuning and RL run through its forward-deployed team for now. @lucataco, who works on models at Cloudflare, showed Clef-flash running locally on an M5 Max and said Replicate’s Cog work now powers fine-tuning and redeploys on Workers AI.
The numbers (UnverifiedClaim). On Cloudflare’s mirror of the Decision Index 0.2.1, Clef scores 61.21 and Clef-flash 57.07, against the official 57.91 for Jev. The official Space has not been regenerated since 28/09 and does not list Clef. The “~13× faster than Jev” line compares Clef-flash’s 38.8 ms GPU time with Jev’s 524 ms hosted-API round trip, which Cloudflare’s own board calls not comparable. On the same day Fastino claimed 64.81 for GLiDE, also self-run.
How we list it. Catalog (one page for both sizes; the quadrant point is the 27B), 6.7 maturity / 7.7 capability / 26 adoption.
