Clef / Clef-flash
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity7.2/10
Apache-2.0 weights (27B, 9B) plus a priced, self-serve Workers AI endpoint with full schema docs. Within a day: official Ollama library entries on /v1/systemone, an Ollaya package and llama.cpp support (PR merged 2 Oct). No GA or model-specific SLA.
- Capability7.5/10
Noul / Choice / Score plus up to 4 images in one pass. Calibration method published (Brier, label smoothing), no ECE. Cloudflare's GPU run: 209 ms median. First independent test of the hosted API: 638–717 ms median, ~1 s p95. DI 61.21 is self-run, not on the official board.
- Adoption37/100
610 points on Hacker News, a 1,245-like launch post from @CloudflareDev, 706 + 237 Hugging Face likes, 33 quantizations (ggml-org among them), Ollama and Ollaya packages and about 15 community repos in 36 hours. No external production case published yet.
Vendor claims
- Decision Index 0.2.1 in Cloudflare's own run: Clef 61.21, Clef-flash 57.07, Jev 1.13 57.91 (official board value)[Vendor claim — not independently verified]
- Clef areas: Knowledge 51.2, Language 61.2, Retrieval 62.5, Tools 81.2, Arts 47.8 (Cloudflare's run)[Vendor claim — not independently verified]
- Latency median / p95: Clef 209.3 / 238.6 ms, Clef-flash 38.8 / 122.4 ms (Cloudflare's run, one RTX PRO 6000)[Vendor claim — not independently verified]
- Clef is currently the leader on the Decision Index and beats Jev on 3 of 4 TypeSafe workflow evals[Vendor claim — not independently verified]
- Across 43 benchmark runs, Clef is 2.5× faster than Jev at the median and Clef-flash 13× faster; a Clef model scores highest on 7 of 10 decision benchmarks[Vendor claim — not independently verified]
Read this first
- “Leader on the Decision Index” is Cloudflare’s own run. The official Decision Index still carries data from 28/09 and does not list Clef. On the same day Fastino claimed 64.81 for GLiDE, also self-run.
- The “13× faster than Jev” line (in Cloudflare’s changelog and from @lucataco, who works on models at Cloudflare) mixes two setups. Clef-flash’s 38.8 ms is single-process GPU time; Jev’s 524 ms is a hosted-API round trip, and Cloudflare’s own leaderboard says the two are “not comparable”. The first independent test of the hosted endpoints (below) found the opposite order: Jev about 100–170 ms, Clef about 640–720 ms.
- Price is per input token only; no output price is listed (decisions generate no text). Per input token, Clef costs about 5.7× Jev ($0.24 vs $0.042 per M) and Clef-flash about 2.1× ($0.09). Clef’s 65,536-token window is double the 32K that OpenRouter lists for Jev.
Clef is a family of decision models from Cloudflare, announced on the Cloudflare blog on 1 October 2026. The weights went up on Hugging Face the evening before (30/09, 18:15 BRT) under Apache-2.0: Clef (27B, from Qwen3.8-27B) and Clef-flash (9B, from Qwen3.5-9B). Both are also served on Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash, and since 2 October they are in the official Ollama library.
The model reads a state (text, JSON, images or video) and a schema of typed questions, and returns a probability for every allowed option in a single forward pass, with no text generation. Cloudflare says the API is “fully compatible with Jev and SystemOne”.
Specs
| Attribute | Clef | Clef-flash |
|---|---|---|
| Base | Qwen/Qwen3.8-27B (vision encoder kept) | Qwen/Qwen3.5-9B |
| Training | Frozen backbone + rank-256 LoRA (stored merged), separate joint schema head; label-smoothed cross-entropy + Brier loss, RLCD as a secondary objective; internal synthetic data | same recipe |
| Workers AI price | $0.24 / M input tokens | $0.09 / M input tokens |
| Context | 65,536 tokens | 65,536 tokens |
| Questions per call | 1–64 (Noul, Choice, Score) | 1–64 |
| Images | up to 4 per request (PNG/JPEG/WebP, embedded) | up to 4 |
| Run locally | Ollama clef (18 GB); Ollaya; llama.cpp (text only) |
Ollama clef-flash (11 GB); Ollaya clef:flash; llama.cpp (text only) |
| License | Apache-2.0 | Apache-2.0 |
| Launch | same |
The Hugging Face release ships the backbone as sharded safetensors plus joint_head.safetensors and a joint_schema_model.py loader with a systemone() helper. On Workers AI the call goes to /ai/run/@cf/cloudflare/clef, not to a /v1/systemone path; Ollama serves the same models at /v1/systemone.
Running it yourself
- Ollama.
clefandclef-flashare in the official library, with native support merged in ollama#18741 on 1 October (19:52 BRT). Requests go to/v1/systemoneand may carry base64 images. Two details on that page do not line up yet: it says “requires Ollama 0.35.1 or later”, but the support was merged after the 0.35.1 tag, and the tag list shows a 256K context while the readme says 64K. - llama.cpp. Support was merged in #29831 on 2 October (21:50 BRT): text only, one sequence per batch, on top of the
/v1/systemoneendpoint. Theggml-org/Clef-GGUFfiles need a build that includes it. Image input is not supported there yet. - Ollaya.
ollaya-dev/clefpackages Clef-flash as ONNX graphs that reference Cloudflare’s own weight files. The packager reports identical decisions to Cloudflare’s code on 571 questions (logits within 4.3e-5; UnverifiedClaim). - Plain GGUF in a chat runtime is not the same model. A GGUF holds only the backbone. Run as a chat model in Ollama 0.35.0, a third-party test got constrained JSON, not the joint schema head’s probabilities.
Evidence (UnverifiedClaim)
From Cloudflare’s Decision Index 0.2.1 mirror (38 benchmarks, upstream board of 28/09, Clef rows added on 01/10):
| Model | Index | Knowledge | Language | Retrieval | Tools | Arts | Median / p95 |
|---|---|---|---|---|---|---|---|
| Clef | 61.21 | 51.2 | 61.2 | 62.5 | 81.2 | 47.8 | 209 / 239 ms |
| Jev 1.13 (official row) | 57.91 | 51.4 | 62.0 | 55.4 | 75.1 | 37.7 | 524 / 536 ms (hosted API) |
| Clef-flash | 57.07 | 50.4 | 52.6 | 52.3 | 82.0 | 49.9 | 39 / 122 ms |
ECE and Brier are blank on that board. The blog also reports TypeSafe workflow evals in which Clef beats Jev on 3 of 4 tasks.
Independent results
The first outside measurements came in within a day. Both are small and task-specific, so they are evidence about those tasks, not a general ranking.
- Nicia, knowledge-base write admission (repo; 85 synthetic cases, hosted APIs, runs 1–2 October). Clef was close to Jev on quality (recall 0.98 / 1.00, routine admits 0.97) but two of its decisions changed when the records were reordered. Clef-flash admitted only about two thirds of harmless writes. Round-trip latency, one request at a time: Jev 98–167 ms median, Clef-flash 429–533 ms, Clef 638–717 ms (p95 0.9–1.1 s), with a slow tail of 9–26 s. A few hours after launch about a fifth of Clef calls at six in flight returned HTTP 429; by the later runs there were none. Clef returned identical probabilities for identical requests; Jev moved by up to 0.12.
- Reid Marlow, coding-agent decisions (DEV; 42 real cases from one agent). Jev 71.4%, Clef-flash 66.7%; the two agreed on 35 of 42.
Fit / anti-fit
Fit: teams already on Cloudflare Workers who want a hosted decision call close to their app; multimodal checks (screenshots, documents with images); anyone who wants the same weights hosted, in Ollama and self-hosted.
Anti-fit: workloads that need an SLA today or a verified score; latency-critical hosted calls (independent round trips are several hundred ms); drop-in use of the TypeSafe SDK against the Workers AI endpoint without an adapter.
Limits
- No ECE or reliability figure published.
- Answers can depend on the order of the records in the state (Nicia’s test).
- “Still early”: the fine-tuning and RL service runs through Cloudflare’s forward-deployed team for now; self-serve comes later.
- Training data is internal and synthetic, so it cannot be audited.
Why it is in the catalog
Clef trains its own decision components (LoRA plus a new joint schema head), publishes the weights and serves them through a documented typed API that returns probabilities. Clef-flash shares the page: same recipe and API, smaller backbone. On the quadrant the point is the 27B Clef.
JevBench v1.5.7 (Benchmark Heaven, checked 5 Oct 2026)
On v1.5.7, Clef-Flash is #26 at official 55.1 and Clef is #63 at 17.0 (text-only measurement of multimodal models). See news. No catalog score change.
Links
- Cloudflare blog: Clef decision models · changelog · Hacker News thread
- Workers AI docs: clef · clef-flash
- Hugging Face: Cloudflare/clef · Cloudflare/clef-flash
- Ollama: clef · clef-flash · llama.cpp #29831 (merged)
- Cloudflare’s Decision Index mirror
- Nicia admission eval · Reid Marlow on DEV
- Launch post (@CloudflareDev) · local demo on an M5 Max (@lucataco, Cloudflare)
- News: launch · Ollama, llama.cpp and the first independent tests
- Contribute / discuss
