Clef / Clef-flash

Last updated:

Open39–239ms$0.24/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity7.2/10

    Apache-2.0 weights (27B, 9B) plus a priced, self-serve Workers AI endpoint with full schema docs. Within a day: official Ollama library entries on /v1/systemone, an Ollaya package and llama.cpp support (PR merged 2 Oct). No GA or model-specific SLA.

  • Capability7.5/10

    Noul / Choice / Score plus up to 4 images in one pass. Calibration method published (Brier, label smoothing), no ECE. Cloudflare's GPU run: 209 ms median. First independent test of the hosted API: 638–717 ms median, ~1 s p95. DI 61.21 is self-run, not on the official board.

  • Adoption37/100

    610 points on Hacker News, a 1,245-like launch post from @CloudflareDev, 706 + 237 Hugging Face likes, 33 quantizations (ggml-org among them), Ollama and Ollaya packages and about 15 community repos in 36 hours. No external production case published yet.

Vendor claims

Read this first

  1. “Leader on the Decision Index” is Cloudflare’s own run. The official Decision Index still carries data from 28/09 and does not list Clef. On the same day Fastino claimed 64.81 for GLiDE, also self-run.
  2. The “13× faster than Jev” line (in Cloudflare’s changelog and from @lucataco, who works on models at Cloudflare) mixes two setups. Clef-flash’s 38.8 ms is single-process GPU time; Jev’s 524 ms is a hosted-API round trip, and Cloudflare’s own leaderboard says the two are “not comparable”. The first independent test of the hosted endpoints (below) found the opposite order: Jev about 100–170 ms, Clef about 640–720 ms.
  3. Price is per input token only; no output price is listed (decisions generate no text). Per input token, Clef costs about 5.7× Jev ($0.24 vs $0.042 per M) and Clef-flash about 2.1× ($0.09). Clef’s 65,536-token window is double the 32K that OpenRouter lists for Jev.

Clef is a family of decision models from Cloudflare, announced on the Cloudflare blog on 1 October 2026. The weights went up on Hugging Face the evening before (30/09, 18:15 BRT) under Apache-2.0: Clef (27B, from Qwen3.8-27B) and Clef-flash (9B, from Qwen3.5-9B). Both are also served on Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash, and since 2 October they are in the official Ollama library.

The model reads a state (text, JSON, images or video) and a schema of typed questions, and returns a probability for every allowed option in a single forward pass, with no text generation. Cloudflare says the API is “fully compatible with Jev and SystemOne”.

Specs

Attribute Clef Clef-flash
Base Qwen/Qwen3.8-27B (vision encoder kept) Qwen/Qwen3.5-9B
Training Frozen backbone + rank-256 LoRA (stored merged), separate joint schema head; label-smoothed cross-entropy + Brier loss, RLCD as a secondary objective; internal synthetic data same recipe
Workers AI price $0.24 / M input tokens $0.09 / M input tokens
Context 65,536 tokens 65,536 tokens
Questions per call 1–64 (Noul, Choice, Score) 1–64
Images up to 4 per request (PNG/JPEG/WebP, embedded) up to 4
Run locally Ollama clef (18 GB); Ollaya; llama.cpp (text only) Ollama clef-flash (11 GB); Ollaya clef:flash; llama.cpp (text only)
License Apache-2.0 Apache-2.0
Launch same

The Hugging Face release ships the backbone as sharded safetensors plus joint_head.safetensors and a joint_schema_model.py loader with a systemone() helper. On Workers AI the call goes to /ai/run/@cf/cloudflare/clef, not to a /v1/systemone path; Ollama serves the same models at /v1/systemone.

Running it yourself

  • Ollama. clef and clef-flash are in the official library, with native support merged in ollama#18741 on 1 October (19:52 BRT). Requests go to /v1/systemone and may carry base64 images. Two details on that page do not line up yet: it says “requires Ollama 0.35.1 or later”, but the support was merged after the 0.35.1 tag, and the tag list shows a 256K context while the readme says 64K.
  • llama.cpp. Support was merged in #29831 on 2 October (21:50 BRT): text only, one sequence per batch, on top of the /v1/systemone endpoint. The ggml-org/Clef-GGUF files need a build that includes it. Image input is not supported there yet.
  • Ollaya. ollaya-dev/clef packages Clef-flash as ONNX graphs that reference Cloudflare’s own weight files. The packager reports identical decisions to Cloudflare’s code on 571 questions (logits within 4.3e-5; UnverifiedClaim).
  • Plain GGUF in a chat runtime is not the same model. A GGUF holds only the backbone. Run as a chat model in Ollama 0.35.0, a third-party test got constrained JSON, not the joint schema head’s probabilities.

Evidence (UnverifiedClaim)

From Cloudflare’s Decision Index 0.2.1 mirror (38 benchmarks, upstream board of 28/09, Clef rows added on 01/10):

Model Index Knowledge Language Retrieval Tools Arts Median / p95
Clef 61.21 51.2 61.2 62.5 81.2 47.8 209 / 239 ms
Jev 1.13 (official row) 57.91 51.4 62.0 55.4 75.1 37.7 524 / 536 ms (hosted API)
Clef-flash 57.07 50.4 52.6 52.3 82.0 49.9 39 / 122 ms

ECE and Brier are blank on that board. The blog also reports TypeSafe workflow evals in which Clef beats Jev on 3 of 4 tasks.

Independent results

The first outside measurements came in within a day. Both are small and task-specific, so they are evidence about those tasks, not a general ranking.

  • Nicia, knowledge-base write admission (repo; 85 synthetic cases, hosted APIs, runs 1–2 October). Clef was close to Jev on quality (recall 0.98 / 1.00, routine admits 0.97) but two of its decisions changed when the records were reordered. Clef-flash admitted only about two thirds of harmless writes. Round-trip latency, one request at a time: Jev 98–167 ms median, Clef-flash 429–533 ms, Clef 638–717 ms (p95 0.9–1.1 s), with a slow tail of 9–26 s. A few hours after launch about a fifth of Clef calls at six in flight returned HTTP 429; by the later runs there were none. Clef returned identical probabilities for identical requests; Jev moved by up to 0.12.
  • Reid Marlow, coding-agent decisions (DEV; 42 real cases from one agent). Jev 71.4%, Clef-flash 66.7%; the two agreed on 35 of 42.

Fit / anti-fit

Fit: teams already on Cloudflare Workers who want a hosted decision call close to their app; multimodal checks (screenshots, documents with images); anyone who wants the same weights hosted, in Ollama and self-hosted.

Anti-fit: workloads that need an SLA today or a verified score; latency-critical hosted calls (independent round trips are several hundred ms); drop-in use of the TypeSafe SDK against the Workers AI endpoint without an adapter.

Limits

  • No ECE or reliability figure published.
  • Answers can depend on the order of the records in the state (Nicia’s test).
  • “Still early”: the fine-tuning and RL service runs through Cloudflare’s forward-deployed team for now; self-serve comes later.
  • Training data is internal and synthetic, so it cannot be audited.

Why it is in the catalog

Clef trains its own decision components (LoRA plus a new joint schema head), publishes the weights and serves them through a documented typed API that returns probabilities. Clef-flash shares the page: same recipe and API, smaller backbone. On the quadrant the point is the 27B Clef.

JevBench v1.5.7 (Benchmark Heaven, checked 5 Oct 2026)

On v1.5.7, Clef-Flash is #26 at official 55.1 and Clef is #63 at 17.0 (text-only measurement of multimodal models). See news. No catalog score change.