Winnow-12B

Last updated:

Open22–53ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity6.6/10

    Apache-2.0 merged GGUF fine-tune of Gemma-4-12B-it on HF with a dedicated llama.cpp-based server that exposes POST /v1/systemone plus chat and vision. Extremely detailed card, BENCHMARKS.md and Quickstart. Independent author, no hosted API or SLA.

  • Capability8.4/10

    Noul, Choice and Score from shared-state logit readout. Author-run ECE/Brier tables without temperature fitted on those evals (JevBench public ECE 0.074). Warm B1 median 22.1 ms Q8 on an RTX PRO 5000 [vendor]. Official JevBench v1.5.7 #2 at 73.2 (capability 79.3) verified on Benchmark Heaven — separate from the…

  • Adoption11/100

    About 30k HF downloads and 48 likes, and 22 GitHub stars on winnow-inference (5 Oct). Own inference server, GGUF presets (Q8/BF16/NVFP4) and a board page. No independent production case study beyond a Benchmark Heaven note that Winnow powers System1 Models s1-pro.

Vendor claims

Read this first

  1. Two different JevBench numbers. The card’s 85.71% is accuracy on the 231-item public subset, measured by Eldan Ring against hosted Jev. The 73.2 on Benchmark Heaven v1.5.7 is the official composite score (#2, Q8). Do not mix them. The subset result is an UnverifiedClaim; the board score we verified live.
  2. Use the Winnow server, not a generic llama.cpp snippet. Generic Hub snippets do not expose /v1/systemone. Decisions are native logit readouts; chat and vision share the same loaded GGUF. Direct decisions are the default — optional same-model “reasoning” and MTP are experimental client workflows.
  3. Training data is private. The author says earlier Kev-v4 panels were used during development; treat author-run suites as non-held-out for those families. The official v1.5.7 board run is a separate measurement.

What it is

Winnow-12B is an independent LoRA fine-tune of Gemma 4 12B IT, merged and shipped as GGUF (Q8_0 recommended, also BF16 and NVFP4) under Apache-2.0. The matching winnow-inference server (MIT, llama.cpp-based) serves typed decisions on POST /v1/systemone and ordinary chat/vision on /v1/chat/completions from one loaded model.

Type How it works
Choice / Noul / Score Shared-state prefill, forked question branches, answer-token logits — no answer text generated
Chat + vision Same weights; optional mmproj-Winnow-12B.gguf projector for images

Point of this page = Winnow-12B Q8 (the board row and the author’s recommended 64K vision setup). BF16 and NVFP4 are the same fine-tune at other precisions.

Specs

Attribute Value
Author Eldan Ring (independent)
Base Gemma-4-12B-it + merged LoRA
License Apache-2.0 (NOTICE cites Gemma 4 Apache upstream + Winnow modifications)
Weights GGUF only (no safetensors) — HF
Runtime EldanRing/winnow-inference — CUDA and Apple Silicon
Launch (HF)
Latency Warm B1 22.1 ms Q8 median (RTX PRO 5000) — UnverifiedClaim
Pricing Self-host (free weights)
Context Q8 tested at 64K with vision on 16 GB (5070 Ti profile)

Numbers

Official board (verified by us, 5 Oct 2026)

Edition System Official score Capability Rank
JevBench v1.5.7 Winnow-12B Q8 73.2 79.3 #2

Statistical co-leader with Cygnet (73.7) per the site’s own note. Capability is a separate headline from official score.

Author-run panels (UnverifiedClaim)

Suite Winnow Q8 Winnow BF16 Hosted Jev 1.13 (OpenRouter)
JevBench public 231 85.71% 85.28% 85.71%
Kev-v9 clean 1,046 81.55% 81.45% 87.00%
Typed teacher agreement 2,000 70.00% 70.20% 73.80%
JevBench public ECE (↓ better) 0.0742 0.0655 0.0474

Source: docs/BENCHMARKS.md (measured 21 Sep 2026 on RTX PRO 5000). No temperature was fitted on those evaluation outputs.

Fit / anti-fit

Fit when you need: a strong open local System One on a single consumer/pro GPU, TypeSafe-compatible /v1/systemone, and optional chat/vision from the same weights.

Anti-fit when you need: a hosted API with SLA, safetensors/vLLM-first serving, or a guarantee that author benches are held-out from every related family.

Alerts

  • GGUF-first: the Hub auto-snippet may point at the small MTP assistant file — pick the explicit Winnow target GGUF (Q8/BF16/NVFP4).
  • NVFP4 / reasoning / MTP: smaller footprint and optional adaptive recipes trade quality; direct native decisions are the default contract.
  • s1-pro: Benchmark Heaven notes Winnow as the model behind System1 Models s1-pro; we have not audited that product — treat as a lead, not a production proof for this page.

Score working (5 Oct 2026)

  • Maturity = availability 8 × 0.30 + docs 9 × 0.25 + integrations 7 × 0.25 + support 1 × 0.20 = 2.40 + 2.25 + 1.75 + 0.20 = 6.60
  • Capability = types 10 × 0.30 + calibration 7 × 0.25 + latency 8 × 0.25 [vendor] + accuracy 8 × 0.20 = 3.00 + 1.75 + 2.00 + 1.60 = 8.35 → 8.4
  • Adoption = engagement 20 × 0.35 + ecosystem 10 × 0.35 + production 2 × 0.30 = 7.00 + 3.50 + 0.60 = 11.10 → 11

Accuracy 8: official public board (#2 at 73.2), not the author’s public-subset %. Calibration stays 7 (published ECE/method, author-run).