Winnow-12B
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity6.6/10
Apache-2.0 merged GGUF fine-tune of Gemma-4-12B-it on HF with a dedicated llama.cpp-based server that exposes POST /v1/systemone plus chat and vision. Extremely detailed card, BENCHMARKS.md and Quickstart. Independent author, no hosted API or SLA.
- Capability8.4/10
Noul, Choice and Score from shared-state logit readout. Author-run ECE/Brier tables without temperature fitted on those evals (JevBench public ECE 0.074). Warm B1 median 22.1 ms Q8 on an RTX PRO 5000 [vendor]. Official JevBench v1.5.7 #2 at 73.2 (capability 79.3) verified on Benchmark Heaven — separate from the…
- Adoption11/100
About 30k HF downloads and 48 likes, and 22 GitHub stars on winnow-inference (5 Oct). Own inference server, GGUF presets (Q8/BF16/NVFP4) and a board page. No independent production case study beyond a Benchmark Heaven note that Winnow powers System1 Models s1-pro.
Vendor claims
- JevBench public subset (231 items): Winnow-12B Q8 198/231 = 85.71%, matching hosted Jev 1.13 via OpenRouter on the author's localhost comparison; BF16 85.28% — public-subset accuracy, not the official composite score[Vendor claim — not independently verified]
- Kev-v9 clean (1,046): Q8 81.55%, BF16 81.45%; Typed teacher agreement (2,000): Q8 70.00%, BF16 70.20% — author-run on RTX PRO 5000[Vendor claim — not independently verified]
- Calibration (author-run, no temperature fitted on these outputs): JevBench public ECE 0.0742 (Q8) / 0.0655 (BF16); Kev-v9 ECE 0.0980 / 0.0989[Vendor claim — not independently verified]
- Warm short-state B1 median latency Q8 22.1 ms, cold 52.7 ms on RTX PRO 5000 Blackwell (localhost); 64K+vision Q8 peak 15.01 GiB on RTX 5070 Ti 16 GB[Vendor claim — not independently verified]
- Official JevBench v1.5.7 (Benchmark Heaven): Winnow-12B Q8 official score 73.2 (#2), capability 79.3 — independent board, not author-run[Vendor claim — not independently verified]
Read this first
- Two different JevBench numbers. The card’s 85.71% is accuracy on the 231-item public subset, measured by Eldan Ring against hosted Jev. The 73.2 on Benchmark Heaven v1.5.7 is the official composite score (#2, Q8). Do not mix them. The subset result is an UnverifiedClaim; the board score we verified live.
- Use the Winnow server, not a generic llama.cpp snippet. Generic Hub snippets do not expose
/v1/systemone. Decisions are native logit readouts; chat and vision share the same loaded GGUF. Direct decisions are the default — optional same-model “reasoning” and MTP are experimental client workflows.- Training data is private. The author says earlier Kev-v4 panels were used during development; treat author-run suites as non-held-out for those families. The official v1.5.7 board run is a separate measurement.
What it is
Winnow-12B is an independent LoRA fine-tune of Gemma 4 12B IT, merged and shipped as GGUF (Q8_0 recommended, also BF16 and NVFP4) under Apache-2.0. The matching winnow-inference server (MIT, llama.cpp-based) serves typed decisions on POST /v1/systemone and ordinary chat/vision on /v1/chat/completions from one loaded model.
| Type | How it works |
|---|---|
| Choice / Noul / Score | Shared-state prefill, forked question branches, answer-token logits — no answer text generated |
| Chat + vision | Same weights; optional mmproj-Winnow-12B.gguf projector for images |
Point of this page = Winnow-12B Q8 (the board row and the author’s recommended 64K vision setup). BF16 and NVFP4 are the same fine-tune at other precisions.
Specs
| Attribute | Value |
|---|---|
| Author | Eldan Ring (independent) |
| Base | Gemma-4-12B-it + merged LoRA |
| License | Apache-2.0 (NOTICE cites Gemma 4 Apache upstream + Winnow modifications) |
| Weights | GGUF only (no safetensors) — HF |
| Runtime | EldanRing/winnow-inference — CUDA and Apple Silicon |
| Launch | (HF) |
| Latency | Warm B1 22.1 ms Q8 median (RTX PRO 5000) — UnverifiedClaim |
| Pricing | Self-host (free weights) |
| Context | Q8 tested at 64K with vision on 16 GB (5070 Ti profile) |
Numbers
Official board (verified by us, 5 Oct 2026)
| Edition | System | Official score | Capability | Rank |
|---|---|---|---|---|
| JevBench v1.5.7 | Winnow-12B Q8 | 73.2 | 79.3 | #2 |
Statistical co-leader with Cygnet (73.7) per the site’s own note. Capability is a separate headline from official score.
Author-run panels (UnverifiedClaim)
| Suite | Winnow Q8 | Winnow BF16 | Hosted Jev 1.13 (OpenRouter) |
|---|---|---|---|
| JevBench public 231 | 85.71% | 85.28% | 85.71% |
| Kev-v9 clean 1,046 | 81.55% | 81.45% | 87.00% |
| Typed teacher agreement 2,000 | 70.00% | 70.20% | 73.80% |
| JevBench public ECE (↓ better) | 0.0742 | 0.0655 | 0.0474 |
Source: docs/BENCHMARKS.md (measured 21 Sep 2026 on RTX PRO 5000). No temperature was fitted on those evaluation outputs.
Fit / anti-fit
Fit when you need: a strong open local System One on a single consumer/pro GPU, TypeSafe-compatible /v1/systemone, and optional chat/vision from the same weights.
Anti-fit when you need: a hosted API with SLA, safetensors/vLLM-first serving, or a guarantee that author benches are held-out from every related family.
Alerts
- GGUF-first: the Hub auto-snippet may point at the small MTP assistant file — pick the explicit Winnow target GGUF (Q8/BF16/NVFP4).
- NVFP4 / reasoning / MTP: smaller footprint and optional adaptive recipes trade quality; direct native decisions are the default contract.
- s1-pro: Benchmark Heaven notes Winnow as the model behind System1 Models s1-pro; we have not audited that product — treat as a lead, not a production proof for this page.
Score working (5 Oct 2026)
- Maturity = availability 8 × 0.30 + docs 9 × 0.25 + integrations 7 × 0.25 + support 1 × 0.20 = 2.40 + 2.25 + 1.75 + 0.20 = 6.60
- Capability = types 10 × 0.30 + calibration 7 × 0.25 + latency 8 × 0.25
[vendor]+ accuracy 8 × 0.20 = 3.00 + 1.75 + 2.00 + 1.60 = 8.35 → 8.4 - Adoption = engagement 20 × 0.35 + ecosystem 10 × 0.35 + production 2 × 0.30 = 7.00 + 3.50 + 0.60 = 11.10 → 11
Accuracy 8: official public board (#2 at 73.2), not the author’s public-subset %. Calibration stays 7 (published ECE/method, author-run).
