JevBench v1.5.7 on Benchmark Heaven: Cygnet and Winnow lead; Jev is #3 official and leads the new capability headline
Benchmark Heaven is showing JevBench v1.5.7 with 111 ranked systems (page checked 5 Oct 2026). The site’s own blurb calls Cygnet and Winnow-12B Q8 the statistical co-leaders, measured across 1,624 decisions per system. We read the live board and confirmed the numbers below.
Official score (top 5)
| Rank | System | Official score | Capability (new headline) |
|---|---|---|---|
| 1 | Cygnet (blockbrain, frozen Gemma-4-12B-it) | 73.7 | 79.0 |
| 2 | Winnow-12B Q8 (EldanRing/Winnow-12B) | 73.2 | 79.3 |
| 3 | Jev 1.13.0 (TypeSafe AI) | 72.1 | 80.0 |
| 4 | JevK5 v0.3 (4B) | 71.9 | 72.3 |
| 5 | Vansa-3.4 (hosted System One API) | 71.6 | 72.8 |
Capability is a separate ranking. On that headline Jev sits at 80.0, tied with SimpleJev Qwen3.8-27B in the broader list we scraped, and ahead of Cygnet (79.0) and Winnow (79.3). Official score and capability do not sort the same way.
Catalog models on this board
| Our page | Board row | Official | Notes |
|---|---|---|---|
| Jev | Jev 1.13.0 | 72.1 (#3) | Capability 80.0 |
| Winnow-12B | Winnow-12B Q8 | 73.2 (#2) | Capability 79.3; new catalog page in this radar |
| JevK5 | JevK5 v0.3 (4B) | 71.9 (#4) | Was #3 on v1.4.2 at 62.04 — different edition |
| ImaJev-4B | Imajev-4B (RTX 5090) | 70.4 (#10) | |
| Lev | lev (Interfaze) | 69.1 (#12) | |
| Hopper | Hopper | 67.5 (#17) | Distinct from Hopper (G) on the Decision Index |
| Surogate Rune | Surogate Rune 26B-A4B v3 | 66.5 (#19) | |
| Clef | Clef-Flash / Clef | 55.1 (#26) / 17.0 (#63) | Multimodal models measured on text |
| AutoJev-27B | AutoJev-27B | 19.5 (#55) | |
| Laya | Laya variants | 0.0 (#91–#111) | Board shows zero official score on these rows |
| Instinct family (not the tuned card) | Instinct Dual 4B / Instinct 27B | 47.0 (#30) / 18.3 (#60) | Frozen readouts; Instinct Tuned 4B is not on the board yet (issue #151) |
Names we are not cataloging yet
| Name | Class for now | Why |
|---|---|---|
| Cygnet | notmodels / method (watch) | Frozen Gemma-4-12B-it recipe (blockbrain-ai/cygnet-recipe); already flagged as method when it appeared as an Ollaya package. #1 official score without trained decision weights. |
| Vansa-3.4 | watch | Hosted System One API at vansa.org; no open weights checked. |
| Malkuth-4B / 2B | watch | Kev post-trains (newfull5/malkuth); #18 at 66.8 for the 4B. |
| Manchego v2.1 | watch | #13 at 68.8. |
What changes here
Evidence sections on the catalog pages above now cite v1.5.7. No score changes in this batch — the board edition moved, and capability is a new axis we have not folded into the rubric yet. Three new catalog pages in the same radar: Instinct Tuned 4B, Gero-4B, and Winnow-12B (the #2 board row).
