JevBench v1.5.7 on Benchmark Heaven: Cygnet and Winnow lead; Jev is #3 official and leads the new capability headline

Last updated:

Source

Benchmark Heaven is showing JevBench v1.5.7 with 111 ranked systems (page checked 5 Oct 2026). The site’s own blurb calls Cygnet and Winnow-12B Q8 the statistical co-leaders, measured across 1,624 decisions per system. We read the live board and confirmed the numbers below.

Official score (top 5)

Rank System Official score Capability (new headline)
1 Cygnet (blockbrain, frozen Gemma-4-12B-it) 73.7 79.0
2 Winnow-12B Q8 (EldanRing/Winnow-12B) 73.2 79.3
3 Jev 1.13.0 (TypeSafe AI) 72.1 80.0
4 JevK5 v0.3 (4B) 71.9 72.3
5 Vansa-3.4 (hosted System One API) 71.6 72.8

Capability is a separate ranking. On that headline Jev sits at 80.0, tied with SimpleJev Qwen3.8-27B in the broader list we scraped, and ahead of Cygnet (79.0) and Winnow (79.3). Official score and capability do not sort the same way.

Catalog models on this board

Our page Board row Official Notes
Jev Jev 1.13.0 72.1 (#3) Capability 80.0
Winnow-12B Winnow-12B Q8 73.2 (#2) Capability 79.3; new catalog page in this radar
JevK5 JevK5 v0.3 (4B) 71.9 (#4) Was #3 on v1.4.2 at 62.04 — different edition
ImaJev-4B Imajev-4B (RTX 5090) 70.4 (#10)
Lev lev (Interfaze) 69.1 (#12)
Hopper Hopper 67.5 (#17) Distinct from Hopper (G) on the Decision Index
Surogate Rune Surogate Rune 26B-A4B v3 66.5 (#19)
Clef Clef-Flash / Clef 55.1 (#26) / 17.0 (#63) Multimodal models measured on text
AutoJev-27B AutoJev-27B 19.5 (#55)
Laya Laya variants 0.0 (#91–#111) Board shows zero official score on these rows
Instinct family (not the tuned card) Instinct Dual 4B / Instinct 27B 47.0 (#30) / 18.3 (#60) Frozen readouts; Instinct Tuned 4B is not on the board yet (issue #151)

Names we are not cataloging yet

Name Class for now Why
Cygnet notmodels / method (watch) Frozen Gemma-4-12B-it recipe (blockbrain-ai/cygnet-recipe); already flagged as method when it appeared as an Ollaya package. #1 official score without trained decision weights.
Vansa-3.4 watch Hosted System One API at vansa.org; no open weights checked.
Malkuth-4B / 2B watch Kev post-trains (newfull5/malkuth); #18 at 66.8 for the 4B.
Manchego v2.1 watch #13 at 68.8.

What changes here

Evidence sections on the catalog pages above now cite v1.5.7. No score changes in this batch — the board edition moved, and capability is a new axis we have not folded into the rubric yet. Three new catalog pages in the same radar: Instinct Tuned 4B, Gero-4B, and Winnow-12B (the #2 board row).