AutoTrust GEV-26B-Decide
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity6.3/10
Apache-2.0 adapter, head and calibration files on Gemma-4-26B-A4B (Gemma 4 terms). Very detailed card with reports, two calibration tables and a training-data disclosure. The own POST /v1/decide server needs a patched vLLM dev build and an 80 GB GPU; no /v1/systemone. Company, no SLA.
- Capability7.8/10
Noul, Choice (2–256) and Score (0–5) over text and images. System 1: 45 ms median on a B200, no p99. ECE 0.035 against gold labels on 1,754 questions outside the DI. The DI 62.48 is AutoTrust's own scoring: adaptive mode breaks the board's latency rule, and some DI training splits were used.
- Adoption10/100
7 likes on the new repo and 13 on the JEV-Gemma4 repo it republishes; about 156 likes across the AutoTrust family, an HF blog, a press release and a third-party GGUF of JEV-27B. Download counts look anomalous (310k in two days) and are not counted. No production use reported.
Vendor claims
- Decision Index 0.2.1 62.48 (balanced raw 70.66, breadth 62.00) vs Jev 1.13 57.91: own scoring, Knowledge & Reasoning with adaptive thinking (0.602), other four areas System 1; not a board entry[Vendor claim — not independently verified]
- Same weights, System 1 only (published 29 Sep as JEV-Gemma4-26B-A4B): Decision Index 0.2.1 58.05, own complete run[Vendor claim — not independently verified]
- 1,754 questions from six public sets outside the Decision Index: System 1 73.3% → adaptive 83.4%, ECE 0.035 for both[Vendor claim — not independently verified]
- System 1: 45 ms median per request and 257 decisions/s with 64 clients on one B200 (vLLM)[Vendor claim — not independently verified]
- Computer use: 95% of 60 multi-step browser tasks at ≈85 ms per click (15% without element text); robot arm (MuJoCo): 40% of 20 scenes at 61 ms per decision[Vendor claim — not independently verified]
- One-move puzzles, System 1 → adaptive: Minesweeper 21 → 86%, Wordle 52 → 100%, Connect Four 54 → 99.5%, Sudoku 76 → 99.5%, chess mate-in-one 46.5 → 78%[Vendor claim — not independently verified]
- Zero-shot images: VL-RewardBench 78.4% (JEV-27B-VL 78.3%); long context 10/10 at 4K–128K; 255-option intent sets 89.0% / 75.2%[Vendor claim — not independently verified]
Read this first
- New name, same weights, not an HF rename. AutoTrust published
autotrust/GEV-26B-Decideas a new repository on 2 Oct 2026 (01:20 BRT). The card says it “was previously published asautotrust/JEV-Gemma4-26B-A4B; the weights are the same”. We checked: the adapter, the decision head and the base shards have identical LFS sha256 hashes. The old repo is still online and unchanged since 29 Sep. GEV adds a vLLM LoRA (adapter_vllm/), a/v1/decideserver with adaptive thinking, a vLLM patch, new reports and demo videos.- Distilled lineage. System 1 “was trained on teacher distributions and ground-truth decision data”, and “where the teacher is wrong, System 1 often is too”. The GEV card does not name the teacher. For the JEV family it is the closed TypeSafe Jev 1.13, via
SargeDev/jev-distill-corpus-v3(labels “distilled from Jev 1.13 via OpenRouter”). We have not verified whether this use of Jev outputs complies with TypeSafe’s terms.- The 62.48 is not a board score. AutoTrust scored it with the Decision Index kit itself. Knowledge & Reasoning runs with adaptive thinking (median 13.4 s per request, which breaks the board’s ≤ 1,000 ms median rule), and the other four areas run System 1 only. The training data also includes the training splits of benchmarks whose test splits the index uses (BANKING77 and CLINC150 among them); AutoTrust says it gave the list to the index maintainers. As of the 28 Sep snapshot, no AutoTrust model is on the official Decision Index.
- “JEV”/“GEV” is not TypeSafe’s product. AutoTrust’s card says it is “not affiliated with, endorsed by, or a product of TypeSafe AI”. For TypeSafe’s hosted model, see Jev.
What it is
GEV-26B-Decide follows the same recipe as AutoTrust’s other models, which it calls Blocks of Experts. The base, google/gemma-4-26B-A4B-it (26B parameters, about 4B active per token), stays frozen. A trained LoRA plus a 24-slot fp32 decision head adds a typed-decision skill.
| Mode | What it does | Output |
|---|---|---|
| System 1 | Noul (yes/no), Choice over 2–256 options, Score 0–5, over text and images; prompts up to 256K tokens | A probability per option, in one forward pass |
Adaptive thinking (opt-in, thinking: "auto") |
System 1 first. Below 0.8 confidence, the base model thinks and its answer is mixed in: p = ½ p1 + ½ p2 | Probabilities |
| System 2 | The unmodified Gemma-4-26B-A4B-it, with or without thinking | Text |
Read-out: the last token’s final-norm hidden state (2,816) goes through the linear head to 24 slots, soft-capped at 30. Inactive slots are masked, and the logits are divided by a per-kind temperature. Up to 16 options are read in one pass. Longer lists run as a tournament in groups of 16. Two temperature tables ship: calibration.json (1.003 / 1.017 / 0.999, used in the index run) and calibration_gold.json (1.214 / 1.098 / 1.000, fitted on gold answers; AutoTrust recommends it when gating actions on confidence).
Serving
- vLLM:
serve.shrunsserve_decide.py, the standard vLLM OpenAI server plus aPOST /v1/decideroute. Plain requests go to System 2, and requests for the LoRA modulejev-decisiongo to System 1. It needs a vLLM build with Gemma 4 support, the bundledpatches/vllm-gemma4-lm-head-lora.patchand one GPU with 80 GB or more. AutoTrust tested it on a vLLM development build from September 2026. - transformers + peft: System 1 with the adapter merged in memory, up to 16 options per pass.
- Not
/v1/systemone, so the TypeSafe SDK does not apply. No GGUF or Ollama/Ollaya package for GEV as of 4 Oct.
AutoTrust’s numbers (UnverifiedClaim)
All numbers below are AutoTrust’s own runs on one B200, taken from the card (lastModified 3 Oct, 08:00 BRT).
Decision Index 0.2.1 (own scoring).
| Balanced skill | Balanced raw | Breadth | |
|---|---|---|---|
| GEV-26B-Decide, adaptive on K&R | 62.48 | 70.66 | 62.00 |
| Same weights, System 1 only (29 Sep run, as JEV-Gemma4-26B-A4B) | 58.05 | 67.35 | 56.98 |
| TypeSafe Jev 1.13 (board, 28 Sep) | 57.91 | — | — |
Areas: Knowledge & Reasoning 0.602 (System 1: 0.429), Language 0.636, Retrieval & Classification 0.679, Tools & Automation 0.697, Arts & Taste 0.415. Raw results are in autotrust/jev-decision-index-results (runs/jev-gemma4-26b-a4b).
Knowledge & Reasoning, System 1 → adaptive (accuracy %). GPQA Diamond 42.9 → 78.6 · CRUXEval 67.5 → 90.7 · CLadder 71.0 → 86.6 · MMLU-Pro 65.0 → 84.6 · BBH 75.0 → 92.0 · GSM8K 97.6 → 99.1. HLE: 8.4 → 17.8, below chance (16.4) without thinking and back to chance level with it. Of the HLE questions answered without thinking, only 1.6% are right. ChessBench gains nothing.
Outside K&R: thinking helps little (BFCL 94.4 → 96.0, CLINC150 93.3 → 94.4) and hurts BANKING77 (88.0 → 85.0), a set whose training split was in the training data.
Validation set (1,754 questions from six public sets outside the index): System 1 73.3% → adaptive 83.4%, ECE 0.035 for both. It thinks on 48% of the questions. This is the only calibration figure measured against gold labels rather than against the teacher.
Latency: System 1 median 45 ms per request, 257 decisions/s with 64 clients. Adaptive mode on K&R: median 13.4 s, p90 33 s (7.7 s / 19 s with speculative decoding). No p99 is published for System 1.
Demos (3 Oct, System 1).
| GEV-26B-Decide | JEV-27B-VL | JEV-9B | |
|---|---|---|---|
| Computer use, 60 browser tasks (numbered boxes + element text) | 95% | 95% | 95% |
| Time per click | ≈ 85 ms | ≈ 260 ms | ≈ 200 ms |
| Robot arm pick-and-place, 20 MuJoCo scenes | 40% | 75% | 50% |
| Time per robot decision | 61 ms | 239 ms | 163 ms |
Without element text, computer use falls to 15%. Asked to choose one of 8 motor commands directly, the robot arm completed 0 of 10 scenes. These are small, self-built simulations.
Games (200 one-move puzzles each, System 1 → adaptive): Minesweeper 21 → 86 · Wordle 52 → 100 · Connect Four 54 → 99.5 · Sudoku 76 → 99.5 · 24 game 89 → 100 · Maze 48.5 → 68 · chess mate-in-one 46.5 → 78. Thinking does not help whole games (2048, Flappy Bird) or sketch recognition.
Other: zero-shot images on VL-RewardBench 78.4% (the decision head was trained on text). Long context 10/10 at 4K–128K (above 128K untested). Many options: MASSIVE 59 options 91.2%, BANKING77 81.5%, CLINC150 95.5%, 255 merged intents 89.0% / 75.2%.
Family
The other AutoTrust repos are still online and were not renamed (4 Oct). They share the recipe and the caveats above.
| Repo | Base | Created (BRT) | Notes |
|---|---|---|---|
| autotrust/JEV-Gemma4-26B-A4B | gemma-4-26B-A4B-it | 29 Sep, 09:34 | Same weights as GEV; System 1 only; own DI 58.05; unchanged since 29 Sep |
| autotrust/JEV-27B | Qwen3.8-27B | 25 Sep, 09:58 | 108.9M trained params; KL ≈ 0.017 to Jev; 137 ms median; own DI 53.30 |
| autotrust/JEV-9B | Qwen3.5-9B | 23 Sep, 09:08 | First generation; ≈ 90 ms; hosts the demo code |
| autotrust/JEV-27B-VL | Qwen3.8-27B (multimodal) | 30 Sep, 02:02 | Same adapter and head as JEV-27B; image input |
JEV-27B evidence, carried over (UnverifiedClaim).
- Six-benchmark table: 84.07% vs Jev 1.13’s 83.85%. AutoTrust ran both models, and the other rows were copied from the NeoHorse-Jev-4B card.
- Fidelity, not accuracy: KL ≈ 0.017, ECE 0.0009 and 90.5% top-1 agreement are measured against the teacher, on a test split of the distillation corpus.
- Base model untouched: HumanEval stays at 78.0% with and without the block, and all 164 completions are byte-identical to the base.
- Third-party GGUF:
prithivMLmods/JEV-27B-GGUF. - Index gap: its own index run (53.30) sits 4.6 points below Jev, which does not fit the “+0.22 over Jev” headline.
Fit / anti-fit
Fit when you need: typed decisions on your own large GPU with a fast System 1 and an opt-in reasoning path for logic, maths and science questions. Also fits one engine for decisions plus normal Gemma generation, or long contexts and choice lists beyond 16 options.
Anti-fit when you need: sub-second answers with thinking on, the /v1/systemone contract, a small GPU (80 GB for the vLLM path), training data independent of a closed vendor’s outputs, or independently verified accuracy. Avoid thinking on classification and routing: it lowered BANKING77 and only returned HLE to chance.
Red flags we track
- Downloads: 310,028 for the new repo within two days, with 7 likes. The old repo has 1,071 downloads and 13 likes. The same pattern shows on JEV-27B-VL (≈ 726k) and JEV-9B (≈ 132k). We do not count these in adoption.
- License tag: HF metadata says
apache-2.0, but only the adapter, head and calibration files are Apache-2.0. The base weights in the repo are under the Gemma 4 terms. - Two repos, one model: benchmarks and links may cite either name.
Why it is in the catalog
The decision block was trained for typed decisions and ships openly. Its documented serving path returns a probability per option. That passes the catalog criteria as a lora-adapter, because the backbone is untouched. Since 4 Oct this page is AutoTrust’s catalog entry. It replaces the former “AutoTrust JEV” family page (scored 6.1 / 7.5 / 9 with JEV-27B as the point), which is now a rename note.
Score working (4 Oct 2026)
-
Maturity = availability 8 × 0.30 + docs 9 × 0.25 + integrations 5 × 0.25 + support 2 × 0.20 = 2.40 + 2.25 + 1.25 + 0.40 = 6.30
-
Capability = types 10 × 0.30 + calibration 7 × 0.25 + latency 8 × 0.25
[vendor]+ accuracy 5 × 0.20 = 3.00 + 1.75 + 2.00 + 1.00 = 7.75 → 7.8 -
Adoption = engagement 20 × 0.35 + ecosystem 8 × 0.35 + production 2 × 0.30 = 7.00 + 2.80 + 0.60 = 10.40 → 10
-
Why these grades:
- Calibration 7: published temperatures plus ECE 0.035 against gold labels; it is a self-run, and HLE shows confident errors.
- Latency 8: 45 ms median, but on a B200 and with no p99.
- Accuracy 5: one point below our usual 6 for a self-run Decision Index score, because of the training-split overlap with the index and a headline that uses a mode the board would reject.
- Adoption: likes and ecosystem only; the download counts are anomalous.
NVFP4 build (5 Oct 2026)
AutoTrust published GEV-26B-Decide-NVFP4 (about 15k downloads / 150 likes by 7 Oct). Only the 3,840 routed experts are in NVFP4 (ModelOpt layout); everything else stays bf16, including the adapter, the decision head and the temperatures. In vLLM it takes 17.1 GiB instead of 51.1 GiB (−66%), so it fits a 24 GB GPU.
AutoTrust’s check (UnverifiedClaim): System 1 only, on GPQA Diamond (43.9% → 44.9%) and HLE (8.4% → 9.7%), within noise; adaptive thinking was not re-evaluated. Native W4A4 was tested only on B200, and the build needs the same vLLM patch as the bf16 model. Scores on this page are unchanged.
Links
- Hugging Face: autotrust/GEV-26B-Decide · JEV-Gemma4-26B-A4B · JEV-27B · JEV-9B · JEV-27B-VL
- Own Decision Index runs (dataset)
- HF blog: JEV-27B · PR Newswire release (29 Sep 2026)
- Training corpus (JEV family): SargeDev/jev-distill-corpus-v3
- News: AutoTrust republishes its Gemma model as GEV-26B-Decide · Five open decision models from HF trending
- Compare: Jev (TypeSafe AI) · Decision 2.0 · AutoJev-27B (unrelated)
