AutoTrust republishes its Gemma decision model as GEV-26B-Decide, with adaptive thinking
AutoTrust AI published autotrust/GEV-26B-Decide on 2 October 2026 at 01:20 (Brasília time). It added demo results on 3 October. The card says the model “was previously published as autotrust/JEV-Gemma4-26B-A4B; the weights are the same.” We compared the files: the adapter, the decision head and the base shards have identical hashes. It is a new repository, not a rename on Hugging Face. The old repo is still online and has not changed since 29 September, and JEV-27B, JEV-9B and JEV-27B-VL keep their names.
What is new. The decision weights are the same. What changed is everything around them:
- Adaptive thinking (opt-in). System 1 answers first, in one pass (about 45 ms on a B200). If its top option is below 0.8 confidence, the Gemma base model thinks and its answer is averaged in.
- A vLLM server with a
POST /v1/decideroute. It needs a patched vLLM development build and an 80 GB GPU. - Two demos: computer use (pick the next click from a screenshot) and a simulated robot arm.
AutoTrust’s numbers (UnverifiedClaim).
| Claim | Result | Note |
|---|---|---|
| Decision Index 0.2.1 | 62.48 vs Jev 57.91 | Own scoring. Knowledge & Reasoning uses adaptive thinking (median 13.4 s per request; the board requires ≤ 1 s). The same weights scored 58.05 with System 1 only |
| Validation, 1,754 questions outside the index | 73.3% → 83.4% with thinking, ECE 0.035 | Thinks on 48% of questions |
| Reasoning benchmarks | GPQA Diamond 42.9 → 78.6, MMLU-Pro 65.0 → 84.6, BBH 75.0 → 92.0 | HLE only rises from below chance to chance (8.4 → 17.8) |
| Computer use, 60 browser tasks | 95% at ≈ 85 ms per click | 15% without element text; JEV-27B-VL also 95%, at 260 ms |
| Robot arm, 20 simulated scenes | 40% at 61 ms per decision | JEV-27B-VL 75%; 0/10 when choosing motor commands directly |
| Puzzles | Minesweeper 21 → 86%, Wordle 52 → 100%, Sudoku 76 → 99.5% | Thinking does not help whole games or sketches |
What to keep in mind.
- Training-data overlap. The card discloses that training used the training splits of benchmarks whose test splits are in the Decision Index (BANKING77 and CLINC150 among them). Thinking lowered BANKING77 (88.0 → 85.0).
- Teacher distillation. System 1 learned from “teacher distributions”, and for AutoTrust’s JEV family the teacher is the closed Jev 1.13. The demos are small simulations AutoTrust built itself.
- Downloads. The new repo shows 310,028 downloads in two days against 7 likes. We do not count that number.
- License. The license tag says Apache-2.0, but that only covers the adapter, head and calibration files. The base weights are under the Gemma 4 terms.
- Not TypeSafe. AutoTrust is not affiliated with TypeSafe.
What changes here. AutoTrust’s catalog entry is now GEV-26B-Decide, scored 6.3 / 7.8 / 10 (maturity / capability / adoption). It replaces the family page, which was scored 6.1 / 7.5 / 9 with JEV-27B as the point. The JEV models and their caveats carried over to the new page. The old URL, /models/autotrust-jev, still works as a rename note, but it is no longer listed in the catalog or on the quadrant.
