AutoTrust republishes its Gemma decision model as GEV-26B-Decide, with adaptive thinking

Last updated:

Source

AutoTrust AI published autotrust/GEV-26B-Decide on 2 October 2026 at 01:20 (Brasília time). It added demo results on 3 October. The card says the model “was previously published as autotrust/JEV-Gemma4-26B-A4B; the weights are the same.” We compared the files: the adapter, the decision head and the base shards have identical hashes. It is a new repository, not a rename on Hugging Face. The old repo is still online and has not changed since 29 September, and JEV-27B, JEV-9B and JEV-27B-VL keep their names.

What is new. The decision weights are the same. What changed is everything around them:

  • Adaptive thinking (opt-in). System 1 answers first, in one pass (about 45 ms on a B200). If its top option is below 0.8 confidence, the Gemma base model thinks and its answer is averaged in.
  • A vLLM server with a POST /v1/decide route. It needs a patched vLLM development build and an 80 GB GPU.
  • Two demos: computer use (pick the next click from a screenshot) and a simulated robot arm.

AutoTrust’s numbers (UnverifiedClaim).

Claim Result Note
Decision Index 0.2.1 62.48 vs Jev 57.91 Own scoring. Knowledge & Reasoning uses adaptive thinking (median 13.4 s per request; the board requires ≤ 1 s). The same weights scored 58.05 with System 1 only
Validation, 1,754 questions outside the index 73.3% → 83.4% with thinking, ECE 0.035 Thinks on 48% of questions
Reasoning benchmarks GPQA Diamond 42.9 → 78.6, MMLU-Pro 65.0 → 84.6, BBH 75.0 → 92.0 HLE only rises from below chance to chance (8.4 → 17.8)
Computer use, 60 browser tasks 95% at ≈ 85 ms per click 15% without element text; JEV-27B-VL also 95%, at 260 ms
Robot arm, 20 simulated scenes 40% at 61 ms per decision JEV-27B-VL 75%; 0/10 when choosing motor commands directly
Puzzles Minesweeper 21 → 86%, Wordle 52 → 100%, Sudoku 76 → 99.5% Thinking does not help whole games or sketches

What to keep in mind.

  • Training-data overlap. The card discloses that training used the training splits of benchmarks whose test splits are in the Decision Index (BANKING77 and CLINC150 among them). Thinking lowered BANKING77 (88.0 → 85.0).
  • Teacher distillation. System 1 learned from “teacher distributions”, and for AutoTrust’s JEV family the teacher is the closed Jev 1.13. The demos are small simulations AutoTrust built itself.
  • Downloads. The new repo shows 310,028 downloads in two days against 7 likes. We do not count that number.
  • License. The license tag says Apache-2.0, but that only covers the adapter, head and calibration files. The base weights are under the Gemma 4 terms.
  • Not TypeSafe. AutoTrust is not affiliated with TypeSafe.

What changes here. AutoTrust’s catalog entry is now GEV-26B-Decide, scored 6.3 / 7.8 / 10 (maturity / capability / adoption). It replaces the family page, which was scored 6.1 / 7.5 / 9 with JEV-27B as the point. The JEV models and their caveats carried over to the new page. The old URL, /models/autotrust-jev, still works as a rename note, but it is no longer listed in the catalog or on the quadrant.