AutoTrust GEV-26B-Decide

Last updated:

Open45–85ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity6.3/10

    Apache-2.0 adapter, head and calibration files on Gemma-4-26B-A4B (Gemma 4 terms). Very detailed card with reports, two calibration tables and a training-data disclosure. The own POST /v1/decide server needs a patched vLLM dev build and an 80 GB GPU; no /v1/systemone. Company, no SLA.

  • Capability7.8/10

    Noul, Choice (2–256) and Score (0–5) over text and images. System 1: 45 ms median on a B200, no p99. ECE 0.035 against gold labels on 1,754 questions outside the DI. The DI 62.48 is AutoTrust's own scoring: adaptive mode breaks the board's latency rule, and some DI training splits were used.

  • Adoption10/100

    7 likes on the new repo and 13 on the JEV-Gemma4 repo it republishes; about 156 likes across the AutoTrust family, an HF blog, a press release and a third-party GGUF of JEV-27B. Download counts look anomalous (310k in two days) and are not counted. No production use reported.

Vendor claims

Read this first

  1. New name, same weights, not an HF rename. AutoTrust published autotrust/GEV-26B-Decide as a new repository on 2 Oct 2026 (01:20 BRT). The card says it “was previously published as autotrust/JEV-Gemma4-26B-A4B; the weights are the same”. We checked: the adapter, the decision head and the base shards have identical LFS sha256 hashes. The old repo is still online and unchanged since 29 Sep. GEV adds a vLLM LoRA (adapter_vllm/), a /v1/decide server with adaptive thinking, a vLLM patch, new reports and demo videos.
  2. Distilled lineage. System 1 “was trained on teacher distributions and ground-truth decision data”, and “where the teacher is wrong, System 1 often is too”. The GEV card does not name the teacher. For the JEV family it is the closed TypeSafe Jev 1.13, via SargeDev/jev-distill-corpus-v3 (labels “distilled from Jev 1.13 via OpenRouter”). We have not verified whether this use of Jev outputs complies with TypeSafe’s terms.
  3. The 62.48 is not a board score. AutoTrust scored it with the Decision Index kit itself. Knowledge & Reasoning runs with adaptive thinking (median 13.4 s per request, which breaks the board’s ≤ 1,000 ms median rule), and the other four areas run System 1 only. The training data also includes the training splits of benchmarks whose test splits the index uses (BANKING77 and CLINC150 among them); AutoTrust says it gave the list to the index maintainers. As of the 28 Sep snapshot, no AutoTrust model is on the official Decision Index.
  4. “JEV”/“GEV” is not TypeSafe’s product. AutoTrust’s card says it is “not affiliated with, endorsed by, or a product of TypeSafe AI”. For TypeSafe’s hosted model, see Jev.

What it is

GEV-26B-Decide follows the same recipe as AutoTrust’s other models, which it calls Blocks of Experts. The base, google/gemma-4-26B-A4B-it (26B parameters, about 4B active per token), stays frozen. A trained LoRA plus a 24-slot fp32 decision head adds a typed-decision skill.

Mode What it does Output
System 1 Noul (yes/no), Choice over 2–256 options, Score 0–5, over text and images; prompts up to 256K tokens A probability per option, in one forward pass
Adaptive thinking (opt-in, thinking: "auto") System 1 first. Below 0.8 confidence, the base model thinks and its answer is mixed in: p = ½ p1 + ½ p2 Probabilities
System 2 The unmodified Gemma-4-26B-A4B-it, with or without thinking Text

Read-out: the last token’s final-norm hidden state (2,816) goes through the linear head to 24 slots, soft-capped at 30. Inactive slots are masked, and the logits are divided by a per-kind temperature. Up to 16 options are read in one pass. Longer lists run as a tournament in groups of 16. Two temperature tables ship: calibration.json (1.003 / 1.017 / 0.999, used in the index run) and calibration_gold.json (1.214 / 1.098 / 1.000, fitted on gold answers; AutoTrust recommends it when gating actions on confidence).

Serving

  • vLLM: serve.sh runs serve_decide.py, the standard vLLM OpenAI server plus a POST /v1/decide route. Plain requests go to System 2, and requests for the LoRA module jev-decision go to System 1. It needs a vLLM build with Gemma 4 support, the bundled patches/vllm-gemma4-lm-head-lora.patch and one GPU with 80 GB or more. AutoTrust tested it on a vLLM development build from September 2026.
  • transformers + peft: System 1 with the adapter merged in memory, up to 16 options per pass.
  • Not /v1/systemone, so the TypeSafe SDK does not apply. No GGUF or Ollama/Ollaya package for GEV as of 4 Oct.

AutoTrust’s numbers (UnverifiedClaim)

All numbers below are AutoTrust’s own runs on one B200, taken from the card (lastModified 3 Oct, 08:00 BRT).

Decision Index 0.2.1 (own scoring).

Balanced skill Balanced raw Breadth
GEV-26B-Decide, adaptive on K&R 62.48 70.66 62.00
Same weights, System 1 only (29 Sep run, as JEV-Gemma4-26B-A4B) 58.05 67.35 56.98
TypeSafe Jev 1.13 (board, 28 Sep) 57.91 — —

Areas: Knowledge & Reasoning 0.602 (System 1: 0.429), Language 0.636, Retrieval & Classification 0.679, Tools & Automation 0.697, Arts & Taste 0.415. Raw results are in autotrust/jev-decision-index-results (runs/jev-gemma4-26b-a4b).

Knowledge & Reasoning, System 1 → adaptive (accuracy %). GPQA Diamond 42.9 → 78.6 · CRUXEval 67.5 → 90.7 · CLadder 71.0 → 86.6 · MMLU-Pro 65.0 → 84.6 · BBH 75.0 → 92.0 · GSM8K 97.6 → 99.1. HLE: 8.4 → 17.8, below chance (16.4) without thinking and back to chance level with it. Of the HLE questions answered without thinking, only 1.6% are right. ChessBench gains nothing.

Outside K&R: thinking helps little (BFCL 94.4 → 96.0, CLINC150 93.3 → 94.4) and hurts BANKING77 (88.0 → 85.0), a set whose training split was in the training data.

Validation set (1,754 questions from six public sets outside the index): System 1 73.3% → adaptive 83.4%, ECE 0.035 for both. It thinks on 48% of the questions. This is the only calibration figure measured against gold labels rather than against the teacher.

Latency: System 1 median 45 ms per request, 257 decisions/s with 64 clients. Adaptive mode on K&R: median 13.4 s, p90 33 s (7.7 s / 19 s with speculative decoding). No p99 is published for System 1.

Demos (3 Oct, System 1).

GEV-26B-Decide JEV-27B-VL JEV-9B
Computer use, 60 browser tasks (numbered boxes + element text) 95% 95% 95%
Time per click ≈ 85 ms ≈ 260 ms ≈ 200 ms
Robot arm pick-and-place, 20 MuJoCo scenes 40% 75% 50%
Time per robot decision 61 ms 239 ms 163 ms

Without element text, computer use falls to 15%. Asked to choose one of 8 motor commands directly, the robot arm completed 0 of 10 scenes. These are small, self-built simulations.

Games (200 one-move puzzles each, System 1 → adaptive): Minesweeper 21 → 86 · Wordle 52 → 100 · Connect Four 54 → 99.5 · Sudoku 76 → 99.5 · 24 game 89 → 100 · Maze 48.5 → 68 · chess mate-in-one 46.5 → 78. Thinking does not help whole games (2048, Flappy Bird) or sketch recognition.

Other: zero-shot images on VL-RewardBench 78.4% (the decision head was trained on text). Long context 10/10 at 4K–128K (above 128K untested). Many options: MASSIVE 59 options 91.2%, BANKING77 81.5%, CLINC150 95.5%, 255 merged intents 89.0% / 75.2%.

Family

The other AutoTrust repos are still online and were not renamed (4 Oct). They share the recipe and the caveats above.

Repo Base Created (BRT) Notes
autotrust/JEV-Gemma4-26B-A4B gemma-4-26B-A4B-it 29 Sep, 09:34 Same weights as GEV; System 1 only; own DI 58.05; unchanged since 29 Sep
autotrust/JEV-27B Qwen3.8-27B 25 Sep, 09:58 108.9M trained params; KL ≈ 0.017 to Jev; 137 ms median; own DI 53.30
autotrust/JEV-9B Qwen3.5-9B 23 Sep, 09:08 First generation; ≈ 90 ms; hosts the demo code
autotrust/JEV-27B-VL Qwen3.8-27B (multimodal) 30 Sep, 02:02 Same adapter and head as JEV-27B; image input

JEV-27B evidence, carried over (UnverifiedClaim).

  • Six-benchmark table: 84.07% vs Jev 1.13’s 83.85%. AutoTrust ran both models, and the other rows were copied from the NeoHorse-Jev-4B card.
  • Fidelity, not accuracy: KL ≈ 0.017, ECE 0.0009 and 90.5% top-1 agreement are measured against the teacher, on a test split of the distillation corpus.
  • Base model untouched: HumanEval stays at 78.0% with and without the block, and all 164 completions are byte-identical to the base.
  • Third-party GGUF: prithivMLmods/JEV-27B-GGUF.
  • Index gap: its own index run (53.30) sits 4.6 points below Jev, which does not fit the “+0.22 over Jev” headline.

Fit / anti-fit

Fit when you need: typed decisions on your own large GPU with a fast System 1 and an opt-in reasoning path for logic, maths and science questions. Also fits one engine for decisions plus normal Gemma generation, or long contexts and choice lists beyond 16 options.

Anti-fit when you need: sub-second answers with thinking on, the /v1/systemone contract, a small GPU (80 GB for the vLLM path), training data independent of a closed vendor’s outputs, or independently verified accuracy. Avoid thinking on classification and routing: it lowered BANKING77 and only returned HLE to chance.

Red flags we track

  • Downloads: 310,028 for the new repo within two days, with 7 likes. The old repo has 1,071 downloads and 13 likes. The same pattern shows on JEV-27B-VL (≈ 726k) and JEV-9B (≈ 132k). We do not count these in adoption.
  • License tag: HF metadata says apache-2.0, but only the adapter, head and calibration files are Apache-2.0. The base weights in the repo are under the Gemma 4 terms.
  • Two repos, one model: benchmarks and links may cite either name.

Why it is in the catalog

The decision block was trained for typed decisions and ships openly. Its documented serving path returns a probability per option. That passes the catalog criteria as a lora-adapter, because the backbone is untouched. Since 4 Oct this page is AutoTrust’s catalog entry. It replaces the former “AutoTrust JEV” family page (scored 6.1 / 7.5 / 9 with JEV-27B as the point), which is now a rename note.

Score working (4 Oct 2026)

  • Maturity = availability 8 × 0.30 + docs 9 × 0.25 + integrations 5 × 0.25 + support 2 × 0.20 = 2.40 + 2.25 + 1.25 + 0.40 = 6.30

  • Capability = types 10 × 0.30 + calibration 7 × 0.25 + latency 8 × 0.25 [vendor] + accuracy 5 × 0.20 = 3.00 + 1.75 + 2.00 + 1.00 = 7.75 → 7.8

  • Adoption = engagement 20 × 0.35 + ecosystem 8 × 0.35 + production 2 × 0.30 = 7.00 + 2.80 + 0.60 = 10.40 → 10

  • Why these grades:

    • Calibration 7: published temperatures plus ECE 0.035 against gold labels; it is a self-run, and HLE shows confident errors.
    • Latency 8: 45 ms median, but on a B200 and with no p99.
    • Accuracy 5: one point below our usual 6 for a self-run Decision Index score, because of the training-split overlap with the index and a headline that uses a mode the board would reject.
    • Adoption: likes and ecosystem only; the download counts are anomalous.

NVFP4 build (5 Oct 2026)

AutoTrust published GEV-26B-Decide-NVFP4 (about 15k downloads / 150 likes by 7 Oct). Only the 3,840 routed experts are in NVFP4 (ModelOpt layout); everything else stays bf16, including the adapter, the decision head and the temperatures. In vLLM it takes 17.1 GiB instead of 51.1 GiB (−66%), so it fits a 24 GB GPU.

AutoTrust’s check (UnverifiedClaim): System 1 only, on GPQA Diamond (43.9% → 44.9%) and HLE (8.4% → 9.7%), within noise; adaptive thinking was not re-evaluated. Native W4A4 was tested only on B200, and the build needs the same vLLM patch as the bf16 model. Scores on this page are unchanged.