AutoTrust JEV (Blocks of Experts)
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity6.1/10
Apache-2.0 weights for 9B and 27B, a long card with evaluation reports, an HF blog and a vLLM quickstart: one engine serves decisions via a LoRA module and plain generation from the frozen base. No /v1/systemone endpoint, so the TypeSafe SDK doesn't apply. Company, no SLA.
- Capability7.5/10
Noul, Choice (2–16) and Score (0–5). A distillation of closed Jev: KL 0.017 and ECE 0.0009 are fidelity to the teacher, in-distribution. 137 ms median per decision on a B200, own run. Not on the official Decision Index; its own index run gives JEV-27B 53.30 vs Jev 57.91.
- Adoption9/100
About 31 HF likes and 1.25k downloads across 9B and 27B, an HF blog, a PR Newswire press release and a third-party GGUF. No integrations beyond vLLM and no production use reported.
Vendor claims
- Six public decision benchmarks: JEV-27B 84.07% vs TypeSafe Jev 1.13 83.85% (AutoTrust ran both; other rows copied from NeoHorse)[Vendor claim — not independently verified]
- Mean KL ≈ 0.017 to Jev 1.13's output distributions on 25,376 held-out rows; ECE 0.0009; choice top-1 agreement with Jev 90.5%[Vendor claim — not independently verified]
- Single decision 137 ms median on one B200 (JEV-9B ≈ 90 ms); 4.2 ms per decision batched[Vendor claim — not independently verified]
- HumanEval 78.0% with and without the decision block; all 164 completions byte-identical to Qwen3.8-27B[Vendor claim — not independently verified]
- Own Decision Index 0.2.1 kit runs: JEV-Gemma4-26B-A4B 58.05, JEV-27B 53.30 (not on the official leaderboard)[Vendor claim — not independently verified]
- JEV-27B-VL: 78.3% on VL-RewardBench, 'highest on the official leaderboard' (self-run vs a board last updated May 2025)[Vendor claim — not independently verified]
Read this first
- This is a distillation of the closed TypeSafe Jev 1.13. AutoTrust trained the decision block on
SargeDev/jev-distill-corpus-v3, a third-party dataset whose largest stream (yuri_v3, 498k rows) has labels “distilled from Jev 1.13 (TypeSafe) via OpenRouter”, per the dataset card. 25,376 of the 29,955 test targets are Jev output distributions. Its headline fidelity metrics (KL ≈ 0.017, ECE 0.0009, 90.5% top-1 agreement) measure how closely it copies the teacher, on a test split of that same corpus. They are not accuracy against ground truth. We have not verified whether this use of Jev outputs complies with TypeSafe’s terms.- “JEV” here is not TypeSafe’s product. AutoTrust’s own card says the model “is not affiliated with, endorsed by, or a product of TypeSafe AI, and shares no weights or code with it.” For TypeSafe’s hosted model, see Jev. This page is titled “AutoTrust JEV” to keep the two apart.
- Not on the official leaderboards. As of the 28 Sep 2026 snapshot, the JEV models are not on the official Decision Index or the JevBench v1.4.2 leaderboard. AutoTrust’s own run of the index kit gives JEV-27B 53.30, 4.6 points below Jev’s 57.91. That sits oddly next to the “+0.22 over Jev” headline from its six-benchmark table.
What it is
AutoTrust AI (Singapore) calls its recipe Blocks of Experts. A strong pretrained model stays frozen as one “expert block”, and a small detachable block is trained for a single skill. In the JEV models:
- System 1 (typed decisions) is a trained LoRA plus a 24-slot decision head. It answers
noul,choiceover 2–16 options andscoreon a fixed 0–5 scale in one prefill pass, with a probability per option. - System 2 (ordinary generation and reasoning) is the base model, bit-identical to the release. HumanEval stays at 78.0% with and without the decision block, and all 164 completions are byte-identical, per the card.
One vLLM engine serves both. Requests addressed to the LoRA module jev-decision go through the decision head as a one-token completion constrained to the option tokens; everything else is plain Qwen. The API is vLLM’s OpenAI-compatible one, not /v1/systemone.
Family
| Repo | Base | Trained block | Created (BRT) | Notes |
|---|---|---|---|---|
| autotrust/JEV-27B (main) | Qwen3.8-27B | 108.9M params (0.4%), ≈ 9.2 B200-hours, ≈ 640k rows | 25 Sep, 09:58 | KL ≈ 0.017; 137 ms median per decision |
| autotrust/JEV-9B | Qwen3.5-9B | 40.2M params, ≈ 3 B200-hours | 23 Sep, 09:08 | First generation; KL ≈ 0.019; ≈ 90 ms; HumanEval 70.7% |
| autotrust/JEV-27B-VL | Qwen3.8-27B (multimodal) | Same adapter and head as JEV-27B | 30 Sep, 02:02 | Adds image input |
| autotrust/JEV-Gemma4-26B-A4B | gemma-4-26B-A4B-it | LoRA + 24-slot head | 29 Sep, 09:34 | Own index run: 58.05 |
Third-party conversion: prithivMLmods/JEV-27B-GGUF (28 Sep; BF16 and Q3_K_M–Q5_K_M). AutoTrust’s documented decision path is vLLM with adapter_vllm/; the GGUF card documents llama.cpp generally.
AutoTrust’s numbers (UnverifiedClaim)
Six public text-decision benchmarks. AutoTrust ran JEV-27B and the hosted Jev 1.13 itself (26–27 Sep). The other rows are copied from the NeoHorse-Jev-4B card, not re-run.
| Model | JevBench | Kev | OpenJev text | Nimble | VitaminC | MASSIVE-en | Mean |
|---|---|---|---|---|---|---|---|
| JEV-27B | 88.70 | 83.75 | 73.89 | 92.91 | 77.46 | 87.71 | 84.07 |
| Jev 1.13 (AutoTrust’s run) | 87.18 | 85.52 | 72.96 | 91.84 | 78.46 | 87.14 | 83.85 |
| NeoHorse-Jev-4B (as reported) | 75.73 | 81.92 | 58.74 | 87.23 | 77.13 | 85.43 | 77.70 |
Other claims: on the third-party decision-models-under-pressure benchmark (human gold labels, 16 options) JEV-27B reaches 96% of the teacher’s accuracy (0.740 vs 0.769), self-run. The card compares its 137 ms median with 238–301 ms measured by others for the hosted API; the hosted numbers include network and queueing. JEV-27B-VL claims 78.3% on VL-RewardBench, above a leaderboard last updated in May 2025.
Fit / anti-fit
Fit when you need: Jev-like behaviour on your own GPU, one engine for both typed decisions and normal generation, or a vLLM-native setup.
Anti-fit when you need: a model whose training data does not depend on a closed vendor’s outputs, independently verified accuracy, the /v1/systemone contract, choice lists longer than 16 or custom score scales, or a small GPU (the 27B bundle is about 54 GB).
Limits
- The teacher’s mistakes are inherited: the card says fidelity “extends to the teacher’s mistakes”.
- Probabilities are learned from Jev’s outputs, which Jev rounds to two decimals.
- Score is fixed at 0–5; choice tops out at 16 options per group.
- Everything above is AutoTrust’s own measurement; nothing is on an official leaderboard yet.
Why it is in the catalog
The decision block’s weights were actually trained for typed decisions (LoRA plus head), and they ship openly with a documented serving path that returns probabilities. That passes the catalog criteria, with artifactKind: lora-adapter because the backbone is untouched. The caveats above stay on the page until the training data question and the leaderboard gap are settled.
Links
- Hugging Face: autotrust/JEV-27B · JEV-9B · JEV-27B-VL · JEV-Gemma4-26B-A4B
- HF blog: JEV-27B
- PR Newswire release (29 Sep 2026)
- Training corpus: SargeDev/jev-distill-corpus-v3
- Own Decision Index runs (dataset)
- News: Five open decision models from HF trending
- Compare: Jev (TypeSafe AI) · AutoJev-27B (unrelated)
