AutoJev-27B
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity5.1/10
Public Apache-2.0 weights plus MIT code with training scripts, a /v1/systemone server and a browser playground, so the TypeSafe SDK works via base_url. Short card, training corpus not bundled, self-host only.
- Capability7.2/10
Noul, Choice and Score in one pass. #4 on the official Decision Index 0.2.1 (56.40 vs Jev 57.91), with a measured index ECE of 0.018, better than the author's own 0.043. Temperature scaling is documented. No latency figures published.
- Adoption11/100
About 47 HF likes and 1k downloads, 110 GitHub stars and 10 forks, plus a community GGUF with ~3k downloads that needs a llama.cpp fork. No production use reported.
Vendor claims
- Overall accuracy 84.60% vs Jev 82.79% and Qwen3.8-27B 69.83%; ECE 0.043 vs Jev 0.053 (author's own held-out set)[Vendor claim — not independently verified]
- One H200, full-weight SFT, 73,000 unique training examples, 286 updates; released checkpoint is step 200[Vendor claim — not independently verified]
- Research, data generation, training, evaluation and deployment were executed by autonomous agents[Vendor claim — not independently verified]
AutoJev-27B is an open decision model published by the Hugging Face / GitHub user denis-pplx on 19 September 2026. It is a full-weight fine-tune of Qwen/Qwen3.8-27B that returns probabilities over the options you supply, in one forward pass per question. The repo ships a server that speaks POST /v1/systemone with choice, noul and score, a browser playground and optional base64 images.
The author describes it as an “independent implementation inspired by Jev”. It is not a TypeSafe product and is not trained on Jev outputs as far as the card says. It is also unrelated to AutoTrust JEV, which uses a similar-sounding name.
Specs
| Attribute | Value |
|---|---|
| Author | denis-pplx (HF + GitHub) |
| Release | 19 Sep 2026 (HF repo 16:54 BRT; GitHub repo created the same day) |
| Base model | Qwen/Qwen3.8-27B |
| Parameters | 26.1B (BF16, ~49 GiB of weights) |
| Training | Full-weight SFT on one H200, 73,000 unique examples, 286 updates; checkpoint 200 released; scalar temperature fitted separately |
| Training data | Generated and curated by the author’s agent pipeline; the exact corpus is not bundled |
| API | POST /v1/systemone (choice, noul, score, optional images); FastAPI docs at /docs |
| License | Weights Apache-2.0, code MIT |
| Latency | No figures published (0–0ms in the frontmatter means “no data”) |
Decision Index 0.2.1 (official Space)
Measured by the official Decision Index Space (data snapshot 28 Sep 2026, 71 systems):
| System | Index (balanced skill) | Rank |
|---|---|---|
| Jev 1.13 | 57.91 | #1 |
| Surogate Rune 26B-A4B v3 | 57.44 | #2 |
| AutoJev-27B | 56.40 | #4 |
| simple-jev (Qwen3.8-27B, runtime) | 55.74 | #5 |
The index also measures calibration: for AutoJev-27B, accuracy 0.730, mean confidence 0.738, ECE 0.018 and Brier 0.361. That is tighter than the author’s own ECE figure.
Note the gap to simple-jev on the same Qwen3.8-27B base: 0.66 points. Most of AutoJev’s accuracy comes from the base model; the fine-tune adds a trained decision readout and better calibration.
Author’s own results (UnverifiedClaim)
| Model | Overall accuracy | ECE | Brier |
|---|---|---|---|
| Qwen3.8-27B | 69.83% | 0.0648 | 0.408 |
| AutoJev-27B | 84.60% | 0.0428 | 0.220 |
| Jev | 82.79% | 0.0527 | 0.254 |
These come from the author’s curated held-out set (holdout-v3), which is not public. Beating Jev on your own test set is a different claim from matching Jev across public suites; on the official index AutoJev sits 1.5 points below Jev.
Community GGUF
thomasgauthier/autojev-27b-GGUF ships a Q4_K_M quant and a vision projector (~3k downloads). It does not run on mainline llama.cpp or LM Studio: you need the jev.cpp fork, which adds the AutoJev classifier head and a --system-one server flag. Another copy, arafathusayn/autojev-27b, is a plain re-upload.
Fit / anti-fit
Fit when you need: a local, TypeSafe-compatible endpoint with good measured calibration, agent gating and routing decisions, or a base to study (training code is public).
Anti-fit when you need: a small GPU (about 49 GiB of BF16 weights plus overhead), published latency, natural-image accuracy (the author says image support is not an accuracy claim), or the training corpus.
Limits
- The training corpus is not published; the headline accuracy is on a private held-out set.
- The card still says “while the weights are private, authenticate”; the repo was public and ungated when we checked on 30 Sep 2026.
- “Built by autonomous agents” describes the author’s process; it does not change how the model is evaluated here.
- No latency or throughput numbers.
Why it is in the catalog
The weights were fully fine-tuned for typed decisions and ship with an open System One-compatible server that returns probabilities. That passes the catalog criteria. The official index places it among the top open entries.
