OneJev (0.8B / 4B / 9B / 27B)
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity5.4/10
Apache-2.0 weights in four sizes (+27B-FP8), training data and recipe released, demo Space and site. The qev server speaks System One, so TypeSafe SDK works via base URL. qev installs from git only (the PyPI name belongs to another project). No hosted API or SLA.
- Capability7.3/10
Noul / Choice / Score plus images and video in one prefill pass. Brier loss and a per-type temperature table, but no ECE published. Latency 31–324 ms is the authors' H200 run, without p99. Accuracy table is author-run on own and public sets; not on the Decision Index.
- Adoption10/100
Four days old: 106 GitHub stars, ~2.3k first-party HF downloads, and day-one community quants (bartowski, mradermacher, onnx-community) with several thousand downloads. A companion server (OpenDecisions) from the same author. No production cases yet.
Vendor claims
- OneJev-27B: 77.4 on the held-out OneJev test set vs 60.8 for Jev-Omni and 71.6 for Qwen3.8-27B thinking (authors' runs)[Vendor claim — not independently verified]
- DecisionBench medium 87.9 vs Jev 90.5; DecisionBench hard 73.1 vs Jev 65.3; TypeSafe evals 89.3 vs 89.0 (Jev = published text-only scores)[Vendor claim — not independently verified]
- MMBench 92.3 and MMStar 72.2 for the 27B[Vendor claim — not independently verified]
- Latency on one H200 with a 1280×720 screenshot, 1 / 10 questions: 0.8B 31 / 51 ms, 4B 64 / 104 ms, 9B 81 / 131 ms, 27B 189 / 324 ms[Vendor claim — not independently verified]
Read this first
- Every number below is the authors’ own run (UnverifiedClaim). OneJev is not on the official Decision Index, whose data was last generated on 28/09.
- Jev rows in the results table use TypeSafe’s published text-only scores; the Jev-Omni and Qwen rows were run by the authors. The comparison is not apples to apples.
pip install qevinstalls a different, unrelated package. The OneJev package is installed from the GitHub repo (pip install git+https://github.com/OmniJev/OneJev).
OneJev is an open family of multimodal decision models from the OmniJev organisation, released on Hugging Face on 27 September 2026 (around 13:15 BRT) in four sizes: 0.8B, 4B, 9B and 27B, plus a 27B FP8 build. Each one is a full fine-tune of a Qwen vision-language model (Qwen3.5-0.8B / 4B / 9B and Qwen3.8-27B) with the vision tower frozen. Like Jev, it answers typed questions in one prefill pass and returns probabilities, not text. Unlike Jev, the state can include images and video next to text.
Specs
| Attribute | Value |
|---|---|
| Maker | OmniJev (maintainer: unikcc) |
| Sizes | 0.8B, 4B, 9B, 27B (+ 27B-FP8) |
| Base | Qwen3.5-0.8B / 4B / 9B, Qwen3.8-27B (vision tower frozen) |
| Training | 1 epoch, lr 5e-6, 16,384-token limit, cross-entropy + Brier loss over answers |
| Data | 99,193 questions; 94,707 released as OneJev-Data (17.7 GB) |
| Question types | Noul, Choice, Score |
| Input | Text/JSON state plus a media field (images, video) |
| Serving | qev serve (System One endpoint, TypeSafe SDK compatible), llama.cpp via GGUF, OpenDecisions |
| License | Apache-2.0 (code and weights) |
| Launch |
Evidence (UnverifiedClaim)
| Benchmark (27B) | OneJev | Jev 1.13 | Jev-Omni | Qwen3.8-27B thinking |
|---|---|---|---|---|
| OneJev test (held out) | 77.4 | — | 60.8 | 71.6 |
| DecisionBench medium | 87.9 | 90.5 | — | — |
| DecisionBench hard | 73.1 | 65.3 | — | — |
| TypeSafe evals | 89.3 | 89.0 | — | — |
| MMBench | 92.3 | — | — | — |
| MMStar | 72.2 | — | — | — |
Latency (one H200, 1280×720 screenshot, 1 / 10 questions): 0.8B 31 / 51 ms, 4B 64 / 104 ms, 9B 81 / 131 ms, 27B 189 / 324 ms. No p99 is given and there is no third-party run yet.
Fit / anti-fit
Fit: decisions about screenshots, UI states, photos or short video (agent step checks, visual triage, moderation queues) where you want open weights and local serving; teams that want to fine-tune further (data and configs are public).
Anti-fit: anything that needs a hosted SLA or an independently verified score; commercial use that cannot accept the Qwen licence chain; pure-text workloads where a smaller text model is already measured on the Decision Index.
Limits
- No calibration metric (ECE) is published, although the training uses Brier loss and
qevapplies a temperature table per question type. - 16,384-token training limit; behaviour beyond it is not documented.
- The repo is four days old, with one maintainer.
Why it is in the catalog
OneJev trains its own decision weights (full fine-tunes with a decision loss), ships them under Apache-2.0 and serves them through a System One endpoint that returns probabilities. That meets the catalog bar. The multimodal input is the distinguishing feature: only Jev-Omni and Cloudflare Clef also take images.
