ImaJev-4B
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity5.6/10
Open Apache-2.0 LoRA family (2B/4B/9B) with a live HF Space, website, technical report and a local /v1/systemone server (MLX or PyTorch). Single maintainer, no hosted API or SLA.
- Capability7.9/10
Choice / noul / score plus image inputs and a trained unknown (can't-tell) output. One fitted temperature with ECE reported per benchmark. p50 96 ms on one H100, 350 ms with four option orders. Board ranks are author-cited → UnverifiedClaim.
- Adoption15/100
229 GitHub stars, 43 HF likes and ~1.2k downloads on the 4B (2 Oct), plus a r/LocalLLaMA post with ~198 score / 66 comments (28 Sep). No third-party integrations or production use reported.
Vendor claims
- #1 of 91 on JevBench v1.4.2.2 (maintainer-scored 27 Sep 2026): JevBench Score 67.37 vs Jev 1.13.0 63.29 (Plumb-4B 65.8, decider-4b v2 64.1)[Vendor claim — not independently verified]
- #1 of 49 on Image JevBench v0.1.3, composite 76.39 vs Jev-Omni 73.10 (board released 28 Sep 2026)[Vendor claim — not independently verified]
- DecisionBench 1.0 (eng): 79.7% primary, #3 behind the benchmark team's own Bosun models (banner says #3 of 56; results table says #3 of 60 records)[Vendor claim — not independently verified]
- p50 96 ms per question on one H100 (serial, JevBench hard item); 350 ms with four option orders averaged and calibration[Vendor claim — not independently verified]
- 83.9% on the author's ImajevBench v2.0-lite test (279 items), ahead of imajev-9b at 82.1%[Vendor claim — not independently verified]
ImaJev-4B is an open System One peer that adds images to the TypeSafe /v1/systemone contract. You send app state, typed questions and one or two photos; it returns a probability for each option, plus a trained unknown (can’t-tell) probability so the app can abstain instead of guessing. The recommended default is the 4B LoRA on Qwen3.5-4B (Apache-2.0); 2B and 9B adapters exist but are still the previous generation. Code, results and a technical report live at github.com/mohit67890/imajev (229★ on 2 Oct 2026), and there is a live Space.
Specs
| Attribute | Value |
|---|---|
| Artifact | LoRA adapter + decision readout (artifactKind: lora-adapter) |
| Base | Qwen3.5-4B (flagship); also 2B / 9B |
| Readout | one linear layer: 255 option codes + unknown |
| Inputs | Text/JSON state (up to 32 KB) + images (photo-vs-record, two photos) |
| Types | choice, noul, score (+ unknown_probability / abstained) |
| Calibration | one fitted temperature (1.305) on 150 template items, none from JevBench |
| License | Apache-2.0 |
| Released | 4B card 24 Sep 2026; 2B / 9B from 23 Sep |
| Serving | Local MLX (Mac) or PyTorch; POST /v1/systemone |
Author-cited results
Every rank and latency below comes from the model card. The JevBench rank is from the maintainer-run board at benchmarkheaven.com and the DecisionBench record is in that benchmark’s registry, but ModelSystem.One has not re-checked either → UnverifiedClaim.
| Signal | Value | Note |
|---|---|---|
| JevBench v1.4.2.2 | 67.37 (#1/91) | Scored 27 Sep 2026; Jev 1.13.0 at 63.29 |
| Image JevBench v0.1.3 | 76.39 (#1/49) | Ahead of Jev-Omni at 73.10 |
| DecisionBench 1.0 (eng) | 79.7% (#3) | Behind the benchmark team’s own Bosun models; banner says “of 56”, table “of 60 records” |
| ImajevBench (author suite) | 83.9% (4B) | 4B ahead of the 9B’s 82.1% |
| Latency | p50 96 ms / 350 ms | One H100, serial; 350 ms averages four option orders |
The author is open about trade-offs: this release abstains on 11 of 14 “can’t tell” test items where the previous release caught all 14 (that ship gate was overridden), and JevBench hard stayed flat within noise. The r/LocalLLaMA write-up on 28 Sep (~198 score) is the main social spike.
Fit / anti-fit
Fit: local multimodal checks (listing vs photo, shipped vs returned), abstention when unsure, Jev-compatible clients that need images.
Anti-fit: hosted SLA, text-only frontier parity without your own eval, or treating board ranks you haven’t checked as audited fact.
Why it is in the catalog
Trained decision adapter (not a frozen-logit readout), public weights, typed API with probabilities, and first-class image inputs. That makes it a catalog peer next to Jev-Omni and Valen, not a runtime.
JevBench v1.5.7 (Benchmark Heaven, checked 5 Oct 2026)
Imajev-4B (RTX 5090) is #10 at official score 70.4 (capability 70.8) on v1.5.7. That sits beside the author-cited v1.4.2.2 #1 claim already on this page — different edition, different scorer run. See news. No catalog score change yet.
Links
- HF imajev-4b · GitHub · Space · Report
- News: ImaJev open multimodal peer
- Compare: Jev-Omni · Valen · OneJev · Jev
