Jev-Omni
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity4.8/10
Public Apache-2.0 weights + loader/example; GPU-class ops (~50GB FP32), custom head stack, no hosted SLA, early release.
- Capability6.9/10
JevBench #7 (51.3) validates multimodal decision quality — not 'on par with Jev' (#2) as claimed, but solid top-10 performance for a community model.
- Adoption38/100
Viral X launch (~2k likes / ~170k views on author post) + ~107 HF likes; still niche outside open-Jev wave.
Vendor claims
- DecisionBench Medium 87.57% / micro 86.01% (80 scenarios / 293 questions) — author card[Vendor claim — not independently verified]
- Author JevBench: 86.15% matched (self-reported) vs independent JevBench v1.4.2: #7 overall (51.3)[Vendor claim — not independently verified]
- MMAU 63.10% (1000 questions); MVBench 53.10% (14 tasks / 2786 questions) — author card[Vendor claim — not independently verified]
- Warm H200: 83 ms (~2k text), 26 ms image, 31 ms 13s audio, 504 ms 16-frame video — author card[Vendor claim — not independently verified]
- Card: 30,000-question fine-tune; decision_config recipe size 24000; tweet cites 30k examples / <100ms on 1 H100 — numbers conflict across sources[Vendor claim — not independently verified]
- X launch post: on par with Jev on typed benchmarks; first-multimodal framing is marketing — catalog treats four-modality ship as open peer fact only[Vendor claim — not independently verified]
JevBench v1.4.2 (independent)
| Metric | Score | Rank |
|---|---|---|
| Overall | 51.3 | #7 |
| Intelligence | 44.8 | - |
| Calibration | 71.2 | - |
| Speed | 76.9 | - |
| Cost | 52.4 | - |
The independent benchmark places Jev-Omni at #7 overall — a strong result for a community multimodal model, though below the author’s “on par with Jev” framing (Jev is #2 at 63.29). The 12B Gemma backbone handles decision quality well; speed (76.9) reflects the heavier compute.
Jev-Omni is an independent open multimodal decision classifier — Gemma 4 12B IT plus a custom head (head.pt) and loader (jev_omni.py) — not a TypeSafe product and not trained on Jev output (author card disclaimer).
State + question + options in; {prediction, prediction_index, confidence, probabilities} out in one forward pass with no text generation. Modalities: text | image | audio | video via media= path. Card: Hugging Face akhilaaa3/Jev-Omni (Apache-2.0). Dataset mentioned: akhilaaa3/decision-bench.
Catalog decisionTypes list choice and boolean (noul via Yes/No options). The card footer also names “score”; the public loader is an option classifier (2–256 slots, best ≤20) — treat score-as-buckets only if you encode buckets yourself. Not a pure transformers AutoModel drop-in: load through the shipped helper.
Among open catalog peers this is the first with all four modalities shipped (text/image/audio/video). openjev already covers image/video/text (audio flagged as coming); Prosodia is audio-specialist and stays off-quadrant.
Public intro: author X post 21 Sep 2026 (@Akhila_988); HF created ~20 Sep 2026. Likes ~107 as of 23 Sep 2026 (downloads API showed 0 at check).
Limits
- Independent JevBench v1.4.2 places Jev-Omni at #7 (51.3) — solid, but a 12-point gap vs Jev (#2 at 63.29) contradicts “on par” framing
- DecisionBench, MMAU, MVBench and H200 latency rows remain self-reported → UnverifiedClaim, not ModelSystem.One numbers
- Card “30k” FT vs
decision_configrecipesize: 24000vs tweet “30k examples / <100ms on 1 H100” — source tension; prefer card tables carefully and keep tweet claims unverified - Launch X video is polished motion-graphics packaging (not raw live UI). On-screen latencies (e.g. text ~17 ms, image ~36 ms, audio ~11 ms / 13s, video ~201–211 ms / 8 frames) differ from the HF H200 table (83 / 26 / 31 / 504 ms) — treat both as UnverifiedClaim
- Demo hygiene: animated bars can show option probs that do not sum to 1 (e.g. YES 53.3% / NO 6.5% on an audio clip) — packaging, not a catalog calibration proof
- CUDA / GPU-class ops (~50 GB FP32 before overhead; BF16 autocast at inference); custom head stack; no hosted SLA
- Best at ≤20 options; head accepts 256 but quality above 20 is not established
- Not affiliated with TypeSafe; “Jev” in the name is the interface shape only
Why it is in the catalog
Trained multimodal decision weights + open inference that returns typed option probabilities across four modalities (peer class of openjev / Tiny-Jev, heavier GPU footprint than mini open peers; distinct from hosted Jev). Radar: 2026-09-23.
