Jev-Omni

Last updated:

Open80–500ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity4.8/10

    Public Apache-2.0 weights + loader/example; GPU-class ops (~50GB FP32), custom head stack, no hosted SLA, early release.

  • Capability6.9/10

    JevBench #7 (51.3) validates multimodal decision quality — not 'on par with Jev' (#2) as claimed, but solid top-10 performance for a community model.

  • Adoption38/100

    Viral X launch (~2k likes / ~170k views on author post) + ~107 HF likes; still niche outside open-Jev wave.

Vendor claims

JevBench v1.4.2 (independent)

Metric Score Rank
Overall 51.3 #7
Intelligence 44.8 -
Calibration 71.2 -
Speed 76.9 -
Cost 52.4 -

The independent benchmark places Jev-Omni at #7 overall — a strong result for a community multimodal model, though below the author’s “on par with Jev” framing (Jev is #2 at 63.29). The 12B Gemma backbone handles decision quality well; speed (76.9) reflects the heavier compute.

Jev-Omni is an independent open multimodal decision classifier — Gemma 4 12B IT plus a custom head (head.pt) and loader (jev_omni.py) — not a TypeSafe product and not trained on Jev output (author card disclaimer).

State + question + options in; {prediction, prediction_index, confidence, probabilities} out in one forward pass with no text generation. Modalities: text | image | audio | video via media= path. Card: Hugging Face akhilaaa3/Jev-Omni (Apache-2.0). Dataset mentioned: akhilaaa3/decision-bench.

Catalog decisionTypes list choice and boolean (noul via Yes/No options). The card footer also names “score”; the public loader is an option classifier (2–256 slots, best ≤20) — treat score-as-buckets only if you encode buckets yourself. Not a pure transformers AutoModel drop-in: load through the shipped helper.

Among open catalog peers this is the first with all four modalities shipped (text/image/audio/video). openjev already covers image/video/text (audio flagged as coming); Prosodia is audio-specialist and stays off-quadrant.

Public intro: author X post 21 Sep 2026 (@Akhila_988); HF created ~20 Sep 2026. Likes ~107 as of 23 Sep 2026 (downloads API showed 0 at check).

Limits

  • Independent JevBench v1.4.2 places Jev-Omni at #7 (51.3) — solid, but a 12-point gap vs Jev (#2 at 63.29) contradicts “on par” framing
  • DecisionBench, MMAU, MVBench and H200 latency rows remain self-reported → UnverifiedClaim, not ModelSystem.One numbers
  • Card “30k” FT vs decision_config recipe size: 24000 vs tweet “30k examples / <100ms on 1 H100” — source tension; prefer card tables carefully and keep tweet claims unverified
  • Launch X video is polished motion-graphics packaging (not raw live UI). On-screen latencies (e.g. text ~17 ms, image ~36 ms, audio ~11 ms / 13s, video ~201–211 ms / 8 frames) differ from the HF H200 table (83 / 26 / 31 / 504 ms) — treat both as UnverifiedClaim
  • Demo hygiene: animated bars can show option probs that do not sum to 1 (e.g. YES 53.3% / NO 6.5% on an audio clip) — packaging, not a catalog calibration proof
  • CUDA / GPU-class ops (~50 GB FP32 before overhead; BF16 autocast at inference); custom head stack; no hosted SLA
  • Best at ≤20 options; head accepts 256 but quality above 20 is not established
  • Not affiliated with TypeSafe; “Jev” in the name is the interface shape only

Why it is in the catalog

Trained multimodal decision weights + open inference that returns typed option probabilities across four modalities (peer class of openjev / Tiny-Jev, heavier GPU footprint than mini open peers; distinct from hosted Jev). Radar: 2026-09-23.