JevK5

Last updated:

Open9–50ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity3.5/10

    Apache-2.0 weights available; runtime published; JevBench v1.4.2 scores independently verified; no hosted inference or SLA.

  • Capability7.4/10

    JevBench v1.4.2 verified: #3 overall (62.04), #1 pure Apache-2.0 open-weight. Choice+Score via jevk5 runtime.

  • Adoption12/100

    Brand-new release (23 Sep 2026). HF page up; community uptake unknown.

Vendor claims

JevK5 is an open-weight Jev alternative based on Qwen3.5-4B with LoRA merged. It uses a one-pass softmax readout — it does not generate text, so there is nothing to parse and nothing to hallucinate. Released under Apache-2.0.

JevBench v1.4.2 confirms #3 overall (62.04 score) and #1 among pure Apache-2.0 open-weight models — Decider-4B is also open but uses mixed licensing for some components.

A smaller JevK5-2B variant (Qwen3.5-2B) ships in the same family for lighter deployments (~3.5GB bf16).

Specs

Attribute Value
Author alibiserikbay (independent)
Base model Qwen3.5-4B + LoRA merged
License Apache-2.0
Launch
Status Open weights (Hugging Face)
Latency ~9ms on H100 (UnverifiedClaim)
Pricing Self-host only (free weights)
Decision types Choice, Score (via jevk5 runtime)
Context 32K (inherited from Qwen3.5)
Smaller variant JevK5-2B — Qwen3.5-2B base
GGUF JevK5-GGUF

Architecture

JevK5 follows the one-pass softmax pattern: given state and typed questions, it produces probabilities over options without generating tokens. The jevk5 Python runtime handles the inference loop and exposes Choice, Score, and Noul question types (verify Noul support in the runtime docs).

JevBench v1.4.2 (verified)

Metric JevK5 JevK5-2B Rank
Overall 62.04 56.48 #3 / #9
Intelligence 48.6 46.2 -
Calibration 77.0 71.3 -
Speed 86.1 89.7 -
Cost 59.7 68.2 -

JevK5 holds #3 overall and #1 among pure Apache-2.0 open-weight models (Decider-4B is also open but uses mixed licensing for some components).

Fit / anti-fit

Fit when you need: self-hosted typed decisions, Apache-2.0 deployment, latency-sensitive routing on capable hardware.

Anti-fit when you need: hosted inference, production SLA, generation, or workflows that require free-text output.