NeoHorse-Jev-4B

Last updated:

Open—$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity6.1/10

    Apache-2.0 bundle on HF and ModelScope (backbone, decision head, runtime wheel), official GGUF repo, DEPLOYMENT guide, vLLM and SGLang adapters, and a /v1/systemone endpoint in the native server. Backed by a company, but no SLA.

  • Capability5.7/10

    Noul, Choice and Score, text plus a single image. The card reports no calibration or latency. The official Decision Index 0.2.1 puts it at 36.75 (#30 of 71, ECE 0.10), far from its self-reported 77.70 six-group average.

  • Adoption9/100

    About 20 HF likes and 2.2k downloads, plus 2.4k on the GGUF repo. The parent NeoHorse-1 project is popular (1.3k GitHub stars), but that is not this model. Cited as a baseline by AutoTrust; no production use reported.

Vendor claims

NeoHorse-Jev-4B is a 4B decision model from TokenRhythm, released on 23 September 2026. It is built on the company’s NeoHorse-1-4B, an agentic post-train of Qwen3.5-4B. Given a state and questions your application defines, it returns decisions with probabilities, using prefill-only inference and three types: Choice, Noul and Score. Requests can also carry a single image.

It is not a TypeSafe product. “Jev” in the name refers to the interface shape.

Lineage: Kev’s architecture, its own training

The inference code vendors model.py and schema.py from Jared Palmer’s Kev (Apache-2.0, credited in the bundle’s NOTICE.md and the card’s acknowledgements), and model_manifest.json lists runtime_origin: github.com/jaredpalmer/kev. The design is Kev’s: a backbone with a merged LoRA plus a separate pointer head (q/k projections, head dimension 256).

It is not a re-upload of Kev. We compared the four tensors of NeoHorse’s pointer_head.safetensors with jaredpalmer/kev-4b’s head.pt on 30 Sep 2026, and none of them match. The backbone is also different (NeoHorse-1-4B instead of Qwen3.5-4B-Base). So it gets its own entry, the same way Jev-Omni and this-that-model did after adapting Decider-2B.

TokenRhythm does not publish the training data or recipe for the decision head and LoRA. The vision tower comes unchanged from Qwen3.5-4B (vision_finetuning: false).

Specs

Attribute Value
Company TokenRhythm (X: @opensquilla)
Release 23 Sep 2026 (HF repo 12:42 BRT); official GGUF repo on 24 Sep
Base model NeoHorse-1-4B (upstream Qwen3.5-4B)
Architecture Backbone with merged LoRA + separate decision pointer head (Kev design)
Inputs Text, or one image plus text
Question types choice, noul (catalog boolean), score
Serving vLLM 0.28.0 or SGLang 0.5.17 adapters (GitHub), or the native neohorse_decision runtime with /v1/decision and /v1/systemone
Weights ~9.08 GB backbone, Apache-2.0, on HF and ModelScope
Calibration / latency Not reported by the author

Self-reported results vs the official index

The card’s headline is a six-group text average: 77.70 for NeoHorse-Jev-4B, ahead of Open-Jev-9B (75.67), Kev-4B (74.25) and Laya English (58.24). All runs are TokenRhythm’s own, with benchmark subsets it selected (for example 282 of 324 Nimble examples). UnverifiedClaim.

The official Decision Index Space (data snapshot 28 Sep 2026, 71 systems) measures it much lower:

System Index (balanced skill) Rank ECE
Jev 1.13 57.91 #1 —
Bespoke Nimble 9B v2 39.57 #24 —
NeoHorse-Jev-4B 36.75 #30 0.104
Kev 4B 34.64 #32 —

So the model is slightly ahead of Kev 4B, which fits its lineage, but the index does not support the “best open model” reading of the card. The index lists its base as Qwen3.5-4B-Base and its kind as “head / adapter”; the card says NeoHorse-1-4B.

Fit / anti-fit

Fit when you need: a 4B decision model you can serve on vLLM or SGLang, agent and game-loop decisions, or a Kev-style model with image input.

Anti-fit when you need: calibrated probabilities out of the box (the index measures ECE 0.10; the author says to set thresholds on your own data), published latency, or accuracy close to Jev.

Limits

  • The author reports no NLL, Brier or ECE, and no latency or memory comparison.
  • Training data and recipe are not published.
  • Multiple questions in one request do not share a single forward pass, per the card.
  • Multilingual and long-input behaviour is not evaluated.

Why it is in the catalog

The decision head and LoRA were trained (the head differs from Kev’s) and ship with an open runtime that returns typed decisions with probabilities. That passes the catalog criteria. Its point on the quadrant uses the official index, not the self-reported average.