Decider-4B

Last updated:

Open5–35ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity6.8/10

    Open weights under Apache-2.0, full HF documentation, TypeSafe API-compatible (POST /v1/systemone), reproducible training docs, per-type temperature maps — no production SLA keeps it below commercial tiers.

  • Capability8.2/10

    JevBench public score 64.13 (rank #1 above Jev); all decision types (choice/score/boolean); 75.0 calibration; speed 92.9 and cost 60.9 on JevBench v1.4.2 — wins on efficiency, not raw intelligence.

  • Adoption24/100

    Inherits momentum from decider-2b wave; v2.1 release went viral as first open model to beat Jev; still niche outside S1-focused builders.

Vendor claims

Decider-4B is the first open model to beat Jev on JevBench — a 4.2B dense model from Mapika with no reinforcement-learning stage, achieving rank #1 via superior speed (92.9) and cost efficiency (60.9), not raw intelligence (49.4 vs Jev’s higher reasoning scores).

Built on Qwen3.5-4B-Base (32 layers, 8 full attention + 24 gated delta-net linear attention), it reads shared state plus typed questions (choice, score, noul → catalog boolean) and returns calibrated distributions in one forward pass. Wire format targets TypeSafe POST /v1/systemone.

JevBench v1.4.2 Performance

Metric Score Notes
Overall 64.13 Rank #1 — first open to beat Jev
Intelligence 49.4 Below Jev on raw reasoning
Calibration 75.0 Good probability calibration
Speed 92.9 5.2ms with CUDA graphs
Cost ~60.9 ~$0.03/1k decisions

How it won: Decider-4B beats Jev on the composite JevBench score by being dramatically faster and cheaper to run, not by being smarter. On hard reasoning items, Jev still leads — but for production workloads where speed and cost matter, Decider-4B offers a compelling open alternative.

Specs

Attribute Value
Base model Qwen/Qwen3.5-4B-Base
Parameters 4.2B (dense)
Architecture 32 layers: 8 full attention + 24 gated delta-net linear attention
Weights 8.4 GB bf16
Current version v2.1 (2026-09-24)
License Apache-2.0
Stage 1 Cross-entropy on mixture v2 (742M tokens)
Stage 2 LoRA rank 64 + 29,325 harder decisions (merged)
RL stage None (unlike decider-2b)
Latency 5.2ms (CUDA graphs) / 32ms (eager) on B300
API POST /v1/systemone (TypeSafe-compatible)
Decision types choice, score, noul → boolean

Training (two stages, no RL)

Stage 1 (v1): One supervised pass over mixture v2 — ~95 public decision datasets, agent trajectories, Mind2Web, 26 additional public datasets, and 10 programmatic families with verifiable gold. 1.89M items, 742M tokens. AdamW on bf16 parameters directly (no FP32 master copy — this preserves more base-model knowledge).

Stage 2 (v2.1): LoRA rank 64 (alpha 128) on attention and MLP weights, 2 epochs over 29,325 rows of harder decisions. Key innovation: replay rows trained toward v1’s own answer distribution (KL divergence loss), not hard labels. This preserves sampled-play performance that v2 lost. Merged into bf16 weights after training.

No reinforcement-learning stage — unlike decider-2b, the 4B model skips RLCD entirely and still achieves top JevBench scores via efficient supervised training.

The Decider Family

Model Base Weights Use case
decider-0.8b Qwen3.5-0.8B 1.4 GB Smallest: routing, yes/no, short-state lookups
decider-2b Qwen3.5-2B 3.8 GB Default: routing, classification, browser agents
decider-4b (this) Qwen3.5-4B 8.4 GB Middle point: harder judgments, no RL stage
decider-35b-a3b Qwen3.5-35B-A3B (3B active) 65 GB When accuracy is worth 3–4× cost
decider-2b-vision Qwen3.5-2B VL 4.1 GB Decisions from images

All share one interface (decider.infer.Decider, POST /v1/systemone) and one readout mechanism.

Usage

from decider.infer import Decider
d = Decider("Mapika/decider-4b")

d.decide(
    "My card was charged twice for the same purchase.",
    [
        {"question": "Which department should handle this?", 
         "options": ["billing", "technical support", "sales"]},
        {"question": "Does this need a refund action?", 
         "options": ["no", "yes"]}
    ]
)
# [{'choice': 'billing', 'confidence': ..., 'probs': {...}}, 
#  {'choice': 'yes', 'confidence': ..., 'probs': {...}}]

Requirements: torch, transformers>=5, flash-linear-attention (Triton kernels for Qwen3.5 linear-attention layers). The per-type temperatures need decider-ai 1.4.0 or later.

Version History

Version Date Changes
v2.1 2026-09-24 Replay trained toward v1’s distribution; per-type temperature map
v2 2026-09-24 Stage 2 LoRA on harder decisions; better on hard judgments, worse on sampled play
v1 2026-09-22 First release: mixture v2 on Qwen3.5-4B-Base

Which version to use:

  • v2.1 (default): sampled play, games, browser agents, hard decisions
  • v2: if you need better-calibrated confidence on hard multi-step items
  • v1: form filling, BabyAI-GoTo navigation, greedy bag-draw play

Limitations

  1. No RL stage: Stated beliefs about action outcomes were not trained against exact laws. On live browser tasks v2.1 is 93.2% sampled; decider-2b v10 with RL is 93.2% but 91.7% on held-out tasks.

  2. Overconfident on hard items: Calibration error 0.147 on held-out generated families (v2: 0.046) and 0.184 on JevBench hard tier (v2: 0.104). Don’t read 0.8 confidence as 80% accuracy on hard questions.

  3. Known regressions vs v1: Issue #9 form case still wrong; BabyAI-GoTo 0.19 vs v1’s 0.54; greedy bag-draw 9.4 points under v1.

  4. Below the 35B: 2.4/2.6 points under decider-35b-a3b on regression set; 2.7 points on JevBench hard tier.

  5. English-centric: Multilingual rows (XNLI, PAWS-X, MASSIVE, Belebele, XCOPA) are a small share of training data.

  6. Schema cache disabled: v2.1’s LoRA was trained only in plain state-first layout. DECIDER_SCHEMA_CACHE=1 won’t activate for this model.

  7. Per-type temps need 1.4.0: Older decider-ai versions use global temperature 1.099 for all answers (same argmax, different probabilities).

Evaluation Summary (v2.1)

Benchmark Result
Regression set (in-task / held-out) 0.831 / 0.784
JevBench public (easy / standard / hard) 1.000 / 0.986 / 0.649
Live MiniWoB++ (sampled / greedy) 93.2% / 93.8%
Bespoke suite (macro / micro) 0.756 / 0.765
Zero-shot games (sampled win rate) 26.9%
CliffWalking −13 (teacher optimal)

Full evaluation tables with confidence intervals are in the HuggingFace model card.