Decider-4B
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity6.8/10
Open weights under Apache-2.0, full HF documentation, TypeSafe API-compatible (POST /v1/systemone), reproducible training docs, per-type temperature maps — no production SLA keeps it below commercial tiers.
- Capability8.2/10
JevBench public score 64.13 (rank #1 above Jev); all decision types (choice/score/boolean); 75.0 calibration; speed 92.9 and cost 60.9 on JevBench v1.4.2 — wins on efficiency, not raw intelligence.
- Adoption24/100
Inherits momentum from decider-2b wave; v2.1 release went viral as first open model to beat Jev; still niche outside S1-focused builders.
Vendor claims
- Regression set 0.831 in-task / 0.784 held-out accuracy[Vendor claim — not independently verified]
- Live MiniWoB++ 93.2% sampled success (v2.1)[Vendor claim — not independently verified]
- 5.2ms latency with CUDA graphs on B300[Vendor claim — not independently verified]
- TypeSafe wire format POST /v1/systemone; noul maps to catalog boolean[Vendor claim — not independently verified]
Decider-4B is the first open model to beat Jev on JevBench — a 4.2B dense model from Mapika with no reinforcement-learning stage, achieving rank #1 via superior speed (92.9) and cost efficiency (60.9), not raw intelligence (49.4 vs Jev’s higher reasoning scores).
Built on Qwen3.5-4B-Base (32 layers, 8 full attention + 24 gated delta-net linear attention), it reads shared state plus typed questions (choice, score, noul → catalog boolean) and returns calibrated distributions in one forward pass. Wire format targets TypeSafe POST /v1/systemone.
JevBench v1.4.2 Performance
| Metric | Score | Notes |
|---|---|---|
| Overall | 64.13 | Rank #1 — first open to beat Jev |
| Intelligence | 49.4 | Below Jev on raw reasoning |
| Calibration | 75.0 | Good probability calibration |
| Speed | 92.9 | 5.2ms with CUDA graphs |
| Cost | ~60.9 | ~$0.03/1k decisions |
How it won: Decider-4B beats Jev on the composite JevBench score by being dramatically faster and cheaper to run, not by being smarter. On hard reasoning items, Jev still leads — but for production workloads where speed and cost matter, Decider-4B offers a compelling open alternative.
Specs
| Attribute | Value |
|---|---|
| Base model | Qwen/Qwen3.5-4B-Base |
| Parameters | 4.2B (dense) |
| Architecture | 32 layers: 8 full attention + 24 gated delta-net linear attention |
| Weights | 8.4 GB bf16 |
| Current version | v2.1 (2026-09-24) |
| License | Apache-2.0 |
| Stage 1 | Cross-entropy on mixture v2 (742M tokens) |
| Stage 2 | LoRA rank 64 + 29,325 harder decisions (merged) |
| RL stage | None (unlike decider-2b) |
| Latency | 5.2ms (CUDA graphs) / 32ms (eager) on B300 |
| API | POST /v1/systemone (TypeSafe-compatible) |
| Decision types | choice, score, noul → boolean |
Training (two stages, no RL)
Stage 1 (v1): One supervised pass over mixture v2 — ~95 public decision datasets, agent trajectories, Mind2Web, 26 additional public datasets, and 10 programmatic families with verifiable gold. 1.89M items, 742M tokens. AdamW on bf16 parameters directly (no FP32 master copy — this preserves more base-model knowledge).
Stage 2 (v2.1): LoRA rank 64 (alpha 128) on attention and MLP weights, 2 epochs over 29,325 rows of harder decisions. Key innovation: replay rows trained toward v1’s own answer distribution (KL divergence loss), not hard labels. This preserves sampled-play performance that v2 lost. Merged into bf16 weights after training.
No reinforcement-learning stage — unlike decider-2b, the 4B model skips RLCD entirely and still achieves top JevBench scores via efficient supervised training.
The Decider Family
| Model | Base | Weights | Use case |
|---|---|---|---|
| decider-0.8b | Qwen3.5-0.8B | 1.4 GB | Smallest: routing, yes/no, short-state lookups |
| decider-2b | Qwen3.5-2B | 3.8 GB | Default: routing, classification, browser agents |
| decider-4b (this) | Qwen3.5-4B | 8.4 GB | Middle point: harder judgments, no RL stage |
| decider-35b-a3b | Qwen3.5-35B-A3B (3B active) | 65 GB | When accuracy is worth 3–4× cost |
| decider-2b-vision | Qwen3.5-2B VL | 4.1 GB | Decisions from images |
All share one interface (decider.infer.Decider, POST /v1/systemone) and one readout mechanism.
Usage
from decider.infer import Decider
d = Decider("Mapika/decider-4b")
d.decide(
"My card was charged twice for the same purchase.",
[
{"question": "Which department should handle this?",
"options": ["billing", "technical support", "sales"]},
{"question": "Does this need a refund action?",
"options": ["no", "yes"]}
]
)
# [{'choice': 'billing', 'confidence': ..., 'probs': {...}},
# {'choice': 'yes', 'confidence': ..., 'probs': {...}}]
Requirements: torch, transformers>=5, flash-linear-attention (Triton kernels for Qwen3.5 linear-attention layers). The per-type temperatures need decider-ai 1.4.0 or later.
Version History
| Version | Date | Changes |
|---|---|---|
| v2.1 | 2026-09-24 | Replay trained toward v1’s distribution; per-type temperature map |
| v2 | 2026-09-24 | Stage 2 LoRA on harder decisions; better on hard judgments, worse on sampled play |
| v1 | 2026-09-22 | First release: mixture v2 on Qwen3.5-4B-Base |
Which version to use:
- v2.1 (default): sampled play, games, browser agents, hard decisions
- v2: if you need better-calibrated confidence on hard multi-step items
- v1: form filling, BabyAI-GoTo navigation, greedy bag-draw play
Limitations
-
No RL stage: Stated beliefs about action outcomes were not trained against exact laws. On live browser tasks v2.1 is 93.2% sampled; decider-2b v10 with RL is 93.2% but 91.7% on held-out tasks.
-
Overconfident on hard items: Calibration error 0.147 on held-out generated families (v2: 0.046) and 0.184 on JevBench hard tier (v2: 0.104). Don’t read 0.8 confidence as 80% accuracy on hard questions.
-
Known regressions vs v1: Issue #9 form case still wrong; BabyAI-GoTo 0.19 vs v1’s 0.54; greedy bag-draw 9.4 points under v1.
-
Below the 35B: 2.4/2.6 points under decider-35b-a3b on regression set; 2.7 points on JevBench hard tier.
-
English-centric: Multilingual rows (XNLI, PAWS-X, MASSIVE, Belebele, XCOPA) are a small share of training data.
-
Schema cache disabled: v2.1’s LoRA was trained only in plain state-first layout.
DECIDER_SCHEMA_CACHE=1won’t activate for this model. -
Per-type temps need 1.4.0: Older
decider-aiversions use global temperature 1.099 for all answers (same argmax, different probabilities).
Evaluation Summary (v2.1)
| Benchmark | Result |
|---|---|
| Regression set (in-task / held-out) | 0.831 / 0.784 |
| JevBench public (easy / standard / hard) | 1.000 / 0.986 / 0.649 |
| Live MiniWoB++ (sampled / greedy) | 93.2% / 93.8% |
| Bespoke suite (macro / micro) | 0.756 / 0.765 |
| Zero-shot games (sampled win rate) | 26.9% |
| CliffWalking | −13 (teacher optimal) |
Full evaluation tables with confidence intervals are in the HuggingFace model card.
Links
- Mapika/decider-4b on HuggingFace
- GitHub: Mapika/decider
- Decider-2B — the 2B sibling with RL stage
- Jev — the TypeSafe commercial model it beats on JevBench
- JevBench — the benchmark where it achieved rank #1
