Decider-2B
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity6.4/10
Full train/serve stack, HF cards, and System One API docs put Decider near the top of open maturity; self-host only (no SLA) keeps it under Laya’s packaging lead.
- Capability5.2/10
JevBench v1.4.2 rank #34 (30.7) — solid calibration (69.7) and speed (70.9) but mid-tier intelligence (40.9); see Decider-4B (#1, 64.13) for stronger variant.
- Adoption16/100
Visible in the open-Jev / S1Bench wave with a serious repo and HF cards; still niche outside builders already hunting replicas.
Vendor claims
- Live MiniWoB++ click tasks ~93% sampled success on v10 (author; v8 ~83%)[Vendor claim — not independently verified]
- TypeSafe wire format POST /v1/systemone; noul maps to catalog boolean[Vendor claim — not independently verified]
- Also ships Mapika/decider-0.8b sibling with the same supervised recipe[Vendor claim — not independently verified]
Family update: The newer Decider-4B (v2.1, September 2026) is now #1 on JevBench with 64.13 — the first open model to beat hosted Jev. This 2B variant remains useful for constrained deployments but sits at #34 (30.7).
Decider-2B is an independent open System One–style decision model — full fine-tune of Qwen3.5-2B-Base plus calibration-aware RL (author v10) — not a TypeSafe product and not a LoRA-only adapter.
It reads shared state plus typed questions (choice, score, noul → catalog boolean) and returns calibrated distributions in one forward pass. Wire format targets TypeSafe POST /v1/systemone. Weights: Mapika/decider-2b; code: github.com/Mapika/decider. Smaller sibling: decider-0.8b.
JevBench v1.4.2
| Metric | Decider-2B | Decider-4B | Rank |
|---|---|---|---|
| Overall | 30.7 | 64.13 | #34 / #1 |
| Intelligence | 40.9 | 49.4 | - |
| Calibration | 69.7 | 75.0 | - |
| Speed | 70.9 | 92.9 | - |
| Cost | 64.4 | 60.9 | - |
The 4B variant’s jump to #1 came from speed and cost optimization, not raw intelligence — both share similar training recipes but the larger backbone handles calibration better.
Limits (read before you ship)
- Author benches (browser, games, public JevBench items, Bespoke suite) are self-reported — not ModelSystem.One evals; comps vs Jev / peers are UnverifiedClaim.
- English-centric; knowledge-heavy / multi-hop items stay weak at 2B without splitting reasoning.
- Schema-cache layouts trade some accuracy for speed; fit temperature on your labels.
- Catalog latency band 50–400ms is a placeholder envelope, not a measured SLA.
Why it is in the catalog
Public typed decision I/O, open weights + reproducible train/serve docs, and System One API compatibility. Stronger open full-model peer next to Laya and NanoJev; sits below hosted Jev on adoption.
