Decider-2B

Last updated:

Open50–400ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity6.4/10

    Full train/serve stack, HF cards, and System One API docs put Decider near the top of open maturity; self-host only (no SLA) keeps it under Laya’s packaging lead.

  • Capability5.2/10

    JevBench v1.4.2 rank #34 (30.7) — solid calibration (69.7) and speed (70.9) but mid-tier intelligence (40.9); see Decider-4B (#1, 64.13) for stronger variant.

  • Adoption16/100

    Visible in the open-Jev / S1Bench wave with a serious repo and HF cards; still niche outside builders already hunting replicas.

Vendor claims

Family update: The newer Decider-4B (v2.1, September 2026) is now #1 on JevBench with 64.13 — the first open model to beat hosted Jev. This 2B variant remains useful for constrained deployments but sits at #34 (30.7).

Decider-2B is an independent open System One–style decision model — full fine-tune of Qwen3.5-2B-Base plus calibration-aware RL (author v10) — not a TypeSafe product and not a LoRA-only adapter.

It reads shared state plus typed questions (choice, score, noul → catalog boolean) and returns calibrated distributions in one forward pass. Wire format targets TypeSafe POST /v1/systemone. Weights: Mapika/decider-2b; code: github.com/Mapika/decider. Smaller sibling: decider-0.8b.

JevBench v1.4.2

Metric Decider-2B Decider-4B Rank
Overall 30.7 64.13 #34 / #1
Intelligence 40.9 49.4 -
Calibration 69.7 75.0 -
Speed 70.9 92.9 -
Cost 64.4 60.9 -

The 4B variant’s jump to #1 came from speed and cost optimization, not raw intelligence — both share similar training recipes but the larger backbone handles calibration better.

Limits (read before you ship)

  • Author benches (browser, games, public JevBench items, Bespoke suite) are self-reported — not ModelSystem.One evals; comps vs Jev / peers are UnverifiedClaim.
  • English-centric; knowledge-heavy / multi-hop items stay weak at 2B without splitting reasoning.
  • Schema-cache layouts trade some accuracy for speed; fit temperature on your labels.
  • Catalog latency band 50–400ms is a placeholder envelope, not a measured SLA.

Why it is in the catalog

Public typed decision I/O, open weights + reproducible train/serve docs, and System One API compatibility. Stronger open full-model peer next to Laya and NanoJev; sits below hosted Jev on adoption.