bev-decider-0.4B

Last updated:

Open—$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity5.1/10

    Open weights, a PyPI package (bev-decider 0.2.1) with a local /v1/systemone server and a public training dataset. The weights are CC-BY-NC-4.0, so commercial use is excluded. Single maintainer, no hosted API or SLA.

  • Capability6.3/10

    Noul / Choice / Score in one forward pass, choice-order invariant by construction. ECE 0.065 is claimed without a published method. No latency figure. Accuracy is author-run on its own split and public sets; not on the Decision Index.

  • Adoption3/100

    Three days old: ~115 HF downloads, 10 likes, a new GitHub repo with no stars. No integrations or production use reported.

Vendor claims

Read this first

  1. Non-commercial weights. The model is released under CC-BY-NC-4.0 because some training sources are non-commercial. The code (GitHub, PyPI) is Apache-2.0, but that does not cover the weights. Do not ship it in a commercial product.
  2. The headline comparison with Jev is on the author’s own split (same distribution as the training data). On sysone-bench, Jev’s published score is 90.7% against 70.8% for bev-decider. All numbers are UnverifiedClaim.

bev-decider-0.4B is a small open decision model by avbiswas, published on Hugging Face on 28 September 2026 (12:58 BRT). It keeps the first 20 of 28 layers of Qwen3-0.6B, merges a rank-8 LoRA into them, and adds a small two-layer attention decision head that compares the options. Every option starts at the same position and cannot see the others inside the backbone, so shuffling the options does not change the answer. It answers Jev-format questions (Noul, Choice, Score) in one forward pass and runs on a laptop CPU.

Specs

Attribute Value
Maker avbiswas (independent)
Base Qwen/Qwen3-0.6B, first 20 of 28 layers, LoRA r=8 merged
Head 2-layer self-attention head, 512-dim, task-type embedding
Parameters ≈0.4B per the card (477.6M counted by Hugging Face, embeddings included)
Training data ~940K questions; dataset avbiswas/bev-decision
Context Trained on states up to 1,024 tokens; 2,048-token default limit; 64 tokens per option
Question types Noul, Choice, Score
Serving pip install "bev-decider[serve]" → bev-decider serve exposes /v1/systemone
License CC-BY-NC-4.0 (weights); Apache-2.0 (code)
Launch

Evidence (UnverifiedClaim)

Benchmark bev-decider-0.4B Jev 1.13 Kev 0.8B Laya
Own held-out split (5,000 q) 74.7 78.0 56.6 50.4
sysone-bench v2 eval 70.8 90.7 (published) — 68.6 (published)
JevBench (public items) 65.8 — — 58.4

The author ran Kev and bev-decider locally; the Jev and Laya rows on sysone-bench are published numbers. ECE on sysone-bench: 0.065.

Fit / anti-fit

Fit: research, hobby and internal prototypes that need a very small local decider (support routing, simple policy checks) with order-independent choices.

Anti-fit: any commercial deployment (licence); long documents (trained at 1,024 tokens); answer verification and temporal reasoning, where the author’s own table shows a gap of 12–21 points to Jev.

Limits

  • No latency figure is published.
  • English only.
  • Calibration method is not described; only the ECE number is given.

Why it is in the catalog

It trains its own decision weights (merged LoRA plus a new head), returns probabilities in the System One format and ships a local /v1/systemone server. The non-commercial licence limits who can use it but not whether it is a decision model, so it is listed with the caveat at the top.