Decision 2.0 (Kai-0.6B to Vega-27B)

Last updated:

Open5–71ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity4.9/10

    Six Apache-2.0 sizes with trained decision heads; system_one() and a pipeline in Transformers via trust_remote_code. Vega-27B is a LoRA that needs Qwen3.8-27B. Cards are short: quickstart and one eval table, no methods, data or calibration notes. Community MLX / ONNX ports. No API, no SLA.

  • Capability6.7/10

    Choice (up to 255 options) / Noul / Score in one pass. The manifest marks calibration as unset, with no ECE. Median 4.9–71.4 ms per single-question request on one unnamed GPU (authors). Vega's DI 56.5 vs Jev 57.91 is the team's own run with the official kit, off the board.

  • Adoption7/100

    Launch post with ~740 likes / 39k views / 680 bookmarks; collection 24 upvotes; ~560 downloads across six repos in about a day; six community MLX / ONNX conversions (incl. onnx-community). No integrations or production use reported.

Vendor claims

  • Decision-2.0-Vega-27B: Decision Index 0.2.1 56.5 (Jev 57.91 on the public board), JevArena 74.0 vs AutoJev-27B 72.1, human-labelled transfer 58.7; 'independent reproduction with the official 0.2.1 kit' run by the authors; '+11.2 over Lux 2.0' (published scores differ by 10.2)[Vendor claim — not independently verified]
  • Decision Index 1.0 → 2.0: Lux-9B 43.5 → 46.3, Nox-4B 34.4 → 43.8, Sol-2B 25.3 → 29.5, Eos-0.8B 18.4 → 20.1, Kai-0.6B 6.5 → 16.3 (1.0 values from the public board, 2.0 values self-run)[Vendor claim — not independently verified]
  • Median latency per single-question request on a single GPU: Kai 4.9 ms, Eos 6.0, Sol 7.2, Nox 12.9, Lux 18.4, Vega 71.4[Vendor claim — not independently verified]

Read this first

  1. Every 2.0 index score is the team’s own run. The cards call it an “independent reproduction with the official 0.2.1 kit”, but the authors ran it; the public Decision Index still shows the 1.0 models from 28 September. The Vega card’s “+11.2 over Lux 2.0” does not match its own table (56.5 − 46.3 = 10.2).
  2. No calibration step. The release manifest lists calibration as unset, and the cards give no ECE. The probabilities are the head’s raw output.
  3. The 1.0 pages stay. Decision 1.0 Lux-9B and Kai keep their pages because their scores rest on the official index. This page covers the 2.0 family as a whole.

Decision 2.0 is the second open decision family from the vLLM Semantic Router project, announced on 2 October 2026 (news). The weights went up on 28–29 September. There are six sizes, all Apache-2.0: Kai-0.6B, Eos-0.8B, Sol-2B, Nox-4B, Lux-9B and a new top model, Vega-27B. Each one takes a state and typed questions and returns a probability for every option without generating text, through the same system_one() call as 1.0.

The six sizes

Model Parameters Built from Decision Index (1.0 → 2.0) Median latency
Kai-0.6B 0.60B Qwen3-0.6B-Base (1.0 Kai was an mmBERT-lineage encoder) 6.5 → 16.3 4.9 ms
Eos-0.8B 0.75B Decision 1.0 Eos 18.4 → 20.1 6.0 ms
Sol-2B 1.88B Decision 1.0 Sol 25.3 → 29.5 7.2 ms
Nox-4B 4.21B Qwen3.5-4B-Base 34.4 → 43.8 12.9 ms
Lux-9B 7.94B Decision 1.0 Lux 43.5 → 46.3 18.4 ms
Vega-27B 29.37B Qwen3.8-27B + rank-512 LoRA (adapter only; base downloaded separately) — → 56.5 71.4 ms

All six share one design, per their configs: a shared bilinear-MLP decision head that scores each option against a global query, up to 255 options per question. Context is 16,384 tokens for most sizes, 8,192 for Kai and 32,768 for Vega.

Evidence (UnverifiedClaim)

From the model cards:

Model JevArena Human-labelled transfer Compared with
Vega-27B 74.0 58.7 AutoJev-27B 72.1 / 58.7, Eikos-27B 69.3, Jebadiah-27B 65.5
Lux-9B 68.1 56.2 Decision 1.0 Lux 65.8, Nimble v2 62.1
Nox-4B 63.6 52.3 Decider 4B 61.9 / 55.5, Jet v6.2 60.4
Sol-2B 52.1 51.3 Decider 2B 49.5, This-That 1.2 46.1
Eos-0.8B 53.9 50.3 Intern-Decision-0.8B 43.5, Kev-0.8B 43.2
Kai-0.6B 48.6 45.9 Decision 1.0 Kai 35.9

JevArena is the team’s own comparison of same-size open models (frozen prompts, invalid answers count as errors). Each card leads its size on that table, but Nox-4B trails Decider 4B on the human-labelled transfer column. The cards say the training data was “audited at row level against all Index test items”.

Launch claims as amplified (3 Oct)

Xunzhuo Liu’s LinkedIn launch post and a widely shared summary by David Hendrickson (@TeksEdge, 3 Oct, 16:00 BRT; about 108 likes and 5.8k views) add claims that are not in the cards (UnverifiedClaim):

  • “#1 at its size on the Jev Decision Index” at 0.6B, 0.8B, 2B and 4B, and “the 27B is #3 overall”. These compare the team’s own 2.0 runs with the 28 Sep board.
    • Nox-4B leads JPT-4B by 0.8 points (43.8 vs 43.04), and Sol-2B leads Decider 2B by 0.5 (29.5 vs 28.97). A re-run could erase margins that small.
    • Eos-0.8B (20.1 vs JPT-0.8B 19.22) and Kai-0.6B (16.3 vs Bosun 0.6B 14.32) lead more clearly.
    • Vega-27B (56.5) would be #3 among open systems, behind Surogate Rune (57.44) and Decider chat Gemma-4-31B (57.33). Counting Jev (57.91), it would be #4.
  • “64 questions about one request, answered in 63 ms on a single GPU.” The post names no model size or GPU. The cards only say that many questions about one input are answered in one pass.
  • “20× faster than the Jev API for a single decision.” No size or setup is given. The hosted Jev median on the index is 524 ms, and that includes the network.
  • “Up to 2.5× the Decision 1.0 score at the same size.” This matches Kai: 6.5 → 16.3.

No score change: accuracy stays at 6 until the 2.0 models are on the public index. By 4 Oct the collection had 49 upvotes and about 1,960 downloads across the six repos. A seventh repo, Decision-2.0-Sol-2B-Reasoning, had appeared with no card numbers yet.

Fit / anti-fit

Fit: self-hosted routing and checks where you want to pick a size for your latency budget (5–70 ms medians on one GPU); trying a 27B open model that the team puts close to Jev, with the 1.0 pages as the third-party baseline.

Anti-fit: anything that needs calibrated probabilities out of the box, a hosted API or a /v1/systemone server; teams that need methods and training data documented before adopting (the 2.0 cards have neither).

Why it is in the catalog

Each 2.0 model ships trained decision weights (a full fine-tune, or a LoRA for Vega) plus its own decision head, returns a probability per option, and covers Choice, Noul and Score. That is the catalog test. One family page, as with StartLux-Decision, because the six share one recipe and one card format. Accuracy stays at 6 until the 2.0 models show up on the public index or in someone else’s run.

GGUF builds (8 Oct 2026)

vLLM Semantic Router published GGUF builds of all six sizes (for example Vega-27B-GGUF and Kai-0.6B-GGUF). On the official Decision Index 0.3, Vega 27B scores 55.88 (rank 13) and Nox 4B 44.95 (rank 38). A “System One Auto” mode that runs Kai first and falls back to Vega was announced on X on 10 Oct; we could not open a primary source, so its cost and latency figures are not listed here. Scores unchanged.