Decision 1.0 Lux-9B

Last updated:

Open51–347ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity5.1/10

    Apache-2.0, ungated, detailed card with methods, 54-task evaluation, diagnostics and attribution; runs in stock Transformers via trust_remote_code with system_one(). Serving needs the separately distributed vLLM Semantic Router Decision runtime. No hosted API or SLA.

  • Capability7.7/10

    Choice (2–255), Noul and Score (2–10) with a shared candidate head. One temperature fitted; official Decision Index ECE 0.076. Median 51 ms / p95 347 ms on the index. Index skill 43.49, #14 of 70: the strongest open entry of its family, well below Jev. 16,384-token input limit.

  • Adoption4/100

    112 downloads and 13 likes. Listed on the official Decision Index with the rest of the Decision 1.0 family. No integrations or production cases reported.

Vendor claims

Read this first

  1. The third-party number is solid but not Jev-class. The official Decision Index scores Lux-9B 43.49 (#14 of 70), ECE 0.076, against Jev’s 57.91. The index run is from 28 September; we assume it used the current weights (last changed 22 September).
  2. Option order moves some answers. The authors’ own test flips 8.33% of answers when options are reordered (Jev: 0%).
  3. The organisation was renamed. The Hugging Face org moved from llm-semantic-router to vllm-sr on 2 October; old links redirect.
  4. There is a successor. Decision 2.0 Lux-9B is fine-tuned from these weights; the team reports 46.3 on the index (+2.8) and 18.4 ms median, in its own run. This page stays on 1.0 because 43.49 is the official third-party number. See the Decision 2.0 news.

Decision 1.0 Lux-9B is the largest open model in the vLLM Semantic Router project’s Decision 1.0 family, published on 22 September 2026. It adapts the Qwen3.5-9B text backbone with full-parameter training, removes the generation head and adds a shared candidate head that returns one probability per supplied answer. It sits alongside the family’s small encoder, Decision 1.0 Kai.

Specs

Attribute Value
Maker vllm-sr (vLLM Semantic Router), formerly llm-semantic-router
Base Qwen/Qwen3.5-9B (text backbone; vision tower and LM head removed)
Parameters 7.94B (backbone 7.937B + 4.2M candidate head)
Question types Choice (2–255 options), Noul, Score (2–10 levels)
Context 16,384 tokens per question including candidates and state; longer fails the request (max_length_exceeded)
Calibration One temperature fitted on 1,814 held-out examples
Usage AutoModel.from_pretrained(..., trust_remote_code=True).system_one(...); pipeline("decision"); or the vLLM Semantic Router Decision runtime (/v1/systemone)
License Apache-2.0
Launch

Evidence

Source Number
Official Decision Index 0.2.1 (28/09) 43.49 skill, #14 of 70; ECE 0.076; median 51 ms / p95 347 ms
Authors’ 54-task suite (UnverifiedClaim) 77.40 vs Kev-9B 71.89, Nox-4B 73.09, Jev 81.05
Authors’ diagnostics (UnverifiedClaim) Brier 0.291, ECE 4.79%; order flip 8.33%

The authors note that their latency chart was measured with earlier weights on an AMD GPU.

Fit / anti-fit

Fit: self-hosted routing and checks on a single GPU where you want open weights, long inputs (up to 16k tokens) and probabilities you can threshold.

Anti-fit: CPU-only deployments (FP32 9B); cases that need Jev-level accuracy or answers that never change with option order; anyone who needs a hosted API.

Why it is in the catalog

Lux is trained for decisions (full-parameter adaptation plus its own candidate head and calibration), released openly and answers in the System One format with probabilities. It was a watch item on the Kai page and gets its own page now that its third-party index result is the family’s strongest.