Decily
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity5.2/10
Apache-2.0, excellent docs (DESIGN.md, TRAINING.md, TRAIN-REPORT.md), four builds (bf16, MLX, MLX-4bit, ONNX-int8), but day-zero (1 star, 7 downloads), no /v1/systemone server, no SLA.
- Capability7/10
In-task accuracy 0.863 beats Decider-2B (0.811) on shared tasks with fewer params (1.7B vs 2B). Zero-shot gap (0.654 vs 0.700) attributed to task coverage (24 vs 95). Calibrated enough to threshold: 93.3% accuracy on top 5% confident predictions.
- Adoption5/100
Day-zero launch: 1 GitHub star, 7 HF downloads, one casual mention in a reply thread. No integrations or production use reported.
Vendor claims
- In-task shared: 0.863 acc / 0.390 NLL / 0.034 ECE (vs Decider-2B 0.811 / 0.453 / 0.032)[Vendor claim — not independently verified]
- Zero-shot fair suite (8 tasks × 300): 0.654 acc / 0.86 NLL / 0.092 ECE[Vendor claim — not independently verified]
- Route B explicit candidate-scorer: +10.6 pt held-out over letter-logits route in controlled comparison[Vendor claim — not independently verified]
- ~75 ms per sentence on M-series chips with MLX 4-bit build[Vendor claim — not independently verified]
- Full fine-tune: 3000 steps ≈ 1.5h on a single RTX 5090 32GB[Vendor claim — not independently verified]
Decily is a 1.7B decision model that scores arbitrary candidates in one forward pass. It uses an explicit candidate-scorer architecture (the author calls it “Route B”) instead of reading letter logits from a frozen LM. You give it a state, a question and any set of options; it returns a calibrated probability distribution over them.
It is a full fine-tune of Qwen3-1.7B-Base with RLCD post-training for calibration. The weights are open (Apache-2.0), the training recipe is documented command-by-command, and the author published four export formats for different deployment targets.
Weights: Hugging Face alexzhang0118/Decily-1.7B (+ MLX, MLX-4bit, ONNX-int8). Repo: github.com/arczhi/decily. Mention: reply to @digitalix (26 Sep 2026).
Specs
| Attribute | Value |
|---|---|
| Author | arczhi / Alex Zhang |
| Base model | Qwen3-1.7B-Base |
| Parameters | 1.7B |
| License | Apache-2.0 |
| HF repo created | |
| Status | Open weights (Hugging Face) |
| Decision types | choice (arbitrary candidate sets) |
| Architecture | Explicit candidate-scorer (Route B) — each candidate scored directly, not via letter logits |
Builds
| Variant | Format | Size | Target |
|---|---|---|---|
| Decily-1.7B | PyTorch bf16 safetensors | ~3.4 GB | server / GPU |
| Decily-MLX | MLX bf16 safetensors | ~3.4 GB | Apple Silicon |
| Decily-MLX-4bit | MLX 4-bit (group 64) | 968 MB | on-device / low memory |
| Decily-ONNX-int8 | ONNX int8 | 1.66 GB | pure CPU / Windows / edge |
The 4-bit build matches bf16 probabilities to 0.003 at the same speed. The 1.7B backbone is standard Qwen3 (not Qwen3.5’s linear attention), so it exports cleanly to ONNX for CPU deployment.
How it works
Decily uses Route B — an explicit candidate-scorer that scores each option directly — instead of the letter-logits approach (Route A) used by Decider and most Jev alternatives. The author claims +10.6 pt held-out over Route A in a controlled same-backbone comparison.
Training pipeline (all documented in TRAINING.md):
- 24 task families + 15% belief data — synthetic stochastic processes with known probability laws teach calibration, not just accuracy
- RLCD post-training — belief proper scoring and confidence ordering for usable selective prediction / abstention
- Probability-space ensemble → distillation — four diverse models (NLL 1.455) distilled into one (NLL 1.451) at 1× inference cost
- Consumer hardware — the v5 full fine-tune is 3000 steps ≈ 1.5h on a single RTX 5090 32GB
Author benchmarks
Everything in this section is author-reported → UnverifiedClaim. All numbers are reported after temperature fitting (raw T=1 logits are over-confident).
| Evaluation | Decily (1.7B, 24 tasks) | Decider-2B (2B, 95 tasks) |
|---|---|---|
| In-task, shared training tasks | 0.863 / 0.390 NLL / 0.034 ECE | 0.811 / 0.453 / 0.032 ECE |
| Zero-shot fair suite (8 tasks × 300) | 0.654 / 0.86 NLL / 0.092 ECE | 0.700 / 0.71 NLL / 0.047 ECE |
| Unseen label sets (60/77-class intents) | 0.581 / 1.451 NLL / 0.068 ECE | — |
How to read this:
- +5.2 pt on shared tasks with a smaller model (1.7B vs 2B). The author calls this “same-domain winner at a quarter of the task count.”
- Zero-shot gap of ~4.6 pt (0.654 vs 0.700). The author attributes this to task coverage (24 vs 95), not model capacity, and expects the gap to close if re-run with more tasks. This is a hypothesis, not a tested result.
- Calibration for abstention: The most-confident 5% of predictions on unseen label sets are 93.3% accurate (vs 79% for the SFT baseline).
The comparison with Decider-2B is done only on tasks held out for both models. TRAIN-REPORT.md §9.23 explains why a naive held-out number would be misleading.
Latency (author-measured)
| Device | Latency |
|---|---|
| Apple Silicon (M-series), MLX 4-bit | ~75 ms per sentence |
All latency numbers are UnverifiedClaim.
Limits
- Day zero. 1 GitHub star, 7 HF downloads at check time. No production use reported.
- 24 tasks only. The zero-shot gap vs Decider-2B (95 tasks) is acknowledged. Scaling claim (“re-run with 60–95 tasks”) is untested.
- No server. No
/v1/systemoneendpoint or TypeSafe API compatibility out of the box — you call the Python API directly. - Choice only. The model scores arbitrary candidates but doesn’t expose
scoreorbooleanas first-class types in the API. - Self-promotion. The author mentioned this in a reply thread; the model is real and well-documented, but independent verification is pending.
Documentation
Unusually thorough for a day-one open model:
DESIGN.md— architecture decisionsTRAINING.md— complete command-level recipe, reproducibleTRAIN-REPORT.md— experiment log with section references (§9.23–§9.25)MODEL_CARD.md— Hugging Face model card
Why it is in the catalog
Trained decision weights (full fine-tune + RLCD), calibrated probability output, explicit candidate-scorer architecture, and public training recipe. It opens a Route B alternative to the letter-logits approach used by most Jev alternatives. The documentation makes it easy to verify or reproduce.
