Decily

Last updated:

Open75–75ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity5.2/10

    Apache-2.0, excellent docs (DESIGN.md, TRAINING.md, TRAIN-REPORT.md), four builds (bf16, MLX, MLX-4bit, ONNX-int8), but day-zero (1 star, 7 downloads), no /v1/systemone server, no SLA.

  • Capability7/10

    In-task accuracy 0.863 beats Decider-2B (0.811) on shared tasks with fewer params (1.7B vs 2B). Zero-shot gap (0.654 vs 0.700) attributed to task coverage (24 vs 95). Calibrated enough to threshold: 93.3% accuracy on top 5% confident predictions.

  • Adoption5/100

    Day-zero launch: 1 GitHub star, 7 HF downloads, one casual mention in a reply thread. No integrations or production use reported.

Vendor claims

Decily is a 1.7B decision model that scores arbitrary candidates in one forward pass. It uses an explicit candidate-scorer architecture (the author calls it “Route B”) instead of reading letter logits from a frozen LM. You give it a state, a question and any set of options; it returns a calibrated probability distribution over them.

It is a full fine-tune of Qwen3-1.7B-Base with RLCD post-training for calibration. The weights are open (Apache-2.0), the training recipe is documented command-by-command, and the author published four export formats for different deployment targets.

Weights: Hugging Face alexzhang0118/Decily-1.7B (+ MLX, MLX-4bit, ONNX-int8). Repo: github.com/arczhi/decily. Mention: reply to @digitalix (26 Sep 2026).

Specs

Attribute Value
Author arczhi / Alex Zhang
Base model Qwen3-1.7B-Base
Parameters 1.7B
License Apache-2.0
HF repo created
Status Open weights (Hugging Face)
Decision types choice (arbitrary candidate sets)
Architecture Explicit candidate-scorer (Route B) — each candidate scored directly, not via letter logits

Builds

Variant Format Size Target
Decily-1.7B PyTorch bf16 safetensors ~3.4 GB server / GPU
Decily-MLX MLX bf16 safetensors ~3.4 GB Apple Silicon
Decily-MLX-4bit MLX 4-bit (group 64) 968 MB on-device / low memory
Decily-ONNX-int8 ONNX int8 1.66 GB pure CPU / Windows / edge

The 4-bit build matches bf16 probabilities to 0.003 at the same speed. The 1.7B backbone is standard Qwen3 (not Qwen3.5’s linear attention), so it exports cleanly to ONNX for CPU deployment.

How it works

Decily uses Route B — an explicit candidate-scorer that scores each option directly — instead of the letter-logits approach (Route A) used by Decider and most Jev alternatives. The author claims +10.6 pt held-out over Route A in a controlled same-backbone comparison.

Training pipeline (all documented in TRAINING.md):

  1. 24 task families + 15% belief data — synthetic stochastic processes with known probability laws teach calibration, not just accuracy
  2. RLCD post-training — belief proper scoring and confidence ordering for usable selective prediction / abstention
  3. Probability-space ensemble → distillation — four diverse models (NLL 1.455) distilled into one (NLL 1.451) at 1× inference cost
  4. Consumer hardware — the v5 full fine-tune is 3000 steps ≈ 1.5h on a single RTX 5090 32GB

Author benchmarks

Everything in this section is author-reported → UnverifiedClaim. All numbers are reported after temperature fitting (raw T=1 logits are over-confident).

Evaluation Decily (1.7B, 24 tasks) Decider-2B (2B, 95 tasks)
In-task, shared training tasks 0.863 / 0.390 NLL / 0.034 ECE 0.811 / 0.453 / 0.032 ECE
Zero-shot fair suite (8 tasks × 300) 0.654 / 0.86 NLL / 0.092 ECE 0.700 / 0.71 NLL / 0.047 ECE
Unseen label sets (60/77-class intents) 0.581 / 1.451 NLL / 0.068 ECE —

How to read this:

  • +5.2 pt on shared tasks with a smaller model (1.7B vs 2B). The author calls this “same-domain winner at a quarter of the task count.”
  • Zero-shot gap of ~4.6 pt (0.654 vs 0.700). The author attributes this to task coverage (24 vs 95), not model capacity, and expects the gap to close if re-run with more tasks. This is a hypothesis, not a tested result.
  • Calibration for abstention: The most-confident 5% of predictions on unseen label sets are 93.3% accurate (vs 79% for the SFT baseline).

The comparison with Decider-2B is done only on tasks held out for both models. TRAIN-REPORT.md §9.23 explains why a naive held-out number would be misleading.

Latency (author-measured)

Device Latency
Apple Silicon (M-series), MLX 4-bit ~75 ms per sentence

All latency numbers are UnverifiedClaim.

Limits

  • Day zero. 1 GitHub star, 7 HF downloads at check time. No production use reported.
  • 24 tasks only. The zero-shot gap vs Decider-2B (95 tasks) is acknowledged. Scaling claim (“re-run with 60–95 tasks”) is untested.
  • No server. No /v1/systemone endpoint or TypeSafe API compatibility out of the box — you call the Python API directly.
  • Choice only. The model scores arbitrary candidates but doesn’t expose score or boolean as first-class types in the API.
  • Self-promotion. The author mentioned this in a reply thread; the model is real and well-documented, but independent verification is pending.

Documentation

Unusually thorough for a day-one open model:

  • DESIGN.md — architecture decisions
  • TRAINING.md — complete command-level recipe, reproducible
  • TRAIN-REPORT.md — experiment log with section references (§9.23–§9.25)
  • MODEL_CARD.md — Hugging Face model card

Why it is in the catalog

Trained decision weights (full fine-tune + RLCD), calibrated probability output, explicit candidate-scorer architecture, and public training recipe. It opens a Route B alternative to the letter-logits approach used by most Jev alternatives. The documentation makes it easy to verify or reproduce.