Choosing an open Decision Model
Start with the workload, not a leaderboard
There is no defensible general claim that an open model is “10× cheaper and faster” than Jev. A hosted API price, a local one-question demo, and a production self-hosted deployment measure different things. Compare models on the same labelled cases, schema, concurrency, hardware, and escalation policy.
Use a Decision Model only when the output is a bounded decision. Keep generation, arithmetic, date comparison, and multi-hop reasoning in code or a System Two model.
What the catalog can support today
| Option | Shape | Best reason to evaluate it | Do not infer |
|---|---|---|---|
| Jev | Hosted typed-decision API | Lowest operational burden and full choice / score / boolean contract |
Vendor speed, cost, or accuracy claims are independent benchmarks |
| Laya | Open ModernBERT full model | Small encoder-style local path | Generic zero-shot quality or calibrated probabilities; its published typed-decision zero-shot result is weak |
| Decider | Open 2B full model | Jev-like wire format and all three primitives | Author browser and Jev comparisons are a production guarantee |
| Nimble | LoRA on Qwen3.5-9B | Open recipe for a schema-shaped classifier | “Free” inference; it needs the base model and suitable hardware |
| NanoJev | Open 0.6B full model | Small research stack with structured heads | Its game demos transfer to your semantic workload |
All open entries above have no per-token API fee in the catalog. That means marginal API price is zero, not that total cost is zero: include GPU/CPU, memory, serving, observability, uptime, and the engineer operating it.
A comparison that is useful
- Freeze a labelled evaluation set from your real decision boundary. Keep a holdout set untouched.
- Give each candidate the same state, criteria, choices, option order variants, timeout, and concurrency.
- Record accuracy (or F1), calibration error, p50/p95 latency, serialized throughput, failures, and total cost assumptions.
- Run the reversed-option variant. A model that changes its answer when candidates move is not ready for an automatic branch.
- Choose the smallest operating burden that meets your accuracy and calibration threshold; escalate low-confidence cases instead of forcing a winner.
The repository includes a local evaluator at scripts/evaluate-decision-models.mjs. It consumes normalized JSONL observations and produces a reproducible report; see the benchmark protocol. It intentionally does not publish a ranking from vendor or community screenshots.
Cost worksheet
For a hosted model, estimate input tokens × posted price plus any gateway and escalation cost. For self-hosting, estimate:
monthly infrastructure + operations labor + observability
--------------------------------------------------------- = cost per decision
successful decisions in the month
Report this beside, rather than in place of, the model’s marginal API price. It makes a low-volume hosted API and a high-volume dedicated GPU comparable without pretending their economics are identical.
Decision rule
Choose Jev when managed availability and its typed contract matter more than local control. Evaluate a portable open model when data residency, offline operation, or high sustained volume justifies owning the serving path. Keep an open model in a shadow or escalation-safe path until its holdout accuracy, calibration, and option-order stability meet the threshold for the decision’s risk.
