Jev-Style-0.8B-Decision-v3
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity5.6/10
Open weights in torch, MLX and GGUF, PyPI package (jev-style 0.2.0), a /v1/systemone-shaped local server, MCP tools, a Claude Code guard and six agent skills. All first-party; the runtime repo is a day old. Self-host only, no SLA.
- Capability6.8/10
First-class noul / choice / score, per-group temperatures fitted on held-out rows, 25.6k-token inputs. Strong, carefully documented evals, but self-run; wins are against Laya, and the author says hosted Jev is ahead on every set where Jev has numbers.
- Adoption12/100
Family downloads are real (the 2B v1 GGUF alone ~6k), but v3 is two days old: ~790 downloads and 12 likes across its builds, 3 GitHub stars, no launch post found. No production use reported.
Vendor claims
- Banking77 (77 intents, never trained): 68.2% vs 49.2% for the best official Laya checkpoint[Vendor claim — not independently verified]
- MASSIVE intent, 37 held-out locales: 65.5% vs 36.1% for Laya multilingual; 71.7% average across 51 locales[Vendor claim — not independently verified]
- JevBench v1.4.1 public items, zero-shot, self-run with the official harness: 64.1% (148/231, 95% CI 57.7–70.0%); not a leaderboard entry[Vendor claim — not independently verified]
- Typed decisions 79.2% (1,583/2,000), in-domain (trained on the train split); +2.6 points over Laya's typed checkpoint[Vendor claim — not independently verified]
- tweet_topic zero-shot: 75.5% vs 79.3% for Jev 1.13 (third-party study numbers); ECE 0.027 vs 0.063[Vendor claim — not independently verified]
- Up to 25,600 input tokens; 98.3% on 1,280 real 24K-token items; preregistered long-context claim passed[Vendor claim — not independently verified]
- About 0.15–0.2 s per short request with MLX on an Apple M1 Max, after the first call[Vendor claim — not independently verified]
- Claude Code guard agrees with 77.6% of 49 hand-labelled tool calls; no deny-labelled call was allowed[Vendor claim — not independently verified]
Jev-Style-0.8B-Decision-v3 is a small open decision model built by fine-tuning every weight of Qwen3.5-0.8B. Send text or JSON with typed questions and it returns a probability for every option, in one forward pass, with no text generation. It is the third generation of an independent series (two 2B LoRA versions came first) and the first one worth cataloguing on its own: smaller, full fine-tune, no cap on the number of options, and a proper local toolchain around it.
It is not TypeSafe, and the author is explicit that no Jev weights, code or outputs were used. “Jev-Style” describes the kind of model.
Weights: Hugging Face chaoliangUNSW/Jev-Style-0.8B-Decision-v3 (Apache-2.0), plus MLX and GGUF builds. Runtime and agent tooling: github.com/lawrence3699/jev-style, linked from the card as the project’s GitHub. Demo: HF Space.
Specs
| Attribute | Value |
|---|---|
| Author | chaoliangUNSW (independent) |
| Base model | Qwen/Qwen3.5-0.8B (text-only; vision tower and MTP head removed) |
| Training | Full fine-tune in bf16, one H100, ~96 min, 131.4M tokens, 321,756-row pool in 19 languages |
| Parameters | 752,393,024 (HF API) |
| Readout | “Verdict slot” per option: logit(" yes") − logit(" no") at the end of each option line, using the tied embeddings. No new parameters |
| Calibration | 20 group temperatures plus a global 0.880, fitted on 15,655 held-out rows |
| Decision types | noul (yes/no), choice (any number of options; 77 tested in one pass), score (2–10 levels) |
| Context | Up to 25,600 input tokens, head up to 2,048; over-budget inputs are rejected, never truncated |
| Builds | safetensors bf16 1.50 GB · MLX bf16 / 8-bit · GGUF F16 / Q8_0 / Q4_K_M (0.53 GB) |
| License | Apache-2.0 (weights and code); see the training-data caveat below |
| Released | (HF) |
| Runtime | pip install "jev-style[torch]" or [mlx] (PyPI 0.2.0); jev-style serve on port 8765 |
Toolchain
The runtime repo is what sets this apart from most small peers:
jev-style serveis a local server that follows the public/v1/systemonerequest shape, with a Playground. It picks MLX on Apple Silicon and PyTorch elsewhere; llama.cpp is optional.- MCP server with
decide,noul,choiceandscoretools. - Claude Code guard, a
PreToolUsehook that returns allow / ask / deny. On its 49 bundled hand-labelled calls it agrees 77.6% of the time and never allowed a call labelled deny (UnverifiedClaim). The author frames it as a second line of defence, not a sandbox. jev-style evalreports accuracy, Brier, ECE and how much you could automate at a 1, 5 or 10% error budget, on your own labels.- Six agent skills installable with
npx skills add lawrence3699/jev-style.
Author benchmarks
All numbers come from the model card, which gives protocols, confidence intervals and data files. They are author-reported → UnverifiedClaim.
| Benchmark | Jev-Style v3 | Comparison | Protocol note |
|---|---|---|---|
| Banking77, 77 intents | 68.2% | Laya best 49.2% | Never trained on Banking77; Laya re-run by the author |
| MASSIVE intent, 37 held-out locales | 65.5% | Laya multilingual 36.1% | 14 other locales were in training |
| tweet_topic, zero-shot | 75.5% | Jev 1.13: 79.3% | Jev number from a third-party study, not re-run |
| JevBench v1.4.1 public (231 items) | 64.1% | Laya 58.4% (board) | Self-run; Laya sits inside v3’s CI |
| Typed decisions (2,000) | 79.2% | Laya typed 76.6% | In-domain for both; Jev’s 72.7% is zero-shot |
The card is unusually candid about what these do and don’t show. Hosted Jev has higher accuracy than v3 on every set where Jev has published numbers, and the author calls v3 “the small local option, not a replacement for the hosted model”. The typed-decisions number is in-domain, and the card explains that gold labels there come from a ~4B teacher, so high scores partly measure agreement with that teacher.
Limits
- Self-run evals. Carefully done, with preregistration for the long-context claim, but nobody outside has rerun them.
- Below Jev on every published comparison, by the author’s own account.
- Training-data licences. The card lists datasets with research-only or unclear terms (DAIR Emotion, AG News, SST-5, MNLI and others) and training data written or labelled by OpenAI and Anthropic models. The weights are Apache-2.0, but check those terms before commercial use.
- Long option lists get chunked. Choice lists that don’t fit the 2,048-token head are scored in chunks with one softmax over all options (runtime update of 26 Sep).
- Batched mode trades exactness for speed. The GGUF
many_mode="batched"path can flip near-tied answers; the default mode matches one call per question. - Early. The runtime repo was created on 25 Sep and has 3 stars; APIs may move.
- No hosted option or SLA.
Fit / anti-fit
Fit when you need: a local decision model on a laptop, gating agent tool calls, many questions over one long document, multilingual intent or routing, a local server for code written against the /v1/systemone shape.
Anti-fit when you need: the accuracy of hosted Jev, a clean commercial licence chain for the training data, or a hosted SLA.
Why it is in the catalog
It passes the trained-weights test: a full fine-tune for typed decisions, with its own readout and fitted temperatures, not a logit wrapper on a frozen model. It returns typed decisions, it is public, and its API is documented. It sits in the sub-1B local slot with Tiny-Jev, Lumma-fev and Laya, and brings the most complete agent toolchain of that group.
