Jev-Style-0.8B-Decision-v3

Last updated:

Open150–200ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity5.6/10

    Open weights in torch, MLX and GGUF, PyPI package (jev-style 0.2.0), a /v1/systemone-shaped local server, MCP tools, a Claude Code guard and six agent skills. All first-party; the runtime repo is a day old. Self-host only, no SLA.

  • Capability6.8/10

    First-class noul / choice / score, per-group temperatures fitted on held-out rows, 25.6k-token inputs. Strong, carefully documented evals, but self-run; wins are against Laya, and the author says hosted Jev is ahead on every set where Jev has numbers.

  • Adoption12/100

    Family downloads are real (the 2B v1 GGUF alone ~6k), but v3 is two days old: ~790 downloads and 12 likes across its builds, 3 GitHub stars, no launch post found. No production use reported.

Vendor claims

Jev-Style-0.8B-Decision-v3 is a small open decision model built by fine-tuning every weight of Qwen3.5-0.8B. Send text or JSON with typed questions and it returns a probability for every option, in one forward pass, with no text generation. It is the third generation of an independent series (two 2B LoRA versions came first) and the first one worth cataloguing on its own: smaller, full fine-tune, no cap on the number of options, and a proper local toolchain around it.

It is not TypeSafe, and the author is explicit that no Jev weights, code or outputs were used. “Jev-Style” describes the kind of model.

Weights: Hugging Face chaoliangUNSW/Jev-Style-0.8B-Decision-v3 (Apache-2.0), plus MLX and GGUF builds. Runtime and agent tooling: github.com/lawrence3699/jev-style, linked from the card as the project’s GitHub. Demo: HF Space.

Specs

Attribute Value
Author chaoliangUNSW (independent)
Base model Qwen/Qwen3.5-0.8B (text-only; vision tower and MTP head removed)
Training Full fine-tune in bf16, one H100, ~96 min, 131.4M tokens, 321,756-row pool in 19 languages
Parameters 752,393,024 (HF API)
Readout “Verdict slot” per option: logit(" yes") − logit(" no") at the end of each option line, using the tied embeddings. No new parameters
Calibration 20 group temperatures plus a global 0.880, fitted on 15,655 held-out rows
Decision types noul (yes/no), choice (any number of options; 77 tested in one pass), score (2–10 levels)
Context Up to 25,600 input tokens, head up to 2,048; over-budget inputs are rejected, never truncated
Builds safetensors bf16 1.50 GB · MLX bf16 / 8-bit · GGUF F16 / Q8_0 / Q4_K_M (0.53 GB)
License Apache-2.0 (weights and code); see the training-data caveat below
Released (HF)
Runtime pip install "jev-style[torch]" or [mlx] (PyPI 0.2.0); jev-style serve on port 8765

Toolchain

The runtime repo is what sets this apart from most small peers:

  • jev-style serve is a local server that follows the public /v1/systemone request shape, with a Playground. It picks MLX on Apple Silicon and PyTorch elsewhere; llama.cpp is optional.
  • MCP server with decide, noul, choice and score tools.
  • Claude Code guard, a PreToolUse hook that returns allow / ask / deny. On its 49 bundled hand-labelled calls it agrees 77.6% of the time and never allowed a call labelled deny (UnverifiedClaim). The author frames it as a second line of defence, not a sandbox.
  • jev-style eval reports accuracy, Brier, ECE and how much you could automate at a 1, 5 or 10% error budget, on your own labels.
  • Six agent skills installable with npx skills add lawrence3699/jev-style.

Author benchmarks

All numbers come from the model card, which gives protocols, confidence intervals and data files. They are author-reported → UnverifiedClaim.

Benchmark Jev-Style v3 Comparison Protocol note
Banking77, 77 intents 68.2% Laya best 49.2% Never trained on Banking77; Laya re-run by the author
MASSIVE intent, 37 held-out locales 65.5% Laya multilingual 36.1% 14 other locales were in training
tweet_topic, zero-shot 75.5% Jev 1.13: 79.3% Jev number from a third-party study, not re-run
JevBench v1.4.1 public (231 items) 64.1% Laya 58.4% (board) Self-run; Laya sits inside v3’s CI
Typed decisions (2,000) 79.2% Laya typed 76.6% In-domain for both; Jev’s 72.7% is zero-shot

The card is unusually candid about what these do and don’t show. Hosted Jev has higher accuracy than v3 on every set where Jev has published numbers, and the author calls v3 “the small local option, not a replacement for the hosted model”. The typed-decisions number is in-domain, and the card explains that gold labels there come from a ~4B teacher, so high scores partly measure agreement with that teacher.

Limits

  • Self-run evals. Carefully done, with preregistration for the long-context claim, but nobody outside has rerun them.
  • Below Jev on every published comparison, by the author’s own account.
  • Training-data licences. The card lists datasets with research-only or unclear terms (DAIR Emotion, AG News, SST-5, MNLI and others) and training data written or labelled by OpenAI and Anthropic models. The weights are Apache-2.0, but check those terms before commercial use.
  • Long option lists get chunked. Choice lists that don’t fit the 2,048-token head are scored in chunks with one softmax over all options (runtime update of 26 Sep).
  • Batched mode trades exactness for speed. The GGUF many_mode="batched" path can flip near-tied answers; the default mode matches one call per question.
  • Early. The runtime repo was created on 25 Sep and has 3 stars; APIs may move.
  • No hosted option or SLA.

Fit / anti-fit

Fit when you need: a local decision model on a laptop, gating agent tool calls, many questions over one long document, multilingual intent or routing, a local server for code written against the /v1/systemone shape.

Anti-fit when you need: the accuracy of hosted Jev, a clean commercial licence chain for the training data, or a hosted SLA.

Why it is in the catalog

It passes the trained-weights test: a full fine-tune for typed decisions, with its own readout and fitted temperatures, not a logit wrapper on a frozen model. It returns typed decisions, it is public, and its API is documented. It sits in the sub-1B local slot with Tiny-Jev, Lumma-fev and Laya, and brings the most complete agent toolchain of that group.