Open Alternative to Jev

Last updated:

RuntimeApache-2.0~54 ★

⚠️ Why not a model

Reads option-letter logits from stock open LLMs (Qwen) in one packed forward pass. No decision training; calibration is an optional temperature you fit yourself.

Technical specs

Base LLMs
Qwen3.6-27B, Qwen3.5-4B, Qwen3.5-2B, Qwen3-1.7B, Qwen3-0.6B, Qwen2.5-0.5B, any ChatML open LLM
Decision types
choice, boolean, score
Features
python-library, hf-transformers, vllm, packed-readout, temperature-scaling, option-permutation, no-server
License
Apache-2.0

Open Alternative to Jev is a Python library, imported as so1, that gets typed decisions out of an ordinary open-weights LLM. It renders each question as a chat turn, packs all the questions about one state into a single token sequence, runs one forward pass and reads the next-token distribution over the option letters at each answer position. Same family as SemIf and simple-jev.

How it works

  1. Each question becomes a normal ChatML user turn with lettered options (A–Z).
  2. Turns are concatenated with a fixed placeholder answer between them, so the state is written once.
  3. One forward pass; logits are read only at the readout positions (logits_to_keep on Transformers, prompt_logprobs on vLLM).
  4. A softmax over the option letters gives the distribution. Nothing outside your options can win.
  5. Optionally, a temperature you fit on your own labels (TemperatureScaler) calibrates it.

mode="separate" runs one sequence per question instead, which avoids cross-question interference. permutations=2 asks each question with the options in both orders and averages the result, to cancel position bias.

The public API is Choice, plus yes_no and scale helpers that build a Choice with yes/no or integer options. Up to 26 options per question. There is no HTTP server: the README says it replaces the library, not the endpoint, so the TypeSafe SDK can’t point at it.

Why it is not a model

No weights are trained or changed. The library works with checkpoints you already have (tested on Qwen2.5-0.5B, Qwen3.5-4B and Qwen3.6-27B; any ChatML template works, others need a ChatFormat). Quality and calibration come from the base LLM, and raw probabilities are over-confident until you fit a temperature. The README says it plainly: “Not a reproduction of Jev.”

Author results (UnverifiedClaim)

On LocalLLaMA/typed-decisions (400 cases, 2,000 decisions), a stock Qwen3.6-27B in 8-bit through the library scores 73.7% accuracy, ECE 0.020, Brier 0.113, 582 ms per case on one H200 MIG slice. The author compares that with Jev 1.13.0 at 72.7% / ECE 0.144 / Brier 0.148 / 710 ms, as measured by the benchmark’s authors through TypeSafe’s API on 18 Sep 2026. With permutations=2 the 27B reaches 75.5% at ECE 0.0075.

Read it with the author’s own caveats and one of ours:

  • A one-point gap on 2,000 decisions is within noise, the author says so, and the score sits on the teacher self-agreement ceiling (73.5%).
  • Latency conditions differ: local GPU forward time vs an API round-trip.
  • It only holds at 4B and up. Qwen3.5-4B scores 59.3% on the same set, Qwen3-0.6B 29.1%, and packing hurts accuracy below about 4B.
  • The Jev ECE of 0.144 is contested. Independent DecisionEval measured Jev 1.13.0 on the same split on 20 Sep 2026 at ECE 0.045 (Brier 0.148, accuracy 0.740) and could not reproduce 0.144. Against that run, the calibration gap largely disappears and the accuracy gap reverses. Details in Jev’s independent evaluation.
  • Answers move with their neighbours. Packing changes 6–9% of individual answers vs one-at-a-time scoring, and rotating question order changes 8% of MMLU answers. Aggregate accuracy holds; per-decision stability doesn’t.

Caveats

  • Not on PyPI yet. The README says pip install open-alternative-jev, but the package was not on PyPI when checked on 26 Sep 2026; install from the repo.
  • Position bias on smaller models: reversing options moved Qwen3.5-4B’s yes/no accuracy by 13.5 points.
  • Pin transformers. Readout positions come from the chat template; a template change can shift them.
  • All benchmarks are author-run → UnverifiedClaim.