Bandits · Your own Jev recipe

Last updated:

MethodNone (no LICENSE file)~28 ★

⚠️ Why not a model

Ships a training recipe, not a model: LoRA on Qwen3.5-4B-Base, local training UI, scorecard vs Jev via /v1/systemone. No released weights, no license.

Technical specs

Base LLMs
Qwen/Qwen3.5-4B-Base
Decision types
choice
Features
training-recipe, lora, local-training-ui, temperature-scaling, scorecard-vs-jev, agent-trace-judge, no-released-weights, no-server
License
None (no LICENSE file)

Read this first

  1. Nothing to download yet. The repo ships the recipe and a local training UI. No trained checkpoint is on Hugging Face (searches for bandr-ai, bandits and the author return nothing, 30 Sep 2026), and the repo has no LICENSE file, so “completely open source” isn’t true in the legal sense.
  2. The headline comparison is in-distribution. Their 4B trained on 135 AgentProcessBench tasks and was tested on 36 held-out tasks from the same benchmark. TypeSafe Jev saw the benchmark cold, through a prompt written for the 4B.
  3. The launch thread claims more than the repo does. See “Thread vs repo” below.
  4. “Jev 4B” is not TypeSafe’s Jev. A thread image labels bandr.ai’s own model “Jev 4B”. For TypeSafe’s hosted model, see Jev.

Bandits is bandr.ai’s toolkit for mining agent traces for post-training. Its recipes/jev package, launched on 30 September 2026 as “Releasing Your Own Jev” (13:41 BRT), turns labelled agent steps into a small Jev-style decision model that you train and run yourself. The thread credits Nimble, Kev and AutoJev for proving the recipe.

How it works

  1. Labels. Each agent step gets a label: success (+1), unclear (0) or failure (−1). Labels come from Bandits’ verifier votes or from your own JSONL (state, question, 2–26 options, target, optional group_id to keep a trace in one split).
  2. Training. LoRA rank 16 on all linear layers of Qwen/Qwen3.5-4B-Base, on one 48 GB GPU (an L40S). The best checkpoint is picked on the dev split.
  3. Readout. One forward pass over a fixed prompt ending in Answer:. A softmax over the option-letter logits gives a probability per option.
  4. Calibration. One temperature is fitted on a calibration split.
  5. Scorecard. jev report compares the verifier, a majority baseline, the untrained base, the trained and calibrated model, and TypeSafe Jev (called through POST /v1/systemone with jev score-api), with paired bootstrap intervals, cost and latency.

jev ui starts a local browser UI for uploading data, training and reading the report. There is no inference server and no /v1/systemone endpoint of its own.

Why it’s here and not in the catalog

It isn’t a frozen-logit wrapper like SemIf or open-alternative-jev: the recipe does train decision weights. But what it ships is the method (code, prompt, training loop and scorecard), not a model you can download or call. Every other open peer in the catalog, such as Nimble, Kev, AutoJev and NeoHorse-Jev, publishes licensed weights. The schema has no “recipe” kind, so this entry uses method.

Promotion trigger: if bandr.ai publishes the trained adapter on Hugging Face with a license, it moves to the catalog as a lora-adapter, with the in-distribution caveat at the top.

Repo results (UnverifiedClaim)

From the repo’s locked test results (28 Sep 2026, one seed, Qwen3.5-4B-Base + LoRA). The test set is 1,920 held-out steps from 36 AgentProcessBench tasks, labelled by humans.

System Accuracy Macro F1 ECE p50 latency
Majority label 60.8% 25.2% 0.013 —
Qwen3.5-4B-Base, untrained 60.7% 35.9% 0.106 0.092 s
Recipe (4B + LoRA) 79.1% 52.6% 0.012 0.135 s
TypeSafe Jev 1.13.0 (API) 66.8% 52.7% 0.073 0.449 s

The gap is +12.3 points (95% CI +7.3 to +17.3). The repo’s own caveats:

  • In-distribution: trained on the same benchmark’s other tasks, while Jev saw it cold.
  • Prompt: both systems got the same input, written for the 4B. Jev may do better with inputs written for it.
  • Macro F1 is tied: the 4B never answers “unclear”, which is 5% of the human labels.
  • Latency isn’t like for like: Jev’s includes the network, the 4B’s is measured on the GPU.
  • Cost: Jev is cheaper, about $0.058 vs $0.077 per 1,000 decisions.
  • TRAIL: neither system beats flagging every step as an error (F1 0.451 for Jev, 0.444 for the 4B, 0.558 for “all errors”). The GAIA split is contaminated with TRAIL questions.

Thread vs repo

Launch thread / video Repo
79.7% vs Jev 66.3% (+13.4) 79.1% vs 66.8% (+12.3)
Beats DeepSeek-V4.1-Flash (763B, 65.3%) by 14 points No DeepSeek row in the results
Customer service 80.3% τ²-bench 78.8%
Post-trained on a 30B verifier; matches it (77.9% vs its 75.6% self-agreement), 36× faster Listed under “Not claimed”: the verifier looped at temperature 0 and needed re-runs with changed settings on 51 of 172 test steps
4B / 8B / 27B Only the 4B was measured; the launch plan lists 27B+ as out of scope
“Completely open source: recipe, data, training, evals” No LICENSE file; raw predictions and reports are gitignored