Bandits publishes a recipe to post-train your own Jev-style trace judge. It beats Jev on one benchmark, and no weights are released

Last updated:

Source

On 30 September 2026 (13:41 BRT), Laxman of bandr.ai posted “Releasing Your Own Jev”: a recipe inside the Bandits trace-mining toolkit for post-training a small model on your own agent’s traces. The model judges each agent step as success, unclear or failure in one forward pass, with a probability for each. The thread credits Nimble, Kev and AutoJev for “proving the recipe”.

What it is. The recipes/jev package LoRA-tunes (rank 16) Qwen/Qwen3.5-4B-Base on labelled decisions, either Bandits’ verifier votes or your own JSONL. It reads the option-letter logits at the Answer: cue, fits one temperature on a calibration split and prints a scorecard against a majority baseline, the untrained model and TypeSafe Jev through its API. Questions are Choice with 2–26 options. There is a local training UI, but no inference server and no /v1/systemone endpoint.

What’s missing. No trained checkpoint on Hugging Face (searches for bandr-ai, bandits and the author return nothing), and the repo has no license file, so “completely open source” isn’t true in the legal sense yet. The thread says 4B/8B/27B, but the published results cover only the 4B, and the repo’s launch plan lists 27B+ as out of scope.

The numbers (UnverifiedClaim). The repo’s locked test results (28 Sep 2026, one seed) cover 1,920 held-out steps from 36 AgentProcessBench tasks. They report 79.1% agreement with human labels vs 66.8% for Jev 1.13, +12.3 points (95% CI +7.3 to +17.3). The model scores ECE 0.012 and runs at 0.135 s p50 on one L40S. The thread and its video say 79.7% vs 66.3% (+13.4) and add DeepSeek-V4.1-Flash at 65.3%; the repo has no DeepSeek row. Per-source numbers also differ: the thread’s “customer service 80.3%” is 78.8% (τ²-bench) in the repo.

The repo’s own caveats are worth reading:

  • The 4B trained on 135 AgentProcessBench tasks and Jev saw the benchmark cold.
  • Both got a prompt written for the 4B.
  • Jev’s latency includes the network.
  • Jev is cheaper per 1,000 decisions ($0.058 vs $0.077).
  • On TRAIL, neither beats flagging every step.

The thread’s “matches a 30B verifier, 36× faster” result (77.9% vs the verifier’s 75.6% self-agreement) is listed in the repo under “Not claimed”: the verifier looped at temperature 0 and needed re-runs with changed settings on 51 of 172 test steps.

How we list it. It’s news, not a catalog entry. It trains real decision weights, like Nimble, Kev or AutoJev, but ships no checkpoint, no license and no way to serve a model. The benchmark result is also a task-specific judge measured in-distribution. It’s listed under Runtimes as a method (a training recipe with no released model) and stays on our watch list. If the adapter appears on Hugging Face with a license, it moves to the catalog as a lora-adapter.