RSI-Jev v3.0

Last updated:

Open10–10ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity5.3/10

    Open weights, MIT research code with versioned experiment logs, RL write-up, and a Jev-shaped /v1/systemone server. Custom load_release loader; no PyPI or hosted SLA.

  • Capability8/10

    First-class choice / noul / score, refit confidence head (suite ECE 0.066 author-reported), ~10 ms per decision claimed. v3.0 adds listwise RL; suite/typed numbers remain self-run and partly in-domain.

  • Adoption9/100

    About 20 GitHub stars after v3.0; HF likes still near zero. Research-scale traction, no public production reports.

Vendor claims

RSI-Jev v3.0 is a 2B open decision model trained by a recursively self-improving research loop — and the first release where a reinforcement-learning stage earned its place. You give it a document and typed questions (choice, noul, score); a single forward pass returns a calibrated probability for every option. It does not generate text.

It is a full fine-tune of Qwen/Qwen3.5-2B-Base with a trained option-scoring head and a confidence head (oof_head_scorefloor) that rescales answers without changing which option wins. v3.0 adds a wider supervised corpus (data-scale-2), an MLP readout combine, and listwise ranking RL (Plackett-Luce / NDCG@5) aimed at memory reranking. Research system: next version of AutoScientists.

Weights: Hugging Face shgao/rsi-jev-v3.0-qwen3.5-2b (Apache-2.0). Code and version records: github.com/Shanghua-Gao/RSI-Jev (MIT, ~20★). Earlier open checkpoints: v2.1, v2.0, v1.0-2B, v1.0-0.8B — history on this card, not separate catalog entries.

Specs

Attribute Value
Author Shanghua Gao (independent; AutoScientists lineage)
Base model Qwen/Qwen3.5-2B-Base
Parameters ~2B (tower + trained option-scoring head)
Readout Trained option-scoring head + MLP xattn combine (v3.0)
Calibration Confidence head oof_head_scorefloor; ships calibration files
Decision types noul, choice, score
License Apache-2.0 (weights); MIT (code)
Released (v3.0 on HF; record dated same day)
Runtime Custom load_release / serve script → local POST /v1/systemone

Serving

Jev-compatible HTTP server (python scripts/serve.py --ckpt …). Routes: /v1/systemone, /v1/models, /v1/limits, /health. jev-latest accepted as an alias. Wire-shape compatibility is the author’s claim, not a TypeSafe endorsement. load_release applies shipped calibration; skipping those files keeps the same top option with different probabilities.

Author benchmarks

Numbers from the v3.0 version record and model card. Author-reported → UnverifiedClaim.

Signal v3.0 Note
Suite mean (15 benches) 0.756 Was 0.736 for v2.1 on this suite
Typed-decisions 0.791 In-domain relative to train coverage — not zero-shot Jev/Laya
Suite ECE 0.066 Confidence head refit after RL
MMLU-Pro 1k 0.364
hippo R@1 (listwise RL) 0.308 vs 0.192 supervised parent; p = 2.4×10⁻¹³
Held-out vs v2.1 +0.016 p = 0.011
Per-decision latency ~10 ms Author claim; no independent p50/p99

v1.0 remains the cleaner zero-shot reference. Prefer v1.0/0.8B if you need a train-split-free baseline.

Limits

  • Most suite / typed-decisions numbers are in-domain to some degree — read the version record before comparing to zero-shot Jev or Laya.
  • Self-run evals with verification hashes; no independent suite rerun yet.
  • Custom loader — not AutoModel.from_pretrained.
  • No PyPI, no hosted SLA.
  • Not TypeSafe.

Fit / anti-fit

Fit when you need: a local 2B decision model with a documented RSI research trail, listwise-RL reranking signal, or a Jev-shaped HTTP server.

Anti-fit when you need: drop-in transformers loading, a hosted SLA, or a headline number you can treat as zero-shot without reading the card.

Why it is in the catalog

Trained tower + option-scoring head (+ confidence head) for typed decisions — not a logit wrapper. Public weights/code and a /v1/systemone surface. Near Jev-Style, Decider, Decily, Eikos, OpenDecider.

Newer release: v6.0-VL (6 Oct 2026)

shgao/rsi-jev-v6.0-vl-4b (Apache-2.0, 4.69B) is a full fine-tune of Qwen3.5-4B-Base, not a LoRA. It has early exits at layers 16, 20 and 32, chosen with an effort setting (low / medium / high / auto) and reported back as usage.depth. It takes text plus up to 4 images; image requests always run all 32 layers.

Author numbers (UnverifiedClaim): Decision Index 0.2.1 full run 46.24 (own run; v5.0-VL 38.38), held-out 0.698, ECE 0.036 by default / 0.024 on auto, median 23 / 27 / 40 ms on one H200 by effort level.

Caveats: five of the image training sources are non-commercial or research-only while the weights ship as Apache-2.0, and it is not settled whether those terms carry over. The v5.0-VL errata disclose that about 1,000 Decision Index test rows were in training (author estimates an effect of 0.04 points or less). Not on the official Decision Index or JevBench. Scores on this page are unchanged.