RSI-Jev v3.0
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity5.3/10
Open weights, MIT research code with versioned experiment logs, RL write-up, and a Jev-shaped /v1/systemone server. Custom load_release loader; no PyPI or hosted SLA.
- Capability8/10
First-class choice / noul / score, refit confidence head (suite ECE 0.066 author-reported), ~10 ms per decision claimed. v3.0 adds listwise RL; suite/typed numbers remain self-run and partly in-domain.
- Adoption9/100
About 20 GitHub stars after v3.0; HF likes still near zero. Research-scale traction, no public production reports.
Vendor claims
- v3.0 typed-decisions 0.791; suite mean 0.756 (was 0.736 on v2.1 under the fifteen-benchmark suite); suite ECE 0.066[Vendor claim — not independently verified]
- First release where reinforcement learning works: listwise ranking reward lifts hippo R@1 0.192 → 0.308 (p = 2.4e-13) vs supervised parent[Vendor claim — not independently verified]
- Held-out set vs v2.1: +0.016 (p = 0.011); data-scale-2 corpus 266,131 questions from 36 sources[Vendor claim — not independently verified]
- About 10 ms per decision in one forward pass[Vendor claim — not independently verified]
- Checkpoint reproduces its training-run predictions at 1.0000 after reload[Vendor claim — not independently verified]
RSI-Jev v3.0 is a 2B open decision model trained by a recursively self-improving research loop — and the first release where a reinforcement-learning stage earned its place. You give it a document and typed questions (choice, noul, score); a single forward pass returns a calibrated probability for every option. It does not generate text.
It is a full fine-tune of Qwen/Qwen3.5-2B-Base with a trained option-scoring head and a confidence head (oof_head_scorefloor) that rescales answers without changing which option wins. v3.0 adds a wider supervised corpus (data-scale-2), an MLP readout combine, and listwise ranking RL (Plackett-Luce / NDCG@5) aimed at memory reranking. Research system: next version of AutoScientists.
Weights: Hugging Face shgao/rsi-jev-v3.0-qwen3.5-2b (Apache-2.0). Code and version records: github.com/Shanghua-Gao/RSI-Jev (MIT, ~20★). Earlier open checkpoints: v2.1, v2.0, v1.0-2B, v1.0-0.8B — history on this card, not separate catalog entries.
Specs
| Attribute | Value |
|---|---|
| Author | Shanghua Gao (independent; AutoScientists lineage) |
| Base model | Qwen/Qwen3.5-2B-Base |
| Parameters | ~2B (tower + trained option-scoring head) |
| Readout | Trained option-scoring head + MLP xattn combine (v3.0) |
| Calibration | Confidence head oof_head_scorefloor; ships calibration files |
| Decision types | noul, choice, score |
| License | Apache-2.0 (weights); MIT (code) |
| Released | (v3.0 on HF; record dated same day) |
| Runtime | Custom load_release / serve script → local POST /v1/systemone |
Serving
Jev-compatible HTTP server (python scripts/serve.py --ckpt …). Routes: /v1/systemone, /v1/models, /v1/limits, /health. jev-latest accepted as an alias. Wire-shape compatibility is the author’s claim, not a TypeSafe endorsement. load_release applies shipped calibration; skipping those files keeps the same top option with different probabilities.
Author benchmarks
Numbers from the v3.0 version record and model card. Author-reported → UnverifiedClaim.
| Signal | v3.0 | Note |
|---|---|---|
| Suite mean (15 benches) | 0.756 | Was 0.736 for v2.1 on this suite |
| Typed-decisions | 0.791 | In-domain relative to train coverage — not zero-shot Jev/Laya |
| Suite ECE | 0.066 | Confidence head refit after RL |
| MMLU-Pro 1k | 0.364 | |
| hippo R@1 (listwise RL) | 0.308 | vs 0.192 supervised parent; p = 2.4×10⁻¹³ |
| Held-out vs v2.1 | +0.016 | p = 0.011 |
| Per-decision latency | ~10 ms | Author claim; no independent p50/p99 |
v1.0 remains the cleaner zero-shot reference. Prefer v1.0/0.8B if you need a train-split-free baseline.
Limits
- Most suite / typed-decisions numbers are in-domain to some degree — read the version record before comparing to zero-shot Jev or Laya.
- Self-run evals with verification hashes; no independent suite rerun yet.
- Custom loader — not
AutoModel.from_pretrained. - No PyPI, no hosted SLA.
- Not TypeSafe.
Fit / anti-fit
Fit when you need: a local 2B decision model with a documented RSI research trail, listwise-RL reranking signal, or a Jev-shaped HTTP server.
Anti-fit when you need: drop-in transformers loading, a hosted SLA, or a headline number you can treat as zero-shot without reading the card.
Why it is in the catalog
Trained tower + option-scoring head (+ confidence head) for typed decisions — not a logit wrapper. Public weights/code and a /v1/systemone surface. Near Jev-Style, Decider, Decily, Eikos, OpenDecider.
Newer release: v6.0-VL (6 Oct 2026)
shgao/rsi-jev-v6.0-vl-4b (Apache-2.0, 4.69B) is a full fine-tune of Qwen3.5-4B-Base, not a LoRA. It has early exits at layers 16, 20 and 32, chosen with an effort setting (low / medium / high / auto) and reported back as usage.depth. It takes text plus up to 4 images; image requests always run all 32 layers.
Author numbers (UnverifiedClaim): Decision Index 0.2.1 full run 46.24 (own run; v5.0-VL 38.38), held-out 0.698, ECE 0.036 by default / 0.024 on auto, median 23 / 27 / 40 ms on one H200 by effort level.
Caveats: five of the image training sources are non-commercial or research-only while the weights ship as Apache-2.0, and it is not settled whether those terms carry over. The v5.0-VL errata disclose that about 1,000 Decision Index test rows were in training (author estimates an effect of 0.04 points or less). Not on the official Decision Index or JevBench. Scores on this page are unchanged.
