Surogate Rune 26B-A4B v3
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity5.8/10
Apache-2.0 bf16 weights (auto-approved gate on HF), detailed card, launch blog and the open surogate engine that trained and serves it. Its endpoint is /api/alpha/decisions, not /v1/systemone, so the TypeSafe SDK can't point at it. Company channels and a hosted API are mentioned, but no SLA.
- Capability7.9/10
Noul, Choice and Score, text and images. #2 on the official Decision Index 0.2.1 (57.44 vs Jev 57.91), a run the maintainers reproduced. Overconfident at the default temperature (index ECE 0.12); the author fits T=2. Latency 90–180 ms per decision is the author's own figure.
- Adoption11/100
About 40 HF likes and 2.5k downloads in the first month, a LinkedIn launch with a dozen reactions, and a third-party mirror. The surogate engine's 838 GitHub stars belong to the engine, not the model. No production use reported.
Vendor claims
- Decision Index 0.2.1 with thinking on: about 59.2, ahead of Jev 1.13 (our estimate, not on the leaderboard)[Vendor claim — not independently verified]
- ECE 12.5% at decision temperature 1, 2.2% at temperature 2, on 33 Decision Index benchmarks[Vendor claim — not independently verified]
- One decision in 0.09–0.18 s median on one RTX PRO 6000; a question that thinks takes about 4.9 s[Vendor claim — not independently verified]
- 83.4% on 1,196 image items built in the Image JevBench v0.1 format (own sample, sealed items not used)[Vendor claim — not independently verified]
- 42 ms per decision on a single RTX 5090; '24B MoE' (LinkedIn launch post, conflicts with the card)[Vendor claim — not independently verified]
Surogate Rune is an open-weights decision model from Invergent, a European company that also builds the surogate training and serving engine. It is a full fine-tune of google/gemma-4-26B-A4B-it: a 26.5B-parameter mixture of experts with about 4B active per token, a 262k context and the vision tower kept. You send a state (text, structured data or an image with text) plus typed questions, and get back one option per question with a probability over all of them, in a single forward pass.
It is not a TypeSafe product and does not claim to be trained on Jev output. Invergent calls the three question types choice, noul and score, the same names Jev uses.
Specs
| Attribute | Value |
|---|---|
| Company | Invergent (HF org surogate) |
| Release | HF repo created 21 Sep 2026 (14:58 BRT) with v1 GGUF builds; v3 bf16 weights and Decision Index results by 26 Sep |
| Base model | google/gemma-4-26B-A4B-it (MoE, 8 of 128 experts active) |
| Training | Full fine-tune with the surogate engine; data and recipe not published |
| Weights | bf16 safetensors, 51.6 GB, Apache-2.0. The HF repo has an auto-approved access gate |
| Inputs | Text, JSON state, images (--vision); 262k context |
| Question types | choice, noul (catalog boolean), score |
| Serving | surogate serve → POST /api/alpha/decisions; also loads as a standard Gemma4ForConditionalGeneration in transformers |
| Thinking | Opt-in per request: questions below 0.7 confidence reason for up to 512 tokens before answering |
| Hosted API | Mentioned in the launch blog; no public pricing or docs for it yet |
Repository naming and history
The canonical repo is surogate/rune-26b-a4b-GGUF, despite the name. It first held the v1 GGUF builds (llama.cpp-compatible, JD-Q3_K_M to Q8_0 plus vision projectors). Those files now live only in its history at revision 2a15504. The current files are v3 in bf16, and the card says there are no GGUF builds of v3 yet. There is no separate non-GGUF repo from Invergent.
A third-party mirror, michaelfeil/rune-26b-a4b, re-uploads the same v3 files without the gate (created 29 Sep 2026, 20:23 BRT). It is not an official release.
Decision Index 0.2.1 (official Space)
Measured by the official Decision Index Space (data snapshot 28 Sep 2026, 71 systems). Invergent submitted the run and says the maintainers’ own reproduction matched it bit for bit.
| System | Index (balanced skill) | Rank |
|---|---|---|
| Jev 1.13 | 57.91 | #1 |
| Surogate Rune 26B-A4B v3 | 57.44 | #2 |
| AutoJev-27B | 56.40 | #4 |
Area scores on the card (leaderboard values): Knowledge & Reasoning 43.4, Language 63.1, Retrieval & Classification 63.5, Tools & Automation 71.2, Arts & Human Taste 41.9. Rune is ahead of Jev on language, retrieval and arts, and clearly behind on knowledge and reasoning (Jev 51.3).
On calibration, the index’s own numbers for Rune are accuracy 0.748, mean confidence 0.867, ECE 0.12 and Brier 0.370, which confirms the author’s note that the model is overconfident at the default temperature.
UnverifiedClaim. The “59.2 with thinking” figure is Invergent’s estimate from its own full run with the 0.2 scorer; it is not on the leaderboard. The 2.2% ECE at temperature 2, the image results and all latency numbers are the author’s own measurements.
Discrepancies between sources
- LinkedIn vs card. The launch post on LinkedIn (23 Sep) says “42ms to make a decision on a single RTX 5090 GPU” and “a 24B MoE with 4B activated per token”. The card and blog say 26.5B parameters and 0.09–0.18 s median per decision on an RTX PRO 6000 (0.2 s p90). We keep both as claims and do not pick one.
- Jev baseline. The card lists Jev 1.13 at 57.89; the blog and the official Space say 57.91.
- Earlier edition. The v1 card reported Rune at 57.24 vs Jev 59.51, from Invergent’s own run of a JD-Q6_K build on an older edition of the index. The current numbers are for v3 on edition 0.2.1.
Fit / anti-fit
Fit when you need: an open decision model close to Jev’s accuracy that you can run on your own hardware, image decisions (screenshots, documents, charts), long inputs, or data that must stay in your environment.
Anti-fit when you need: a drop-in for the TypeSafe SDK (different endpoint), a small GPU (51.6 GB of weights; Invergent serves it on 96 GB cards), knowledge-heavy questions, or calibrated probabilities without setting the decision temperature.
Limits
- Trained data and recipe are not published.
- Default probabilities are overconfident; use
--decision-temperature 2, which never changes the chosen option. - surogate 1.5.3 can crash under sustained load on a MoE router value; the card says to use a build with fix #217.
- Thinking does not work together with images yet.
Why it is in the catalog
Invergent trained the weights for typed decisions (full fine-tune, decision protocol) and ships them openly with a documented request format that returns probabilities. That passes the catalog criteria. It is the strongest open entry on the official Decision Index as of 28 Sep 2026.
