Surogate Rune 26B-A4B v3

Last updated:

Open90–180ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity5.8/10

    Apache-2.0 bf16 weights (auto-approved gate on HF), detailed card, launch blog and the open surogate engine that trained and serves it. Its endpoint is /api/alpha/decisions, not /v1/systemone, so the TypeSafe SDK can't point at it. Company channels and a hosted API are mentioned, but no SLA.

  • Capability7.9/10

    Noul, Choice and Score, text and images. #2 on the official Decision Index 0.2.1 (57.44 vs Jev 57.91), a run the maintainers reproduced. Overconfident at the default temperature (index ECE 0.12); the author fits T=2. Latency 90–180 ms per decision is the author's own figure.

  • Adoption11/100

    About 40 HF likes and 2.5k downloads in the first month, a LinkedIn launch with a dozen reactions, and a third-party mirror. The surogate engine's 838 GitHub stars belong to the engine, not the model. No production use reported.

Vendor claims

Surogate Rune is an open-weights decision model from Invergent, a European company that also builds the surogate training and serving engine. It is a full fine-tune of google/gemma-4-26B-A4B-it: a 26.5B-parameter mixture of experts with about 4B active per token, a 262k context and the vision tower kept. You send a state (text, structured data or an image with text) plus typed questions, and get back one option per question with a probability over all of them, in a single forward pass.

It is not a TypeSafe product and does not claim to be trained on Jev output. Invergent calls the three question types choice, noul and score, the same names Jev uses.

Specs

Attribute Value
Company Invergent (HF org surogate)
Release HF repo created 21 Sep 2026 (14:58 BRT) with v1 GGUF builds; v3 bf16 weights and Decision Index results by 26 Sep
Base model google/gemma-4-26B-A4B-it (MoE, 8 of 128 experts active)
Training Full fine-tune with the surogate engine; data and recipe not published
Weights bf16 safetensors, 51.6 GB, Apache-2.0. The HF repo has an auto-approved access gate
Inputs Text, JSON state, images (--vision); 262k context
Question types choice, noul (catalog boolean), score
Serving surogate serve → POST /api/alpha/decisions; also loads as a standard Gemma4ForConditionalGeneration in transformers
Thinking Opt-in per request: questions below 0.7 confidence reason for up to 512 tokens before answering
Hosted API Mentioned in the launch blog; no public pricing or docs for it yet

Repository naming and history

The canonical repo is surogate/rune-26b-a4b-GGUF, despite the name. It first held the v1 GGUF builds (llama.cpp-compatible, JD-Q3_K_M to Q8_0 plus vision projectors). Those files now live only in its history at revision 2a15504. The current files are v3 in bf16, and the card says there are no GGUF builds of v3 yet. There is no separate non-GGUF repo from Invergent.

A third-party mirror, michaelfeil/rune-26b-a4b, re-uploads the same v3 files without the gate (created 29 Sep 2026, 20:23 BRT). It is not an official release.

Decision Index 0.2.1 (official Space)

Measured by the official Decision Index Space (data snapshot 28 Sep 2026, 71 systems). Invergent submitted the run and says the maintainers’ own reproduction matched it bit for bit.

System Index (balanced skill) Rank
Jev 1.13 57.91 #1
Surogate Rune 26B-A4B v3 57.44 #2
AutoJev-27B 56.40 #4

Area scores on the card (leaderboard values): Knowledge & Reasoning 43.4, Language 63.1, Retrieval & Classification 63.5, Tools & Automation 71.2, Arts & Human Taste 41.9. Rune is ahead of Jev on language, retrieval and arts, and clearly behind on knowledge and reasoning (Jev 51.3).

On calibration, the index’s own numbers for Rune are accuracy 0.748, mean confidence 0.867, ECE 0.12 and Brier 0.370, which confirms the author’s note that the model is overconfident at the default temperature.

UnverifiedClaim. The “59.2 with thinking” figure is Invergent’s estimate from its own full run with the 0.2 scorer; it is not on the leaderboard. The 2.2% ECE at temperature 2, the image results and all latency numbers are the author’s own measurements.

Discrepancies between sources

  • LinkedIn vs card. The launch post on LinkedIn (23 Sep) says “42ms to make a decision on a single RTX 5090 GPU” and “a 24B MoE with 4B activated per token”. The card and blog say 26.5B parameters and 0.09–0.18 s median per decision on an RTX PRO 6000 (0.2 s p90). We keep both as claims and do not pick one.
  • Jev baseline. The card lists Jev 1.13 at 57.89; the blog and the official Space say 57.91.
  • Earlier edition. The v1 card reported Rune at 57.24 vs Jev 59.51, from Invergent’s own run of a JD-Q6_K build on an older edition of the index. The current numbers are for v3 on edition 0.2.1.

Fit / anti-fit

Fit when you need: an open decision model close to Jev’s accuracy that you can run on your own hardware, image decisions (screenshots, documents, charts), long inputs, or data that must stay in your environment.

Anti-fit when you need: a drop-in for the TypeSafe SDK (different endpoint), a small GPU (51.6 GB of weights; Invergent serves it on 96 GB cards), knowledge-heavy questions, or calibrated probabilities without setting the decision temperature.

Limits

  • Trained data and recipe are not published.
  • Default probabilities are overconfident; use --decision-temperature 2, which never changes the chosen option.
  • surogate 1.5.3 can crash under sustained load on a MoE router value; the card says to use a build with fix #217.
  • Thinking does not work together with images yet.

Why it is in the catalog

Invergent trained the weights for typed decisions (full fine-tune, decision protocol) and ships them openly with a documented request format that returns probabilities. That passes the catalog criteria. It is the strongest open entry on the official Decision Index as of 28 Sep 2026.