Gero-4B
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity4.6/10
Apache-2.0 fine-tune of Qwen3-4B on HF with a long card and copy-paste transformers helpers. No dedicated runtime, no /v1/systemone server, no hosted API. Single independent author, no SLA.
- Capability6.1/10
Choice (up to 256 options), Score (2–10 levels) and yes/no, with no position bias by design. Calibration via a described RL stage with a proper scoring rule, but no ECE or temperatures published. No latency figures and no public benchmark numbers.
- Adoption2/100
About 65 HF downloads and 6 likes (5 Oct). No GitHub runtime, no integrations and no production use reported.
Vendor claims
- Branching cross-encoder: shared state/question prefix, one linear scorer per option branch, softmax across branches; option order does not change any probability[Vendor claim — not independently verified]
- Training stage 3 uses reinforcement learning for calibrated decisions with a proper scoring rule over sampled outcomes from the true answer distribution[Vendor claim — not independently verified]
- Supports up to 256 choice options and ordered score scales of 2–10 levels; English is the training language[Vendor claim — not independently verified]
Read this first
- No published numbers. The card describes architecture and a three-stage training recipe, including reinforcement learning for calibration, but it does not report accuracy, ECE, latency or any public benchmark. Everything about quality is an architectural claim until someone measures it.
- Not
/v1/systemone. Inference is viaAutoModelForSequenceClassificationhelpers in the card (one forward pass per option, then softmax). There is no TypeSafe-compatible server in the repo.- Latency fields are 0–0 ms on this page because none are published — that is the catalog sentinel for “unknown”, not a claim of zero cost.
What it is
Gero-4B is the first model in an independent “Gero” family (named after Gerolamo Cardano). It starts from Qwen/Qwen3-4B and is rebuilt into a scorer: the language-model head is removed and a single learned linear scorer reads the final token of each option.
| Type | Goal | Returns |
|---|---|---|
| Choice | Pick one of up to 256 options | Selected label, probabilities, confidence |
| Score | Place the state on 2–10 ordered levels | Expected score, per-level probabilities, confidence |
| Yes/no | Whether a statement holds | P(yes) |
Branching cross-encoder. The state and the question form a shared prefix. Each option is its own branch that attends to that prefix; options never see one another, and the prefix never sees the options. A softmax across branches turns the scores into one distribution. Consequences the author emphasises: no position bias, any option count with one shared scorer, and (with prefix caching) the state is encoded once.
Training (as disclosed)
- Readout — branching architecture and shared scorer.
- Instruction and structure — purpose-built data for the three question types, option sets from 1 to 256, ordered scales, negation and unbound attributes.
- Reinforcement learning for calibrated decisions — outcomes sampled from each item’s true answer distribution; reward combines probability of the sampled outcome with the calibration of the top answer’s confidence (described as a proper scoring rule).
No dataset, compute budget, ECE or temperature table is published. English is the training language.
Fit / anti-fit
Fit when you need: an open scorer that natively handles long choice lists (up to 256) without a reserved slot per option, and you are willing to measure calibration yourself.
Anti-fit when you need: verified accuracy, a /v1/systemone endpoint, published latency, or anything beyond English without your own eval.
Score working (5 Oct 2026)
- Maturity = availability 8 × 0.30 + docs 6 × 0.25 + integrations 2 × 0.25 + support 1 × 0.20 = 2.40 + 1.50 + 0.50 + 0.20 = 4.60
- Capability = types 10 × 0.30 + calibration 5 × 0.25 + latency 4 × 0.25 + accuracy 4 × 0.20 = 3.00 + 1.25 + 1.00 + 0.80 = 6.05 → 6.1
- Adoption = engagement 5 × 0.35 + ecosystem 2 × 0.35 + production 0 × 0.30 = 1.75 + 0.70 + 0 = 2.45 → 2
Calibration 5: RL method is described (“claim + method” band starts at 7 when numbers exist); without ECE it stays in the lower published-method band. Latency 4 and accuracy 4 follow the “no public numbers” precedents (d1 latency 4; accuracy claims without numbers = 3–5).
