Hopper

Last updated:

Open130–570ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity4.8/10

    Open weights (research/demo license due to RACE data), serving code Apache-2.0, calibration map published; extensive disclosure on development process and limitations.

  • Capability7.2/10

    JevBench #5 overall (Score 59.43), strong calibration (79.1), supports choice/score/boolean via one-pass softmax with per-type temperature calibration.

  • Adoption20/100

    Top 5 JevBench ranking generates community interest; HF discussion activity on long-menu support; self-host only limits broader adoption.

Vendor claims

Hopper is an open LoRA adapter (rank 16) on Qwen3.5-4B designed for the JevBench setting: document + policy + question → probability distribution over options. One forward pass, no text generation. Research and demo use only due to RACE training data license terms.

The author provides unusually transparent disclosure on benchmark-directed development: 26 distinct model/prompt configurations were tested, plus more than twenty calibration-map variants. Half of the public JevBench items (115) served as a development gate; the other half (116) was reserved. This transparency is valuable — not a red flag — but users should expect held-out performance closer to the frozen base model on unseen distributions.

JevBench v1.4.2 Results

Metric Score
Overall 59.43 (Rank #5)
Intelligence 48.0
Calibration 79.1
Speed 86.8
Cost 58.7 (~$0.024/1k decisions)

JevBench maintainer’s official run. Author’s local numbers on the 231 public items: easy 100%, standard 94.4%, hard 68.5%.

Decision Index 0.2.1 (official board)

On the official Decision Index 0.2.1 board (snapshot 28 Sep 2026), Hopper (G) 1.2 scores 40.77, 18th of 70 open entries. Hopper (G) is a general-purpose sibling adapter (HopitAI/hopper-g) trained on from Hopper 1.0, under the same research-and-demo license. It is not the exact checkpoint behind the JevBench number above. The row comes from the authors’ complete run with the official scorer. The maintainers measured a median of 23.3 ms on an RTX PRO 6000 and an ECE of 0.093. For comparison on the same board: Jev 57.91, Kev 9B 38.48, Decider 4B 40.70. Fabio Akita’s post (3 Oct 2026, in Portuguese) highlights it as the 4B option with “41 points at 23 milliseconds”. The capability score (7.2) still rests on Hopper 1.0’s JevBench result and has not been re-scored against the index. Hopper (G) 1.3 is not on the board yet.

Architecture

Hopper uses a one-pass softmax readout over option letters (A, B, C, …) with thinking disabled. No tokens are generated — the answer is the softmax over next-token logits restricted to option letters.

A calibration map (hopper.json) rescales the distribution with one temperature per answer type:

  • Choice: 0.790
  • Noul (boolean): 0.753
  • Score: 0.900

The calibration map was fitted only on HopitAI’s own held-out JevBench-style items — never on actual JevBench items. It cannot change the top answer (dividing log-probabilities by a positive number preserves ordering); it only adjusts confidences.

Specs

Attribute Value
Author HopitAI
Base model Qwen/Qwen3.5-4B (revision 851bf6e)
Adapter LoRA rank 16
License Research/demo only (adapter); Apache-2.0 (serving code)
Launch
Status Open weights (Hugging Face)
Latency ~130ms (RunPod GPU)
Pricing Self-host only (free weights)
Decision types Choice, Score, Boolean (noul)
Context 32K (inherited from Qwen3.5)
Serving code github.com/hopit-ai/hopper
Wire format JevBench /v1/systemone compatible

Training Data

Mixed training from three sources:

  1. Synthetic decision families — LLM-generated, labels computed in code
  2. JevBench-style synthetic set — items kept only where independent LLM solvers agreed; checked against all public JevBench questions/states (normalized identity, 8-word sequence matching) with matches dropped
  3. Public human-labelled datasets — ARC, CommonsenseQA, MMLU aux (includes RACE → license restriction), SNLI, MultiNLI, VitaminC, BoolQ, SQuAD v2, CLINC OOS, DBpedia-14, HelpSteer2

⚠️ License note: The MMLU auxiliary set bundles RACE, which its authors release for non-commercial research only with terms that extend to derived data. This is why Hopper is research/demo only. A version trained without RACE is in development.

Benchmark-Directed Development (Disclosed)

The author is transparent about extensive development against public JevBench items:

  • 26 model/prompt configurations tested
  • 20+ calibration-map variants explored
  • Development half (115 items): used as a gate many times; hard tier scores 0.709
  • Reserved half (116 items): scored only in aggregate; hard tier scores 0.661 (level with frozen base at 0.643)

The JevBench-style training set’s style sheet was written by reading the development half, and some training items target behaviors observed to fail on hard items.

“Expect the held-out hard items to score below the public ones, and expect the judge tier, which we have never seen, to be the least predictable part.” — Model card

This is not contamination — no JevBench item or paraphrase was used in training, verified via hash-based checking. It is honest benchmark-directed development with full disclosure. The 0 state/instruction matches from JevBench’s scan of 17 released files vs 231 public tasks supports this.

Long Menus (>26 options) — v1.1.1

Since the adapter uses letter logits (A–Z), menus over 26 options require a two-stage approach (from v1.1.1):

  1. First stage: Tournament deals menu into ceil(n/26) chunks, reads each with single pass, keeps top 10 per chunk
  2. Final pass: Ordinary single pass decides among survivors

Every option gets a probability: final distribution mixed with 5% uniform over whole menu (eliminated options get 0.05/n, never zero).

Benchmark Options Top-1 Accuracy First Stage Recall@10 Latency p50
BANKING77 77 0.673 0.887 321ms
CLINC150 (in-scope) 150 0.863 0.977 569ms

Measured on A10G, 300 items each. On these long menus, Hopper is level with the frozen base within sampling error.

Known Limitations

  • One pass of a 4B model — no step-by-step reasoning; any problem needing intermediate results is decided in one forward pass
  • Dates and multi-step arithmetic are weak — date differences, deadlines, chained calculations fail often
  • Long documents needing several hops are weak — accuracy drops when answers need facts from distant parts
  • Calibration fitted on author’s data — confidences may be off on different distributions; v1.1.0 simplified to three temperatures because richer maps didn’t transfer
  • Dev half flatters it — on reserved half, level with frozen base on hard accuracy
  • English only — no other languages tested
  • Long menus overconfident — mean top probability 0.78 against 0.67 accuracy on BANKING77

How to Use

With the serving package:

from hopper_decisions import Decider
decider = Decider(adapter="HopitAI/hopper")
decider.decide({
    "state": "The customer wants a refund for order 12.",
    "questions": {
        "decision": {
            "type": "choice",
            "instructions": "Route the ticket.",
            "criteria": {"refund": "money back", "track": "where is it"}
        }
    }
})

Or with transformers + peft directly (see model card for full example).

Dependencies: flash-linear-attention==0.5.2, causal-conv1d 1.7.0, torch==2.8.0, transformers==5.17.0, peft==0.21.0, accelerate==1.15.0. Without the linear-attention kernels, inference is >10× slower.

Fit / Anti-fit

Fit when you need: self-hosted typed decisions, transparent development disclosure, strong calibration, research/demo use cases, JevBench-compatible wire format.

Anti-fit when you need: commercial deployment (RACE license restriction), production SLA, hosted inference, tasks requiring multi-hop reasoning or date arithmetic, non-English languages.

JevBench v1.5.7 (Benchmark Heaven, checked 5 Oct 2026)

Hopper (the JevBench checkpoint, not Hopper G) is #17 at official score 67.5 (capability 68.9) on v1.5.7. See news. No catalog score change yet.