Instinct Tuned 4B

Last updated:

Open62–110ms$0.01/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity6.3/10

    Apache-2.0 merged weights on HF, an open reference runtime with POST /v1/systemone, and a hosted API (free preview on api.zoowork.ai; paid ZooData at $0.01/M input). Card, API docs and GitHub runtime are detailed. Company product, no public SLA.

  • Capability7.8/10

    Noul, Choice (2–16) and Score (2–16). Published temperatures (T=2.80 choice/score, T=0.50 noul) and self-run ECE 0.04–0.11 on hard. p50 62 ms, p95 ≈110 ms estimated [vendor]. Self-reported JevBench public 198/231 (85.71%) with the checkpoint selected on that set — not held-out, not on the official board yet.

  • Adoption6/100

    About 55 HF downloads and 4 likes, and 9 GitHub stars on the runtime (5 Oct). A JevBench bench-request issue is open. No third-party integrations or production use reported beyond ZooWork's own API.

Vendor claims

Read this first

  1. The 85.71% is not a held-out score. ZooWork ran the 231 public JevBench items with the official client and got 198/231. The card says those same public items were used during development for model selection. Treat the table as a reference point, not an independent result. The model is not on the official JevBench v1.5.7 board yet; a bench request asks maintainers to score the production API.
  2. Do not call generate(). The decision is read from the label-token logits with a fixed prompt and temperature. Use the reference runtime or the hosted /v1/systemone API.
  3. Family, not only this checkpoint. ZooWork also hosts instinct (frozen Qwen3.8-27B readout, no fine-tune) and instinct-dual-4b (frozen Qwen3.5-4B with two option orders averaged). Those two already appear on JevBench v1.5.7 (#60 at 18.3 and #30 at 47.0). This page covers the fine-tuned open weights.

What it is

Instinct Tuned 4B is ZooWork’s open decision model: a LoRA fine-tune of Qwen/Qwen3.5-4B, merged into the base weights and released under Apache-2.0. You send a shared state and typed questions; it returns a probability for every candidate from one forward pass and generates no text.

Type Range Output
Choice 2–16 described options Probability per option, argmax, confidence
Noul Yes/no proposition P(yes)
Score 2–16 ordered levels Probability per level and expected level

Read-out. The request is rendered with a fixed prompt (instinct.prompt.v1), run through the chat template with thinking off, and scored from the last hidden state at full depth (layer 32) via the LM-head rows of the candidate label tokens. Softmax uses serving temperatures from decision_config.json: T = 2.80 for choice and score, T = 0.50 for noul (sharpened because JevBench treats P(yes) in (0.2, 0.8) as an abstention).

Serving

  • Open runtime: SerendipityOneInc/instinct (Apache-2.0, ~9★). CLI instinct-decide, Python InstinctModel, and a local POST /v1/systemone reference server. Needs a CUDA GPU with BF16 (Ampere+); about 8.5 GB download, runs on a single 24 GB GPU. Pins transformers==5.16.1.
  • Hosted API: same schema on two routes — free evaluation preview at https://api.zoowork.ai/v1/systemone, paid ZooData at https://api.zoodata.ai/v1/systemone ($0.01/M input for this model; output free). Compatible with the TypeSafe typesafe adapter per the bench-request issue.
  • Self-hosting the weights does not use ZooData billing.

ZooWork’s numbers (UnverifiedClaim)

Claim Figure Note
JevBench public, all 231 198/231 (85.71%) Official client, full depth, bf16, original option order; used for selection
Easy / standard / hard 48/48 · 69/72 · 81/111 Same run
ECE hard (after T) 0.04–0.11 Depends on option order; measured before the separate noul T
Latency p50 62 ms, p95 ≈110 ms (est.) Warmed serial, direct-upstream; no internet/TLS

Training disclosure. About 9.4k decision items (base ~5.9k plus ~900 synthetic hard examples upsampled 4×). Decontamination: every training item has <20% 13-gram overlap with the JevBench public set — that check covers training data only, not the selection use of the public set. Training data is not released yet. T = 2.80 was fitted on a held-out development set from the same task distribution as the benchmark.

Limits

  • Option-order sensitivity: accuracy stays stable, but hard-split ECE moves by 0.04–0.07 between orders.
  • Text only, 8,192-token limit (longer inputs are rejected). Vision tower weights are carried over from the base but unused.
  • Mostly English evaluation. The model decides; it does not explain.

Fit / anti-fit

Fit when you need: a small open decision model with a real /v1/systemone path, optional hosted inference at $0.01/M, and published temperatures.

Anti-fit when you need: independently verified accuracy, choice lists longer than 16, image inputs, or a model that was not tuned against the public JevBench set.

Score working (5 Oct 2026)

  • Maturity = availability 8 × 0.30 + docs 8 × 0.25 + integrations 6 × 0.25 + support 2 × 0.20 = 2.40 + 2.00 + 1.50 + 0.40 = 6.30
  • Capability = types 10 × 0.30 + calibration 7 × 0.25 + latency 8 × 0.25 [vendor] + accuracy 5 × 0.20 = 3.00 + 1.75 + 2.00 + 1.00 = 7.75 → 7.8
  • Adoption = engagement 8 × 0.35 + ecosystem 8 × 0.35 + production 0 × 0.30 = 2.80 + 2.80 + 0 = 5.60 → 6

Accuracy stays at 5 (below the usual 6 for a self-run board score) because the public set was used for checkpoint selection and the model is not on the official board yet.