Instinct Tuned 4B
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity6.3/10
Apache-2.0 merged weights on HF, an open reference runtime with POST /v1/systemone, and a hosted API (free preview on api.zoowork.ai; paid ZooData at $0.01/M input). Card, API docs and GitHub runtime are detailed. Company product, no public SLA.
- Capability7.8/10
Noul, Choice (2–16) and Score (2–16). Published temperatures (T=2.80 choice/score, T=0.50 noul) and self-run ECE 0.04–0.11 on hard. p50 62 ms, p95 ≈110 ms estimated [vendor]. Self-reported JevBench public 198/231 (85.71%) with the checkpoint selected on that set — not held-out, not on the official board yet.
- Adoption6/100
About 55 HF downloads and 4 likes, and 9 GitHub stars on the runtime (5 Oct). A JevBench bench-request issue is open. No third-party integrations or production use reported beyond ZooWork's own API.
Vendor claims
- JevBench public (231 items): 198/231 = 85.71% (easy 48/48, standard 69/72, hard 81/111); checkpoint also used for model selection — not a held-out score[Vendor claim — not independently verified]
- Direct-upstream serving latency p50 62 ms; p95 ≈110 ms estimated (warmed, serial; excludes internet/TLS/gateway)[Vendor claim — not independently verified]
- ECE on the hard split after temperature scaling: 0.04–0.11 depending on option order (measured at T=2.80 for every type, before the separate noul temperature)[Vendor claim — not independently verified]
- ZooData paid API: instinct-tuned-4b at US$0.01 per million input tokens; free evaluation preview at api.zoowork.ai[Vendor claim — not independently verified]
Read this first
- The 85.71% is not a held-out score. ZooWork ran the 231 public JevBench items with the official client and got 198/231. The card says those same public items were used during development for model selection. Treat the table as a reference point, not an independent result. The model is not on the official JevBench v1.5.7 board yet; a bench request asks maintainers to score the production API.
- Do not call
generate(). The decision is read from the label-token logits with a fixed prompt and temperature. Use the reference runtime or the hosted/v1/systemoneAPI.- Family, not only this checkpoint. ZooWork also hosts
instinct(frozen Qwen3.8-27B readout, no fine-tune) andinstinct-dual-4b(frozen Qwen3.5-4B with two option orders averaged). Those two already appear on JevBench v1.5.7 (#60 at 18.3 and #30 at 47.0). This page covers the fine-tuned open weights.
What it is
Instinct Tuned 4B is ZooWork’s open decision model: a LoRA fine-tune of Qwen/Qwen3.5-4B, merged into the base weights and released under Apache-2.0. You send a shared state and typed questions; it returns a probability for every candidate from one forward pass and generates no text.
| Type | Range | Output |
|---|---|---|
| Choice | 2–16 described options | Probability per option, argmax, confidence |
| Noul | Yes/no proposition | P(yes) |
| Score | 2–16 ordered levels | Probability per level and expected level |
Read-out. The request is rendered with a fixed prompt (instinct.prompt.v1), run through the chat template with thinking off, and scored from the last hidden state at full depth (layer 32) via the LM-head rows of the candidate label tokens. Softmax uses serving temperatures from decision_config.json: T = 2.80 for choice and score, T = 0.50 for noul (sharpened because JevBench treats P(yes) in (0.2, 0.8) as an abstention).
Serving
- Open runtime:
SerendipityOneInc/instinct(Apache-2.0, ~9★). CLIinstinct-decide, PythonInstinctModel, and a localPOST /v1/systemonereference server. Needs a CUDA GPU with BF16 (Ampere+); about 8.5 GB download, runs on a single 24 GB GPU. Pinstransformers==5.16.1. - Hosted API: same schema on two routes — free evaluation preview at
https://api.zoowork.ai/v1/systemone, paid ZooData athttps://api.zoodata.ai/v1/systemone($0.01/M input for this model; output free). Compatible with the TypeSafetypesafeadapter per the bench-request issue. - Self-hosting the weights does not use ZooData billing.
ZooWork’s numbers (UnverifiedClaim)
| Claim | Figure | Note |
|---|---|---|
| JevBench public, all 231 | 198/231 (85.71%) | Official client, full depth, bf16, original option order; used for selection |
| Easy / standard / hard | 48/48 · 69/72 · 81/111 | Same run |
| ECE hard (after T) | 0.04–0.11 | Depends on option order; measured before the separate noul T |
| Latency | p50 62 ms, p95 ≈110 ms (est.) | Warmed serial, direct-upstream; no internet/TLS |
Training disclosure. About 9.4k decision items (base ~5.9k plus ~900 synthetic hard examples upsampled 4×). Decontamination: every training item has <20% 13-gram overlap with the JevBench public set — that check covers training data only, not the selection use of the public set. Training data is not released yet. T = 2.80 was fitted on a held-out development set from the same task distribution as the benchmark.
Limits
- Option-order sensitivity: accuracy stays stable, but hard-split ECE moves by 0.04–0.07 between orders.
- Text only, 8,192-token limit (longer inputs are rejected). Vision tower weights are carried over from the base but unused.
- Mostly English evaluation. The model decides; it does not explain.
Fit / anti-fit
Fit when you need: a small open decision model with a real /v1/systemone path, optional hosted inference at $0.01/M, and published temperatures.
Anti-fit when you need: independently verified accuracy, choice lists longer than 16, image inputs, or a model that was not tuned against the public JevBench set.
Score working (5 Oct 2026)
- Maturity = availability 8 × 0.30 + docs 8 × 0.25 + integrations 6 × 0.25 + support 2 × 0.20 = 2.40 + 2.00 + 1.50 + 0.40 = 6.30
- Capability = types 10 × 0.30 + calibration 7 × 0.25 + latency 8 × 0.25
[vendor]+ accuracy 5 × 0.20 = 3.00 + 1.75 + 2.00 + 1.00 = 7.75 → 7.8 - Adoption = engagement 8 × 0.35 + ecosystem 8 × 0.35 + production 0 × 0.30 = 2.80 + 2.80 + 0 = 5.60 → 6
Accuracy stays at 5 (below the usual 6 for a self-run board score) because the public set was used for checkpoint selection and the model is not on the official board yet.
