Mica v0.1 4B
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity5.4/10
Apache-2.0 merged BF16 weights plus six GGUF quants, a Dockerfile, llama.cpp build scripts and a /v1/systemone server that JevBench's unchanged typesafe adapter runs against. Detailed card with a quant table. Self-host only, no SLA.
- Capability7.3/10
Noul, Choice (2–255 options) and Score (2–10 levels). Temperature 1.124 fitted on a calibration set, held-out ECE 5.4%. p50 54 ms / p95 552 ms on an RTX 3090 and JevBench public per-item outputs, all author-run. Not on the Decision Index.
- Adoption8/100
About 37 HF likes and 4k downloads (GGUF included) five days after release, 27 GitHub stars. No launch post found, no integrations or production use reported.
Vendor claims
- JevBench public 231 items via the llama.cpp server: easy 1.000, original 1.000, hard 0.649, ECE 0.064, p50 54 ms, p95 552 ms (RTX 3090, BF16 GGUF)[Vendor claim — not independently verified]
- Held-out set (7,328 items, EN + KO): Mica 67.4 vs JEV 1.13 74.7, Kev 4B 56.4, Nimble 9B 53.9[Vendor claim — not independently verified]
- Planted wrong-option note in the state: Mica 69.1% correct vs JEV 17.5%, Kev 31.4%[Vendor claim — not independently verified]
- No JevBench item in the 77,732 training rows (exact-match and 8-gram checks)[Vendor claim — not independently verified]
- Tetris demo: 223 lines over three seeds vs 55 for Kev 4B and 17 for Laya[Vendor claim — not independently verified]
Mica v0.1 4B is a small open decision model published on Hugging Face by sky7350 on 25 September 2026, with code and scripts on GitHub under akivet. You give it a state, a question and the allowed answers; it returns a probability for each answer, with no generated text. A decision costs one prefill, and several questions about the same state share it.
It speaks TypeSafe’s /v1/systemone format, so Jev clients work unchanged. It was trained on English and Korean. It is not a TypeSafe product, not distilled from Jev, and not a derivative of another catalogued model: the card uses Kev, Laya, JevK5 and Nimble only as baselines.
Specs
| Attribute | Value |
|---|---|
| Author | sky7350 (HF) / akivet (GitHub) |
| Release | 25 Sep 2026 (HF repo 04:44 BRT) |
| Base model | Qwen/Qwen3.5-4B (revision 851bf6e), shape unchanged, no new heads |
| Training | Rank-16 LoRA on attention and Gated DeltaNet projections, merged; one epoch of cross-entropy on verified answers; best of three seeds |
| Data | About 34k source decisions / 78k rows in 12 areas (coding agents, code review, computer use, policies, routing, games, general knowledge…), checked by running code where possible |
| Readout | Softmax over 255 fixed single-token labels (yes/no uses “No”/“Yes”), divided by a fitted temperature of 1.124 |
| Question types | noul (catalog boolean), choice (2–255 options), score (2–10 levels) |
| Files | Merged BF16 safetensors; GGUF in BF16, Q8_0, Q6_K, Q5_K_M, Q4_K_M, Q4_0 |
| Serving | Docker image or llama.cpp b11010 build → http://127.0.0.1:8010/v1/systemone; inputs over 8,192 tokens get HTTP 400 |
| License | Apache-2.0 |
Author’s results (UnverifiedClaim)
JevBench public set (231 items, through JevBench’s own runner and unchanged typesafe adapter, per-item output published): easy 1.000, original 1.000, hard 0.649, ECE 0.064. The author’s PyTorch evaluation gives 69.5 on the hard tier; the served GGUF gives 64.9. Half of the public hard items were used to compare recipes during development.
| Set | n | Mica | JEV 1.13 | Kev 4B | Nimble 9B |
|---|---|---|---|---|---|
| SemIf | 252 | 94.4 | 98.4 | 89.3 | 93.2 |
| Kev transfer v9 | 1,264 | 69.2 | 82.0 | 73.5 | – |
| MMLU-Pro | 10,032 | 53.0 | 82.3 | 49.7 | – |
| Held-out (EN + KO) | 7,328 | 67.4 | 74.7 | 56.4 | 53.9 |
| Perturbed inputs | 2,752 | 77.3 | 73.3 | 48.7 | 44.5 |
The held-out set was written after the training data was frozen. Some of the author’s own sets (long inputs, chat style, knowledge) are close to the training distribution, and the card says to read them as in-distribution. Held-out calibration: ECE 5.4%, 2.5% of answers wrong at confidence ≥ 0.9 (JEV: 3.8% and 2.0%).
Latency on one RTX 3090, one request at a time: p50 54 ms / p90 526 ms (BF16), 47 ms / 466 ms (Q4_K_M). The p90 comes from hard items of up to ~3.7k tokens.
Quantization: on the author’s 1,402-item calibration set, Q5_K_M keeps 96.1% of BF16 answers and is the suggested pick for an 8 GB GPU.
Fit / anti-fit
Fit when you need: a small local decision model on a consumer GPU, Korean or mixed English/Korean inputs, agent and code-review decisions, or robustness to instructions planted inside the state.
Anti-fit when you need: knowledge-heavy questions (53.0 on MMLU-Pro vs 82.3 for Jev), long English policy documents (weakest public set), inputs over 8k tokens, or independently verified numbers.
Limits
- Every number is author-run; Mica is not on the Decision Index or the JevBench v1.4.2 leaderboard.
- Planted notes in the state still move answers somewhat.
- Single release (v0.1) from an individual author, no hosted option.
Why it is in the catalog
The LoRA was trained for typed decisions and merged into published weights, with a System One-compatible server that returns probabilities. That passes the catalog criteria.
