Tev1-4B-experimental
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity5.8/10
Open weights, full training recipe, hosted serverless endpoint, and public inference examples; still experimental, generative, and without a published latency benchmark or SLA.
- Capability2.3/10
Choice-only structured output from a 2–24 option prompt. The model keeps Qwen's autoregressive head and exposes token logprobs as preferences, not calibrated probabilities; development results are self-reported.
- Adoption19/100
4B: ~2.4k HF downloads and 24 likes (was 741 / 12); 0.8B: ~800 downloads; GitHub recipe 177 stars. Since 30 Sep both sizes ship in Ollama (~2.7k pulls in ~14h). Still an experimental niche release.
Vendor claims
- 88.0% on 1,000 main development decisions and 300/300 policy-transfer decisions[Vendor claim — not independently verified]
- Training the recipe costs about $17[Vendor claim — not independently verified]
- Serverless pricing of $0.042 per million input tokens and $0 per million output tokens[Vendor claim — not independently verified]
Tev1-4B-experimental is an open, Jev-inspired decision model from Together AI. It is a supervised fine-tune of Qwen3.5-4B trained to choose one option from a structured state, question, and list of 2–24 choices.
The model card is explicit about the boundary: this is not a non-autoregressive Jev runtime. It keeps Qwen’s standard language-model head, and the application maps the returned option letter back to a semantic key. For that reason, the catalog entry is limited to choice; it does not claim Jev’s calibrated Score or Boolean interfaces.
Training recipe
The public repository includes the data recipe and a training example. The new v1 run combines 37,840 unique training examples with 4,568 validation examples across language classification, policy decisions, routing, and synthetic research classification. The training example uses LoRA-style supervised fine-tuning on Qwen3.5-4B.
The repository warns that the saved evaluation mixture influenced model development. The reported 88.0% result on 1,000 main decisions and 300/300 policy-transfer decisions are bring-up results, not an independent benchmark. The public model card also says that logprobs are model preferences, not calibrated confidence.
The 0.8B sibling, Tev1-0.8B-experimental, went public on Hugging Face on 25 September 2026. Same recipe and interface on Qwen3.5-0.8B; its card says the weight licence “is being finalized”. It is covered here rather than as a separate entry.
On Ollama (30 Sep 2026)
Both sizes are launch models for Ollama’s /v1/systemone endpoint: ollama pull tev1 (4B, 4.5 GB Q8_0, 4.21B params) and ollama pull tev1:0.8b (812 MB Q8_0, 752M params). Ollama’s library page lists three question types (choice, noul, score) and advises staying within 2–24 options, the range Tev1 was trained on.
The catalog keeps decisionTypes at choice. On Ollama, the noul and score shapes come from Ollama’s candidate scoring, not from Tev1’s training, and Together’s card still describes a one-letter output with uncalibrated logprobs.
UnverifiedClaim: Ollama’s run on Bespoke Labs’ 13 public human-labelled datasets (3,880 decisions) gives 73.3% for Tev1 4B and 63.5% for Tev1 0.8B, against 76.0% for Jev 1.13 (Jev figure from Bespoke’s published run). (blog · library)
Fit / anti-fit
Fit when you need: an open Jev-inspired classifier, a reproducible fine-tuning recipe, or a controlled choice decision behind a typed application schema.
Anti-fit when you need: a non-autoregressive System One runtime, calibrated probabilities, Boolean/Noul or Score semantics, or production decisions without your own held-out evaluation.
Limits
- No public p50/p99 latency benchmark. The
0–0msfrontmatter value is an unavailable-data sentinel. - The 88% and 100% figures are development results and UnverifiedClaim.
- The model is still a generative Qwen checkpoint and may produce prose outside the constrained interface.
- The
$0.042/Minput and$0/Moutput prices are vendor claims for Together’s serverless endpoint.
