Span-01
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity3.5/10
Announced behind a waitlist with no public API contract, no request/response schema and no latency numbers. Respan's own platform docs and SDKs are strong, but the model itself is not yet in the gateway catalog.
- Capability4.3/10
Direct parallel probabilities for present/absent/not_observable in one forward pass, plus a public benchmark with a downloadable dataset. One decision type only, no published p50/p99, and no calibration study for Span-01 itself.
- Adoption30/100
Strong launch-day signal for a waitlist product (29.4k views, 632 likes, 400 bookmarks, YC company) but no public integration, no open weights and no reported production use of the model.
Vendor claims
- Overall behavior F1 of 84.3, against 81.5 for GPT-6 Luna and 71.5 for Jev 1.13.0[Vendor claim — not independently verified]
- 2× cheaper and 18% better than Jev; 700× cheaper and 4% better than GPT-6 Luna[Vendor claim — not independently verified]
- Production behavior benchmark overall 0.806, against Jev 0.716, Sonnet 5 0.719 and GPT-6 Sol 0.885[Vendor claim — not independently verified]
- Span-01 Lite scores better than Jev and Sonnet 5 and is free[Vendor claim — not independently verified]
- Pricing of $0.02 per million input tokens with free output; Span-01 Lite free on both[Vendor claim — not independently verified]
Span-01 is a proprietary reasoning classifier from Respan AI (Keywords AI, Inc., a YC company), announced on 24 September 2026. It is the first direct commercial competitor to Jev in the category: same shape of product — a dedicated classifier that returns probabilities instead of text — built for a narrower and more operationally specific job.
You hand it a trace and one or more behavior definitions written in plain language. It returns a direct probability for each definition in parallel, in a single forward pass, with no token-by-token generation.
The idea
The hard part of production classification is not emitting a label. It is applying a definition the model has never seen to a trace it has never seen. Span-01 is trained for exactly that: Respan describes general classification reasoning trained with RLAIF, then specialization for behavior detection, so the model executes new definitions instead of memorizing a fixed taxonomy.
Three architectural details are claimed:
- Hyper-parallel definition branches — several unseen behaviors are evaluated over the same trace context at once.
- Hybrid attention — used to keep that reasoning intact across long production traces.
- One true forward pass — direct
present,absentandnot_observableprobabilities, no generation.
The not_observable output is the interface detail worth borrowing. Most decision models force a verdict even when the input does not contain the evidence. Span-01 has a first-class third state, which is functionally the same job as Jev’s noul, and it shows up in the benchmark’s own label distribution.
Specs
| Attribute | Value |
|---|---|
| Company | Respan AI / Keywords AI, Inc. (YC) |
| Announcement author | Frank Chen |
| Launch | |
| Status | Early access, waitlist only |
| Variants | Span-01, Span-01 Lite |
| Decision output | present / absent / not_observable probabilities |
| Architecture | Non-autoregressive, hyper-parallel definition branches, hybrid attention |
| Training | RLAIF, then behavior-detection specialization (vendor term) |
| Latency | Not published |
| Pricing | $0.02/M input, output free · Span-01 Lite free on both |
| Weights | Closed — no model repository on Hugging Face |
| Benchmark | respanai/behavior-benchmark |
Benchmarks
Behavior detection (Respan’s own suite)
| System | Core | Multilingual | All |
|---|---|---|---|
| Span-01 | 0.784 | 0.902 | 0.843 |
| GPT-5.6-terra | 0.756 | 0.917 | 0.837 |
| deepseek-v4-flash | 0.743 | 0.885 | 0.814 |
| GPT-6-luna | 0.728 | 0.901 | 0.815 |
| Span-01 Lite | 0.673 | 0.848 | 0.761 |
| qwen3-235b | 0.584 | 0.772 | 0.678 |
| Sonnet 5 | 0.576 | 0.858 | 0.610 |
| jev | 0.574 | 0.856 | 0.715 |
| laya | 0.196 | 0.250 | 0.223 |
All figures are UnverifiedClaim. GPT-5.6-terra actually beats Span-01 on the multilingual split, and the header slide that markets “84.3 vs 81.5 vs 71.5” omits GPT-6 Sol, which scores higher in the vendor’s own domain table.
Per-domain production behavior scores
| Behavior domain | Span-01 | Jev | Sonnet 5 | GPT-6 Sol |
|---|---|---|---|---|
| Jailbreak and prompt injection | 0.779 | 0.752 | 0.709 | 0.878 |
| Safety and refusals | 0.803 | 0.751 | 0.785 | 0.911 |
| Privacy and secrets | 1.000 | 0.948 | 0.901 | 0.935 |
| Hallucination and grounding | 0.796 | 0.673 | 0.756 | 0.903 |
| Agent and tool reliability | 0.845 | 0.691 | 0.677 | 0.861 |
| Task and instruction following | 0.702 | 0.671 | 0.621 | 0.821 |
| Response quality | 0.771 | 0.691 | 0.722 | 0.956 |
| Overall | 0.806 | 0.716 | 0.719 | 0.885 |
This is the number to look at. Span-01 beats Jev and Sonnet 5, and a frontier model still sets the ceiling. The overall figures here (0.806) and the F1 figures in the table above (0.843) are different metrics and should not be mixed.
The benchmark has no ground truth — read this before trusting the gap
The dataset card is unusually candid, which is a point in Respan’s favor and a hard limit on the headline claim. Labels come from three methods:
| method | core |
multilingual |
|---|---|---|
gpt-5.6-sol and claude-opus-5 agree |
24,074 | 5,541 |
Tie-break: teachers disagreed, gpt-6-astra decided |
2,421 | — |
| Construction: set by how the trace was built | — | 2,851 |
The card says it directly: “These are model labels, not ground truth”, and the teachers “can’t be fairly evaluated on this benchmark”.
So “18% better than Jev” measures agreement with frontier teacher models on teacher labels, not intrinsic decision quality. A classifier post-trained to approximate those teachers has a structural advantage over a competitor that was not. The dataset is still genuinely useful — 1,990 traces, 34,885 trace–behavior pairs, 19 languages, documented evaluation protocol, per-row upstream licensing — but it is a vendor benchmark.
Two more limits: positive labels are 10.4% of rows and the protocol reports F1 on present without publishing recall or a confusion matrix, and the cost Pareto chart normalizes to $0.003/event while excluding a $299/month base fee.
Span-01 judged 11 decision models
Respan inverted its own benchmark and used Span-01 as the evaluation layer for 11 emerging decision models. Span-01 does not appear in the ranking — it produced the signal.
| Model | Accuracy | Paired flip rate | Injection ASR | ECE |
|---|---|---|---|---|
| Jev 1.13.0 | 0.932 | 0.021 | 0.063 | 0.045 |
| Tev1 4B | 0.901 | 0.052 | 0.042 | 0.031 |
| Kev 4B | 0.833 | 0.060 | 0.078 | 0.070 |
| Decider 2B v10 | 0.747 | 0.128 | 0.208 | 0.074 |
| AlexWortega OpenJev 4B v5 | 0.744 | 0.116 | 0.089 | 0.118 |
| Bosun v3.1 1.7B | 0.680 | 0.208 | 0.323 | 0.089 |
| Tiny-Jev 0.6B | 0.646 | 0.128 | 0.224 | 0.267 |
| Laya | 0.542 | 0.298 | 0.396 | 0.092 |
| MoJev 0.85B | 0.501 | 0.150 | 0.307 | 0.129 |
| Open-Jev DeBERTa v3 large | 0.472 | 0.194 | 0.385 | 0.147 |
| Von 1.2 | 0.422 | 0.242 | 0.297 | 0.071 |
Paired flip rate measures consistency under equivalent variants. Injection ASR is attack success rate. ECE is expected calibration error.
This is the first external evaluation of the Jev line that publishes calibration error, and the hosted Jev comes out with the second-lowest flip rate and a competitive ECE — the best result in that column among the larger models. Treat it as a third-party claim: Span-01 is judging competitors in a domain it was specialized for, and no prompt or version is published. Nine of these eleven models are catalogued on this site; we have not folded these numbers into their entries.
Fit / anti-fit
Fit when you need: a behaviour-detection layer over LLM traces, judge-style evaluation of agent output, prompt-injection or privacy screens, or a cheap replacement for an LLM-as-a-judge call on a narrow taxonomy — once the waitlist opens.
Anti-fit when you need: open weights, self-hosting, a documented API contract, a general decision model outside behaviour detection, calibrated Score semantics, or any decision you cannot waitlist for.
Limits
- No public API contract. No endpoint, request schema or response shape is published — only the price table and a rate-limited playground (10 comparisons per IP per 10 minutes).
- No latency data. “At classifier speed” is the entire argument and it carries no p50/p99 number. The
0–0msfrontmatter value is an unavailable-data sentinel. - No weights, no license, no repo.
author=respanaireturns zero models on Hugging Face. This is not a peer of the open catalog entries. - One decision type. No
choiceover arbitrary options, noscorerubric.not_observableis a product idea worth copying, not a broad decision interface. - The judge is not audited. Respan publishes ECE for the 11 models it evaluated and none for Span-01 itself.
- The vendor benchmark is the only benchmark. Teacher labels, 10.4% positives, F1 on
present, and a cost Pareto that excludes the base fee. - A frontier model still wins. GPT-6 Sol scores 0.885 against Span-01’s 0.806. Span-01 replaces the mid-tier judge, not the frontier ceiling.
