Span-01 challenges Jev on behavior detection, then judges 11 decision models
On 24 September 2026, Respan AI (the YC company formerly known as Keywords AI) announced Span-01, a proprietary reasoning classifier, and Span-01 Lite, a free variant. On 25 September the project published the launch post on X, pinned to the company profile. Access is waitlist-only at the time of writing: no weights, no published API contract, no latency numbers.
Two things happened in one release, and the second one matters more for this catalog.
1. A commercial peer that claims to beat Jev
Span-01 returns direct present / absent / not_observable probabilities for plain-language behavior definitions, in one non-autoregressive forward pass. Respan trained general classification reasoning with RLAIF, then specialized for behavior detection, and uses hyper-parallel definition branches plus hybrid attention to evaluate several unseen behaviors over one long trace.
On Respan’s own suite it reports 84.3 overall F1, against 71.5 for Jev 1.13.0, 81.5 for GPT-6 Luna and 83.7 for GPT-5.6-terra. The per-domain table is the more honest read: Span-01 scores 0.806 overall, ahead of Jev (0.716) and Sonnet 5 (0.719), but behind GPT-6 Sol (0.885). The “first hyper-parallel reasoning classifier” headline is a marketing claim; the category now has two commercial vendors and a working incumbent.
The benchmark is unusually well documented, which is rare in this category. respanai/behavior-benchmark ships 1,990 traces, 34,885 trace–behavior pairs and 19 languages with core and multilingual splits, per-row upstream licensing, a stated evaluation protocol (F1 on present, bootstrap over task_id) and an explicit limitations section. All benchmark numbers remain UnverifiedClaim — and the card is frank that the labels are model labels, not ground truth, produced mostly by agreement between gpt-5.6-sol and claude-opus-5.
2. The inverted benchmark: 11 decision models judged
Then Respan turned the classifier around. Span-01 was used as the evaluation layer for 11 emerging decision models, across accuracy, consistency, injection resistance and calibration. Span-01 does not appear in its own ranking.
| Model | Accuracy | Paired flip rate | Injection ASR | ECE |
|---|---|---|---|---|
| Jev 1.13.0 | 0.932 | 0.021 | 0.063 | 0.045 |
| Tev1 4B | 0.901 | 0.052 | 0.042 | 0.031 |
| Kev 4B | 0.833 | 0.060 | 0.078 | 0.070 |
| Decider 2B v10 | 0.747 | 0.128 | 0.208 | 0.074 |
| AlexWortega OpenJev 4B v5 | 0.744 | 0.116 | 0.089 | 0.118 |
| Bosun v3.1 1.7B | 0.680 | 0.208 | 0.323 | 0.089 |
| Tiny-Jev 0.6B | 0.646 | 0.128 | 0.224 | 0.267 |
| Laya | 0.542 | 0.298 | 0.396 | 0.092 |
| MoJev 0.85B | 0.501 | 0.150 | 0.307 | 0.129 |
| Open-Jev DeBERTa v3 large | 0.472 | 0.194 | 0.385 | 0.147 |
| Von 1.2 | 0.422 | 0.242 | 0.297 | 0.071 |
Read it with care. This is the first external evaluation of the Jev line that publishes expected calibration error, and the hosted Jev holds up well: best-but-one consistency under paired variants, second-best accuracy, and a competitive ECE against much smaller peers. Tiny-Jev’s 0.267 ECE and Open-Jev DeBERTa’s 0.147 are useful calibration warnings for anyone shipping those locally.
It is also a vendor judging competitors in a domain it was specialized for, with no prompt, no version pinning and no third-party audit. The judge publishes ECE for all eleven and none for itself.
Why this is a category-level event
Three things changed in one announcement:
- The category got its second commercial vendor, and the incumbent is now measured against someone rather than against LLMs only.
- A decision model became an evaluation layer. Span-01 is positioned as a replacement for LLM-as-a-judge — the exact pattern that produced pi-jev v0.6.0 and the LangChain harness on the open side.
- The
not_observabledecision state is now a product pattern, not an internal detail. A first-class “the input does not contain enough to decide” output is a better interface than forcing a binary verdict on incomplete evidence.
