Span-01 challenges Jev on behavior detection, then judges 11 decision models

Last updated:

Source

On 24 September 2026, Respan AI (the YC company formerly known as Keywords AI) announced Span-01, a proprietary reasoning classifier, and Span-01 Lite, a free variant. On 25 September the project published the launch post on X, pinned to the company profile. Access is waitlist-only at the time of writing: no weights, no published API contract, no latency numbers.

Two things happened in one release, and the second one matters more for this catalog.

1. A commercial peer that claims to beat Jev

Span-01 returns direct present / absent / not_observable probabilities for plain-language behavior definitions, in one non-autoregressive forward pass. Respan trained general classification reasoning with RLAIF, then specialized for behavior detection, and uses hyper-parallel definition branches plus hybrid attention to evaluate several unseen behaviors over one long trace.

On Respan’s own suite it reports 84.3 overall F1, against 71.5 for Jev 1.13.0, 81.5 for GPT-6 Luna and 83.7 for GPT-5.6-terra. The per-domain table is the more honest read: Span-01 scores 0.806 overall, ahead of Jev (0.716) and Sonnet 5 (0.719), but behind GPT-6 Sol (0.885). The “first hyper-parallel reasoning classifier” headline is a marketing claim; the category now has two commercial vendors and a working incumbent.

The benchmark is unusually well documented, which is rare in this category. respanai/behavior-benchmark ships 1,990 traces, 34,885 trace–behavior pairs and 19 languages with core and multilingual splits, per-row upstream licensing, a stated evaluation protocol (F1 on present, bootstrap over task_id) and an explicit limitations section. All benchmark numbers remain UnverifiedClaim — and the card is frank that the labels are model labels, not ground truth, produced mostly by agreement between gpt-5.6-sol and claude-opus-5.

2. The inverted benchmark: 11 decision models judged

Then Respan turned the classifier around. Span-01 was used as the evaluation layer for 11 emerging decision models, across accuracy, consistency, injection resistance and calibration. Span-01 does not appear in its own ranking.

Model Accuracy Paired flip rate Injection ASR ECE
Jev 1.13.0 0.932 0.021 0.063 0.045
Tev1 4B 0.901 0.052 0.042 0.031
Kev 4B 0.833 0.060 0.078 0.070
Decider 2B v10 0.747 0.128 0.208 0.074
AlexWortega OpenJev 4B v5 0.744 0.116 0.089 0.118
Bosun v3.1 1.7B 0.680 0.208 0.323 0.089
Tiny-Jev 0.6B 0.646 0.128 0.224 0.267
Laya 0.542 0.298 0.396 0.092
MoJev 0.85B 0.501 0.150 0.307 0.129
Open-Jev DeBERTa v3 large 0.472 0.194 0.385 0.147
Von 1.2 0.422 0.242 0.297 0.071

Read it with care. This is the first external evaluation of the Jev line that publishes expected calibration error, and the hosted Jev holds up well: best-but-one consistency under paired variants, second-best accuracy, and a competitive ECE against much smaller peers. Tiny-Jev’s 0.267 ECE and Open-Jev DeBERTa’s 0.147 are useful calibration warnings for anyone shipping those locally.

It is also a vendor judging competitors in a domain it was specialized for, with no prompt, no version pinning and no third-party audit. The judge publishes ECE for all eleven and none for itself.

Why this is a category-level event

Three things changed in one announcement:

  • The category got its second commercial vendor, and the incumbent is now measured against someone rather than against LLMs only.
  • A decision model became an evaluation layer. Span-01 is positioned as a replacement for LLM-as-a-judge — the exact pattern that produced pi-jev v0.6.0 and the LangChain harness on the open side.
  • The not_observable decision state is now a product pattern, not an internal detail. A first-class “the input does not contain enough to decide” output is a better interface than forcing a binary verdict on incomplete evidence.

Sources