Span-01

Last updated:

Waitlist—$0.02/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity3.5/10

    Announced behind a waitlist with no public API contract, no request/response schema and no latency numbers. Respan's own platform docs and SDKs are strong, but the model itself is not yet in the gateway catalog.

  • Capability4.3/10

    Direct parallel probabilities for present/absent/not_observable in one forward pass, plus a public benchmark with a downloadable dataset. One decision type only, no published p50/p99, and no calibration study for Span-01 itself.

  • Adoption30/100

    Strong launch-day signal for a waitlist product (29.4k views, 632 likes, 400 bookmarks, YC company) but no public integration, no open weights and no reported production use of the model.

Vendor claims

Span-01 is a proprietary reasoning classifier from Respan AI (Keywords AI, Inc., a YC company), announced on 24 September 2026. It is the first direct commercial competitor to Jev in the category: same shape of product — a dedicated classifier that returns probabilities instead of text — built for a narrower and more operationally specific job.

You hand it a trace and one or more behavior definitions written in plain language. It returns a direct probability for each definition in parallel, in a single forward pass, with no token-by-token generation.

The idea

The hard part of production classification is not emitting a label. It is applying a definition the model has never seen to a trace it has never seen. Span-01 is trained for exactly that: Respan describes general classification reasoning trained with RLAIF, then specialization for behavior detection, so the model executes new definitions instead of memorizing a fixed taxonomy.

Three architectural details are claimed:

  • Hyper-parallel definition branches — several unseen behaviors are evaluated over the same trace context at once.
  • Hybrid attention — used to keep that reasoning intact across long production traces.
  • One true forward pass — direct present, absent and not_observable probabilities, no generation.

The not_observable output is the interface detail worth borrowing. Most decision models force a verdict even when the input does not contain the evidence. Span-01 has a first-class third state, which is functionally the same job as Jev’s noul, and it shows up in the benchmark’s own label distribution.

Specs

Attribute Value
Company Respan AI / Keywords AI, Inc. (YC)
Announcement author Frank Chen
Launch
Status Early access, waitlist only
Variants Span-01, Span-01 Lite
Decision output present / absent / not_observable probabilities
Architecture Non-autoregressive, hyper-parallel definition branches, hybrid attention
Training RLAIF, then behavior-detection specialization (vendor term)
Latency Not published
Pricing $0.02/M input, output free · Span-01 Lite free on both
Weights Closed — no model repository on Hugging Face
Benchmark respanai/behavior-benchmark

Benchmarks

Behavior detection (Respan’s own suite)

System Core Multilingual All
Span-01 0.784 0.902 0.843
GPT-5.6-terra 0.756 0.917 0.837
deepseek-v4-flash 0.743 0.885 0.814
GPT-6-luna 0.728 0.901 0.815
Span-01 Lite 0.673 0.848 0.761
qwen3-235b 0.584 0.772 0.678
Sonnet 5 0.576 0.858 0.610
jev 0.574 0.856 0.715
laya 0.196 0.250 0.223

All figures are UnverifiedClaim. GPT-5.6-terra actually beats Span-01 on the multilingual split, and the header slide that markets “84.3 vs 81.5 vs 71.5” omits GPT-6 Sol, which scores higher in the vendor’s own domain table.

Per-domain production behavior scores

Behavior domain Span-01 Jev Sonnet 5 GPT-6 Sol
Jailbreak and prompt injection 0.779 0.752 0.709 0.878
Safety and refusals 0.803 0.751 0.785 0.911
Privacy and secrets 1.000 0.948 0.901 0.935
Hallucination and grounding 0.796 0.673 0.756 0.903
Agent and tool reliability 0.845 0.691 0.677 0.861
Task and instruction following 0.702 0.671 0.621 0.821
Response quality 0.771 0.691 0.722 0.956
Overall 0.806 0.716 0.719 0.885

This is the number to look at. Span-01 beats Jev and Sonnet 5, and a frontier model still sets the ceiling. The overall figures here (0.806) and the F1 figures in the table above (0.843) are different metrics and should not be mixed.

The benchmark has no ground truth — read this before trusting the gap

The dataset card is unusually candid, which is a point in Respan’s favor and a hard limit on the headline claim. Labels come from three methods:

method core multilingual
gpt-5.6-sol and claude-opus-5 agree 24,074 5,541
Tie-break: teachers disagreed, gpt-6-astra decided 2,421 —
Construction: set by how the trace was built — 2,851

The card says it directly: “These are model labels, not ground truth”, and the teachers “can’t be fairly evaluated on this benchmark”.

So “18% better than Jev” measures agreement with frontier teacher models on teacher labels, not intrinsic decision quality. A classifier post-trained to approximate those teachers has a structural advantage over a competitor that was not. The dataset is still genuinely useful — 1,990 traces, 34,885 trace–behavior pairs, 19 languages, documented evaluation protocol, per-row upstream licensing — but it is a vendor benchmark.

Two more limits: positive labels are 10.4% of rows and the protocol reports F1 on present without publishing recall or a confusion matrix, and the cost Pareto chart normalizes to $0.003/event while excluding a $299/month base fee.

Span-01 judged 11 decision models

Respan inverted its own benchmark and used Span-01 as the evaluation layer for 11 emerging decision models. Span-01 does not appear in the ranking — it produced the signal.

Model Accuracy Paired flip rate Injection ASR ECE
Jev 1.13.0 0.932 0.021 0.063 0.045
Tev1 4B 0.901 0.052 0.042 0.031
Kev 4B 0.833 0.060 0.078 0.070
Decider 2B v10 0.747 0.128 0.208 0.074
AlexWortega OpenJev 4B v5 0.744 0.116 0.089 0.118
Bosun v3.1 1.7B 0.680 0.208 0.323 0.089
Tiny-Jev 0.6B 0.646 0.128 0.224 0.267
Laya 0.542 0.298 0.396 0.092
MoJev 0.85B 0.501 0.150 0.307 0.129
Open-Jev DeBERTa v3 large 0.472 0.194 0.385 0.147
Von 1.2 0.422 0.242 0.297 0.071

Paired flip rate measures consistency under equivalent variants. Injection ASR is attack success rate. ECE is expected calibration error.

This is the first external evaluation of the Jev line that publishes calibration error, and the hosted Jev comes out with the second-lowest flip rate and a competitive ECE — the best result in that column among the larger models. Treat it as a third-party claim: Span-01 is judging competitors in a domain it was specialized for, and no prompt or version is published. Nine of these eleven models are catalogued on this site; we have not folded these numbers into their entries.

Fit / anti-fit

Fit when you need: a behaviour-detection layer over LLM traces, judge-style evaluation of agent output, prompt-injection or privacy screens, or a cheap replacement for an LLM-as-a-judge call on a narrow taxonomy — once the waitlist opens.

Anti-fit when you need: open weights, self-hosting, a documented API contract, a general decision model outside behaviour detection, calibrated Score semantics, or any decision you cannot waitlist for.

Limits

  • No public API contract. No endpoint, request schema or response shape is published — only the price table and a rate-limited playground (10 comparisons per IP per 10 minutes).
  • No latency data. “At classifier speed” is the entire argument and it carries no p50/p99 number. The 0–0ms frontmatter value is an unavailable-data sentinel.
  • No weights, no license, no repo. author=respanai returns zero models on Hugging Face. This is not a peer of the open catalog entries.
  • One decision type. No choice over arbitrary options, no score rubric. not_observable is a product idea worth copying, not a broad decision interface.
  • The judge is not audited. Respan publishes ECE for the 11 models it evaluated and none for Span-01 itself.
  • The vendor benchmark is the only benchmark. Teacher labels, 10.4% positives, F1 on present, and a cost Pareto that excludes the base fee.
  • A frontier model still wins. GPT-6 Sol scores 0.885 against Span-01’s 0.806. Span-01 replaces the mid-tier judge, not the frontier ceiling.

Sources