Lev
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity5.2/10
Open weights on HF + GitHub with detailed docs and a TypeSafe-compatible local server; no hosted SLA keeps maturity mid-tier.
- Capability6.8/10
S1Bench 68.9% macro (13 subsets), ECE 0.115, three decision types. Weaker than Jev on minimal-edit pairs and summeval-consistency.
- Adoption12/100
Day-one YC-backed launch with modest traction (6.6k views, 96 likes); no production use reported yet.
Vendor claims
- 68.9% macro accuracy on all 13 S1Bench subsets[Vendor claim — not independently verified]
- 69ms engine compute for a short request on H100[Vendor claim — not independently verified]
- 414–654ms end-to-end on Modal H100 vs 335–346ms Jev API[Vendor claim — not independently verified]
- Mean ECE 0.115 (vs Jev 0.091); better calibrated on 5 of 13 subsets[Vendor claim — not independently verified]
- Training: LoRA r=32, α=64, 3 epochs, 7.8 hours on one H100[Vendor claim — not independently verified]
Lev is a LoRA adapter on Qwen3.5-4B from Interfaze (YC P26), released under Apache-2.0. It speaks the TypeSafe /v1/systemone protocol: code written for the TypeSafe SDK works once you change the base_url.
It reads a state (text, ticket, email, or JSON) and a set of typed questions in a single forward pass — no autoregressive decode, zero output tokens. Question types: noul (yes/no → catalog boolean), choice, and score. Output is calibrated probabilities over exactly the options you supply.
What it is
- Open weights:
interfaze-ai/lev(Apache-2.0 adapter; Qwen base under its own license) - Repo + local server: github.com/Abhinavexists/lev
- API:
POST /v1/systemone— TypeSafe-compatible; official SDK works with a localbase_url - Install:
pip install "lev[serve] @ git+https://github.com/Abhinavexists/lev#subdirectory=packages/lev"
Quickstart:
import lev
model = lev.load("interfaze-ai/lev")
state = "Hi, I was charged twice for my order #4471 and I want a refund."
questions = {
"intent": {
"type": "choice",
"instructions": "What does the customer want?",
"criteria": {"refund": "wants money back", "cancel": "wants to cancel", "track": "wants to know where an order is", "other": "anything else"},
},
"urgent": {"type": "noul", "instructions": "Does this need a human within the hour?"},
}
result = model.system_one(state, questions)
print(result.answers["intent"].choice) # refund
print(result.answers["intent"].probabilities) # {'refund': 0.84, ...}
print(result.usage.output_tokens) # 0
S1Bench results
UnverifiedClaim (source: huggingface.co/interfaze-ai/lev):
| Subset | Task | Lev | Jev | Δ |
|---|---|---|---|---|
| vitaminc-dev | claim verification | 0.668 | 0.801 | −13.3 |
| massive-en-US | intent routing (18 scenarios) | 0.857 | 0.874 | −1.7 |
| massive-de-DE | intent routing (German) | 0.823 | 0.871 | −4.8 |
| boolq | yes/no reading comprehension | 0.827 | 0.893 | −6.6 |
| squad2 | answerability | 0.813 | 0.836 | −2.3 |
| paws | adversarial paraphrase | 0.776 | 0.900 | −12.4 |
| multinli | natural language inference | 0.890 | 0.836 | +5.4 |
| civil_comments | toxicity | 0.760 | 0.803 | −4.3 |
| aegis2 | safety moderation | 0.800 | 0.804 | −0.4 |
| helpsteer2 | helpfulness (5 levels) | 0.386 | 0.341 | +4.5 |
| summeval-relevance | summary relevance (5 levels) | 0.358 | 0.358 | 0.0 |
| summeval-consistency | summary faithfulness (5 levels) | 0.271 | 0.812 | −54.1 |
| pubmedqa | biomedical yes/no/maybe | 0.732 | 0.764 | −3.2 |
| Macro | 0.689 | 0.761 | −7.2 |
Lev and Jev ran through the same harness on all 3,880 S1Bench items. The harness reproduces TypeSafe’s published Jev figures within 0.8pp.
Where Jev is clearly ahead: minimal-edit pairs (paws −12.4pp, vitaminc −13.3pp) and summeval-consistency (−54.1pp), where Lev rates most fully faithful summaries one level low. The untuned Qwen backbone scores 0.826 on that subset — fine-tuning introduced the regression.
Where Lev leads or ties: multinli (+5.4pp) and helpsteer2 (+4.5pp), though the per-subset noise floor is 5–9pp.
How it works
- Label-token readout: each option gets a short code; the answer is read from the next-token logits over those codes. Choices are read in two option orders and averaged to cancel position bias.
- Candidate-path head: past ~68 options, a small learned head matches the state against each option’s text.
- Calibration: temperatures fitted after training, one per question type and readout mode.
Optimizations that mattered (from the card):
- One batched forward instead of prefill-and-fork: 169 → 69 ms
- Label-token readout up to the tokenizer’s limit: +51pp on MASSIVE’s 60 intents
- Skipping codes that split: banking77 0.818 → 0.980
Limits (read before you ship)
- Author caveat: not for production people-affecting decisions without your own eval
- Minimal edits and fine-grained ratings are weak — inputs that differ by one swapped word, and quality ratings over five levels, are where Lev is least accurate and can be confidently wrong
- Calibration is fitted on the training distribution; tasks unlike the training mix may be less well calibrated
- Questions are answered independently — encode a joint decision as one
choice, or ask in stages - English only
- Needs a GPU for real-time use (4B backbone takes seconds on CPU)
Training
UnverifiedClaim (source: huggingface.co/interfaze-ai/lev):
- Data: 200,000 examples from 29 sources built on 26 public HuggingFace datasets (topic, sentiment, emotion, intent, NLI, paraphrase, QA, toxicity, helpfulness)
- Contamination guard: refuses any source that resolves to one of the 13 S1Bench subsets
- Recipe: LoRA r=32, α=64 on q/k/v/o attention and MLP projections; 3 epochs, 18,750 steps, batch size 32, learning rate 5e-5, Qwen chat format, one H100, 7.8 hours
Why it is in the catalog
Public typed decision I/O, documented TypeSafe-compatible API, open weights, and a full S1Bench run against Jev. It sits next to Kev as an open / local point on the quadrant — with artifactKind: lora-adapter so the card does not pretend this is a frontier base model.
Not affiliated with or endorsed by TypeSafe AI.
