Lev

Last updated:

Open69–654ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity5.2/10

    Open weights on HF + GitHub with detailed docs and a TypeSafe-compatible local server; no hosted SLA keeps maturity mid-tier.

  • Capability6.8/10

    S1Bench 68.9% macro (13 subsets), ECE 0.115, three decision types. Weaker than Jev on minimal-edit pairs and summeval-consistency.

  • Adoption12/100

    Day-one YC-backed launch with modest traction (6.6k views, 96 likes); no production use reported yet.

Vendor claims

Lev is a LoRA adapter on Qwen3.5-4B from Interfaze (YC P26), released under Apache-2.0. It speaks the TypeSafe /v1/systemone protocol: code written for the TypeSafe SDK works once you change the base_url.

It reads a state (text, ticket, email, or JSON) and a set of typed questions in a single forward pass — no autoregressive decode, zero output tokens. Question types: noul (yes/no → catalog boolean), choice, and score. Output is calibrated probabilities over exactly the options you supply.

What it is

  • Open weights: interfaze-ai/lev (Apache-2.0 adapter; Qwen base under its own license)
  • Repo + local server: github.com/Abhinavexists/lev
  • API: POST /v1/systemone — TypeSafe-compatible; official SDK works with a local base_url
  • Install: pip install "lev[serve] @ git+https://github.com/Abhinavexists/lev#subdirectory=packages/lev"

Quickstart:

import lev

model = lev.load("interfaze-ai/lev")
state = "Hi, I was charged twice for my order #4471 and I want a refund."
questions = {
    "intent": {
        "type": "choice",
        "instructions": "What does the customer want?",
        "criteria": {"refund": "wants money back", "cancel": "wants to cancel", "track": "wants to know where an order is", "other": "anything else"},
    },
    "urgent": {"type": "noul", "instructions": "Does this need a human within the hour?"},
}
result = model.system_one(state, questions)
print(result.answers["intent"].choice)  # refund
print(result.answers["intent"].probabilities)  # {'refund': 0.84, ...}
print(result.usage.output_tokens)  # 0

S1Bench results

UnverifiedClaim (source: huggingface.co/interfaze-ai/lev):

Subset Task Lev Jev Δ
vitaminc-dev claim verification 0.668 0.801 −13.3
massive-en-US intent routing (18 scenarios) 0.857 0.874 −1.7
massive-de-DE intent routing (German) 0.823 0.871 −4.8
boolq yes/no reading comprehension 0.827 0.893 −6.6
squad2 answerability 0.813 0.836 −2.3
paws adversarial paraphrase 0.776 0.900 −12.4
multinli natural language inference 0.890 0.836 +5.4
civil_comments toxicity 0.760 0.803 −4.3
aegis2 safety moderation 0.800 0.804 −0.4
helpsteer2 helpfulness (5 levels) 0.386 0.341 +4.5
summeval-relevance summary relevance (5 levels) 0.358 0.358 0.0
summeval-consistency summary faithfulness (5 levels) 0.271 0.812 −54.1
pubmedqa biomedical yes/no/maybe 0.732 0.764 −3.2
Macro 0.689 0.761 −7.2

Lev and Jev ran through the same harness on all 3,880 S1Bench items. The harness reproduces TypeSafe’s published Jev figures within 0.8pp.

Where Jev is clearly ahead: minimal-edit pairs (paws −12.4pp, vitaminc −13.3pp) and summeval-consistency (−54.1pp), where Lev rates most fully faithful summaries one level low. The untuned Qwen backbone scores 0.826 on that subset — fine-tuning introduced the regression.

Where Lev leads or ties: multinli (+5.4pp) and helpsteer2 (+4.5pp), though the per-subset noise floor is 5–9pp.

How it works

  1. Label-token readout: each option gets a short code; the answer is read from the next-token logits over those codes. Choices are read in two option orders and averaged to cancel position bias.
  2. Candidate-path head: past ~68 options, a small learned head matches the state against each option’s text.
  3. Calibration: temperatures fitted after training, one per question type and readout mode.

Optimizations that mattered (from the card):

  • One batched forward instead of prefill-and-fork: 169 → 69 ms
  • Label-token readout up to the tokenizer’s limit: +51pp on MASSIVE’s 60 intents
  • Skipping codes that split: banking77 0.818 → 0.980

Limits (read before you ship)

  • Author caveat: not for production people-affecting decisions without your own eval
  • Minimal edits and fine-grained ratings are weak — inputs that differ by one swapped word, and quality ratings over five levels, are where Lev is least accurate and can be confidently wrong
  • Calibration is fitted on the training distribution; tasks unlike the training mix may be less well calibrated
  • Questions are answered independently — encode a joint decision as one choice, or ask in stages
  • English only
  • Needs a GPU for real-time use (4B backbone takes seconds on CPU)

Training

UnverifiedClaim (source: huggingface.co/interfaze-ai/lev):

  • Data: 200,000 examples from 29 sources built on 26 public HuggingFace datasets (topic, sentiment, emotion, intent, NLI, paraphrase, QA, toxicity, helpfulness)
  • Contamination guard: refuses any source that resolves to one of the 13 S1Bench subsets
  • Recipe: LoRA r=32, α=64 on q/k/v/o attention and MLP projections; 3 epochs, 18,750 steps, batch size 32, learning rate 5e-5, Qwen chat format, one H100, 7.8 hours

Why it is in the catalog

Public typed decision I/O, documented TypeSafe-compatible API, open weights, and a full S1Bench run against Jev. It sits next to Kev as an open / local point on the quadrant — with artifactKind: lora-adapter so the card does not pretend this is a frontier base model.

Not affiliated with or endorsed by TypeSafe AI.