TypeLLM

Last updated:

RuntimeApache-2.0~885 ★

⚠️ Why not a model

Constrained decoding layer over existing autoregressive LLMs — does not change architecture or weights.

Technical specs

Base LLMs
Qwen3.8-27B, Qwen3.5-0.8B, Qwen3.5-4B, Qwen3.5-9B, MiniCPM5-1B, Ling-mini-2.0, Ring-mini-2.0
Decision types
choice, boolean, integer, number, string
Features
sglang, thinking-mode, image-input, dag-execution, permutation-averaging, kv-prefix-reuse, hosted-api, playground
License
Apache-2.0

JevBench score

JevBench: 195/231 (no thinking), 228/231 (with thinking)[Author-reported — not independently verified]

TypeLLM brings type-safe generation to existing autoregressive LLMs without changing their architecture or weights. Built on SGLang, it forces models to produce schema-guaranteed outputs through JSON Schema.

How it works

TypeLLM uses constrained decoding to ensure outputs match a defined schema:

  1. You provide context + typed questions (JSON Schema)
  2. TypeLLM constrains the model’s generation to valid values
  3. Returns typed results: string, integer, number, boolean, enum

Unlike Jev (which is non-autoregressive), TypeLLM keeps the base LLM’s autoregressive generation but constrains the output space.

Why it’s not a model

TypeLLM explicitly states: “brings type-safe generation to existing autoregressive LLMs without changing their architecture or weights.”

It’s a runtime layer that:

  • Does not train any new weights
  • Does not modify the base model
  • Uses inference-time constrained decoding

Key features

  • Thinking mode: Optional reasoning before the constrained answer
  • Image input: Vision-language model support (tested with Qwen3.8-27B)
  • DAG execution: Dependency-aware field execution with depends_on
  • Permutation averaging: Reduce option-order bias on enum questions
  • KV prefix reuse: Shared context caching for efficiency

Hosted service (29 Sep 2026)

TypeLLM now also runs as a hosted service. “Not another Jev,” the team posted on 29 September at 11:00 BRT.

  • Playground: open to everyone, with $5 of credit for new users.
  • API: “rolling out to early-access users in the coming days”. Keys are created in the dashboard with an access code. Endpoint POST https://api.typellm.ai/v1/generate; Python client pip install -U typellm. The self-hosted SGLang path is still documented.
  • Output types: string, integer, number, boolean and enum, plus image input, dependent fields (depends_on) and per-question thinking.
Pricing (typellm.ai, 30 Sep 2026) Price
Input tokens (context, questions as JSON, images) $0.05 / M
Thinking tokens (only for questions with "thinking": true) $0.50 / M
Typed outputs Free
New users $5 credit

The pricing page doesn’t say which base model the hosted API serves. The benchmark on the site uses Qwen3.8-27B. An enterprise tier (private deployment, dedicated capacity) is available on request.

It stays in (NOT)Models: a hosted API doesn’t change the fact that TypeLLM is constrained decoding over existing LLMs, with no decision weights of its own. See news.

JevBench results

TypeLLM reports 195/231 (84.42%) without thinking and 228/231 (98.70%) with thinking enabled on the 231 public JevBench tasks, using Qwen3.8-27B. The same table on typellm.ai lists Jev 1.13.0 at 86.58%, Open-Jev 27B v1.1 at 85.28%, and two models without type-safe output: GPT-5.6 Luna (none) at 89.18% and GPT-6 Astra (low) at 100%. These are author-reported figures (UnverifiedClaim).

Comparison with Jev

The TypeLLM README includes a direct comparison:

Feature TypeLLM Jev
Rubric scoring Numeric enum (no dedicated Score API) Score
Thinking mode ✓ —
Image input ✓ Not documented
DAG execution ✓ —