TypeLLM
⚠️ Why not a model
Constrained decoding layer over existing autoregressive LLMs — does not change architecture or weights.
Technical specs
- Base LLMs
- Qwen3.8-27B, Qwen3.5-0.8B, Qwen3.5-4B, Qwen3.5-9B, MiniCPM5-1B, Ling-mini-2.0, Ring-mini-2.0
- Decision types
- choice, boolean, integer, number, string
- Features
- sglang, thinking-mode, image-input, dag-execution, permutation-averaging, kv-prefix-reuse, hosted-api, playground
- License
- Apache-2.0
JevBench score
JevBench: 195/231 (no thinking), 228/231 (with thinking)[Author-reported — not independently verified]Links
TypeLLM brings type-safe generation to existing autoregressive LLMs without changing their architecture or weights. Built on SGLang, it forces models to produce schema-guaranteed outputs through JSON Schema.
How it works
TypeLLM uses constrained decoding to ensure outputs match a defined schema:
- You provide context + typed questions (JSON Schema)
- TypeLLM constrains the model’s generation to valid values
- Returns typed results:
string,integer,number,boolean,enum
Unlike Jev (which is non-autoregressive), TypeLLM keeps the base LLM’s autoregressive generation but constrains the output space.
Why it’s not a model
TypeLLM explicitly states: “brings type-safe generation to existing autoregressive LLMs without changing their architecture or weights.”
It’s a runtime layer that:
- Does not train any new weights
- Does not modify the base model
- Uses inference-time constrained decoding
Key features
- Thinking mode: Optional reasoning before the constrained answer
- Image input: Vision-language model support (tested with Qwen3.8-27B)
- DAG execution: Dependency-aware field execution with
depends_on - Permutation averaging: Reduce option-order bias on enum questions
- KV prefix reuse: Shared context caching for efficiency
Hosted service (29 Sep 2026)
TypeLLM now also runs as a hosted service. “Not another Jev,” the team posted on 29 September at 11:00 BRT.
- Playground: open to everyone, with $5 of credit for new users.
- API: “rolling out to early-access users in the coming days”. Keys are created in the dashboard with an access code. Endpoint
POST https://api.typellm.ai/v1/generate; Python clientpip install -U typellm. The self-hosted SGLang path is still documented. - Output types: string, integer, number, boolean and enum, plus image input, dependent fields (
depends_on) and per-question thinking.
| Pricing (typellm.ai, 30 Sep 2026) | Price |
|---|---|
| Input tokens (context, questions as JSON, images) | $0.05 / M |
Thinking tokens (only for questions with "thinking": true) |
$0.50 / M |
| Typed outputs | Free |
| New users | $5 credit |
The pricing page doesn’t say which base model the hosted API serves. The benchmark on the site uses Qwen3.8-27B. An enterprise tier (private deployment, dedicated capacity) is available on request.
It stays in (NOT)Models: a hosted API doesn’t change the fact that TypeLLM is constrained decoding over existing LLMs, with no decision weights of its own. See news.
JevBench results
TypeLLM reports 195/231 (84.42%) without thinking and 228/231 (98.70%) with thinking enabled on the 231 public JevBench tasks, using Qwen3.8-27B. The same table on typellm.ai lists Jev 1.13.0 at 86.58%, Open-Jev 27B v1.1 at 85.28%, and two models without type-safe output: GPT-5.6 Luna (none) at 89.18% and GPT-6 Astra (low) at 100%. These are author-reported figures (UnverifiedClaim).
Comparison with Jev
The TypeLLM README includes a direct comparison:
| Feature | TypeLLM | Jev |
|---|---|---|
| Rubric scoring | Numeric enum (no dedicated Score API) | Score |
| Thinking mode | ✓ | — |
| Image input | ✓ | Not documented |
| DAG execution | ✓ | — |
