Liquid AI launches d1, a hosted decision model it says beats Jev on HF's Decision Index
Liquid AI, the lab behind the LFM models, announced d1 on 29 September 2026 at 15:35 BRT: “our first decision model” and “the first model to outperform Jev on @huggingface’s Decision Index”. The post says d1 wins on multilingual evals, is more robust against prompt injection and handles longer inputs better. It is live on the Liquid API (console.liquid.ai) and “available on OpenRouter soon”. By the next morning the post had about 1.8k likes and 680k views.
What shipped. A hosted endpoint, POST https://api.liquid.ai/decisions/v1/systemone, with the model id d1:free in the docs. It takes the same three question types as Jev (Noul, Choice, Score) and returns probabilities with output_tokens: 0. Liquid’s quickstart uses TypeSafe’s own SDKs (typesafe-sdk on PyPI, @typesafe-ai/sdk on npm) pointed at https://api.liquid.ai, so code written for Jev moves over by changing the base URL and key. A migration guide covers replacing LLM classification, routing and scoring calls, and the cookbook has a road-decider demo.
What is not public. No weights (nothing named d1 in Liquid’s Hugging Face org), no parameter count, base model or training method, no latency numbers and no pricing beyond the current free tier. GIGAZINE reports pricing will come later, and that Liquid replied “for sure” when asked whether d1 will run locally.
The benchmark (UnverifiedClaim). The launch chart is captioned “Internal reproduction of Decision Index 0.2.1 by Hugging Face”. d1 vs Jev 1.13: Arts 45.5 vs 37.7, Language 67.6 vs 62.0, Retrieval 60.7 vs 55.4, Tools 74.1 vs 75.1, Knowledge 43.3 vs 51.3; Index 58.9 vs 57.9. The Jev numbers match the official Decision Index Space exactly. d1 itself is not on that index, so the one-point lead rests on Liquid’s own run. The chart also shows d1 eight points behind Jev on knowledge and reasoning.
How we list it. d1 is a decision model with a public, documented typed API that returns probabilities, from a lab that presents it as its own model. It goes into the catalog and onto the quadrant at 4.9 maturity / 6.5 capability / 23 adoption. That puts it below Jev (6.0 / 8.2 / 58) for now: calibration and latency are undocumented, and the only benchmark is self-run. That is the difference from the OpenAI Decisions API, announced the same day, which has no docs and is described as a way of using GPT-6 Luna, so it stays under runtimes.
