Valen
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity5.8/10
Apache-2.0 weights and training code, HF Space demo, technical notes, and a native path in vLLM Jev. Preview checkpoint still needs the frozen Qwen3.5-2B base; no hosted SLA.
- Capability7.3/10
First-class Choice / Noul / Score over text and vision without answer-token generation. Author claims RLCD stage and ~122 ms/step on Sokoban; evals are game/VQA-heavy and partly outcome-conditioned.
- Adoption43/100
About 479 GitHub stars within a week of the 23 Sep public repo, HF Preview downloads, and serving demos via vLLM Jev. No public production case studies yet.
Vendor claims
- Valen-Preview-0923: 87.60% accuracy on 500 single-step Sokoban action questions; 38/100 complete games solved (≤200 moves) on a model-outcome-conditioned easy/medium set[Vendor claim — not independently verified]
- Four parallel Sokoban trajectories: 7–10 decisions each, averaging 122–128 ms per step; all four finish within 1.24 seconds[Vendor claim — not independently verified]
- On one Sokoban level with a 10-step cap: Valen solves in 9 decisions / 1.13 s cumulative latency vs Qwen3.8-27B-FP8 at 198.05 s thinking (and fail in no-thinking)[Vendor claim — not independently verified]
- vLLM Jev path on 500 Valen image questions: median latency 231.9 → 77.3 ms (3.0×) with 86.8% target-support accuracy on both paths[Vendor claim — not independently verified]
Valen brings vision to System One decision-making. You give it text, images or video plus typed questions (choice, noul, score); a shared decision head returns probabilities over your candidates. It does not generate answer tokens.
The public preview Valen-Team/Valen-Preview-0923 is a 2B Sokoban-trained checkpoint (same weights as Valen-Sokoban-RLCD-2B): General 100k SFT decision head, then experimental RLCD on 30k Sokoban records. Backbone is frozen Qwen/Qwen3.5-2B (pinned revision) — both downloads are required. Code, data recipes and SFT/RLCD trainers: github.com/Liuziyu77/Valen (Apache-2.0, ~479★ at radar check). Demo: HF Space.
Independent of TypeSafe AI. Tagline on the card: “System One Model, now with vision.”
Specs
| Attribute | Value |
|---|---|
| Author | Valen-Team / Liuziyu77 (independent) |
| Flagship weights | Valen-Team/Valen-Preview-0923 (+ frozen Qwen/Qwen3.5-2B) |
| Modalities | Text, image, video |
| Decision types | Choice, Noul, Score (catalog choice / boolean / score) |
| License | Apache-2.0 |
| Released | (HF Preview + GH) |
| Serving | Project inference scripts; also vllm-jev serve Valen-Team/Valen-Preview-0923 (vLLM Jev) |
Author benchmarks
All numbers below are author-reported → UnverifiedClaim. The Preview card itself flags that the 100-game Sokoban set retained successful 2B RLCD cases before filling — it is outcome-conditioned, not an unbiased full benchmark.
| Signal | Value | Note |
|---|---|---|
| Sokoban single-step (500) | 87.60% | Action questions |
| Sokoban full games (100) | 38/100 | ≤200 moves; easy/medium mix |
| Per-step latency (demo) | ~122–128 ms | Four parallel games |
| vLLM Jev image median | 77.3 ms | vs 231.9 ms baseline path; 500 questions |
General-family panels on the repo (5k VQA-style questions) use separately trained checkpoints — not this Sokoban Preview.
Limits
- Preview is Sokoban-specialised. Do not read game scores as a general multimodal decision leaderboard.
- Needs the frozen base. Checkpoint alone is not enough to run.
- RLCD is labelled experimental in the project badges.
- Custom stack. Prefer the Valen repo or vLLM Jev; not a drop-in chat
transformersgenerate path. - Not TypeSafe. Wire compatibility via community servers is not affiliation.
Fit / anti-fit
Fit when you need: open multimodal typed decisions (especially visual game/UI state), a research stack with SFT+RLCD recipes, or a checkpoint already wired into vLLM Jev.
Anti-fit when you need: a general text-only business router with clean typed-decisions zero-shot numbers, a hosted SLA, or a single-file AutoModel.
Why it is in the catalog
Trained decision weights (SFT + RLCD stage) return Choice/Noul/Score over multimodal state without generating answer text. Public weights, docs and a Space meet the inclusion bar. It sits near Jev-Omni and openjev on the multimodal open side, with a stronger game/vision focus.
