vLLM Jev

Last updated:

RuntimeApache-2.0~94 ★

⚠️ Why not a model

Serving engine for existing Jev-style checkpoints (vLLM on Linux, MLX/MPS on Mac). No decision training — it selects a protocol and runs inference.

Technical specs

Base LLMs
ZefanCai/Open-Jev-2B, OpenJev-0.6B, Tiny-Jev, Laya checkpoints, Valen-Team/Valen-Preview-0923, vjev vision
Decision types
choice, boolean, score
Features
vllm, mlx, apple-silicon, multimodal, http-systemone, batching, kv-cache, laya-head, marker-scores
License
Apache-2.0

vLLM Jev serves compatible open decision checkpoints through vLLM on Linux NVIDIA, with an Apple Silicon preview via MLX or PyTorch MPS. You point it at a Hugging Face model ID; it prepares the checkpoint and exposes Choice, Noul and Score over HTTP. It does not train weights.

uv pip install .
vllm-jev serve ZefanCai/Open-Jev-2B
# → http://127.0.0.1:8795  POST /v1/systemone

On Mac (macOS 15+): source scripts/install_mac.sh then vllm-jev serve Valen-Team/Valen-Preview-0923 for text/image decisions.

Why it is not a model

No decision fine-tune lives in this repo. The runtime selects and verifies a native protocol for a supported checkpoint (scalar candidate branches, marker scores, or Laya’s trained decision head) and runs inference with scheduling, batching, compilation and KV cache from vLLM — or MLX/MPS on Apple Silicon. Quality comes from the checkpoint you serve (Valen, Open-Jev-2B, Laya, Tiny-Jev, …).

Supported surface (author docs)

  • Linux NVIDIA: native vLLM for listed text and multimodal checkpoints.
  • Apple Silicon preview: Open-Jev-2B, OpenJev-0.6B, Tiny-Jev, Laya (MPS), Valen text/image (MLX).
  • Multimodal: Valen and vjev vision models accept text and images on the same System One API.
  • Readouts: scalar candidate branches, marker scores, Laya decision head.

Author performance tables (A800 / M5) claim large latency cuts vs the checkpoints’ own HTTP servers — treat every multiplier and millisecond as UnverifiedClaim. Example from the README: on 500 Valen image questions, median 231.9 → 77.3 ms (3.0×) with 86.8% target-support accuracy on both paths.

Caveats

  • Checkpoint-dependent. Wrong architecture or unsupported ID will not magically become a decision model.
  • Not TypeSafe’s stack. Wire shape /v1/systemone is community compatibility.
  • Mac path is preview and often sequential (concurrency queues).
  • Amplifying posts (e.g. NFT_Chen 29 Sep) repeat author speed claims — still UnverifiedClaim.