OpenJev 27B (openjev)

Last updated:

Open80–227ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity5.8/10

    Open 27B weights under CC BY-NC 4.0 (commercial use needs a separate licence) with FP8, MLX and GGUF builds, each with measured accuracy. Own /v1/systemone helper over vLLM, plus official llama.cpp support (ggml-org GGUF). Detailed README and serving guide. Independent project, no SLA.

  • Capability7/10

    Choice (up to 52 options per pass), Noul and Score with fixed calibration settings; answers flip 2.3% under option shuffling, no ECE published. About 80 ms for short text and 227 ms for a 1.1k-token page on one H100 FP8 (author). 84.0% vs hosted Jev 85.4% on the same 10,000 questions, self-run.

  • Adoption12/100

    98 likes on the main repo and about 18,000 downloads across its five builds, plus 1,200 for ggml-org's GGUF. One of five models in llama.cpp's launch of /v1/systemone. No production use reported.

Vendor claims

Read this first

  1. Non-commercial weights. The weights are CC BY-NC 4.0. Commercial use needs a licence from the authors (the README points to a Loop AI support address). The helper and serving code are Apache-2.0.
  2. Not the same model as openjev. That page is Alex Wortega’s MIT release (4B and other sizes). This one is the openjev organisation’s 27B model, the “OpenJev” that llama.cpp supports.
  3. Every number is the authors’ own run. The 10,000-question comparison with hosted Jev is careful (same questions and option order, failures counted), but it is not on the Decision Index or another public board.

OpenJev is an open decision model published on Hugging Face on 20 September 2026 by the openjev organisation, which describes itself as an independent project not affiliated with TypeSafe. It is a fine-tune of Qwen3.8-27B that reads text, JSON, web pages and screenshots and answers typed questions with a probability per option. On 2 October it was one of the five models in llama.cpp’s launch of /v1/systemone, where it is the only one that reads images.

How it works

The model gives each option a letter and reads the scores of exactly those letters at the first output position, one forward pass per question for up to 52 options. Fixed calibration settings turn the scores into probabilities. Fine-tuning trains that single position to carry the decision and to stay stable when options are reordered. It is a trained decision model, not a readout on a frozen LLM: the README compares it throughout with “the same base model before tuning”.

Specs

Attribute Value
Base Qwen/Qwen3.8-27B (fine-tuned)
Question types Choice (up to 52 options per pass; more in several passes), Noul, Score
Input Text, JSON, DOM, one image per request; up to 16,384 tokens
Languages English, German, French, Hindi, Chinese, Japanese
Serving vLLM + the helper/shim.py decision API (POST /v1/systemone); llama.cpp (llama serve -hf ggml-org/OpenJev-GGUF, text and images)
Builds bf16 (~54 GB), FP8 (~29 GB), MLX 8-bit / 4-bit (text only), GGUF Q4_K_M–Q8_0 (text only)
License Weights CC BY-NC 4.0; helper / serving code Apache-2.0
Launch

Evidence (UnverifiedClaim)

From the model card:

Test OpenJev Hosted Jev Base before tuning
10,000 text questions, 34 sources 84.0% 85.4% 80.4%
Science / facts / claims (1,558) 82.5% 89.0% 81.0%
Ethics / policy judgement (877) 78.1% 75.8% 65.5%
Desktop next action from a screenshot (2,000) 88.0% — 76.5%
MiniWoB, 100 task types 39 39 38
Option-shuffle flips (2,000) 2.3% — 18.5%

The quantized builds report their own accuracy: FP8 84.2% and MLX 8-bit 84.0% on the same 10,000 questions; GGUF Q4_K_M 82.8% vs 83.2% for bf16 on a separate 1,789-question set.

Fit / anti-fit

Fit: browser and desktop agents choosing the next element or action; text routing and policy checks where you want an open model close to Jev’s accuracy; research and other non-commercial use on one 80 GB GPU, or a 24 GB card with the Q4 GGUF.

Anti-fit: commercial products without a separate licence; CPU-only or small-GPU deployments (27B); knowledge-heavy questions, where it trails Jev by several points.

Why it is in the catalog

OpenJev fine-tunes its own weights for typed decisions, publishes them with a Jev-shaped API that returns a probability per option, and covers Choice, Noul and Score. It had no page until now. We had attributed llama.cpp’s “OpenJev” to the openjev page, and this page corrects that.