OpenJev 27B (openjev)
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity5.8/10
Open 27B weights under CC BY-NC 4.0 (commercial use needs a separate licence) with FP8, MLX and GGUF builds, each with measured accuracy. Own /v1/systemone helper over vLLM, plus official llama.cpp support (ggml-org GGUF). Detailed README and serving guide. Independent project, no SLA.
- Capability7/10
Choice (up to 52 options per pass), Noul and Score with fixed calibration settings; answers flip 2.3% under option shuffling, no ECE published. About 80 ms for short text and 227 ms for a 1.1k-token page on one H100 FP8 (author). 84.0% vs hosted Jev 85.4% on the same 10,000 questions, self-run.
- Adoption12/100
98 likes on the main repo and about 18,000 downloads across its five builds, plus 1,200 for ggml-org's GGUF. One of five models in llama.cpp's launch of /v1/systemone. No production use reported.
Vendor claims
- 10,000 text questions from 34 public sources, same questions for every model: OpenJev 84.0% vs hosted Jev 85.4%, its own base before tuning 80.4%, Nimble 9B 75.7% (6,922 fresh questions; gap to Jev 1.2 points on those)[Vendor claim — not independently verified]
- Agent steps: desktop next action from a screenshot 88.0% (base 76.5%), web next action on unseen sites 87.4% (68.5%); MiniWoB 39 of 100 tasks, tied with hosted Jev[Vendor claim — not independently verified]
- One H100, FP8: ~80 ms for a short text decision; first read of a 1.1k-token page 227 ms median (cached page, new question 118 ms); ~210 ms for a web step with ~23 options[Vendor claim — not independently verified]
- Answer changes in 2.3% of cases when options are shuffled (18.5% for the base before tuning)[Vendor claim — not independently verified]
Read this first
- Non-commercial weights. The weights are CC BY-NC 4.0. Commercial use needs a licence from the authors (the README points to a Loop AI support address). The helper and serving code are Apache-2.0.
- Not the same model as openjev. That page is Alex Wortega’s MIT release (4B and other sizes). This one is the
openjevorganisation’s 27B model, the “OpenJev” that llama.cpp supports.- Every number is the authors’ own run. The 10,000-question comparison with hosted Jev is careful (same questions and option order, failures counted), but it is not on the Decision Index or another public board.
OpenJev is an open decision model published on Hugging Face on 20 September 2026 by the openjev organisation, which describes itself as an independent project not affiliated with TypeSafe. It is a fine-tune of Qwen3.8-27B that reads text, JSON, web pages and screenshots and answers typed questions with a probability per option. On 2 October it was one of the five models in llama.cpp’s launch of /v1/systemone, where it is the only one that reads images.
How it works
The model gives each option a letter and reads the scores of exactly those letters at the first output position, one forward pass per question for up to 52 options. Fixed calibration settings turn the scores into probabilities. Fine-tuning trains that single position to carry the decision and to stay stable when options are reordered. It is a trained decision model, not a readout on a frozen LLM: the README compares it throughout with “the same base model before tuning”.
Specs
| Attribute | Value |
|---|---|
| Base | Qwen/Qwen3.8-27B (fine-tuned) |
| Question types | Choice (up to 52 options per pass; more in several passes), Noul, Score |
| Input | Text, JSON, DOM, one image per request; up to 16,384 tokens |
| Languages | English, German, French, Hindi, Chinese, Japanese |
| Serving | vLLM + the helper/shim.py decision API (POST /v1/systemone); llama.cpp (llama serve -hf ggml-org/OpenJev-GGUF, text and images) |
| Builds | bf16 (~54 GB), FP8 (~29 GB), MLX 8-bit / 4-bit (text only), GGUF Q4_K_M–Q8_0 (text only) |
| License | Weights CC BY-NC 4.0; helper / serving code Apache-2.0 |
| Launch |
Evidence (UnverifiedClaim)
From the model card:
| Test | OpenJev | Hosted Jev | Base before tuning |
|---|---|---|---|
| 10,000 text questions, 34 sources | 84.0% | 85.4% | 80.4% |
| Science / facts / claims (1,558) | 82.5% | 89.0% | 81.0% |
| Ethics / policy judgement (877) | 78.1% | 75.8% | 65.5% |
| Desktop next action from a screenshot (2,000) | 88.0% | — | 76.5% |
| MiniWoB, 100 task types | 39 | 39 | 38 |
| Option-shuffle flips (2,000) | 2.3% | — | 18.5% |
The quantized builds report their own accuracy: FP8 84.2% and MLX 8-bit 84.0% on the same 10,000 questions; GGUF Q4_K_M 82.8% vs 83.2% for bf16 on a separate 1,789-question set.
Fit / anti-fit
Fit: browser and desktop agents choosing the next element or action; text routing and policy checks where you want an open model close to Jev’s accuracy; research and other non-commercial use on one 80 GB GPU, or a 24 GB card with the Q4 GGUF.
Anti-fit: commercial products without a separate licence; CPU-only or small-GPU deployments (27B); knowledge-heavy questions, where it trails Jev by several points.
Why it is in the catalog
OpenJev fine-tunes its own weights for typed decisions, publishes them with a Jev-shaped API that returns a probability per option, and covers Choice, Noul and Score. It had no page until now. We had attributed llama.cpp’s “OpenJev” to the openjev page, and this page corrects that.
