JevEmbed
Quadrant scores
See the full quadrantScored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.
- Maturity5.8/10
Public GitHub (47★), multiple Apache-2.0 Hugging Face checkpoints, dataset, CLI and optional HTTP server. Academic project without a hosted SLA; recipes still moving across 0.6B/KaLM/4B releases.
- Capability5.3/10
First-class Choice / Score / Noul via embedding similarity with published temperatures/slopes. Strong author accuracy on the in-domain JevEmbed-Data test; no independent p50/p99 or external suite entry yet.
- Adoption21/100
Repository at 47 GitHub stars with three public fine-tunes (KaLM, Qwen3-Embedding-0.6B, Qwen3-Embedding-4B on 28 Sep) plus a 1.6M-question dataset. No public production case studies.
Vendor claims
- JevEmbed-Qwen3-Embedding-4B: overall hard-label accuracy 85.86% on JevEmbed-Data test (64,110 scored); Choice 90.31%, Score level 73.41%, Noul 95.88% vs base embedding 36.29% overall[Vendor claim — not independently verified]
- One-epoch LoRA (r64) on 1,601,157 training questions; final checkpoint step 3,127; embeddings 2,560-d with last-token pooling[Vendor claim — not independently verified]
- Framework exposes Choice, Score and Noul through a Python API, CLI and optional FastAPI server with Jev-shaped routes[Vendor claim — not independently verified]
JevEmbed turns embedding models into typed decision engines. You pick an embedder (including HIT-TMG’s own fine-tunes), ask Choice / Score / Noul questions over a state, and get probabilities — not chat. It is independent of TypeSafe AI.
The project is both a framework (any compatible embedding model, local or HTTP) and a family of trained weights. Fine-tuned checkpoints are LoRA adapters merged into Sentence Transformers models; adapters also ship under lora/. Flagship as of 28 September 2026: HIT-TMG/JevEmbed-Qwen3-Embedding-4B. Earlier: 0.6B, KaLM-Embedding-V2.5. Training data: HIT-TMG/JevEmbed-Data.
Code: github.com/HITsz-TMG/JevEmbed (47★ at radar check). Org: HIT-TMG / Harbin Institute of Technology (Shenzhen) TMG. First public framework activity around ; KaLM weights ; Qwen3-Embedding-4B .
Specs
| Attribute | Value |
|---|---|
| Author | HIT-TMG (HITsz-TMG) |
| Flagship weights | HIT-TMG/JevEmbed-Qwen3-Embedding-4B (merged LoRA on Qwen3-Embedding-4B) |
| Other public weights | Qwen3-Embedding-0.6B, KaLM-Embedding-V2.5 |
| Decision types | Choice, Score, Noul (mapped to catalog choice / score / boolean) |
| Embedding dim (4B) | 2,560 (last-token pooling, normalized) |
| License | Apache-2.0 on published weight cards; training-data source licenses vary (see dataset report) |
| Runtime | jevembed Python package — local Sentence Transformers, remote /v1/embeddings, optional FastAPI |
| Latency | Not published → catalog sentinel 0 / 0 |
Author benchmarks (4B)
From the 4B model card. Author-reported → UnverifiedClaim. Evaluation is on the project’s own JevEmbed-Data test split (in-domain relative to training), BF16, identical prompts/scoring for base vs LoRA.
| Metric | Base Qwen3-Embedding-4B | Final LoRA | Δ |
|---|---|---|---|
| Overall hard-label acc (64,110) | 36.29% | 85.86% | +49.57 pp |
| Choice acc (17,487) | 38.11% | 90.31% | +52.20 pp |
| Score level acc (24,260) | 30.55% | 73.41% | +42.86 pp |
| Noul binary acc (22,363) | 41.09% | 95.88% | +54.78 pp |
Do not paste 85.86% next to zero-shot typed-decisions or JevBench figures from other cards — different data and protocol.
Limits
- In-domain headline. The strong 4B numbers are on JevEmbed-Data, the same family used for training.
- Embedding geometry, not RLCD. Temperatures/slopes are config knobs; there is no published RLCD-style calibration story.
- Dual nature. Pointing the framework at a frozen generic embedder is closer to a runtime; this catalog entry is for the trained JevEmbed checkpoints.
- Data provenance. Dataset processing report notes mixed source licenses — read before commercial redistribution of derivatives.
- Not TypeSafe. Optional Jev-shaped HTTP routes are community wire compatibility, not an affiliation.
Fit / anti-fit
Fit when you need: local embedding-based typed decisions, a research stack that already speaks Choice/Score/Noul, or a fine-tuned Qwen3/KaLM embedder for those tasks.
Anti-fit when you need: a generative LLM, a hosted TypeSafe SLA, or a zero-shot number on LocalLLaMA/typed-decisions / JevBench without a new run.
Why it is in the catalog
The public LoRA merges are trained decision weights for Choice / Score / Noul, not a prompt wrapper on a frozen chat model. Code, dataset and multiple checkpoints are open. It sits next to encoder-style peers such as Laya, GLiNER2.5-Decide and Julia-1, with similarity scoring instead of a dedicated classification head.
