JevEmbed

Last updated:

Open—$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity5.8/10

    Public GitHub (47★), multiple Apache-2.0 Hugging Face checkpoints, dataset, CLI and optional HTTP server. Academic project without a hosted SLA; recipes still moving across 0.6B/KaLM/4B releases.

  • Capability5.3/10

    First-class Choice / Score / Noul via embedding similarity with published temperatures/slopes. Strong author accuracy on the in-domain JevEmbed-Data test; no independent p50/p99 or external suite entry yet.

  • Adoption21/100

    Repository at 47 GitHub stars with three public fine-tunes (KaLM, Qwen3-Embedding-0.6B, Qwen3-Embedding-4B on 28 Sep) plus a 1.6M-question dataset. No public production case studies.

Vendor claims

JevEmbed turns embedding models into typed decision engines. You pick an embedder (including HIT-TMG’s own fine-tunes), ask Choice / Score / Noul questions over a state, and get probabilities — not chat. It is independent of TypeSafe AI.

The project is both a framework (any compatible embedding model, local or HTTP) and a family of trained weights. Fine-tuned checkpoints are LoRA adapters merged into Sentence Transformers models; adapters also ship under lora/. Flagship as of 28 September 2026: HIT-TMG/JevEmbed-Qwen3-Embedding-4B. Earlier: 0.6B, KaLM-Embedding-V2.5. Training data: HIT-TMG/JevEmbed-Data.

Code: github.com/HITsz-TMG/JevEmbed (47★ at radar check). Org: HIT-TMG / Harbin Institute of Technology (Shenzhen) TMG. First public framework activity around ; KaLM weights ; Qwen3-Embedding-4B .

Specs

Attribute Value
Author HIT-TMG (HITsz-TMG)
Flagship weights HIT-TMG/JevEmbed-Qwen3-Embedding-4B (merged LoRA on Qwen3-Embedding-4B)
Other public weights Qwen3-Embedding-0.6B, KaLM-Embedding-V2.5
Decision types Choice, Score, Noul (mapped to catalog choice / score / boolean)
Embedding dim (4B) 2,560 (last-token pooling, normalized)
License Apache-2.0 on published weight cards; training-data source licenses vary (see dataset report)
Runtime jevembed Python package — local Sentence Transformers, remote /v1/embeddings, optional FastAPI
Latency Not published → catalog sentinel 0 / 0

Author benchmarks (4B)

From the 4B model card. Author-reported → UnverifiedClaim. Evaluation is on the project’s own JevEmbed-Data test split (in-domain relative to training), BF16, identical prompts/scoring for base vs LoRA.

Metric Base Qwen3-Embedding-4B Final LoRA Δ
Overall hard-label acc (64,110) 36.29% 85.86% +49.57 pp
Choice acc (17,487) 38.11% 90.31% +52.20 pp
Score level acc (24,260) 30.55% 73.41% +42.86 pp
Noul binary acc (22,363) 41.09% 95.88% +54.78 pp

Do not paste 85.86% next to zero-shot typed-decisions or JevBench figures from other cards — different data and protocol.

Limits

  • In-domain headline. The strong 4B numbers are on JevEmbed-Data, the same family used for training.
  • Embedding geometry, not RLCD. Temperatures/slopes are config knobs; there is no published RLCD-style calibration story.
  • Dual nature. Pointing the framework at a frozen generic embedder is closer to a runtime; this catalog entry is for the trained JevEmbed checkpoints.
  • Data provenance. Dataset processing report notes mixed source licenses — read before commercial redistribution of derivatives.
  • Not TypeSafe. Optional Jev-shaped HTTP routes are community wire compatibility, not an affiliation.

Fit / anti-fit

Fit when you need: local embedding-based typed decisions, a research stack that already speaks Choice/Score/Noul, or a fine-tuned Qwen3/KaLM embedder for those tasks.

Anti-fit when you need: a generative LLM, a hosted TypeSafe SLA, or a zero-shot number on LocalLLaMA/typed-decisions / JevBench without a new run.

Why it is in the catalog

The public LoRA merges are trained decision weights for Choice / Score / Noul, not a prompt wrapper on a frozen chat model. Code, dataset and multiple checkpoints are open. It sits next to encoder-style peers such as Laya, GLiNER2.5-Decide and Julia-1, with similarity scoring instead of a dedicated classification head.