Julia-1

Last updated:

Open33–295ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity4.8/10

    Apache-2.0 weights, a detailed card with provenance and metrics files, and a typed Python API. Since 2 Oct it runs in llama.cpp's /v1/systemone server (official ggml-org GGUF), so Jev clients can point at it. No PyPI package; hosted API announced but not live. No SLA.

  • Capability6.5/10

    First-class choice / score / noul at 144M params on CPU. Author evals are unusually well documented but self-run; Jev numbers are supplied references, not a new run. No calibration step (calibration: null). Weak on long label lists.

  • Adoption14/100

    Launch post on 26 Sep (161 likes, ~2.7k views a few hours in), 13 HF likes and 0 downloads at check time. No integrations or production use reported yet.

Vendor claims

Julia-1 is a small open decision model from Supersonic Labs, a Brazilian lab. You give it a state, a question and 2 to 20 candidate answers; it scores the candidates and returns one pick plus a probability per option. No text generation, no GPU required.

It is a fine-tune of JHU CLSP’s mmBERT-small, a multilingual ModernBERT encoder, with a decision head on top. The lab says so plainly on its launch page, including the line “Julia 1 is not a fine-tuned Qwen model”. It is not TypeSafe and does not speak the /v1/systemone contract.

Weights: Hugging Face SupersonicLabs/Julia-1 (Apache-2.0). Launch page: supersoniclabs.ia.br/julia-1. Announcement: @supersonicai (26 Sep 2026, 16:00 BRT).

Specs

Attribute Value
Author Supersonic Labs (Brazil)
Base model jhu-clsp/mmBERT-small (multilingual ModernBERT encoder)
Parameters 144.3M (HF API: 144,292,870, FP32)
Weights 550.5 MiB model.safetensors, FP32
License Apache-2.0 (weights and inference code; training pipeline not released)
HF repo created
Public launch
Status Open weights (Hugging Face)
Decision types choice (2–20 options), score (2–20 ordered levels), noul (yes/no)
Context Evaluated at 1,024 tokens (context + question + options); runtime now defaults to 8,192, only smoke-tested at that length
Runtime Python 3.11+, PyTorch on CPU (CUDA optional); install from the HF snapshot, no PyPI package
Other builds Julia-1-ONNX with a WebGPU adapter (created 25 Sep 2026)

How it works

Each question becomes a row of state + question + options. The encoder reads the row and a two-layer decision head scores every option; a softmax turns the scores into probabilities. Questions in one request are scored independently in a batch, so they don’t see each other.

The typed API lives in julia/typed.py:

  • choice takes a mapping of caller IDs to descriptions and returns the winning ID
  • score takes an ordered rubric and returns the expected zero-based level
  • noul takes no criteria and returns P(true)

For label sets larger than 20, the repo ships a hierarchical Router that narrows candidates in groups and reranks the survivors. The card is upfront that this narrowing can drop the right answer, and that the final probabilities only cover the last group, not the whole label set.

Author benchmarks

Everything in this section comes from Supersonic Labs’ card, metrics file and launch page. They are author-reported → UnverifiedClaim, not ModelSystem.One measurements.

Run of 24 Sep 2026, H200, BF16, strict encoding, 1,024-token limit. Protocol pinned to AbdelStark/jev-benchmarks @ 0d610cc; typed cases from LocalLLaMA/typed-decisions.

Benchmark Correct / total Julia-1 Jev reference (supplied)
Typed decisions 1,463 / 2,000 73.15% 72.70%
AG News (4 labels, pilot) 94 / 100 94% 91%
DAIR Emotion (6 labels, pilot) 86 / 100 86% 48%
Banking77 (72-label shortlist, pilot) 64 / 100 64% 87%

Typed breakdown: Choice 428/600, Noul 484/600, Score 551/800. A CPU rerun on 25 Sep landed at 72.55% on typed decisions and 60/100 on Banking77, with 3 abstentions counted as errors. On MASSIVE (18 scenario labels, not intents) the author reports 71.50% macro accuracy across 52 locales, 86.25% for pt-PT and 86.75% for en-US. Brazilian Portuguese has not been evaluated yet.

How to read this:

  • The Jev column is not a new Jev run. The protocol supplies those numbers. Independent DecisionEval measured Jev 1.13.0 at 0.740 (95% CI 0.721–0.759) on the same typed-decisions test split on 20 Sep 2026. Julia’s 73.15% falls inside that interval, so “above Jev” on typed decisions reads better as a tie.
  • Pilots are small and may not be zero-shot. 100 examples each. The provenance file lists AG News, Banking77, DAIR Emotion and MASSIVE among the validation sources for the released checkpoint, which suggests those task families were in the training mix. The Jev references are zero-shot. The card doesn’t say whether the pilot rows themselves were excluded.
  • Banking77 is the known weak spot. It runs through the shortlist Router rather than a native 72-way call, and the lab calls it out as its main failure.
  • Typed-decisions gold labels come from a teacher model, so accuracy there is agreement with that teacher, not ground truth.

Latency (author-measured)

Device and test Median p95
Apple M4, one decision per call (100-word context, 4 options) 33.15 ms 44.23 ms
Apple M4, batch of 16 312.11 ms per batch 313.44 ms
Intel Core i5-1235U, typed decisions 294.81 ms 428.49 ms
Intel Core i5-1235U, Banking77 (with Router narrowing) 3,713.54 ms 5,125.10 ms
Samsung SM-X510 tablet, ONNX Runtime on CPU 203 ms 205 ms

The catalog latency band (33–295 ms) spans the M4 single-call median and the i5 typed-decisions median. The Router path is an order of magnitude slower. All of it is UnverifiedClaim.

Limits

  • Possible acquiescence bias. In an informal test on X, a user ran both Jev and Julia through Political Compass, 16Personalities and Big Five questionnaires in Portuguese. Julia agreed with all 62 Political Compass propositions and never disagreed with any 16Personalities item, landing on “authoritarian right” and ENFJ-T respectively. The tester concluded “o Julia simplesmente concorda com qualquer coisa aparentemente” (Julia simply agrees with anything apparently). This is one data point (n=1, no code shared) but flags a potential systematic tendency to pick the affirmative option. Independent replication would be valuable.
  • Benchmarks are self-run and the Jev comparison uses supplied references. See the notes above before quoting “beats Jev”.
  • No calibration step. inference-policy.json ships "calibration": null, and the card says the probabilities are “not guaranteed certainty”. No ECE or reliability data is published. Measure calibration on your own labels before you threshold on it.
  • 2–20 options per native call. Bigger label sets go through the Router, which can lose the correct answer.
  • Context is only evaluated at 1,024 tokens. The launch page describes a 1,024-token configuration; the card, updated on 26 Sep, says the runtime now defaults to 8,192 tokens, with only a CPU smoke test at that length (about 24 s for one 8,192-token row). Accuracy at 8k is not established.
  • Knowledge and multi-step reasoning are out of scope, by the lab’s own account. The provenance file for the released checkpoint shows 26% on an MMLU sample and 28.5% on ARC-Challenge.
  • Custom runtime. The Python package in julia/ is required; it is not a transformers pipeline and has no /v1/systemone server.
  • Hosted API is a plan. The $0.025/M input price is announced, not live.
  • Day zero. 0 HF downloads and 13 likes when checked on 26 Sep 2026.

Fit / anti-fit

Fit when you need: cheap routing or classification over a short, explicit list of options, CPU or edge inference, multilingual inputs, a small model you can ship inside an app.

Anti-fit when you need: long label lists, calibrated confidence out of the box, long documents, knowledge-heavy or multi-step questions, a drop-in for code written against the TypeSafe API, or a hosted SLA.

Why it is in the catalog

Trained decision weights (fine-tune plus a dedicated decision head), an open typed choice / score / noul interface and public, documented evals. It opens the sub-200M CPU slot next to Lumma-fev and Tiny-Jev, and joins Eikos as a Brazilian-built entry. The 144M size and the published provenance make it an easy model to check independently.

The lab has a published manifesto that explains its priorities: local inference over cloud handoffs, small systems that can be measured against clear outcomes, and building useful infrastructure instead of chasing hype. The tagline is “Less distance. More possibility.” This framing is consistent with the technical documentation’s focus on constraints, honest failure modes and reproducibility.