Eikos

Last updated:

Open23–800ms$0/M input

Quadrant scores

See the full quadrant

Scored with the public quadrant rubric (maturity × capability; bubble = adoption). Revise as evidence lands.

  • Maturity5.4/10

    Weights live on HF since 23 Sep (4B, 27B, FP8, INT4, MLX), full code/data/eval repo, TypeSafe-compatible /v1/systemone server with agent sessions. No third-party integrations yet; self-host only, no SLA.

  • Capability7.2/10

    noul / choice / score with a published calibration story (ECE per build, T=1 rationale, quantization gate). All accuracy numbers come from the author's own harness; the JevBench leaderboard entry is still pending. No p50/p99 latency.

  • Adoption25/100

    PT-BR launch post at ~374 likes / ~28k views; 16 + 11 likes and ~250 downloads on the two main HF repos, ~20 GitHub stars, dataset public. No production use reported.

Vendor claims

Eikos is an open family of typed-decision models from Brazilian developer Caio Vicentino (@0xCVYH). It comes in two sizes, 4B and 27B, and targets rule-heavy decisions in finance, trading and trade finance: apply a stated rule or policy to a case and return a probability for every allowed answer. It does not predict prices.

The weights went up on Hugging Face on the evening of the announcement day (23 Sep 2026, 20:48–20:59 BRT), together with FP8 and INT4 builds; MLX builds for Apple Silicon followed on 24 Sep. The code, data pipeline, training recipes and evaluation harness are on GitHub, and the training set is published as caiovicentino1/eikos-decisions. This page said “release pending” until 26 Sep; that was stale.

The name comes from the Greek εἰκός, “the probable”.

Specs

Attribute Value
Author Caio Vicentino (independent) · HF caiovicentino1 · GitHub caiovicentino
Sizes Eikos-4B (4.66B params, base Qwen/Qwen3.5-4B) · Eikos-27B (27.78B params, base Qwen/Qwen3.8-27B)
Training LoRA rank 64, one epoch, merged and exported as a full checkpoint in the official Qwen layout; distilled from GLM-5.3-Flash on generated, programmatic and long-dossier items. 4B is a soup of two recipes
Readout SemIf prompt format; next-token logits read only over option letters (A–Z, then AA, AB…), softmax at T = 1
Decision types noul (yes/no), choice, score (ordinal)
Options Up to 160 in one pass on the 27B (per-model limit in decision_config.json); a tournament beyond that
Context Trained on inputs up to 32k tokens (4B) and 12k (27B); long-context probe up to 64k; default vLLM max-model-len 16,384
Languages Trained on English and Portuguese; Spanish held out as a zero-shot test
Builds bf16, FP8, INT4 (GPTQ), MLX (4B 8-bit / 4-bit, 27B 4-bit): nine HF repos
Size on disk 27B: 55.6 GB bf16, 19.4 GB INT4, 15.1 GB MLX 4-bit · 4B: 9.3 GB bf16, 2.4 GB MLX 4-bit
Serving vLLM ≥ 0.30 (older builds give wrong answers on batched long requests), PyTorch/transformers, MLX or MPS
API serve.py: TypeSafe-compatible POST /v1/systemone, plus /v1/sessions for agent sessions with an incremental state
License MIT for the fine-tuning deltas and code; Qwen base models Apache-2.0; dataset CC BY 4.0
Weights published
Status Open weights (Hugging Face)

How it works

The state, the question and lettered options go into a SemIf-style prompt. The model is run once and the server reads the logits of the option letters at the answer position, then applies a softmax. All questions in a request share one pass over the state through vLLM’s prefix cache.

That readout is the same trick runtimes like SemIf use on frozen models. The difference is that Eikos trains the weights for it: soft cross-entropy on the option-letter logits against the teacher’s probabilities, with option-order permutation and an auxiliary rationale loss that is switched off at inference. That is why it sits in the catalog and not under Runtimes.

Author benchmarks

All numbers come from the Eikos-27B card and the GitHub README. They are author-reported → UnverifiedClaim, measured with the author’s own harness.

Benchmark (author’s harness) Eikos-4B Eikos-27B Jev (author’s reference)
JevBench public, original / hard 91.7 / 72.1 100.0 / 82.9 98.6 / 73.0
DecisionBench (OOD), medium / hard 77.1 / 66.9 88.4 / 78.5 89.1 / 69.3
General battery (9 human-labeled tasks) 76.0 82.5 84.1
Finance (CUAD, financial sentiment, FinQA-judge) 74.7 85.3 79.9
Trade rules, unseen rules 76.0 88.0 79.4
Central-bank stance (balanced acc.) 38.4 44.5 58.6
Decision hidden in 64k tokens 74.2 88.3 n/a (API limit 32k)
RuleArena NBA (balanced acc.) 0.50 0.50 —

Release builds report ECE between 0.021 and 0.043, and about 2–2.7% error on the decisions each build takes at ≥90% confidence.

How to read this:

  • JevBench here means the public items only, scored by the author, not the official leaderboard. The card says the leaderboard submission is pending.
  • The Jev column was measured by the author through Jev’s API on the author’s batteries. It is not a TypeSafe number and not an independent run.
  • The rules suites come from the same generator family used in training. The author flags this and points to “unseen trade rules” and RuleArena as the transfer tests. On RuleArena both sizes score 0.50: no discrimination.
  • The general battery still favours Jev (84.1 vs 82.5 for the 27B), and central-bank stance is a clear weakness of both sizes.
  • Teacher distillation. Labels come from GLM-5.3-Flash at maximum reasoning effort, filtered for agreement with a gold answer.

Limits

  • Benchmarks are author-run. Code and data are public, so they can be rerun, but nobody outside has done it yet.
  • Single pass, no multi-step reasoning. Long chains of arithmetic across many rules are out of reach; an optional “verify” mode exists but is off and not in any reported number.
  • Calibration ships at T = 1. The author found every fitted temperature made the model over-confident on hard data, so re-fit calib.json on your own labels.
  • Context beyond training degrades. The 64k result is a probe; the models were trained on 32k (4B) and 12k (27B).
  • vLLM version matters. Below 0.30, batched long requests return wrong answers on this hybrid (Gated DeltaNet) architecture.
  • Big footprint for a decision model. The 27B needs a serious GPU; the 4B MLX 4-bit build (2.4 GB) is the laptop option at ~0.4–0.8 s per decision on an M4.
  • Not advice. It applies the rules it is given; it does not know your jurisdiction’s current law.

Fit / anti-fit

Fit when you need: rule and policy checks on financial cases (order limits, KYC/AML, lending policy, documentary credits, VAT), many questions over one long dossier, a self-hosted server that speaks the TypeSafe contract, Portuguese or English inputs, calibrated confidence you can gate on.

Anti-fit when you need: price prediction, long multi-step computation, monetary-policy stance, a small edge model, or a hosted SLA.

Why it is in the catalog

Trained decision weights (LoRA fine-tune merged into full checkpoints) with open typed I/O, a TypeSafe-compatible server, and code, data and evaluation all published. It is the first finance-focused peer in the catalog and, alongside Julia-1, one of its Brazilian-built entries. Next up for it: an independent rerun and the official JevBench entry.