Ollama

Last updated:

RuntimeMIT~182,000 ★

⚠️ Why not a model

Inference server. Since v0.35 it serves decision models trained by others (Nimble, Tev1) on a local /v1/systemone endpoint; it has no decision weights of its own.

Technical specs

Base LLMs
Bespoke-Nimble-9B, Tev1-4B-experimental, Tev1-0.8B-experimental
Decision types
choice, noul, score
Features
local, systemone-endpoint, typesafe-sdk-compatible, gguf, multi-question, keep-alive, no-api-key
License
MIT

Ollama, the widely used local LLM runner, now serves decision models. Its blog post (dated 29 Sep 2026) and the @ollama announcement (30 Sep 2026, 01:25 BRT) introduce a local /v1/systemone endpoint “based on TypeSafe’s Jev API”, available from Ollama 0.35.

ollama pull nimble
curl http://localhost:11434/v1/systemone -d '{"model": "nimble", "state": "...", "questions": {...}}'

What it does

  • Endpoint: POST /v1/systemone on the usual port 11434. Send state (a string, or JSON serialized as text) and 1–64 named questions; get back one answer per question plus token usage.
  • Question types: choice (2–26 options, returns the top option, per-option probabilities and confidence), noul (probability of true) and score (2–26 ordered levels, returns the probability-weighted level, a legend and probabilities).
  • TypeSafe SDK works unchanged: set TYPESAFE_BASE_URL=http://localhost:11434, TYPESAFE_API_KEY=ollama (any value) and TYPESAFE_DEFAULT_MODEL=nimble. Local requests need no key.
  • Scoring, not generation: each question is scored separately against the full state; answers are not passed to later questions. Probabilities are normalized over the supplied candidates.

Models at launch

Model Pull Maker Notes
Nimble 9B ollama pull nimble Bespoke Labs Fine-tuned from Qwen3.5-9B; 9.5 GB; Apache-2.0
Tev1 4B ollama pull tev1 / tev1:4b Together AI Qwen3.5-4B fine-tune; 4.5 GB Q8_0 (4.21B params)
Tev1 0.8B ollama pull tev1:0.8b Together AI Qwen3.5-0.8B fine-tune; 812 MB Q8_0 (752M params)

Ollama says more decision models are coming, “including models served by Ollama’s cloud”. For now the endpoint rejects cloud models.

Limits (from the API reference)

  • Local only. Streaming, images, tools and generation controls are not supported.
  • Request body up to 64 KiB. Each rendered prompt must fit the loaded context; input is never truncated.
  • Needs compatible GGUF weights and a scoring-capable runner. MLX / Safetensors models are not supported yet (MLX speed-ups on Apple Silicon are on the roadmap).
  • confidence is 1 − H(p)/ln(N), a measure of how peaked the distribution is. Ollama’s docs say plainly that it is not calibrated correctness.
  • usage.output_tokens counts tokens generated internally for scoring, so it is not always zero.

Vendor numbers (UnverifiedClaim)

  • Latency: Nimble 9B “averaged 91ms per decision” in a Pac-Man demo on a MacBook Pro M5 Max (blog); “under 100ms” on the same machine (library page). No p50/p99 or other hardware.
  • Accuracy: mean accuracy on Bespoke Labs’ 13 public human-labelled datasets (3,880 decisions): Jev 1.13 76.0%, Nimble 9B 75.7%, Tev1 4B 73.3%, Tev1 0.8B 63.5%. Ollama ran Nimble and Tev1 itself; the Jev number comes from Bespoke Labs’ published run on the same decisions (suite).

Why it’s not a model

Ollama trains no decision weights. It is inference and distribution infrastructure: the decision quality comes entirely from the models it serves. That puts it in the same class as Ollaya and vllm-jev, and it stays off the quadrant.

Tev1 note. Together’s model card says Tev1 was trained to return one option letter and that its logprobs are preferences, not calibrated confidence. On Ollama, the noul and score shapes for Tev1 come from Ollama’s candidate scoring, not from separate training. The catalog keeps Tev1 at choice.

Not the same as Ollaya

Ollaya is an independent project (“Ollama for decision models”, port 11435) that says it is not affiliated with Ollama. This entry is Ollama’s own native support. Both speak TypeSafe’s /v1/systemone shapes, but they have different model registries and different model lists.