Ollama adds a local /v1/systemone endpoint and three Jev-style decision models

Last updated:

Source

Ollama now runs decision models on your own machine. In a blog post dated 29 September and an announcement on X at 01:25 BRT on 30 September, the team shipped a local /v1/systemone endpoint “based on TypeSafe’s Jev API”, available from Ollama 0.35. The pitch: no extra cost and lower latency when the model runs locally, for jobs like ticket triage, model routing and content moderation.

ollama pull nimble

Three models at launch. Nimble 9B from Bespoke Labs (nimble), and Together AI’s Tev1 in 4B (tev1) and 0.8B (tev1:0.8b). Ollama says more decision models are coming, including ones served from Ollama’s cloud. Within about 14 hours, the Nimble and Tev1 library pages showed roughly 2.7k pulls each.

Same contract as Jev. Requests carry state and up to 64 named questions of type choice, noul or score; responses give the choice, per-option probabilities and a confidence value. TypeSafe’s official Python SDK works unchanged when pointed at http://localhost:11434. The API reference is candid about limits: 64 KiB per request, no streaming, images or tools, GGUF models only, and confidence is a measure of how peaked the distribution is, “not calibrated correctness”.

Vendor numbers (UnverifiedClaim). Nimble “averaged 91ms per decision” playing a Pac-Man demo on a MacBook Pro M5 Max. On Bespoke Labs’ 13 public human-labelled datasets (3,880 decisions), Ollama reports mean accuracy of 75.7% for Nimble, 73.3% for Tev1 4B and 63.5% for Tev1 0.8B, against 76.0% for Jev 1.13. Ollama ran Nimble and Tev1 itself; the Jev figure is taken from Bespoke Labs’ published run.

Why it matters. Until now, running open decision models behind the System One API meant a separate tool such as Ollaya, an independent project with no ties to Ollama. With Ollama’s install base, /v1/systemone becomes a default local interface, and open models like Nimble get distribution that GitHub and Hugging Face alone didn’t give them. It also makes TypeSafe’s request shape a de facto standard: Ollama, Ollaya and now Liquid AI’s hosted d1 all accept it.

How we list it. Ollama trains no decision weights, so it goes under runtimes, off the quadrant. The Nimble and Tev1 catalog entries now mention Ollama availability. Nimble’s maturity and adoption scores went up a little; Tev1’s adoption did too. Tev1 stays choice-only in the catalog: the noul and score shapes Ollama exposes for it come from Ollama’s candidate scoring, not from Tev1’s training.