Ollaya

Last updated:

RuntimeApache-2.0~935 ★

⚠️ Why not a model

Serves existing open decision models locally (Laya, Decider, NLI, GLiClass, Kev, Von); no trained weights of its own — it's distribution and inference infra.

Technical specs

Base LLMs
Decision types
choice, score, boolean
Features
local, ollama-style, typesafe-api, desktop-app, cli, docker, nvidia-gpu, apple-silicon, mlx, systemd, vision, images
License
Apache-2.0

Ollaya is the Ollama for decision models — a local runtime that pulls, caches and serves open decision models behind an API compatible with TypeSafe’s /v1/systemone.

What it does

One binary, one command: ollaya run laya. The runtime:

  1. Pulls model weights from a registry (like ollama pull)
  2. Serves them on localhost:11435 with TypeSafe’s request/response shapes
  3. Works with the official TypeSafe Python SDK unchanged — just point TYPESAFE_BASE_URL at the local server

Desktop apps for macOS, Windows and Linux, plus a CLI and Docker image. NVIDIA GPUs (CUDA 13, driver R580+), Apple Silicon (MLX for Laya and NLI) or CPU fallback.

Why it’s not a model

Ollaya does not train any weights. It’s infrastructure that distributes and runs models that others have trained. The decision-making capability comes entirely from the models it serves (Laya, Decider, Kev, etc.), not from Ollaya itself.

Models available

Model Author Sizes Notes
laya Convai Innovations 322m, 421m Fastest; MLX on Apple Silicon
decider Mapika 0.75b, 1.9b, 4.2b Most accurate in Ollaya’s tests
nli Moritz Laurer 396m, 435m Zero-shot entailment classifiers
gliclass Knowledgator 439m Instruction-following zero-shot
qwen3guard Qwen team 0.6b Safety guard, 119 languages
decision vLLM Semantic Router 0.75b Fine-tuned Qwen3.5 + endpoint head
kev Jared Palmer 0.76b, 4.2b, 7.9b LoRA + pointer head
von Victor Hugo Panisa 395m ModernBERT-large, 8k context

Latency (author-reported)

All numbers are UnverifiedClaim — measured by Ollaya on an RTX 4090 via the HTTP API.

Model Median latency
laya:multilingual 8.1 ms
laya:en 9.6 ms
gliclass 14.7 ms
nli 20.4 ms
decider:0.8b 155 ms
decider:2b 190 ms

For comparison, Ollaya cites third-party benchmarks for the hosted Jev API at 236–276 ms. Different setups (local GPU vs network round-trip), so treat this as order-of-magnitude context.

TypeSafe compatibility

The SDK works unchanged:

export TYPESAFE_BASE_URL=http://localhost:11435
export TYPESAFE_API_KEY=local        # any value works
export TYPESAFE_DEFAULT_MODEL=laya

Or call the endpoint directly — same JSON shapes as TypeSafe’s hosted API.

Platforms

Platform Desktop app CLI GPU
macOS (Apple Silicon, 14+) ✓ .dmg ✓ Apple GPU (MLX)
Windows 10/11 x64 ✓ .exe/.msi ✓ PowerShell NVIDIA CUDA 13
Linux x86-64 ✓ AppImage/.deb/.rpm ✓ systemd NVIDIA CUDA 13
Linux ARM64 — ✓ CPU only
WSL 2 — ✓ NVIDIA CUDA 13
Docker (amd64/arm64) — ✓ GHCR NVIDIA :cuda image

Caveats

  • Not affiliated with Ollama or TypeSafe — independent open-source project. Since 30 Sep 2026 Ollama itself has native decision-model support on port 11434 (see Ollama); that is a separate product with its own model list
  • Latency comparisons mix local GPU with hosted API; read as order-of-magnitude
  • Model quality depends on the underlying model, not Ollaya
  • Anonymous maintainers (ollaya-dev org)

Vision (v0.7.5)

From v0.7.5 Ollaya serves decider:2b-vision (Mapika) and accepts an images field (base64 PNG) on the API/CLI. Vision latency figures in the release notes are UnverifiedClaim. See /news/ollaya-vision.