Ollaya
⚠️ Why not a model
Serves existing open decision models locally (Laya, Decider, NLI, GLiClass, Kev, Von); no trained weights of its own — it's distribution and inference infra.
Technical specs
- Base LLMs
- Decision types
- choice, score, boolean
- Features
- local, ollama-style, typesafe-api, desktop-app, cli, docker, nvidia-gpu, apple-silicon, mlx, systemd, vision, images
- License
- Apache-2.0
Links
Ollaya is the Ollama for decision models — a local runtime that pulls, caches and serves open decision models behind an API compatible with TypeSafe’s /v1/systemone.
What it does
One binary, one command: ollaya run laya. The runtime:
- Pulls model weights from a registry (like
ollama pull) - Serves them on
localhost:11435with TypeSafe’s request/response shapes - Works with the official TypeSafe Python SDK unchanged — just point
TYPESAFE_BASE_URLat the local server
Desktop apps for macOS, Windows and Linux, plus a CLI and Docker image. NVIDIA GPUs (CUDA 13, driver R580+), Apple Silicon (MLX for Laya and NLI) or CPU fallback.
Why it’s not a model
Ollaya does not train any weights. It’s infrastructure that distributes and runs models that others have trained. The decision-making capability comes entirely from the models it serves (Laya, Decider, Kev, etc.), not from Ollaya itself.
Models available
| Model | Author | Sizes | Notes |
|---|---|---|---|
| laya | Convai Innovations | 322m, 421m | Fastest; MLX on Apple Silicon |
| decider | Mapika | 0.75b, 1.9b, 4.2b | Most accurate in Ollaya’s tests |
| nli | Moritz Laurer | 396m, 435m | Zero-shot entailment classifiers |
| gliclass | Knowledgator | 439m | Instruction-following zero-shot |
| qwen3guard | Qwen team | 0.6b | Safety guard, 119 languages |
| decision | vLLM Semantic Router | 0.75b | Fine-tuned Qwen3.5 + endpoint head |
| kev | Jared Palmer | 0.76b, 4.2b, 7.9b | LoRA + pointer head |
| von | Victor Hugo Panisa | 395m | ModernBERT-large, 8k context |
Latency (author-reported)
All numbers are UnverifiedClaim — measured by Ollaya on an RTX 4090 via the HTTP API.
| Model | Median latency |
|---|---|
| laya:multilingual | 8.1 ms |
| laya:en | 9.6 ms |
| gliclass | 14.7 ms |
| nli | 20.4 ms |
| decider:0.8b | 155 ms |
| decider:2b | 190 ms |
For comparison, Ollaya cites third-party benchmarks for the hosted Jev API at 236–276 ms. Different setups (local GPU vs network round-trip), so treat this as order-of-magnitude context.
TypeSafe compatibility
The SDK works unchanged:
export TYPESAFE_BASE_URL=http://localhost:11435
export TYPESAFE_API_KEY=local # any value works
export TYPESAFE_DEFAULT_MODEL=laya
Or call the endpoint directly — same JSON shapes as TypeSafe’s hosted API.
Platforms
| Platform | Desktop app | CLI | GPU |
|---|---|---|---|
| macOS (Apple Silicon, 14+) | ✓ .dmg | ✓ | Apple GPU (MLX) |
| Windows 10/11 x64 | ✓ .exe/.msi | ✓ PowerShell | NVIDIA CUDA 13 |
| Linux x86-64 | ✓ AppImage/.deb/.rpm | ✓ systemd | NVIDIA CUDA 13 |
| Linux ARM64 | — | ✓ | CPU only |
| WSL 2 | — | ✓ | NVIDIA CUDA 13 |
| Docker (amd64/arm64) | — | ✓ GHCR | NVIDIA :cuda image |
Caveats
- Not affiliated with Ollama or TypeSafe — independent open-source project. Since 30 Sep 2026 Ollama itself has native decision-model support on port 11434 (see Ollama); that is a separate product with its own model list
- Latency comparisons mix local GPU with hosted API; read as order-of-magnitude
- Model quality depends on the underlying model, not Ollaya
- Anonymous maintainers (ollaya-dev org)
Links
- Website
- GitHub — 446★, Apache-2.0
- Docker image
Vision (v0.7.5)
From v0.7.5 Ollaya serves decider:2b-vision (Mapika) and accepts an images field (base64 PNG) on the API/CLI. Vision latency figures in the release notes are UnverifiedClaim. See /news/ollaya-vision.
