Addis PulseStudio

Ollaya puts TypeSafe decision models on your machine for 10 ms and no fee

A local serving layer for open-weight decision models that drops the TypeSafe round-trip from 236 ms to single-digit milliseconds, with one environment variable and no per-call charge.

4 min read806 words

What happened

Ollaya shipped as a local server that runs open-source decision models the way Ollama runs LLMs: one binary, one port, zero per-token fees. A five-question request to the bundled Laya model returns in roughly 10 ms on an RTX 4090, where the hosted TypeSafe API reports a median of 236–276 ms.

Context

TypeSafe built Jev as a hosted API for structured, typed decisions: pick an option, score a hypothesis, route a job. The official Python SDK 0.7.1 wraps the round-trip cleanly. The catch is the round-trip itself. Every call leaves the machine, hits a metered endpoint, and returns in hundreds of milliseconds. The four models Ollaya ships (Laya, decider, nli, gliclass) all came from open-weight releases by their respective authors, but no local serving layer spoke TypeSafe's API shapes until now. Ollaya fills that gap.

How it works

None of the four bundled models generate tokens. Laya, a 322 m-parameter English model (421 m for 100+ languages) from Convai Innovations, performs a single forward pass and returns a typed answer with per-choice confidence scores and a full probability distribution. The API response reports output_tokens as 0. decider (0.75 b / 1.9 b, built on Qwen3.5 by Mapika) reads its answer from option-letter logits in the same single pass. nli (396 m / 435 m, by Moritz Laurer) scores every option as an entailment hypothesis. gliclass (439 m, by Knowledgator) scores all options in one pass so cost barely grows with option count.

Inference runs through ONNX Runtime on CPU or NVIDIA GPU. Ollaya pulls weights from each author's Hugging Face repository, pins them to a commit, and verifies sha256. It does not re-host weights. The default endpoint is 127.0.0.1:11435, and the TypeSafe SDK redirects there by setting TYPESAFE_BASE_URL. Every model runs on CPU; an NVIDIA GPU on Linux, WSL 2, or Docker reduces latency to milliseconds. Ollaya supports macOS, Windows, Linux, and Docker.

Our read

The number that matters is not the 10 ms. It is the ECE. Ollaya's testing puts Laya's expected calibration error at 0.081 after temperature fitting, versus 0.246 for TypeSafe Jev. That is roughly a 3× gap in how well a model's stated confidence tracks its actual accuracy, and in a pipeline routing call that gap is the difference between a dispatch that is right nine times out of ten and one that is right seven out of ten. A fast model that miscalibrates will misroute more often than a slower one that knows when it is guessing.

What the release does not document is training-data provenance. No source states what corpus produced Laya, decider, nli, or gliclass. For a studio shipping product decisions through a classifier, open-weight is not the same as auditable. The weights are pinned and hash-checked, but the recipe behind them is not published.

The second-order effect matters more than the latency win. Because the TypeSafe SDK redirects to a local port with a single environment variable, the barrier to switching is effectively zero. That means the hosted API's pricing leverage over this class of routing task drops to nothing for anyone with a GPU and a terminal. Vendors who priced Jev on per-call margins now compete against a model that costs electricity.

What this changes

For a studio already calling TypeSafe or Jev for pipeline routing: classifying a brief's intent, tagging assets, dispatching a job to the right generation queue. This is a one-line change. Set TYPESAFE_BASE_URL=http://localhost:11435, run ollaya run laya, and the round-trip drops from 236 ms to single-digit milliseconds on a GPU. No per-call fee, no API key.

For a studio not calling TypeSafe: nothing changes yet. There is no ComfyUI node, no workflow integration, and these models classify. They do not generate video, images, or audio. The dependency is net-new and the near-term use case is narrow. If your routing logic is a 20-line if/else, a 322 m-parameter classifier is not the tool you need this Monday.

License

The Ollaya runtime is Apache-2.0: commercial use, modification, and redistribution are all permitted. The bundled model weights are pulled from their authors' Hugging Face repositories, but the sources do not state a licence for Laya, decider, nli, or gliclass. Check each model card before building anything commercial on them.

Key takeaways

  • Ollaya turns the TypeSafe decision-model API into a local, zero-fee server with a single environment-variable redirect.
  • Laya's 0.081 ECE versus Jev's 0.246 is a larger practical gap than the 10 ms latency difference.
  • All four models run on CPU via ONNX Runtime; an NVIDIA GPU cuts latency to single-digit milliseconds.
  • The Apache-2.0 runtime licence is permissive; the model weights' licences are not stated in any source.
  • Training-data provenance is not documented for any of the bundled models.

Sources

  1. Ollaya – Ollama for open-source, Jev-style decision models — tier 3
inferencedecision-modelslocal-aimachine-learning

How this post was made

Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.

Drafted
Independent sources
1
cluster pair
gemma4:12b
cluster label
gemma4:12b
radar brief
gemma4:12b
research brief
qwen3.8:27b
draft article
qwen3.8:27b
short script
qwen3.8:27b
seo pack
gemma4:12b
Run
editorial-20260925T223947Z