Addis PulseStudio

Ollama v0.40.0 Makes MLX the Default Runtime on Apple Silicon

A routing change that removes a silent-failure class, three underdocumented decision models, and an embedding model that could replace a cloud call in a small studio's tracker.

4 min read781 words

What happened

Ollama v0.40.0 shipped on 6 October 2026. On Apple Silicon, any model architecture the MLX runtime supports now routes to MLX by default β€” no config flag, no workaround. The release also adds gemma4, qwen3.6, and qwen3.5 as runnable models, introduces three decision models (Nimble tev1, clef, clef-flash), and brings the embedding model embeddinggemma-2 onto the MLX path.

Context

Before v0.40.0, an M-series Mac user who wanted MLX acceleration in Ollama had to set it up manually or work around the default CPU path. The release notes frame this as a routing change, not a new feature: the MLX backend existed; now it is the first thing the runtime tries. The full changelog spans v0.35.1 through v0.40.0, meaning several intermediate releases of work landed before this tag. The notes also state that additional models will continue to be tested and enabled, signalling the current MLX-supported architecture list is not final.

How it works

MLX is Apple's machine-learning framework built for its own silicon. In practical terms, Ollama's serving layer now detects whether a model's architecture has a compiled MLX kernel and, if it does, runs inference on the GPU/ANE path instead of falling back to CPU. For a user the effect is that ollama run on an M-series Mac simply gets faster for supported architectures; there is no new CLI flag to learn.

The new model additions split into three groups. Generative text models: gemma4, qwen3.6, qwen3.5. Decision models β€” Nimble tev1, clef, clef-flash β€” placed on the MLX runtime but not further described in the release notes in terms of task, input schema, or parameter count. And embeddinggemma-2, a text-embedding model now available on the MLX path. Ollama remains what it has always been: a model-serving and management layer. It adds no ComfyUI nodes, no video-generation runtime, no diffusion backends.

Our read

The new model names will get the press, but the routing change is the part that actually shifts a workflow. Until now, MLX acceleration on a Mac was a know-the-trick situation: set the right config, verify it took, hope the architecture had a kernel. Making it the default removes a class of silent failures where a user thinks they are on the GPU path and is actually on CPU. For a solo operator doing prompt iteration or asset tagging on a laptop, that is the difference between closing the browser before it OOMs and just typing.

The decision models are the more interesting unknown. Three names β€” Nimble tev1, clef, clef-flash β€” with no task description, no input/output contract, no parameter count in the release notes. They appear to target classification or routing tasks, but that is inference from context, not documentation. Until someone ships a schema, wiring them into a QA gate means exploratory API calls first. The "clef-flash" suffix suggests a speed variant, but that is a guess from the name.

And the line about additional models continuing to be "tested and enabled" is doing more work than it looks. It means today's MLX architecture list is a v1, not a final state. A model that runs on MLX this month may have been disabled last quarter because its kernel was unstable. Treat the support matrix as moving, not fixed.

What this changes

If a team member runs an M-series Mac as a local dev box, the upgrade is ollama update and the next ollama run is on the fast path. No config file to edit. Embeddinggemma-2 is a candidate for a lightweight local semantic-search index over shot lists or asset tags in a production tracker, replacing a cloud API call for that one query. The decision models are a watch-and-test item: if the pipeline already calls an Ollama endpoint, prototype a flag-a-render-for-review step once the model contract is documented. The ComfyUI video-generation stack is untouched. Nothing in this release changes the core render pipeline.

License

No source in the brief states a licence for Ollama v0.40.0, the MLX integration, the newly added models (gemma4, qwen3.6, qwen3.5), the decision models, or embeddinggemma-2. Check each model's card and the Ollama repository before building anything commercial on top of any of them.

Key takeaways

  • MLX is now the default inference backend on Apple Silicon for supported architectures; no configuration change is required.
  • gemma4, qwen3.6, qwen3.5, three decision models, and embeddinggemma-2 are new in v0.40.0.
  • The decision models' task, schema, and parameter count are not described in the release notes.
  • The MLX-supported architecture list is explicitly incomplete; more models are still being tested and enabled.
  • No licence is stated in any source for any component of this release.

Sources

  1. v0.40.0 β€” tier 2
ollamamlxapple siliconllm

How this post was made

Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.

Drafted
Independent sources
1
cluster pair
clef:27b
cluster label
gemma4:12b
radar brief
gemma4:12b
research brief
qwen3.8:27b
draft article
qwen3.8:27b
short script
qwen3.8:27b
seo pack
gemma4:12b
Run
editorial-20261006T185729Z