Ollama v0.34.3: thinking controls become a query, not a guess
A structured reasoning-levels field in /api/show, Nemotron H vision models on Apple Silicon via MLX, and a HuggingFace pull fix round out a small patch with one outsized implication for pipelines that swap models.
What happened
Ollama v0.34.3 shipped on 19 September 2026 with a structured thinking-controls field in its /api/show endpoint and Nemotron H vision-model support on Apple Silicon via the MLX backend. The patch spans v0.34.2 through v0.34.3-rc0 and also bundles a HuggingFace model-pull fix and a macOS window-management correction.
Context
Ollama is a local LLM serving runtime, the layer between application code and model weights. Before this release, a client that needed a model's reasoning-level options had to hard-code them or parse documentation. The thinking field in /api/show is the first time Ollama exposes that dimension as structured, queryable data in the runtime's own API. It sits alongside existing model metadata and changes how a pipeline discovers what a model can do at inference time, without a per-model configuration file.
How it works
The /api/show response now carries a thinking object listing the available reasoning levels and the model's default. The release-notes example uses the identifier 'glm-5.3-flash:cloud', which returns ['low', 'high', 'max'] with a default of 'max'. A client queries the endpoint once, reads the field, and sets the inference parameter accordingly. No per-model hard-coding, no config-file lookup.
Separately, Nemotron H vision models are now runnable on Apple Silicon through the MLX backend, so vision-capable inference on a Mac no longer requires a discrete GPU server. The HuggingFace pull fix repairs a download path for models fetched from HF. The macOS app no longer reopens windows the user closed in the previous session. The full changelog spans v0.34.2 through v0.34.3-rc0, so the window fix and the HF fix land in the same patch cycle rather than as separate hotfixes.
Our read
The thinking-controls field is the story here, and it is smaller than the Nemotron H headline. What it actually does is turn a static configuration problem into a dynamic one. In a ComfyUI pipeline where a custom node calls Ollama for prompt expansion or scene description, the node previously had to assume — by hard-coded constant — that one model wants thinking: "high" and another wants thinking: "max". Now it asks. That is the difference between a pipeline that breaks the day you swap the model and one that adapts without a code change.
The cloud identifier in the example deserves a second look. 'glm-5.3-flash:cloud' is not a local path; it is a cloud-tier tag. Ollama has positioned itself as a local runtime, yet the API is clearly being shaped to treat local and cloud backends as the same surface. Whatever node you build against /api/show today should not need a fork when the model moves between a local shard and a cloud endpoint. That is a design commitment, and it is worth noting because it means the thinking field is likely to be the stable contract even as the models behind it change.
What the release does not tell you: which Nemotron H variants are covered, what hardware floor they require, or whether the thinking field appears for every model or only those that implement a reasoning dimension. Those gaps matter if you are about to wire a new node around the field and want it to degrade gracefully when the answer is absent.
What this changes
If your stack calls Ollama for LLM-assisted steps — prompt expansion, script generation, content tagging — add a /api/show query before the first inference call. Read the thinking field, set the parameter, move on. When you swap the underlying model, the node still works without a config edit.
If you run on an M-series Mac and have been waiting for a vision-capable model to work locally via MLX, Nemotron H is that model. Frame-level analysis or content tagging in the video workflow no longer needs a separate GPU box.
The HuggingFace pull fix is operational relief: models you were pulling from HF should download cleanly now. The macOS window fix is a QoL item with no pipeline impact.
Nothing changes if you are not on Apple Silicon, not using HF pulls, and not calling Ollama's API from custom code.
License
The sources do not state a licence for Ollama v0.34.3 or for the Nemotron H vision models. Before building anything commercial on either, check the model card and the Ollama repository's licence file. Publishing weights does not confirm a licence, and the absence of one in the release notes is not a signal that one does not exist.
Key takeaways
- /api/show now returns a structured thinking field with available levels and default, so pipelines can query a model's reasoning options instead of hard-coding them per model.
- Nemotron H vision models run on Apple Silicon via MLX, bringing local vision inference to M-series Macs without a separate GPU server.
- The release example uses a cloud-tier identifier, suggesting Ollama's API is being unified across local and cloud backends under one contract.
- Supported Nemotron H variants, hardware requirements, and whether the thinking field is universal are not specified in the release notes.
- No licence is stated in the available sources; commercial-use verification is required before shipping anything built on these weights.
Sources
- v0.34.3 — tier 2
How this post was made
Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.
- Drafted
- Independent sources
- 1
- cluster pair
- gemma4:12b
- cluster label
- gemma4:12b
- radar brief
- gemma4:12b
- research brief
- qwen3.8:27b
- draft article
- qwen3.8:27b
- short script
- qwen3.8:27b
- seo pack
- gemma4:12b
- Run
- editorial-20260919T201353Z