Claude Sonnet 5.5: the token-efficiency release that hides a bigger story in the benchmarks
Anthropic's second Claude 5.5-family model holds the per-token price flat and cuts per-task cost through fewer tokens. The Terminal-Bench jump from 10.3% to 70.6% is the number that matters.
What happened
Anthropic released Claude Sonnet 5.5 on 28 September 2026, the second model in its 5.5 family and the tier aimed at fast, well-scoped work. It runs 30%+ faster than Sonnet 5 and costs up to 30% less per task, while the per-token price stays flat at $2 per million input and $10 per million output tokens.
Context
Sonnet 5 launched roughly three months before this release, selling itself on efficient agentic deployment at a lower cost than competitors. The week before Sonnet 5.5 shipped, OpenAI released enhanced versions of its Sol and Luna mid-tier models. The mid-tier price war that accelerated through 2025 has become a weekly cadence, and each new release narrows the gap between "good enough for most tasks" and "the model you need for the hard ones."
How it works
The cost claim is mechanical, not commercial. The per-token price did not move. What moved is token consumption: Anthropic states Sonnet 5.5 needs far fewer tokens to finish the same work, and it batches tool calls more aggressively than Sonnet 5, reducing the number of steps in an agentic loop. The operator-facing control is the effort dial: Low, Medium, High, Max, xhigh. Default is Medium in Claude Code and the claude.ai apps, High on the Claude Platform. At Low or Medium effort, Sonnet 5.5 beats Sonnet 5's best score on several benchmarks at roughly a tenth of the per-task cost. At Max effort it performs comparably to Opus 5.5 on several evaluations. The jump from 10.3% to 70.6% on Terminal-Bench 4.0, an agentic coding benchmark, is the single most striking number in the release.
Our read
The "30% cheaper" framing undersells what actually happened. The per-token price did not move; what moved is how many tokens the model burns per task. A model that finishes a coding loop in a fraction of the tokens is not a discount. It is a different tool at the same price. The Terminal-Bench 4.0 jump from 10.3% to 70.6% is the real headline, because that benchmark measures multi-step, tool-using, terminal-level work, exactly the loop a small studio runs when scripting a ComfyUI pipeline or iterating on workflow JSON.
On Sonnet versus Opus, the sources pull in different directions. TechCrunch reports Anthropic's benchmarks show Sonnet 5.5 outperforming Opus 5.5 on agentic coding, attributed to Sonnet's ability to spawn multiple agents within cost limits. Anthropic's own page says Sonnet 5.5 at Max effort is comparable to Opus on several evaluations, then states Opus remains clearly stronger at complex, open-ended work requiring sustained judgment. Simon Willison's independent testing lands between those: almost as good as Opus on some coding tasks. The honest read: Sonnet 5.5 has closed the gap on scoped, iterative, tool-heavy work and has not closed it on the reasoning Opus was designed for.
One data point from Willison deserves attention. At 'max' thinking effort, Sonnet 5.5 consumed 128,000 tokens and cost $1.28 to fail at generating an SVG. At 'xhigh' it produced the same image for 5.74 cents in 41 seconds. The effort dial is a routing decision, not a linear quality slider.
What this changes
The claude.ai free tier now serves Sonnet 5.5. If your text-planning step, whether shot lists, ComfyUI workflow JSON, or prompt refinement, runs on a free or local model, the free tier is now stronger, and that change lands Monday with zero API cost.
If you are already paying Sonnet 5 API rates, the token-efficiency gain is real: same per-token price, fewer tokens per task. Point your key at Sonnet 5.5 and re-run the same task to see the delta.
If a local model handles your planning layer, nothing here forces a change. No ComfyUI node, no local weights, no on-premises endpoint for Sonnet 5.5 appears in any source. Haiku 5.5 is coming in the coming weeks with no date, and that is the tier that would actually pressure a local setup.
License
No source in the brief states a software or weights licence for Sonnet 5.5. These are API-only models; no open weights have been published. The terms governing use are Anthropic's commercial API agreement, not an open-source licence. Check Anthropic's model card and API terms before shipping anything that depends on the endpoint.
Key takeaways
- The cost advantage of Sonnet 5.5 is entirely token efficiency, not a price cut: $2/$10 per million in/out tokens is unchanged from Sonnet 5.
- Terminal-Bench 4.0 scores jumped from 10.3% to 70.6%, the largest single-benchmark improvement in the release and the one that matters for agentic, tool-using pipelines.
- The effort dial is a routing decision, not a quality slider: Simon Willison's test shows 'max' can cost roughly 22× more than 'xhigh' and produce a worse result.
- The claude.ai free tier now serves Sonnet 5.5, giving a zero-cost LLM for text-planning steps in a video-production pipeline.
- For studios running local open-weight models, same-day community benchmarks suggest the local stack remains competitive at Sonnet 5.5's Low and Medium effort tier.
Sources
How this post was made
Drafted from clustered primary sources by the models below, then read, edited and approved by a human before it was published. The sources are listed in full at the end of the article.
- Drafted
- Independent sources
- 4
- cluster pair
- gemma4:12b
- cluster label
- gemma4:12b
- radar brief
- gemma4:12b
- research brief
- qwen3.8:27b
- draft article
- qwen3.8:27b
- short script
- qwen3.8:27b
- seo pack
- gemma4:12b
- Run
- editorial-20260928T225029Z