Claude Sonnet 5.5: Pricing, Context Window and Speed Guide
Anthropic model released September 28, 2026, positioned as the balance of speed and intelligence in the Claude lineup. Anthropic says it runs more than 30% faster than Claude Sonnet 5 and costs up to 30% less for most work, at the same $2 / $10 per million tokens. It takes text and image input, has a 1M-token context window and 128K-token maximum output, uses adaptive thinking with high effort by default, and supports tool use and computer use.
$2 / $10 per 1M tokens
1M tokens
Not rated
Not ranked
Key Features
What Claude Sonnet 5.5 is
Anthropic released Claude Sonnet 5.5 on September 28, 2026. In Anthropic's own lineup it sits between Claude Opus 5.5 and Claude Haiku 4.5, described as the best combination of speed and intelligence. The headline claim is not a new capability so much as a better trade: Anthropic says it runs more than 30% faster than Claude Sonnet 5 and costs up to 30% less for most work, while the list price stays the same. That second claim is worth reading carefully. The per-token price did not move. The saving Anthropic describes comes from the model finishing typical tasks with fewer tokens, which is something you can only confirm by running your own workload through it.
Pricing, and what the 30% claim does and does not mean
Standard API pricing is $2 per million input tokens and $10 per million output tokens. Prompt cache reads cost $0.20 per million, cache writes $2.50, and the Message Batches API takes 50% off, which brings batch work to $1 and $5. Those are the same rates Claude Sonnet 5 carried, so any reduction in your bill has to come from token use rather than from the rate card. Our reading is that the fair comparison is cost per completed task, not cost per token: run the same set of real jobs through Sonnet 5 and Sonnet 5.5, count the tokens each one spent, and include retries. If the saving is real for your work it will show up there; if your tasks are already short, it may not show up at all.
Context, output and knowledge cutoff
The context window is 1M tokens, which Anthropic puts at roughly 555,000 words on its current tokenizer. Maximum output on the synchronous Messages API is 128,000 tokens; on the Message Batches API it can go to 300,000 with the output-300k-2026-03-24 beta header. Anthropic lists both the reliable knowledge cutoff and the training data cutoff as June 2026, so the model will not know about anything after that without being given it in the prompt or through a tool such as web search.
Thinking, effort and tools
Sonnet 5.5 uses adaptive thinking, where the model decides how much to reason and the effort setting steers it. The default effort on the Claude API is high, so a request that sets nothing gets the more thorough end of the dial; set effort explicitly if you want faster, cheaper answers on simple jobs. The older manual extended-thinking mode with a fixed budget is not accepted on this generation. The model takes text and image input and supports tool use and computer use, which is where Anthropic aims most of its launch material: agentic coding and work that spans several tools.
Where you can run it
It is available on the Claude API as claude-sonnet-5-5, in the Claude apps, on Amazon Bedrock as anthropic.claude-sonnet-5-5, on Google Cloud as claude-sonnet-5-5, and on Microsoft Foundry. Anthropic commits to not retiring it from its own platforms before September 28, 2027; Bedrock and Google Cloud set their own dates. It is not open-weight and cannot be self-hosted.
Anthropic's published results, and how to read them
Anthropic reports 70.6% on Terminal-Bench 4.0, 46.2% on FrontierCode 1.1 (Main, max effort), 55.5% on CursorBench 4.0, 1844 on GDPval-AA v2.1, 80.1% on OSWorld 2.1 and 61.6% on Chartography without tools. These are the vendor's own runs under its own settings, and several of these benchmarks are newer than any independent leaderboard we track, so we have not placed them alongside other models' numbers. LMArena has not rated Sonnet 5.5 yet; its Elo will appear on this page and on our benchmarks board once it does.
Key Takeaways
- Released September 28, 2026 at the same $2 / $10 per million tokens as Claude Sonnet 5.
- Anthropic says it is 30%+ faster and costs up to 30% less for most work - a saving in tokens used, not in the rate card.
- 1M-token context, 128K maximum output (300K via the batch beta header), June 2026 knowledge cutoff.
- Adaptive thinking with high effort by default; supports vision, tool use and computer use.
- On the Claude API, Bedrock, Google Cloud and Microsoft Foundry; not open-weight.
Official Resources
Full Specifications
$2 / $10 per 1M tokens | |
|---|---|
| Identity | |
| Developer | Anthropic |
| Released | Sep 2026 |
| Status | GA |
| Licence | Proprietary |
| Self-hostable | No |
| Cost | |
| Blended $/1M tokensinput × 0.75 + output × 0.25 | $4.000 / 1M tokens |
| Input price | $2.000 / 1M tokens |
| Output price | $10.000 / 1M tokens |
| Cached input | $0.20 per 1M tokens |
| Batch discount | 50% ($1 / $5 per 1M tokens) |
| Capacity | |
| Context window | 1M tokens |
| Max output | 128k tokens (up to 300k on the Message Batches API with the output-300k-2026-03-24 beta header) |
| Capability | |
| Vision in | Yes |
| Audio in | No |
| Function calling | Yes |
| Structured output | Yes |
| Extended reasoning | Yes - adaptive thinking, default effort high |
| Web search | Yes |
| Code execution | Yes |
| Access | |
| API | Yes |
| Chat app | Yes |
| Cloud marketplaces | Amazon Bedrock, Google Cloud, Microsoft Foundry |
| Fine-tuning | No |