Anthropic Publishes Claude Sonnet 5.5 Model Report

Image: anthropic.com
Anthropic has published a model report for Claude Sonnet 5.5 on its transparency hub, dated October 2, 2026, summarising the model's safety evaluations and deployment safeguards. The full detail is in the Claude Sonnet 5.5 system card. Claude Sonnet 5.5 was released in September 2026.
Political even-handedness
Anthropic tested 1,350 pairs of requests on 150 political topics, each pair asking for the same thing from opposite political perspectives. Claude Sonnet 5.5 scored 97.9% through the API and 99.0% on claude.ai. On acknowledging opposing viewpoints, Anthropic reports that through the API Sonnet 5.5 did this slightly less often than Claude Sonnet 5, at 41.4% against 45.7%.
Honesty
Across eight categories of misleading behaviour, Claude Sonnet 5.5 did better than Claude Sonnet 5 on seven. On factual questions it got more questions right than Sonnet 5, but also gave slightly more wrong answers. When it had made hidden changes, it mentioned them 96.2% of the time, about as often as Claude Opus 5.5. Anthropic also ran MASK, a public test of whether a model will say something it believes is false when a user or its instructions pressure it to.
Safeguards and capability thresholds
Anthropic deployed Claude Sonnet 5.5 with the same chemical and biological misuse protections it deployed for Claude Opus 5, treating it as having CB-1 capabilities, and states that autonomy threat model 1 is applicable to the model. Anthropic reports that Sonnet 5.5 is not more capable than its most advanced models and is less capable than Claude Opus 5.5 on most of its tests. On Anthropic's capability index, which combines many tests into one score, Sonnet 5.5 scored 167.93 against 169.12 for Opus 5.5.
Related AI News
Enjoyed this? Get more in your inbox.
Weekly AI breakthroughs, tool reviews, and practical guides.