On September 28, 2026, Anthropic introduced Claude Sonnet 5.5 and called it the second model in the Claude 5.5 family. The page says it is a clear upgrade over Claude Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.
Sonnet 5.5 is described as a faster, lower-cost complement to Claude Opus 5.5. Opus 5.5 is for complex work that needs careful judgment. Sonnet 5.5 is said to be strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets. The page also says it has a sharp eye for design. Claude Haiku 5.5, for high-volume and cost-sensitive applications, is planned to join the family in the coming weeks. No date is given.
[1]
On Terminal-Bench 4.0, an agentic coding evaluation, Sonnet 5.5 scores 70.6% and Sonnet 5 scores 10.3%. The same table lists Opus 5.5 at 66.4%. The footnote says that Opus result is at Xhigh effort, the model’s highest reported score. It is not a same-effort comparison.
On GDPval-AA, the page says Sonnet 5.5 scores two points below Opus 5.5. In the table the order is Sonnet 5.5, Sonnet 5, then Opus 5.5: 1844, 1449, and 1846. Footnote 3 says Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release Claude Platform deployment. Anthropic found a bug there that could degrade requests using structured outputs. The company expects any effect to be small and to understate the score. The bug has since been fixed. This piece does not use the GPT comparison scores. Footnote 4 says some of those comparisons may not yet reflect the latest model.
On OSWorld 2.1 the table is labeled partial: 80.1% for Sonnet 5.5, 57.0% for Sonnet 5, and 81.8% for Opus 5.5. The page also says this is the first Sonnet to beat Pokémon Red using only screenshots.
Anthropic says benchmark scores capture only one facet of capability. In its own testing, and in testing by external testers, Opus 5.5 remains clearly stronger at complex, open-ended work that requires sustained judgment.
[1]The price matches Sonnet 5: $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads. The page says Sonnet 5.5 typically needs far fewer tokens for the same work. In Anthropic’s testing it costs up to 30% less per task than its predecessor. It generates output more than 30% faster than Sonnet 5, which the page calls the fastest Sonnet so far.
Its cybersecurity capabilities are described as comparable to Opus 5, not to Opus 5.5. It is the first Sonnet to launch with cyber safeguards and fallbacks like those on the company’s most capable models. Users can still find and fix bugs as part of routine software development. Higher-risk cybersecurity tasks visibly fall back to Sonnet 5. Biology safeguards are the same as Sonnet 5’s. Both sets target a narrow band of high-risk requests. Routine software development and most life sciences work are unaffected.
On an automated behavioral audit of roughly 1,850 scenarios, Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment, resistance to misuse, and honesty. The page also says no set of evaluations reliably catches every failure, and that Sonnet 5.5 may have tendencies the company has not found.
Like Opus 5.5 and Sonnet 5, Sonnet 5.5 is available with zero data retention. It is available on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. The Claude Platform model id is claude-sonnet-5-5. Anyone running Sonnet with thinking off must switch to between_tools, which keeps up-front thinking off, before moving to Sonnet 5.5.
[1]The page does not present Sonnet 5.5 as ahead of Opus 5.5 across the board. The Terminal-Bench cell above 66.4% compares against Opus at its highest effort. The GDPval-AA and AA-Briefcase scores come from a pre-release deployment, before the structured-output bug was fixed. Anthropic says that bug would understate the score. This piece did not rerun the tests, and it does not use early-tester quotations or the system card.
[1]要点
- Sonnet 5.5 is more than 30% faster than Sonnet 5 and up to 30% cheaper for most work, at the same list price.
- Terminal-Bench 4.0 is 70.6% versus 10.3%. The table’s 66.4% for Opus 5.5 is its highest-effort score.
- GDPval-AA is two points below Opus. A pre-release structured-output bug has been fixed and may have understated the score.
- Higher-risk cybersecurity tasks fall back to Sonnet 5. Haiku 5.5 is only “the coming weeks,” with no date.