Anthropic released Claude Sonnet 5.5 on September 28, 2026. The company positions it as a faster, lower-cost complement to Claude Opus 5.5 for well-scoped everyday work, bug fixes and coding tasks.
What Anthropic announced
- Sonnet 5.5 is available on the Claude Platform, AWS, Google Cloud and Microsoft Azure using the model name
claude-sonnet-5-5. - Anthropic says it produces output 30% or more faster than Sonnet 5 and is priced at $2 per million input tokens and $10 per million output tokens.
- The company says it typically uses fewer tokens per task, claiming up to 30% lower task cost than Sonnet 5 in its testing.
- For agentic terminal coding, Anthropic reports a 70.6% score on Terminal-Bench 4.0, compared with 10.3% for Sonnet 5. Benchmark results are useful signals, not a substitute for a team’s own regression suite.
Why this matters for QA engineers
Faster output and fewer tool calls can shorten the feedback loop for bounded tasks such as generating test data, investigating a failed build or proposing a small test repair. But a model upgrade can also alter JSON schemas, tool-call patterns, retries and assertions. Treat Sonnet 5.5 as a version change: replay representative test-agent tasks, compare pass rates and false-positive rates, and keep approval gates around code or test changes.
Migration checks for automation teams
- Pin the model identifier in a staging environment and run a baseline suite for structured-output, browser, API and CI-triage tasks.
- Measure end-to-end latency, token use, tool failures and human-review findings against the current Sonnet 5 workflow.
- If a workflow disabled thinking, review Anthropic’s extended-thinking guidance: the announcement says teams need the
between_toolssetting before moving to Sonnet 5.5. - Exercise denied and higher-risk security requests. Anthropic says some cybersecurity requests can visibly fall back to Sonnet 5, which is important behavior to capture in integration tests.
Sonnet 5.5 also launches with cybersecurity safeguards and fallbacks similar to those used for Opus 5.5, while routine software development is intended to remain unaffected. That makes explicit fallback and refusal-path testing worthwhile for teams using the model in developer tooling.
