Anthropic Releases Claude Sonnet 5.5: 70.6% on Terminal-Bench 4

Anthropic Releases Claude Sonnet 5.5: 70.6% on Terminal-Bench 4.0 at the Same $2/$10 Price

By Michal Sutter β€” September 28, 2026

Anthropic has released Claude Sonnet 5.5, the second model in the Claude 5.5 family following Claude Opus 5.5. Anthropic positions Sonnet 5.5 as a faster, lower-cost complement to Opus 5.5, targeting well-scoped everyday tasks, bug fixing, and polished documents, slides, and spreadsheets.

Is it deployable? Yes. It is live on the Claude Platform as claude-sonnet-5-5, as well as on AWS, Google Cloud, and Microsoft Azure. It is a closed-weights model, so self-hosting is not an option.

What Changed Versus Sonnet 5

Anthropic reports four main upgrades over Sonnet 5:

  • Speed: Output generation is more than 30% faster, making it the fastest Sonnet to date.
  • Cost per task: Up to 30% lower, because it requires fewer tokens and tool calls.
  • Writing: Clearer prose, with early testers calling it a better collaboration partner.
  • Vision and long-horizon work: It is the first Sonnet to beat PokΓ©mon Red using only screenshots.

Specs from the models overview: 1M-token context, 128K max output, and a June 2026 reliable knowledge cutoff. Adaptive thinking is on by default. Effort runs across five levels: low, medium, high, xhigh, and max.

Benchmarks

All scores below are vendor-reported in the launch post. Methodology is detailed in the Sonnet 5.5 System Card.

  • Terminal-Bench 4.0: 70.6%, versus 10.3% for Sonnet 5 and 66.4% for Opus 5.5 at Xhigh.
  • CursorBench 4.0: 55.5%, about 2 points below Opus 5.5 (57.8%).
  • FrontierCode 1.1: 52.1% at Xhigh and 46.2% at Max. GPT-6 Sol scored 49.3%.
  • GDPval-AA v2.1: 1844, versus 1846 for Opus 5.5 and 1449 for Sonnet 5.
  • [Additional benchmark data truncated in the source β€” insert the remaining vendor-reported scores here if available.]

via MarkTechPost

Related