Anthropic Launches Claude Sonnet 5.5 With Lower Costs and Stronger Coding Scores

Anthropic Launches Claude Sonnet 5.5 With Lower Costs and Stronger Coding Scores

Image to help understand the article

Anthropic unveiled Claude Sonnet 5.5 on Sept. 28, 2026, positioning the artificial intelligence model as a faster and more cost-efficient upgrade to its previous Sonnet offering.

The company said Sonnet 5.5 can surpass Claude Sonnet 5’s best benchmark results while operating at low or medium effort and costing roughly one-tenth as much per task. The model is available under the API identifier claude-sonnet-5-5.

Pricing and availability

Claude Sonnet 5.5 is available through the Claude Platform and major U.S. cloud providers Amazon Web Services, Google Cloud and Microsoft Azure. Anthropic also offers configurations in which customer data is not retained.

API pricing is $2 per million input tokens and $10 per million output tokens. Cached input reads cost 20 cents per million tokens, while cache writes cost $2.50. VentureBeat reported that the pricing matches GPT-6 Sol and is half the rate of Claude Opus 5.5, which costs $4 per million input tokens and $20 per million output tokens.

SiliconANGLE reported that Sonnet 5.5 generates output 30% faster than the previous generation and requires fewer tokens, reducing practical costs by about 30%.

Benchmark gains come with trade-offs

In results released by Anthropic, Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0, compared with 10.3% for Sonnet 5 and 66.4% for Opus 5.5.

On FrontierCode 1.1 Main, Sonnet 5.5 scored 46.2% at its maximum effort setting. That exceeded Sonnet 5’s 42.4% but trailed Opus 5.5 at 54.4% and GPT-6 Sol at 49.3%.

Sonnet 5.5 received a score of 1,844 on GDPval-AA v2.1, nearly matching Opus 5.5’s 1,846 and exceeding scores of 1,449 for Sonnet 5 and 1,487 for GPT-6 Sol. On Humanity’s Last Exam, it scored 64.5%, ahead of Sonnet 5’s 54.9% but behind Opus 5.5’s 67.7%. Its OSWorld 2.1 score of 80.1% followed the same pattern, beating Sonnet 5’s 57% while falling short of Opus 5.5’s 81.8%.

Anthropic’s results showed that Opus 5.5 remained ahead on FrontierCode, CursorBench, GDPval-AA, AA-Briefcase and Humanity’s Last Exam. The company said Opus 5.5 is still clearly stronger on complex, open-ended work requiring sustained judgment.

Independent tests highlight cost and verbosity

Artificial Analysis gave Sonnet 5.5 an Intelligence Index score of 56, placing it third among 216 models and well above the median score of 26. The research group measured a cost of $7.60 per task.

The model generated 410 million output tokens during the evaluation, compared with a median of 88 million. Artificial Analysis characterized it as highly verbose. It measured an output speed of 141.9 tokens per second, ranking 28th, and a time to first token of 353.32 seconds.

In a separate evaluation by Vals AI, Sonnet 5.5 ranked second among 66 models with a Vals Index score of 69.22%, plus or minus 0.96 percentage points. Opus 5.5 led by 0.47 percentage points with a score of 69.69%. Sonnet 5.5 cost $20.80 per test, compared with $32.77 for Opus 5.5.

Performance varied by specialty. Sonnet 5.5 ranked 31st among 41 models on CyberBench with a score of 59.58%. It placed 36th among 70 models on Harvey’s Legal Agent Benchmark, scoring 2.92%.

Early business use and safety limits

Yashodha Bhavnani, a vice president at cloud content management company Box, told VentureBeat that Sonnet 5.5 was 2.4 times faster while using 12% fewer tokens overall. A Zendesk AI director told SiliconANGLE that the model processed customer service tickets 20% faster, helping users receive assistance without waiting.

SiliconANGLE also reported that Sonnet 5.5 became the first Sonnet model to complete the video game Pokémon Red using only screenshots.

Anthropic cautioned that biological safety controls may mistakenly block some microbiology and virology requests. SiliconANGLE reported that safeguards may activate in cybersecurity, biology and sensitive-content tasks, though most software development and life sciences work is not expected to be affected.

Source: Original Korean article - Trendy News Korea

Comments