
Image to help understand the article
Anthropic cuts costs with its first 5.5 model
Anthropic unveiled Claude Opus 5.5 on Sept. 22, introducing the first model in its 5.5 series with lower prices and faster output than its predecessor.
The company said Opus 5.5 delivers performance comparable to Claude Fable 5.1 on most tasks while costing 40% less to run than Opus 5. It also generates output more than 30% faster, according to Anthropic.
API pricing is $4 per million input tokens and $20 per million output tokens, down from $5 and $25, respectively, for Opus 5. Tokens are units of text processed or generated by an artificial intelligence model.
Claude Opus 5.5 is available through Anthropic's Claude platform as well as Amazon Web Services, Google Cloud and Microsoft Azure. Its API model identifier is claude-opus-5-5.
Benchmark results show a close contest
Anthropic's FrontierCode v1.1 results put Opus 5.5 at 54.4%, narrowly ahead of OpenAI's GPT-6 Astra at 53.3%. Anthropic said Opus 5.5 exceeded Astra's best score using its standard medium-effort setting while costing about one-fifth as much per task.
For Terminal-Bench 4.0, Anthropic said Opus 5.5 matched GPT-6 Astra at roughly 40% of the cost. At the highest effort setting shown in the company's table, Opus 5.5 scored 66.4%, compared with Astra's 57.9%.
Astra led in other evaluations. It scored 64.6% on Terminal-Bench-Science 0.1, compared with 58.7% for Opus 5.5. On AutomationBench, Astra scored 41.4% and Opus 5.5 scored 40.0%. Opus 5.5 posted 57.8% on CursorBench 4.0, compared with 41.7% for GPT-5.6 Sol; no Astra score was provided for that test.
The companies' benchmark figures have not been independently reproduced. In direct testing by independent evaluator Vals AI, the two models were nearly even, with Astra slightly ahead. Digital Applied advised treating gaps of less than 5 percentage points as inconclusive without direct testing. It found that Opus 5.5 generally led on price and shared benchmarks, while Astra retained an advantage in automation and science tasks.
Safety gains come with acknowledged limits
Anthropic said Opus 5.5 received the highest score yet in its internal audit of autonomous behavior. Attempts to act outside permitted boundaries fell by about 85% from Opus 5, the company said, while Opus 5.5 and Fable 5.1 recorded the lowest prompt-injection success rates. Safeguards route cybersecurity and biology-related tasks to other models.
The company also acknowledged that AI safety testing remains incomplete, saying that building evaluations capable of reliably identifying every failure before deployment is still an unsolved problem. Anthropic said it also observed signs that Opus 5.5 often suspects when it is being evaluated, a behavior that can complicate assessments of how a model will perform under ordinary conditions.
Release follows call to pace AI development
Opus 5.5 is Anthropic's first model release since CEO Dario Amodei called for managing the pace of AI capability gains so that safeguards have time to catch up. OpenAI gave GPT-6 Astra a limited release Sept. 3, while Anthropic introduced Claude Fable 5.1 and Claude Mythos 5.1 earlier in September.
Comments
Post a Comment