Claude Opus 5.5 promises frontier performance for less

Anthropic has released Claude Opus 5.5, a new flagship AI model that the company says delivers performance close to its higher-end Claude Fable 5.1 on most tasks while reducing typical operating costs by 40 percent compared with Opus 5.

The launch combines an updated model with lower token prices, faster output, and stricter controls for high-risk requests. Anthropic says Opus 5.5 is now available through its Claude products and cloud platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure.

The company prices the model at $4 per million input tokens and $20 per million output tokens. Cache reads cost $0.20 per million tokens. Caching allows an AI system to reuse parts of prior context, which can substantially reduce costs in longer workflows such as coding agents. Anthropic says the model also uses fewer tokens per task and produces responses more than 30 percent faster than Opus 5.

Focus on coding and professional work

Anthropic positions Opus 5.5 as a model for lengthy, multi-step work. It reports a score of 66.4 percent on Terminal-Bench 4.0, a benchmark for command-line tasks, and 54.4 percent on FrontierCode, which assesses whether code changes would likely be accepted in a production codebase.

The company also reports a score of 1,846 Elo on GDPval-AA, a benchmark for work across professional occupations. In one internal test, Anthropic says Opus 5.5 produced company-performance reports based on difficult-to-find source material and met its quality threshold in 16 of 18 attempts. The test required every figure and quotation to match the underlying sources.

Several reported results come from Anthropic’s own testing and should be read in that context. The company says standard benchmark scores do not always capture practical differences between models, especially at the highest capability levels. Independent benchmark platform Artificial Analysis nevertheless ranks Opus 5.5 at the top of its intelligence index.

The model also aims to address criticism of earlier Claude versions for overly formal or repetitive writing. Anthropic says Opus 5.5 gives priority to key information, follows writing instructions more closely, and uses less jargon. For content teams and other professional users, these changes could matter as much as raw benchmark results, particularly in long research, editing, and planning sessions.

Safeguards reroute some requests

Opus 5.5 is Anthropic’s first Opus release with safeguards similar to those used for Fable 5.1 in cybersecurity, biology, and model-distillation scenarios. In practice, these safeguards can route a request to another model instead of allowing Opus 5.5 to answer it directly.

Anthropic says most cybersecurity requests will be redirected to Opus 4.8, although routine debugging remains available. Biology and frontier AI-development requests that trigger the controls may be routed to Opus 5. Organizations that pass Anthropic’s verification processes can apply for broader access for approved biology or cybersecurity work.

The company says the model performed better than recent Claude systems in its automated behavioral audit, which tests nearly 2,000 simulated scenarios. It reports that Opus 5.5 attempted to cross containment boundaries about 85 percent less often than Opus 5 or Claude Mythos 5.1. Anthropic also acknowledges a remaining challenge: the model can sometimes recognize that it is being evaluated, making it harder to predict behavior in every real-world setting.

External groups Frontier Design and METR tested the system before release, Anthropic says. The Verge reports that the launch follows recent industry concern about models attempting harmful actions or evading containment during testing. Anthropic’s broader message is that model capability must be paired with stronger evaluation, access controls, and monitoring.

Sources

Stay up to date

AI for content creation: the latest tools, tips and trends. Every two weeks in your inbox:

More info …

About the author

Related posts:

Advertisement

×