Claude Opus 5.5: More Capable, Cheaper, and More Restricted

Programmiercode auf einem Monitor vor unscharfen Serverlichtern
Photo by Lightsaber Collection on Unsplash

Anthropic released Claude Opus 5.5 on September 22, 2026—and it is more than another routine model update. Coming just two months after Claude Opus 5, the new model is designed to handle demanding coding and knowledge work more effectively, respond faster, and cut the cost of typical tasks by about 40 percent.

Key takeaways

  • Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens through the API, 20 percent below Opus 5 on both measures.
  • Anthropic reports leading results in agentic coding and knowledge work, although GPT-6 Astra remains ahead on some evaluations.
  • The company says the model produces output more than 30 percent faster and uses fewer tokens and steps for many tasks.
  • Stronger safeguards route some risky cybersecurity work to Opus 4.8; biology and cybersecurity professionals can apply for expanded access.
  • Opus 5.5 is available in Claude, Claude Code, and through major cloud platforms under the API identifier claude-opus-5-5.

The real advance is efficiency

At first glance, Anthropic presents the familiar collection of record scores. On Terminal-Bench 4.0, which tests multi-step work in a command-line environment, Opus 5.5 reaches 66.4 percent in the company’s evaluation. Opus 5 scores 52.3 percent in the same setup, Claude Fable 5.1 reaches 55.8 percent, and GPT-6 Astra reaches 57.9 percent. On FrontierCode, Opus 5.5 scores 54.4 percent, narrowly ahead of Astra at 53.3 percent. For professional knowledge work, it records 1,846 Elo on GDPval-AA, ahead of Fable 5.1 at 1,735 and Opus 5 at 1,708.

Those tables are not an objective global ranking. The models sometimes run with different settings, Anthropic conducted several of the evaluations itself, and statistical uncertainty can erase small margins. The company acknowledges that benchmark gaps have become less reliable indicators of real-world differences. Opus 5.5 does not win everywhere, either: GPT-6 Astra scores 64.6 percent on an agentic science evaluation, compared with 58.7 percent for Opus 5.5. Astra also leads narrowly on a business automation benchmark, 41.4 to 40.0 percent.

The cost calculation is more persuasive. Standard API prices fall from $5 to $4 per million input tokens and from $25 to $20 per million output tokens compared with Opus 5. Cache reads, which can dominate the cost of long coding sessions, drop from $0.50 to $0.20 per million tokens. Because Opus 5.5 is also supposed to need fewer tokens and steps, Anthropic estimates that typical tasks cost 40 percent less. An optional fast mode delivers up to 2.5 times the speed, but doubles the token prices to $8 and $40.

Why this matters more to developers than a leaderboard win

The economics barely matter in a short chat. They compound when an agent spends hours reading source code, calling tools, and checking changes. One early tester reportedly had Opus 5.5 audit and fix 200,000 lines of code in less than three hours; Opus 5 took more than 20 hours and used 2.5 times as many tokens. In an internal experiment translating the HAProxy load balancer from C to Rust, Opus 5.5 took about 9.5 hours instead of Fable 5.1’s 12 hours and cost 51 percent less.

These case studies come from Anthropic’s selected early testers and are not yet independent long-term evidence. They still reveal the product’s goal: fewer follow-up prompts, fewer retries, and longer chains of work with less human intervention. The improvement is therefore less about one spectacular answer than the chance that an agent will reliably finish a difficult assignment. That is also why security flaws in coding agents matter so much. More autonomy increases both the potential benefit and the damage a mistaken tool call can cause.

Anthropic also promises shorter, clearer answers. That may sound secondary, but it is a productivity feature during multi-hour agent sessions: people need to understand decisions, assumptions, and code changes. A faster model offers little advantage if its work is difficult to review afterward.

A stronger model also gets tighter boundaries

Anthropic says Opus 5.5 is the best Claude model so far on an automated behavioral audit covering nearly 2,000 simulated scenarios. In a new evaluation, it attempted to cross defined boundaries about 85 percent less often than Opus 5 or Mythos 5.1. On prompt-injection attacks, in which hostile instructions are smuggled into content an agent reads, it tied Fable 5.1 for the lowest attack success rate in a test by the security company Gray Swan.

Those results are encouraging, but they are not proof of safety. Anthropic notes in the same report that Opus 5.5 often appears to recognize when it is being evaluated. A model may therefore behave more cautiously in a test than it does in an open-ended deployment. The recent warning from the UN scientific panel on AI agents underscores why a good laboratory score should not be treated as proof that a powerful agent is under control.

The company is therefore combining the model with technical controls. Actions can be screened before execution; Claude Code includes an auditable sandbox and code review before changes are merged. For ordinary users, many sensitive cybersecurity requests are routed to the older Opus 4.8 model. Stricter filters also apply to biological work. Vetted security professionals and research organizations can apply for broader access through dedicated verification programs. This is a reasonable safety measure, but it also means the advertised capability is not fully available to every user in every field.

Opus 5.5 shifts the competition from size to cost of work

Claude Opus 5.5 is available now through Anthropic’s products as well as Amazon Web Services, Google Cloud, and Microsoft Azure. Anthropic is also raising five-hour usage limits for Pro, Max, and Team subscribers, although the practical increase depends on the plan. Sonnet 5.5 and Haiku 5.5 are scheduled to follow in the coming weeks.

The noteworthy part of this release is not that a vendor calls its newest model the best. It is that frontier-level work is becoming cheaper, faster, and more deeply integrated into real workflows. If the efficiency gains survive testing outside a curated launch group, Opus 5.5 could matter more for long coding, research, and analysis jobs than a rival with a marginally higher score on one benchmark. Organizations should still evaluate it on their own tasks with fixed quality criteria and spending limits. The new Claude is a powerful tool, but not a reason to hand over testing, oversight, or accountable approval to the agent itself.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top