“Opus 5 as your daily driver, the model you hand complex work to and review when it's done,” an Anthropic spokesperson said. “Fable 5 for your most ambitious work, the days-long autonomous projects nothing could take on before... Sonnet 5 for work you run at scale, where speed and cost per call decide what ships. Haiku 4.5 for subagents and instant answers.”
Anthropic says Opus 5 posted leading results on several coding and knowledge-work evaluations. On Frontier-Bench v0.1, the model scored 43.3%, compared with 18.7% for Opus 4.8 and 33.7% for Fable 5. The company also reported that Opus 5 achieved three times the score of the next-best model on ARC-AGI 3 and exceeded Fable 5 on OSWorld 2.0 while operating at slightly more than one-third of the cost.
The company acknowledged that Opus 5 does not lead in every category. Mythos 5 remains ahead on cybersecurity and biology research, while an OpenAI model still holds the top result on one agentic coding benchmark.
Anthropic said Opus 5 performs best on tasks with clearly defined outcomes, while Fable 5 remains better suited to longer projects that require sustained reasoning across many steps.
“The evals where Opus 5 wins are bounded tasks with a specific outcome, which is where it's strongest. What those evals don't measure is duration,” the spokesperson said. “One way to put it: Opus 5 is the best tool for the jobs benchmarks can see, and Fable 5 is what you reach for when the job outruns the benchmark.”
The model includes an adjustable effort setting that allows customers to trade some performance for faster responses and lower token use. Anthropic is emphasizing that balance as enterprises pay closer attention to the cost of running AI systems at scale.
Several early customers reported efficiency gains. Harvey said Opus 5 matched the performance of Opus 4.8 at its highest reasoning setting while using 26% fewer tokens on average. Fundamental Research Lab said the model delivered nine percentage points higher accuracy on difficult financial modeling tasks while requiring about one-third fewer turns and tool calls and 60% less time.
Zapier CEO Wade Foster said Opus 5 completed a churn-prevention workflow with a perfect result without consuming more tokens than earlier Claude models. Scott Wu, CEO of Cognition, said the model approached Fable-level performance at half the cost on FrontierCode 1.1, with particular strength in debugging and identifying root causes.
Anthropic is also highlighting improvements in how Opus 5 checks and revises its own work. In one test, the model created a computer vision pipeline to reconstruct a 3D object from raw pixel data after being given no direct way to inspect the drawing. In another, it identified the cause of a software bug and repaired an edge case that had been missed in an existing community patch.
Cristian Rivera, a staff software engineer at Stripe, said he gave the model “a chief-of-staff role over my dev environments” for a weekend. “It built its own monitor, drove each box, and pulled me in only for the judgment calls.”
The model also arrives with a lighter safety framework than Fable 5. Anthropic said Opus 5 produced the lowest overall rate of misaligned behavior among the models it compared, including Opus 4.8, Sonnet 5 and Fable 5.
Anthropic intentionally limited Opus 5’s training on offensive cybersecurity tasks. The model identified vulnerabilities in 79.4% of cases on the company’s OSS-Fuzz evaluation, close to Mythos 5’s 80%, but completed only four exploit-development challenges compared with 13 for Mythos 5.
The company expects Opus 5’s cybersecurity classifiers to activate about 85% less often than those used with Fable 5. When a request triggers a classifier in Claude.ai, Claude Code or Claude Cowork, the system can route the prompt to Opus 4.8 instead.
“The model it falls back to has lower capability levels making the risk of harmful use lower as well,” the spokesperson said, adding that users are informed when the switch occurs.
Anthropic is also making Automatic Fallbacks available as an optional beta feature for API customers. The setting routes blocked requests to a less capable model so users can receive a response rather than an error.
Opus 5 is not covered by the 30-day data retention policy that applies to Fable 5 and Mythos 5. Anthropic said the model supports customers that require zero data retention.
The company also introduced a Fast mode that runs at roughly 2.5 times the standard speed for twice the base price. Other updates include automatic API fallback routing and the ability to change tools during a conversation without invalidating the prompt cache.
Developers can access the model through the Claude API under the name claude-opus-5.
This analysis is based on reporting from Venture Beat.
Image courtesy of Anthropic.
This article was generated with AI assistance and reviewed for accuracy and quality.