Anthropic Launches Claude Opus 5: Top-Tier Performance at a Lower Cost
The new model matches the level of Anthropic's best model at a significantly lower cost, excelling in programming, scientific research, and business automation tasks.
The new model matches the level of Anthropic's best model at a significantly lower cost, excelling in programming, scientific research, and business automation tasks.
Anthropic today launched Claude Opus 5, a new large language model that, according to internal evaluations, achieves performance close to the top-tier Claude Fable 5 but at half the cost. The model is immediately available on all platforms, becomes the default model on the Claude Max subscription, and is the most powerful model on the Claude Pro tier.
On Frontier-Bench v0.1, an evaluation for software engineering, Opus 5 surpasses all other models and more than doubles the score of its predecessor Opus 4.8 at a lower cost per task. On ARC-AGI 3, a test for solving novel problems, Opus 5 achieves a score three times higher than the next best model. On Zapier AutomationBench, which measures the ability to complete business tasks end-to-end, Opus 5 achieves a pass rate about 1.5 times higher than the best competitor at the same cost. On OSWorld 2.0, a benchmark for computer usage, Opus 5 outperforms Fable 5 at just one-third of the cost.
Anthropic particularly highlights the model's abilities in autonomously solving complex programming problems. In one Frontier-Bench task, Opus 5 was given a drawing of a machine part and the task of writing code to reconstruct it as a 3D FreeCAD model, but deliberately without direct access to the drawing. According to Anthropic's blog description, the model responded by writing its own computer vision pipeline to extract geometry from raw pixels, then reconstructed the entire machine part. No competing model with the same setup could solve the task even after five attempts. The Lovable platform reported a 22% improvement on the most difficult agentic programming tasks compared to Opus 4.7, with significantly less variability in results.
In the domain of science, Opus 5 shows a 10.2 percentage point improvement over Opus 4.8 on organic chemistry tasks, and a 7.7 percentage point improvement on protein function prediction tasks. Users in genomics describe the model as one that "behaves more like a careful scientist than any model we've run," highlighting its ability to choose the right statistical tests and cross-check results. Box reported an 8% improvement in analyzing specialized business content, with 11% gains in data analysis and 17% in due diligence processes compared to Opus 4.8.
One of the key advantages of Opus 5 is its drastically reduced computational resource consumption while achieving better results. Users in financial modeling reported an average of 9 percentage points higher accuracy with one-third fewer turns and tool calls, and 60% less time compared to Opus 4.8. On trading benchmarks, the model uses roughly one-seventh the tokens for reasoning with less than half the latency of its predecessor. Legal users note a 26% reduction in the average number of generated tokens when reasoning at maximum capacity, with retained or improved accuracy.
Anthropic states that Opus 5 is their most aligned model to date, scoring 2.3 on an automated audit of misaligned behavior, the lowest among all recent models. The model shows the lowest rates of deceptive behavior and the lowest risk of misuse compared to Opus 4.8, Sonnet 5, and Fable 5. In cybersecurity, Opus 5 lags behind Mythos 5 on cybersecurity tasks, and the model's pricing remains the same as its predecessor Opus 4.8.
Early access users consistently highlight improved judgment as the most important new feature. JetBrains noted in its evaluation: "It thinks harder before writing a single line, catches its own logical errors during planning instead of after the fact, and considers why an answer is correct, not just whether it is functional." One developer described an instance where Opus 5 challenged a proposed architectural design, did not back down under pressure, but precisely explained the value of the original proposal and offered a compromise solution that retained the benefits while correcting the flaw.