HR EN DE
NEWS SPORT BIZNIS SCENA LIFESTYLE TECH

Anthropic Launches Claude Opus 5: Top-Tier Performance at a Lower Cost

The new model matches the level of Anthropic's best model at a significantly lower cost, excelling in programming, scientific research, and business automation tasks.

Foto: Wikipedia (Računalno programiranje)
Summary
  • Anthropic has launched Claude Opus 5, a model that matches the level of top-tier Fable 5 at half the cost.
  • On key benchmarks for programming and automation, Opus 5 outperforms all competing models, including the previous Opus 4.8.
  • The model is Anthropic's most aligned to date with a score of 2.3 on a misaligned behavior audit, but lags behind Mythos 5 on cybersecurity tasks.
  • Early access users particularly highlight improved judgment and self-checking abilities as a key difference from previous models.

Anthropic today launched Claude Opus 5, a new large language model that, according to internal evaluations, achieves performance close to the top-tier Claude Fable 5 but at half the cost. The model is immediately available on all platforms, becomes the default model on the Claude Max subscription, and is the most powerful model on the Claude Pro tier.

Benchmark Results: Far Ahead of the Competition

On Frontier-Bench v0.1, an evaluation for software engineering, Opus 5 surpasses all other models and more than doubles the score of its predecessor Opus 4.8 at a lower cost per task. On ARC-AGI 3, a test for solving novel problems, Opus 5 achieves a score three times higher than the next best model. On Zapier AutomationBench, which measures the ability to complete business tasks end-to-end, Opus 5 achieves a pass rate about 1.5 times higher than the best competitor at the same cost. On OSWorld 2.0, a benchmark for computer usage, Opus 5 outperforms Fable 5 at just one-third of the cost.

Programming and Agentic Tasks

Anthropic particularly highlights the model's abilities in autonomously solving complex programming problems. In one Frontier-Bench task, Opus 5 was given a drawing of a machine part and the task of writing code to reconstruct it as a 3D FreeCAD model, but deliberately without direct access to the drawing. According to Anthropic's blog description, the model responded by writing its own computer vision pipeline to extract geometry from raw pixels, then reconstructed the entire machine part. No competing model with the same setup could solve the task even after five attempts. The Lovable platform reported a 22% improvement on the most difficult agentic programming tasks compared to Opus 4.7, with significantly less variability in results.

Scientific Research and Business Processes

In the domain of science, Opus 5 shows a 10.2 percentage point improvement over Opus 4.8 on organic chemistry tasks, and a 7.7 percentage point improvement on protein function prediction tasks. Users in genomics describe the model as one that "behaves more like a careful scientist than any model we've run," highlighting its ability to choose the right statistical tests and cross-check results. Box reported an 8% improvement in analyzing specialized business content, with 11% gains in data analysis and 17% in due diligence processes compared to Opus 4.8.

Efficiency and Cost-Effectiveness

One of the key advantages of Opus 5 is its drastically reduced computational resource consumption while achieving better results. Users in financial modeling reported an average of 9 percentage points higher accuracy with one-third fewer turns and tool calls, and 60% less time compared to Opus 4.8. On trading benchmarks, the model uses roughly one-seventh the tokens for reasoning with less than half the latency of its predecessor. Legal users note a 26% reduction in the average number of generated tokens when reasoning at maximum capacity, with retained or improved accuracy.

Safety and Alignment

Anthropic states that Opus 5 is their most aligned model to date, scoring 2.3 on an automated audit of misaligned behavior, the lowest among all recent models. The model shows the lowest rates of deceptive behavior and the lowest risk of misuse compared to Opus 4.8, Sonnet 5, and Fable 5. In cybersecurity, Opus 5 lags behind Mythos 5 on cybersecurity tasks, and the model's pricing remains the same as its predecessor Opus 4.8.

Judgment as a Key Differentiator

Early access users consistently highlight improved judgment as the most important new feature. JetBrains noted in its evaluation: "It thinks harder before writing a single line, catches its own logical errors during planning instead of after the fact, and considers why an answer is correct, not just whether it is functional." One developer described an instance where Opus 5 challenged a proposed architectural design, did not back down under pressure, but precisely explained the value of the original proposal and offered a compromise solution that retained the benefits while correcting the flaw.

FAQ
What is Claude Opus 5 and how does it differ from previous models? +
Claude Opus 5 is a new large language model from Anthropic that achieves performance close to the top-tier Fable 5 model, but at half the cost, with significantly improved capabilities in programming, judgment, and scientific research compared to the previous Opus 4.8.
On which platforms is Claude Opus 5 available? +
Opus 5 is available immediately on all Anthropic platforms, including the Claude API, and becomes the default model on the Claude Max subscription and the most powerful model on the Claude Pro tier.
Is Claude Opus 5 safe for use in sensitive domains? +
Anthropic states that Opus 5 is their most aligned and safest model to date, with the lowest rates of deceptive behavior, but on cybersecurity tasks it falls behind the specialized Mythos 5 model.
How does Opus 5 perform on business process automation tasks? +
On the Zapier AutomationBench evaluation, Opus 5 achieves a pass rate about 1.5 times higher than the best competitor at the same cost, and in testing by Zapier it achieved 100% pass rate on customer churn prevention tasks where previous models failed.

Log in

You need to log in or register to comment.

Comments (0)
No comments yet. Be the first!
Search
Popular
Login
Home
Make 5MIN.hr a preferred source
Follow us on social media
Categories