LIVE
News

Anthropic Unveils Claude Opus 5: Performance and Efficiency Gains

Anthropic has released Claude Opus 5, positioning the model as the new default on Claude Max and the strongest offering on Claude Pro.

Shane Barrett·updated July 29, 2026

Anthropic Unveils Claude Opus 5: Performance and Efficiency Gains

Per Anthropic's announcement, the model delivers state-of-the-art performance on coding and knowledge work evaluations at half the cost of Claude Fable 5, while maintaining the same price point as its predecessor, Opus 4.8. The release shifts the cost-performance frontier on several established benchmarks and introduces a configurable effort-setting mechanism that trades token consumption against task quality.

Benchmark Results

On Frontier-Bench v0.1, Opus 5 surpasses all previously evaluated models and more than doubles Opus 4.8's performance at a lower per-task cost. On CursorBench 3.2 at max effort, the model lands within 0.5 percentage points of Fable 5's peak score while operating at half the cost per task; at high, xhigh, and max effort settings, Opus 5 achieves greater performance at a given cost than any competing model in Anthropic's comparison set. The model also posts a new state-of-the-art on GDPval-AA.

On scientific research tasks, Opus 5 outperforms Opus 4.8 across all of Anthropic's life sciences evaluations. The largest gains appear on organic chemistry—specifically molecular structure inference from spectroscopy data, where the model scores 10.2 percentage points higher than Opus 4.8—and on protein-function prediction from sequence variation, where it scores 7.7 percentage points higher. The model also produces stronger visual outputs and demonstrates greater capacity for self-verification and iterative refinement.

Safety and Alignment

Anthropic's automated behavioral audit identifies Opus 5 as the most aligned model in its lineup to date. It adheres to the Claude Constitution more closely than Opus 4.8, Sonnet 5, or Fable 5; exhibits the lowest rates of deceptive behavior; and is the least susceptible to being manipulated into misuse. The model is also flagged as the safest in avoiding reckless actions with hard-to-reverse side effects.

On dual-use capability evaluations conducted with private-sector and government partners, Opus 5 does not advance the frontier in risky domains. It remains behind Mythos 5 on both biology research and offensive cybersecurity, though Anthropic notes that the deliberate exclusion of cyber-task training—consistent with the policy applied to Opus 4.8—has still yielded substantial improvements from general capability gains, bringing Opus 5 close to Mythos 5 on vulnerability discovery.

Implementation Considerations

The effort-setting parameter functions as a tunable inference knob: customers can select higher effort to maximize intelligence or lower effort to conserve tokens for faster, cheaper runs. Independent verification of the cited benchmarks—particularly the CursorBench 3.2 and Frontier-Bench v0.1 gains—remains outstanding. Practitioners evaluating the model should track third-party reproductions and the forthcoming System Card for full methodology. For broader infrastructure context around model deployment and evaluation, IEEE's virtual training course on large language models outlines the current professional training landscape.