Anthropic's Claude Opus 5 has demonstrated strong performance across multiple benchmarks while maintaining lower costs compared to its competitors. According to Artificial Analysis, Opus 5 achieved an Intelligence Index score of 61, placing it just ahead of Fable 5 at 60. This score combines results from nine tests measuring knowledge work, coding, scientific reasoning, and factual accuracy. Opus 5 also outperformed Fable 5 in coding tasks, achieving first place on the Artificial Analysis Coding Index when paired with Claude Code. On Terminal-Bench v2.1, it scored 89 percent at 'max,' matching the previous leader GPT-5.6 Sol.
Opus 5's performance is notable in its ability to handle complex tasks efficiently. At the 'high' and 'xhigh' performance levels, it outperformed both Opus 4.8 and Sonnet 5 while maintaining lower costs. The model's cost per task dropped by 20 percent to $17.79, down from $22.30 for Fable 5. At the 'xhigh' variant, the cost is $14.26, and at 'high,' it is $10.41, which is less than half of Fable 5's cost. However, the model's hallucination rate climbed to 50 percent, raising concerns about reliability in high-stakes applications.
Artificial Analysis worked with Anthropic to test the model before its public release, providing insights into its capabilities and limitations. The findings highlight the tight competition among frontier models, with no single model able to pull away or claim a clear advantage. This underscores the potential for AI models to become commoditized as the field continues to evolve. Source: thedecoder