Claude Opus 5.5 takes the intelligence crown — and cuts its own price
Anthropic's new flagship posts the highest score Artificial Analysis has ever measured while listing 20 per cent cheaper than the model it replaces — and OpenAI's GPT-6 Sol answers by undercutting it at half the price.
Anthropic released Claude Opus 5.5 on 22 September and, on Artificial Analysis's Intelligence Index, it went straight to the top — a score of 58, the highest the benchmarking outfit says it has measured, and well clear of a field whose median sits at 25. What made the launch unusual was the price tag underneath the score: $4 per million input tokens and $20 per million output, down from the $5 and $25 that Opus 5 charged in July. A new flagship that is both more capable and cheaper than the old one breaks the usual pattern, where the best model is also the dearest.
The competitive pressure is easy to read once the numbers are laid out. OpenAI's GPT-6 Sol lists at $2 / $10 — half the Opus 5.5 rate — and on Artificial Analysis's own cost-per-task figures the gap narrows further: Opus 5.5 scored 51.2 at medium effort for about $1.34 a task, against Sol's 47.5 at maximum effort for $1.06, or 39.8 at its default setting for just $0.25. Grok 4.7, from Elon Musk's xAI, sits on 46. The labs are no longer only racing on capability; they are racing on what a unit of that capability costs to run.
The frontier AI model race has entered its comparison-shopping phase. Ars Technica
For anyone building on these models, that shift is the story. A year ago the calculus was which model was smart enough; increasingly it is which one is smart enough at a price that survives contact with a real workload. Anthropic's answer is to fold in cheaper cache reads — $0.20 per million tokens, down from $0.50 — and to keep a pricier fast-mode tier for those who want lower latency; OpenAI's is to lead with a headline rate that looks almost disposable at the default setting.
The comparisons carry an asterisk worth keeping in view. Artificial Analysis has just moved its Intelligence Index to a new version and reweighted the components, so scores from before the change do not line up cleanly with the ones after it, and rival labs quote different leaderboards depending on which cut they favour. The list prices, though, are unambiguous — and they are all pointing the same way: down.