Start your day with intelligence. Get The OODA Daily Pulse.
Artificial Analysis evaluations reveal that Claude Sonnet 5.5 captures the #2 spot on the Intelligence Index, scoring just two points behind Opus 5.5 when running at maximum reasoning effort. The model achieves near-parity with leading frontier models on complex knowledge work and terminal-based agent tasks, scoring 64% on Terminal-Bench 4.0. To reach these performance levels, Sonnet 5.5 utilizes the highest output token volume measured by the platform averaging roughly 193,000 output tokens per task. Despite its heavy token consumption, Anthropic has kept pricing identical to its predecessor at $2 per million input tokens and $10 per million output tokens. While matching top-tier models in coding and agentic execution, the mid-tier class model still trails Opus 5.5 on broad factual knowledge and advanced scientific reasoning benchmarks.
Full benchmarking : Anthropic’s latest model achieves parity with larger frontier systems on agentic workflows and coding tasks, though risks increased costs due to high token consumption.