Claude Sonnet 5.5 trades token efficiency for agentic performance
Anthropic's new model reaches second place on the Intelligence Index by using significantly more output tokens, matching top-tier agents in terminal tasks while lagging in factual knowledge.
Anthropic has released Claude Sonnet 5.5, a mid-tier model that now ranks second on the Artificial Analysis Intelligence Index. Published on September 28, 2026, this update shows the model closing the gap with flagship competitors in agentic workflows, though it achieves these gains through unusually high token consumption rather than raw architectural efficiency.
What happened
Claude Sonnet 5.5 moves to the number two spot on the Intelligence Index, trailing only Opus 5.5 at maximum effort. The model gains 18 points over its predecessor, Sonnet 5, primarily driven by massive improvements in terminal use and complex knowledge work. In Terminal-Bench 4.0, which evaluates how well models handle command-line interactions, Sonnet 5.5 scores 64 percent. This places it slightly ahead of both Opus 5.5 and GPT-6 Astra, which each score 60 percent.
Despite these performance jumps, the model does not lead in all categories. It remains behind Opus 5.5 in factual knowledge and scientific reasoning. On the AA-Omniscience benchmark for factual accuracy, Sonnet 5.5 scores 54 percent compared to 66 percent for Opus 5.5. However, it does exhibit a lower hallucination rate of 47 percent versus 59 percent for the larger model. The evaluations were conducted on a pre-release version that contained a bug affecting structured outputs, which Anthropic states is fixed for the public release.
How it works
The core mechanism behind Sonnet 5.5’s improved ranking is its "max effort" setting, which triggers an extreme increase in output token usage. To complete tasks in the Intelligence Index, the model generates approximately 193,000 output tokens per task. This is the highest token usage recorded by Artificial Analysis, exceeding Opus 5.5 and Sonnet 5 by about 60 percent and surpassing GPT-6 Astra by roughly seven times. The model essentially brute-forces complex problems by generating extensive chains of thought and verification steps.
Anthropic offers five effort settings: low, medium, high, xhigh, and max. The Intelligence Index evaluations ran across all five levels with default fallbacks enabled. The model fell back to the previous Sonnet 5 version in only 0.1 percent of tasks, mostly within Terminal-Bench 4.0. While the max effort setting drives the high leaderboard position, lower effort settings are less competitive. At low, medium, and high efforts, GPT-6 Sol configurations deliver equivalent or better performance with fewer output tokens.
Key details
- Ranking: Second on the Artificial Analysis Intelligence Index, behind Opus 5.5 (max).
- Token Usage: Uses ~193k output tokens per task at max effort, the highest measured to date.
- Pricing: Identical to Sonnet 5 at $0.2/$2/$10 per 1M cache input/input/output tokens.
- Cost Per Task: Approximately $7.60 per task, which is 50 percent higher than Sonnet 5.
- Terminal Performance: Scores 64% on Terminal-Bench 4.0, outperforming Opus 5.5 and GPT-6 Astra.
- Context Window: Remains unchanged at 1 million tokens for text and image input.
Why it matters
For engineering teams building agentic systems, Sonnet 5.5 presents a trade-off between capability and cost efficiency. The model matches or exceeds leading competitors in practical terminal use and automation benchmarks like AA-Briefcase and AutomationBench-AA. This makes it a strong candidate for workflows requiring deep interaction with development environments or complex data processing pipelines. However, this performance comes with a significant price tag. The cost per task is roughly $7.60, making it substantially more expensive to run than its predecessor for the same volume of work.
The positioning of Sonnet 5.5 also highlights a shift in how mid-tier models compete. Rather than trying to beat flagship models on pure reasoning density, Anthropic has optimized Sonnet 5.5 to spend more compute and tokens to reach similar outcomes. For developers, this means the choice between Sonnet 5.5 and models like GPT-6 Sol depends heavily on the specific task. If token budget is tight, GPT-6 Sol may offer better value at lower effort levels. If the task requires maximum agentic persistence in a terminal environment, Sonnet 5.5 at max effort becomes a viable, albeit costly, alternative to Opus 5.5.
What you can do
- Benchmark your current agentic workflows against Sonnet 5.5 at max effort to see if the 64% Terminal-Bench score translates to real-world reliability.
- Compare the cost per task of $7.60 against your existing budget for Sonnet 5 or GPT-6 Sol to determine financial viability.
- Test the five effort settings individually to find the optimal balance between token usage and accuracy for your specific use case.
- Monitor hallucination rates if factual accuracy is critical, noting that Sonnet 5.5 has a lower rate than Opus 5.5 but also lower overall factual scores.
- Verify structured output handling in your applications, ensuring you are running the public release where the pre-release bug has been fixed.
- Evaluate whether the 1 million token context window meets your needs, as this specification has not changed from the previous version.
