Building with AI

Anthropic releases Claude Haiku 5.5 with new pricing and tokenization changes

Anthropic launched Claude Haiku 5.5, matching OpenAI's GPT-6 Luna pricing for short contexts but introducing a less efficient tokenizer and higher costs for long inputs.

A glass cube with glowing circuits next to priced paper notes
Illustration generated for this article

Anthropic has released Claude Haiku 5.5, a fast and low-cost language model designed to compete directly with recent offerings from OpenAI. The update arrives alongside significant changes to pricing structures, tokenization efficiency, and subscription benefits for API users. This release marks a strategic shift in how Anthropic positions its entry-level model against competitors like GPT-6 Luna.

What happened

The previous iteration, Haiku 4.5, had become relatively expensive compared to market alternatives. It was priced at $1 per million input tokens and $5 per million output tokens. This rate was ten times higher than OpenAI’s GPT-6 Luna, which launched just last month. Haiku 5.5 addresses this gap by exactly matching Luna’s base pricing of $0.10 per million input tokens and $0.50 per million output tokens. However, this competitive rate only applies to contexts up to 100,000 tokens.

Beyond the 100,000-token threshold, the price for Haiku 5.5 increases fivefold to $0.50 for input and $2.50 for output. In contrast, GPT-6 Luna maintains lower rates for longer contexts, only increasing to $0.20/$0.75 after 272,000 tokens. This structure suggests that while Haiku 5.5 is cost-competitive for standard tasks, it becomes significantly more expensive for large-scale data processing or long-context applications compared to its primary rival.

Another critical change involves the underlying tokenizer. Haiku 5.5 uses a new, less generous tokenizer than its predecessor. Testing indicates that the same long prompt consumes approximately 1.25 times more tokens in Haiku 5.5 than it did in Haiku 4.5. This inefficiency acts as a hidden price increase, meaning users may pay more for the same content even within the cheaper tier. For workloads that stay under 100,000 tokens, Haiku reports higher benchmark scores than Luna, but for larger inputs, Luna appears to be the more economical choice.

How it works

Haiku 5.5 introduces reasoning capabilities that cannot be disabled, defaulting to a medium effort level. Users can adjust the thinking effort between low, medium, high, xhigh, and max. Lower effort settings produce faster results with minimal cost, while higher settings engage deeper reasoning traces. For example, generating an SVG image of a pelican on a bicycle took seven seconds and cost roughly 0.09 cents at low effort. The same task at max effort took over five minutes but still cost only about 3.38 cents, demonstrating the trade-off between latency and computational depth.

The model also integrates with updated developer tools. The llm-anthropic plugin now supports dynamic model refreshes, eliminating the need for new software versions for every model release. This allows developers to test new capabilities immediately using command-line interfaces. The reasoning trace feature provides visibility into the model’s decision-making process, which can be useful for debugging complex outputs or understanding benchmark performance.

Key details

  • Haiku 5.5 matches GPT-6 Luna pricing at $0.10/$0.50 per million tokens for inputs up to 100,000 tokens.
  • Prices jump fivefold to $0.50/$2.50 per million tokens for contexts exceeding 100,000 tokens.
  • The new tokenizer is less efficient, using 1.25x more tokens than Haiku 4.5 for the same text.
  • Reasoning cannot be disabled; it defaults to medium effort and supports levels up to max.
  • Anthropic is halving the price of cache reads for Sonnet 5.5.
  • Max and Team subscribers now receive monthly API credits equal to their subscription cost.

Why it matters

For software engineers building AI-powered features, the pricing tier break at 100,000 tokens is a critical architectural constraint. Applications that frequently process long documents or maintain extensive conversation histories will face steep cost penalties with Haiku 5.5 compared to GPT-6 Luna. Developers must carefully monitor context window usage to avoid unexpected billing spikes. The hidden cost of the new tokenizer further complicates budget forecasting, requiring teams to re-evaluate token consumption metrics from previous models.

The introduction of non-disableable reasoning changes the performance profile of the model. While this may improve accuracy for complex logic tasks, it introduces variable latency that can impact user experience in real-time applications. Teams relying on predictable response times will need to test different effort levels to find the right balance between speed and quality. The inability to turn off reasoning means that even simple queries may incur slightly higher computational overhead than in previous versions.

What you can do

  • Audit your current token usage to determine if most requests fall below the 100,000-token threshold.
  • Benchmark Haiku 5.5 against GPT-6 Luna for your specific long-context workflows to compare total costs.
  • Test different reasoning effort levels to identify the optimal setting for your application’s latency requirements.
  • Update your llm-anthropic plugin to the latest version to support dynamic model switching without redeployment.
  • Review your Anthropic subscription tier to claim the new monthly API credits via the Billing settings.
  • Configure auto-reload disabling if you plan to use up your monthly API credits to prevent overage charges.

Tools from the Bytechap store

$89

DocBento

Self-hosted document management that reads every scan and answers with page citations.

Live demo

Keep reading

All stories