Comrades, Anthropic is slashing prices again. This time it's the cheapest model in the Claude family, Haiku.
On October 7, Anthropic officially released Claude Haiku 5.5, claiming it to be its fastest, most capable, and lowest-priced small model to date.
The most outrageous part is the price, as low as $0.1 per million input tokens, a direct 90% price cut compared to the previous generation.
Performance hasn't fallen behind either. According to test results published by Anthropic, Haiku 5.5 surpasses OpenAI's GPT-6 Luna on multiple benchmarks including computer operation, knowledge work, and Agent programming, with some scores even showing a considerable gap.
Moreover, Haiku 5.5 also introduces adjustable reasoning intensity, supporting a 1 million token context window and up to 128,000 token output.
At this point, Anthropic has completed the update of the entire Claude 5.5 family in just 15 days: Opus 5.5 arrived on September 22, Sonnet 5.5 was released on September 28, and now Haiku 5.5 has officially joined.
This time, Anthropic has focused on those AI tasks with massive call volumes and extreme price sensitivity.
Multiple scores surpass GPT-6 Luna, computer operation success rate 72.4%
According to the Benchmark scores published by Anthropic, Haiku 5.5 has achieved significant improvements over the previous generation, and in some tests has even achieved better results than GPT-6 Luna.
The most notable result comes from OSWorld. This is a test measuring an AI Agent's ability to operate a real computer, requiring the model to complete cross-application, multi-step operation tasks.
On the OSWorld 2.1 offline task subset, Haiku 5.5 achieved a 72.4% success rate, while the previous generation Haiku 4.5 only had 15.7%, and GPT-6 Luna had 48.9%.
In terms of knowledge work, Haiku 5.5 scored 1620 in GDPval-AA v2.1, surpassing GPT-6 Luna's 1437. In the Agent programming test Terminal-Bench 4.0, the two scored 39.2% and 16.4% respectively.
However, there is an important detail in these data. Haiku 5.5's highest Benchmark scores often require more reasoning computation.
Haiku 5.5 introduces adjustable reasoning intensity for the first time in the Haiku series, allowing developers to control how much computational resources the model invests through the effort parameter.
Media outlet VentureBeat noticed that in the Terminal-Bench test published by Anthropic, the 39.2% score corresponds to maximum reasoning intensity. After switching to medium reasoning intensity, the score drops to about 20%, and medium intensity is exactly the default setting for Haiku 5.5.
Third-party evaluation agency Artificial Analysis also observed a similar phenomenon. In its Intelligence Index evaluation, Haiku 5.5 scored 43 points at maximum reasoning intensity and 34 points at medium reasoning intensity. Under the same evaluation, the average cost per task was about $0.21 and $0.05 respectively.
In other words, the model can exchange increased reasoning budget for better task performance, but the actual cost may also increase significantly.
In its own Terminal-Bench 4.0 test, Artificial Analysis measured scores of 33% at maximum reasoning intensity and 15% at medium intensity. Due to different evaluation settings, these numbers differ from Anthropic's official results.
This is also why when looking at small model performance, one needs to consider reasoning intensity, task completion rate, and actual cost at the same time.
On more complex Agent programming tasks, Sonnet 5.5 still has a clear advantage: its Terminal-Bench 4.0 score reaches 70.6%.
Anthropic also clearly stated that Sonnet and Opus are still more suitable for complex Agent programming tasks, while Haiku's advantages are concentrated in high-frequency, clearly scoped work.
Price directly cut to one-tenth, competing head-on with GPT-6 Luna
If performance improvements make Haiku 5.5 competitive, then price may be the most important highlight of this release. Anthropic designed two tiers of API pricing for the new model.
For requests with prompt length not exceeding 100,000 tokens, Haiku 5.5's input and output unit prices are both 90% lower than the previous generation. Above this threshold, the price is still 50% lower than Haiku 4.5.
Anthropic stated that about 90% of requests to the previous generation Haiku were within 100,000 tokens. Combining different request lengths and changes in token consumption, the company estimates that Haiku 5.5 completes the average cost of similar tasks drops by about 75%.
There is also an easily overlooked detail here — Haiku 5.5 switched to a new Tokenizer. According to the official developer documentation, the same piece of text may generate about 30% more tokens in the new model than in Haiku 4.5, with the specific increase depending on the content. Therefore, a 90% drop in API unit price does not mean the bill for every task can be reduced by 90%.
Another rival worth watching is GPT-6 Luna. OpenAI's base pricing for Luna is also $0.1 per million input tokens and $0.5 for output, and the cache read price is also consistent with Haiku 5.5's low-price tier.
However, there is a clear difference between the two companies' long-context billing thresholds.
Haiku 5.5 enters a higher price tier after prompts exceed 100,000 tokens; GPT-6 Luna only triggers long-context surcharges after input exceeds 272,000 tokens. This means that for requests between 100,000 and 272,000 tokens, Luna has an advantage in unit token price.
As for which model is ultimately more cost-effective for completing a task, it still needs to be judged based on actual token usage, inference budget, and task success rate.
Looking at the entire market, Anthropic is also putting pressure on other low-price models.
The API prices for the same period listed by VentureBeat show that Google Gemini 3.5 Flash-Lite costs $0.3 per million input tokens and $2.5 for output; Gemini 3.8 Flash costs $0.75 and $3.75 respectively.
Of course, different models have different capability positioning, caching strategies, and task performance. These numbers first reflect competition in API pricing.
The price war among small models is further escalating.
8 million calls a week, enterprises begin to make small models work for large models
What exactly is Haiku 5.5 suitable for?
The scenarios given by Anthropic include document summarization, information classification, database queries, context compression, and serving as a Subagent in complex Agent workflows.
Behind this is a very realistic problem: today's Agents are becoming increasingly complex, and a task often requires multiple model calls. If every step is handed to a large model, costs will accumulate quickly.
Financial AI company Rogo provides an example. In its workflow, a more capable large model is responsible for creating financial presentation decks. When processing a certain slide, the Haiku 5.5 Subagent enters the enterprise 10-K annual report, finds the required business revenue data, and then hands the result to the main model.
For this kind of task, the model needs to complete information search and extraction quickly and accurately. Rogo believes that Haiku 5.5 achieves a balance suitable for large-scale repeated calls among accuracy, speed, and price.
Enterprise content management platform Box also provided early test results. Compared with Haiku 4.5, Haiku 5.5's internal evaluation score increased by 11 points, and latency dropped by about 50%.
Box stated that this makes the new model suitable for analysis work that needs to run at scale, such as cost reports, financial summaries, and periodic reviews. However, Box did not disclose the complete evaluation criteria.
Asana's tests covered Agent tasks such as bug classification, project creation, and large project portfolio search. According to the results it disclosed, compared with the model currently in use, Haiku 5.5 reduced task completion latency by more than 30% and increased single-round Agent inference speed by up to 2.5 times.
Another example that better demonstrates cost value comes from AlphaSense. This company's Ask in Document feature needs to handle about 8 million calls per week, mainly answering specific questions about one or several documents.
Among 400 test queries, Haiku 5.5's internal evaluation score reached 0.84, while Haiku 4.5 was 0.76.
If a model is called millions of times per week, even if it saves only a small amount of money each time, the long-term accumulated cost change may be considerable.
These data come from early customer feedback disclosed by Anthropic. The specific workloads, comparison models, and evaluation standards are not completely consistent, and more independent testing is still needed.
Sonnet also cuts prices, and the Claude 5.5 family begins to compete on cost
In addition to Haiku 5.5, Anthropic also adjusted the pricing of Sonnet 5.5 at the same time. Starting from October 7, Sonnet 5.5's cache read fee will be reduced by 50%, from $0.2 per million tokens to $0.1.
Anthropic estimates that this adjustment can further reduce the cost of most Agent workloads by about 20%, with the actual reduction depending on the extent to which applications reuse cached context.
For Agents that need to repeatedly read long conversation histories, tool descriptions, and task context, cache costs directly affect the system's long-term operating expenses.
Anthropic also launched monthly API credits for Claude Max and Team subscribers: $100 for Max 5x users, $200 for Max 20x users, and up to $500 shared by Team users.
These credits can be used for model calls on the Claude platform, encouraging more subscribers to try building their own Agents and applications.
On the developer side, Haiku 5.5 is already available through platforms such as the Anthropic API, Amazon Bedrock, Google Cloud, and Microsoft Azure, with the API model ID claude-haiku-5-5.
Anthropic also updated the Python and TypeScript SDKs, adding Beta-stage browser operation and computer operation support.
At this point, the product division of labor in the Claude 5.5 series is already quite clear.
Opus 5.5 handles high-difficulty, complex reasoning tasks, Sonnet 5.5 covers daily complex work and programming, and Haiku 5.5 focuses on tasks with high call volume, cost sensitivity, and requirements for fast response.
From September 22 to October 7, Anthropic took 15 days to complete this round of product updates, with all three models emphasizing performance and cost efficiency.
The release of Haiku 5.5 also makes one trend more obvious. As Agents gradually move from demonstrations to practical applications, developers must consider the call costs of the entire system. Which steps in a task require a stronger model, which can be handed to a small model, and how work is allocated among different models will all affect final commercial viability.
For Anthropic, perhaps the most attractive part of Haiku 5.5 is here: it makes more tasks that were originally difficult to scale due to excessively high call costs begin to become economically viable.
Competition among small models has already begun to penetrate every execution link of the Agent system.
Reference materials
https://www.anthropic.com/claude-haiku-5-5
https://venturebeat.com/technology/anthropic-launches-claude-haiku-5-5-with-90-api-price-reduction-matching-gpt-6-luna
https://x.com/claudeai/status/2107894042235166750
This article comes from the WeChat public account "Machine Heart" (ID: almosthuman2014), editor: Spoon-billed Sandpiper









