in

Claude Haiku 5.5 Launches at $0.10 per Million Input Tokens

Claude Haiku 5.5 pricing headline over a developer desk

Anthropic released Claude Haiku 5.5 on October 7, 2026, pricing its smallest model at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. That is about 90% below Haiku 4.5, which charged $1 and $5, according to figures Anthropic published. Prompts above 100,000 tokens move to a second price list of $0.50 input and $2.50 output.

What the price list actually says

Anthropic’s Claude Haiku page lists the two tiers and notes that prompt caching can save up to 90% and batch processing 50%. The model string is claude-haiku-5-5, and it is available in Claude.ai, Claude Code, the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Foundry. It is the first Haiku with an effort control, so you can spend less thinking on easy tasks and more on hard ones.

Hosting 75% off

MarkTechPost’s launch coverage adds the specs Anthropic reported: a 1 million token context window, up to 128,000 output tokens, and text and image input. On OSWorld 2.1 (offline subset) Anthropic reports 72.4%, against 15.7% for Haiku 4.5. Sonnet 5.5 scores higher on every benchmark in the table, and Anthropic positions Haiku as a helper model under Opus 5.5 or Sonnet 5.5 rather than the lead for complex agentic coding.

The 100,000 token cliff

The cheap headline has a catch that matters if you process long documents. A DEV Community analysis points out that the higher tier is not a marginal bracket. Once a prompt crosses 100,000 tokens, the whole request is billed at the higher rates, output included. Using the published prices, its author calculates that a 100,000 token prompt with 2,000 output tokens costs $0.011, while 100,001 tokens costs $0.055. The same piece notes the new tokenizer produces roughly 30% more tokens than Haiku 4.5 for the same text, so your documents hit the line sooner than the old numbers suggest. These are the article’s own calculations, not measurements from the API.

The practical fix is chunking. Splitting a long contract or transcript into pieces that each stay under the line keeps every request on the cheap list.

Why freelancers and small agencies should care

Take a hypothetical job: 1,000 customer messages, each about 5,000 input tokens, with 500 tokens of output per reply. At the short-prompt rates that is 5 million input tokens ($0.50) plus 0.5 million output tokens ($0.25), or $0.75 in total before caching or batch discounts. At Haiku 4.5 prices the same job would come to $7.50. That arithmetic is ours, based on the published rates, and your real bill will vary with how many tokens your prompts use.

For people selling automation, such as support bots, data entry, document sorting or lead qualification, the model cost per task is now small enough that your fee, not the API bill, is the main number. Anthropic describes the model as suited to form filling and data entry, and fast enough for live chat and voice agents. If you are deciding which Claude to build on, our earlier look at Sonnet 5.5 covers the model Anthropic says does the harder reasoning above Haiku.

A crowded price war

MarkTechPost notes that Haiku 5.5’s short-prompt list prices match OpenAI’s GPT-6 Luna, whose higher tier starts above 272,000 tokens, which makes Luna cheaper for very long prompts. Non-default temperature, top_p or top_k values also return a 400 error on Haiku 5.5, so code written for older models may need small changes.

The open question is token appetite. Cheaper per token only means cheaper per job if the model does not use far more tokens to finish it, and with adaptive thinking on by default, builders will need to run their own cost tests on real workloads before promising clients a price.

Hosting 75% off

Written by Fahad Manzur

Google’s new SynthID website can detect AI-created images, video, and audio

Google’s new SynthID website can detect AI-created images, video, and audio