[AINews] Claude Haiku 5.5 — better than GPT-6 Luna at the same pricing

Anthropic shipped Claude Haiku 5.5, the first update to its small model tier since Haiku 4.5 in October 2025. The company called it the cheapest, fastest and most capable small model it has released, and said it costs about 75% less to run on average than Haiku 4.5. It is live on the Claude Platform and in Claude Code. The same day, Anthropic cut prices on Sonnet 5.5 and on its subscription plans, and began including monthly Claude Platform API credits with Max 5x ($100), Max 20x ($200) and Team (up to $500, pooled) plans, usable on any model and in third-party harnesses.
Haiku 5.5 is priced in two tiers: $0.10 per 1M input tokens and $0.50 per 1M output tokens under 100K tokens, and $0.50 / $2.50 above 100K. Cache reads cost $0.01 per 1M under 100K and $0.05 above; 5-minute cache writes cost $0.125 per 1M, or $0.625 above 100K. Context is 1M tokens, up from 200K for Haiku 4.5, with text and image input and text output. Sonnet 5.5 cache reads were halved from $0.20 to $0.10 per 1M tokens, which Anthropic says makes that model about 20% cheaper on most long-running or agentic work.
The Python and TypeScript Claude SDKs gained built-in computer-use and browser-use toolsets that run the action loop and send clicks and keystrokes to drivers from browser_use, Browserbase, E2B or Daytona. Anthropic positions Haiku 5.5 as a subagent alongside Opus 5.5 or Sonnet 5.5 for high-volume, cost-sensitive tasks such as summaries, compactions and database queries. This is the first Haiku with effort settings and adaptive thinking.
Artificial Analysis reported an Intelligence Index of 43 at max effort, up 26 points from the previous Haiku and slightly ahead of GLM-5.3 Flash (42), Gemini 3.8 Flash (41) and GPT-6 Luna (38), while trailing Sonnet 5.5 (56). It measured about 162K output tokens per Index task at max effort, roughly 3x GPT-6 Luna, and noted that prompts over 100K tokens pay 5x more. Terminal-Bench 4.0 came in at 33%, up from 0%. Partner claims include Cursor's 10x lower cost on shorter requests, Devin's 58.4% on FrontierCode 1.1, and GitHub Copilot in VS Code matching Sonnet 5 on many coding tasks. A pre-release safety bug caused over-refusal; Anthropic is working on a fix.
Why it matters: Teams building agent pipelines now have a cheaper small-model option with the same list price as GPT-6 Luna, but its higher token usage and 5x surcharge above 100K tokens mean real savings depend on prompt length and effort setting.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
yay small models