Claude Haiku 4.5
The fastest model Anthropic ships, still on the Claude 4.5 generation with a 200K window and an early-2025 knowledge cutoff. Worth it when response latency is the product requirement; otherwise Sonnet 5 offers more capability and five times the context for twice the price.
Pricing checked against provider documentation on . How we verify
claude-haiku-4-5 reasoning Input
$1.00
/ 1M tokens
Output
$5.00
/ 1M tokens
Context
200K
200,000 tokens
Max output
64K
tokens per request
What it costs in practice
A workload of 10M input and 2M output tokens per month — roughly a moderately busy production assistant — runs $20.00/month on Claude Haiku 4.5. Routing that same volume through the batch API brings it to $10.00. The next cheaper option, DeepSeek V4-Pro , would cost $10.56.
Where Claude Haiku 4.5 fits
- Real-time chat suggestions
- Lightweight extraction and tagging
- Guardrail and routing calls
- High-frequency background jobs
Specifications
- API model ID
- claude-haiku-4-5
- Provider
- Anthropic
- Model family
- Claude 4.5
- Released
- October 2025
- Knowledge cutoff
- February 2025
- Context window
- 200,000 tokens
- Max output
- 64,000 tokens
- Native reasoning
- Yes
- Relative latency
- low
- Input modalities
- text, image
- Output modalities
- text
- Open weights
- No
- Cached input
- $0.10 / 1M tokens
- Batch discount
- 50%
Frequently asked
How much does Claude Haiku 4.5 cost?
Claude Haiku 4.5 costs $1.00 per million input tokens and $5.00 per million output tokens, with cached input at $0.10. Batch processing is 50% cheaper.
What is Claude Haiku 4.5's context window?
Claude Haiku 4.5 accepts up to 200,000 tokens of context and can generate up to 64,000 output tokens per request.
Is Claude Haiku 4.5 a reasoning model?
Yes. Claude Haiku 4.5 performs native chain-of-thought before answering. Those thinking tokens are billed at the output rate, so budget above the sticker price.
What is Claude Haiku 4.5 best for?
Lowest-latency Claude responses. The fastest model Anthropic ships, still on the Claude 4.5 generation with a 200K window and an early-2025 knowledge cutoff. Worth it when response latency is the product requirement; otherwise Sonnet 5 offers more capability and five times the context for twice the price.