Claude Models & API Pricing: Fable, Opus, Sonnet and Haiku
Compare current Claude model pricing, prompt caching, 1M context, Batch and Fast mode, tool charges, data residency, and Sonnet 5 introductory pricing.
Current Claude model family
Capability, composed in four voices
Use Fable for maximum capability, Opus for complex agentic work, Sonnet for balance, or Haiku for speed and cost.
USD per 1M tokens · Claude API
Claude Fable 5
Anthropic’s most capable broadly released model for long-running agents.
- Model ID
claude-fable-5
- Input
- $10
- 5m cache write
- $12.50
- 1h cache write
- $20
- Cache hit
- $1
- Output
- $50
- Context
- 1M
- Max output
- 128K
- Reliable knowledge
- Jan 2026
- Thinking
- Adaptive, always on
- Latency
- Slower
Claude Opus 5
Complex agentic coding and enterprise work with adaptive thinking.
- Model ID
claude-opus-5
- Input
- $5
- 5m cache write
- $6.25
- 1h cache write
- $10
- Cache hit
- $0.50
- Output
- $25
- Context
- 1M
- Max output
- 128K
- Reliable knowledge
- May 2026
- Thinking
- Adaptive
- Latency
- Moderate
Claude Sonnet 5
The strongest balance of speed and intelligence, with introductory pricing.
- Model ID
claude-sonnet-5
- Input
- $2
- 5m cache write
- $2.50
- 1h cache write
- $4
- Cache hit
- $0.20
- Output
- $10
- Context
- 1M
- Max output
- 128K
- Reliable knowledge
- Jan 2026
- Thinking
- Adaptive
- Latency
- Fast
Claude Haiku 4.5
Anthropic’s fastest model with near-frontier intelligence.
- Model ID
claude-haiku-4-5-20251001
- Input
- $1
- 5m cache write
- $1.25
- 1h cache write
- $2
- Cache hit
- $0.10
- Output
- $5
- Context
- 200K
- Max output
- 64K
- Reliable knowledge
- Feb 2025
- Thinking
- Extended thinking
- Latency
- Fastest
Current Claude Model Family
Anthropic’s current core lineup spans four workload tiers. Claude Fable 5 is the most capable broadly released model, Opus 5 targets complex agentic coding and enterprise work, Sonnet 5 balances speed and intelligence, and Haiku 4.5 prioritizes latency and cost.
The visual comparison above uses Anthropic’s official model overview and pricing pages checked on August 10, 2026. Claude Mythos 5 shares Fable 5 specifications and pricing but is limited-access through Project Glasswing, so it is not presented as a self-service default.
Standard Token Pricing
Claude API prices are quoted in USD per million tokens (MTok). Input and output are billed separately, and output tokens are materially more expensive across all four models.
Choose by workload rather than price alone: Fable for maximum broadly available capability, Opus for complex production agents, Sonnet for the general-purpose balance, and Haiku for high-volume execution.
- Fable 5: $10 input · $50 output
- Opus 5: $5 input · $25 output
- Sonnet 5: introductory $2 input · $10 output through August 31, 2026
- Haiku 4.5: $1 input · $5 output
Sonnet 5 Introductory Price Ends August 31
Sonnet 5 has time-limited introductory pricing of $2 per MTok input and $10 per MTok output through August 31, 2026. On September 1, 2026, standard pricing becomes $3 input and $15 output.
The cache prices change at the same time: 5-minute writes move from $2.50 to $3.75, 1-hour writes from $4 to $6, and cache hits from $0.20 to $0.30. Budget September and later workloads with the standard rates now.
Prompt Caching Prices
Anthropic separates cache creation from cache reads. A 5-minute cache write costs 1.25× the base input price, a 1-hour write costs 2×, and cache hits or refreshes cost 0.1× base input.
Caching is valuable when long system instructions, documents, or examples repeat across requests. Calculate whether reuse is frequent enough to offset the initial write premium.
Context Windows, Batch and Fast Mode
Fable 5, Opus 5 and Sonnet 5 include a full 1M-token context window at standard per-token rates and support 128K synchronous output. Haiku 4.5 supports 200K input and 64K output.
Batch API reduces both input and output prices by 50%. Fast mode is a research preview for Opus 5 and Opus 4.8 on the first-party Claude API; Opus 5 Fast mode costs $10 input and $50 output per MTok. Fast mode cannot be combined with Batch.
Data Residency and Tool Charges
For Claude 4.6 and later, US-only inference selected through inference_geo applies a 1.1× multiplier to input, output, cache writes, and cache reads. Default global routing uses standard prices.
Web search costs $10 per 1,000 searches plus normal tokens. Web fetch has no separate tool fee, but fetched content becomes billable input. Code execution includes 1,550 free organization-hours per month and then costs $0.05 per container-hour when it is not free through eligible web search or fetch usage.
Claude Managed Agents and Cost Control
Claude Managed Agents bill model tokens at standard rates plus $0.08 per running session-hour. Idle, rescheduling, and terminated time is not billed as runtime. Batch, Fast mode, data-residency, and partner-cloud pricing modifiers do not apply to Managed Agents sessions.
For cost control, route simple work to Haiku, general production work to Sonnet, complex agents to Opus, and highest-capability workloads to Fable. Add prompt caching for repeated context and Batch for non-interactive jobs.
FAQ
Which Claude model is cheapest?
Claude Haiku 4.5 is the lowest-cost current core model at $1 input and $5 output per MTok.
When does Sonnet 5 pricing change?
The $2/$10 introductory input/output rates apply through August 31, 2026. Standard $3/$15 rates begin September 1, 2026.
Do current Claude models support 1M context?
Fable 5, Opus 5 and Sonnet 5 do. Haiku 4.5 has a 200K context window.
Can prompt caching and Batch be combined?
Yes. Batch and prompt-caching discounts can be combined. Fast mode cannot be combined with Batch.
Related in This Series
Related Providers
Sources
- Anthropic Models & PricingAnthropic · Checked 2026-08-10
- Claude Models OverviewAnthropic · Checked 2026-08-10
- Anthropic 中文定价Anthropic · Checked 2026-08-10