Skip to main content
Comparison12 min readPublished: 2026-08-10Updated: 2026-08-10

Claude Models & API Pricing: Fable, Opus, Sonnet and Haiku

Compare current Claude model pricing, prompt caching, 1M context, Batch and Fast mode, tool charges, data residency, and Sonnet 5 introductory pricing.

By Heizi· Founder & Editor· Published: 2026-08-10Hands-on tested

Current Claude model family

Capability, composed in four voices

Use Fable for maximum capability, Opus for complex agentic work, Sonnet for balance, or Haiku for speed and cost.

USD per 1M tokens · Claude API

Claude Fable 5

Anthropic’s most capable broadly released model for long-running agents.

Model ID
claude-fable-5
Input
$10
5m cache write
$12.50
1h cache write
$20
Cache hit
$1
Output
$50
Context
1M
Max output
128K
Reliable knowledge
Jan 2026
Thinking
Adaptive, always on
Latency
Slower

Claude Opus 5

Complex agentic coding and enterprise work with adaptive thinking.

Model ID
claude-opus-5
Input
$5
5m cache write
$6.25
1h cache write
$10
Cache hit
$0.50
Output
$25
Context
1M
Max output
128K
Reliable knowledge
May 2026
Thinking
Adaptive
Latency
Moderate

Claude Sonnet 5

The strongest balance of speed and intelligence, with introductory pricing.

Model ID
claude-sonnet-5
Input
$2
5m cache write
$2.50
1h cache write
$4
Cache hit
$0.20
Output
$10
Context
1M
Max output
128K
Reliable knowledge
Jan 2026
Thinking
Adaptive
Latency
Fast

Claude Haiku 4.5

Anthropic’s fastest model with near-frontier intelligence.

Model ID
claude-haiku-4-5-20251001
Input
$1
5m cache write
$1.25
1h cache write
$2
Cache hit
$0.10
Output
$5
Context
200K
Max output
64K
Reliable knowledge
Feb 2025
Thinking
Extended thinking
Latency
Fastest
Sonnet 5 introductory input/output pricing is $2/$10 through August 31, 2026. Standard $3/$15 pricing begins September 1, 2026.
Official Anthropic snapshot checked August 10, 2026.

Current Claude Model Family

Anthropic’s current core lineup spans four workload tiers. Claude Fable 5 is the most capable broadly released model, Opus 5 targets complex agentic coding and enterprise work, Sonnet 5 balances speed and intelligence, and Haiku 4.5 prioritizes latency and cost.

The visual comparison above uses Anthropic’s official model overview and pricing pages checked on August 10, 2026. Claude Mythos 5 shares Fable 5 specifications and pricing but is limited-access through Project Glasswing, so it is not presented as a self-service default.

Standard Token Pricing

Claude API prices are quoted in USD per million tokens (MTok). Input and output are billed separately, and output tokens are materially more expensive across all four models.

Choose by workload rather than price alone: Fable for maximum broadly available capability, Opus for complex production agents, Sonnet for the general-purpose balance, and Haiku for high-volume execution.

  • Fable 5: $10 input · $50 output
  • Opus 5: $5 input · $25 output
  • Sonnet 5: introductory $2 input · $10 output through August 31, 2026
  • Haiku 4.5: $1 input · $5 output

Sonnet 5 Introductory Price Ends August 31

Sonnet 5 has time-limited introductory pricing of $2 per MTok input and $10 per MTok output through August 31, 2026. On September 1, 2026, standard pricing becomes $3 input and $15 output.

The cache prices change at the same time: 5-minute writes move from $2.50 to $3.75, 1-hour writes from $4 to $6, and cache hits from $0.20 to $0.30. Budget September and later workloads with the standard rates now.

Prompt Caching Prices

Anthropic separates cache creation from cache reads. A 5-minute cache write costs 1.25× the base input price, a 1-hour write costs 2×, and cache hits or refreshes cost 0.1× base input.

Caching is valuable when long system instructions, documents, or examples repeat across requests. Calculate whether reuse is frequent enough to offset the initial write premium.

Context Windows, Batch and Fast Mode

Fable 5, Opus 5 and Sonnet 5 include a full 1M-token context window at standard per-token rates and support 128K synchronous output. Haiku 4.5 supports 200K input and 64K output.

Batch API reduces both input and output prices by 50%. Fast mode is a research preview for Opus 5 and Opus 4.8 on the first-party Claude API; Opus 5 Fast mode costs $10 input and $50 output per MTok. Fast mode cannot be combined with Batch.

Data Residency and Tool Charges

For Claude 4.6 and later, US-only inference selected through inference_geo applies a 1.1× multiplier to input, output, cache writes, and cache reads. Default global routing uses standard prices.

Web search costs $10 per 1,000 searches plus normal tokens. Web fetch has no separate tool fee, but fetched content becomes billable input. Code execution includes 1,550 free organization-hours per month and then costs $0.05 per container-hour when it is not free through eligible web search or fetch usage.

Claude Managed Agents and Cost Control

Claude Managed Agents bill model tokens at standard rates plus $0.08 per running session-hour. Idle, rescheduling, and terminated time is not billed as runtime. Batch, Fast mode, data-residency, and partner-cloud pricing modifiers do not apply to Managed Agents sessions.

For cost control, route simple work to Haiku, general production work to Sonnet, complex agents to Opus, and highest-capability workloads to Fable. Add prompt caching for repeated context and Batch for non-interactive jobs.

FAQ

Which Claude model is cheapest?

Claude Haiku 4.5 is the lowest-cost current core model at $1 input and $5 output per MTok.

When does Sonnet 5 pricing change?

The $2/$10 introductory input/output rates apply through August 31, 2026. Standard $3/$15 rates begin September 1, 2026.

Do current Claude models support 1M context?

Fable 5, Opus 5 and Sonnet 5 do. Haiku 4.5 has a 200K context window.

Can prompt caching and Batch be combined?

Yes. Batch and prompt-caching discounts can be combined. Fast mode cannot be combined with Batch.

Related Providers

Sources