September 25, 2026. Two price cuts this week left the two largest model vendors' rate cards looking simple and reading complicated. The per-token headline is on every launch page; the cache multipliers, the batch and fast tiers, the long-context rules and the regional uplifts are spread across pricing pages, model pages and caching guides. This page puts the current OpenAI and Anthropic tables side by side as read on September 25, 2026, then costs one identical agent turn on every model so an SMB or agency can see what a workload actually bills. Prices are USD per million tokens unless stated. We will update this page when either vendor moves.
OpenAI models, standard tier
Each entry lists input, cached input, cache writes and output. OpenAI bills any request over 272,000 input tokens at double the input and cache rates and 1.5 times the output rate for the whole request; the second set of figures is that long-context tier.
- GPT-6 Astra (released September 3): $10.00, $1.00, $12.50, $50.00. Over 272K: $20.00, $2.00, $25.00, $75.00. Context 1,050,000 tokens, 922,000 max input, 128,000 max output.
- GPT-6 Sol (September 22): $2.00, $0.20, $2.50, $10.00. Over 272K: $4.00, $0.40, $5.00, $15.00. Same context limits as Astra.
- GPT-6 Luna (September 22): $0.10, $0.01, $0.125, $0.50. Over 272K: $0.20, $0.02, $0.25, $0.75. Same context limits.
- GPT-5.6 Sol: $4.00, $0.40, $5.00, $20.00; over 272K $8.00, $0.80, $10.00, $30.00. OpenAI's page says this promotional price holds at least through November 21, 2026.
- GPT-5.6 Terra: $2.00, $0.20, $2.50, $12.00; over 272K $4.00, $0.40, $5.00, $18.00.
- GPT-5.6 Luna: $0.20, $0.02, $0.25, $1.20; over 272K $0.40, $0.04, $0.50, $1.80.
- GPT-5.5: $5.00, $0.50, no cache write charge, $30.00; over 272K $10.00, $1.00, $45.00. GPT-5.4: $2.50, $0.25, $15.00; over 272K $5.00, $0.50, $22.50.
- Tiers. Batch and Flex processing halve every figure above. Fast mode (renamed from Priority processing on July 30, 2026) doubles them. Regional data residency endpoints add 10% for models released on or after March 5, 2026, and FedRAMP endpoints add 10%.
- Voice and audio. gpt-realtime-2.1 audio is $32.00 input, $0.40 cached and $64.00 output per million audio tokens; gpt-realtime-2.1-mini is $10.00, $0.30 and $20.00. gpt-live-1 is $0.05 per minute, live transcription $0.017 per minute and live translation $0.034 per minute.
Anthropic models
Each entry lists input, output, five-minute cache write, one-hour cache write and cache read. Anthropic bills the full 1M token context window at standard rates for Claude 4.6 and later, so there is no long-context tier: a 900K token request costs the same per token as a 9K one.
- Claude Fable 5.1: $10.00, $50.00, $12.50, $20.00, $0.25 (cache reads at 0.025x input). Context 1M, output 128K.
- Claude Opus 5.5 (September 22): $4.00, $20.00, $5.00, $8.00, $0.20 (cache reads at 0.05x input). Context 1M, output 128K. Fast mode $8.00 input and $40.00 output.
- Claude Sonnet 5: $2.00, $10.00, $2.50, $4.00, $0.20. Context 1M, output 128K.
- Claude Haiku 4.5: $1.00, $5.00, $1.25, $2.00, $0.10. Context 200K, output 64K.
- Claude Opus 5 and Opus 4.8: $5.00, $25.00, cache reads $0.50; Fast mode $10.00 and $50.00. Claude Mythos 5.1 is listed at $10.00 and $50.00 and is available only through Anthropic's trusted access programs.
- Tiers. The Batch API halves input and output. Fast mode, in research preview on the first-party API only, is available for Opus 5.5, Opus 5 and Opus 4.8 and is not available with Batch. Regional and multi-region endpoints carry a 10% premium over global. Cache write multipliers are 1.25x input for five minutes and 2x for one hour; the models released since Claude 4.7 use a tokenizer that produces roughly 30% more tokens for the same text, which Anthropic discloses on the pricing page.
One agent turn, costed on every model
Take a typical automation turn: 50,000 input tokens, of which 40,000 are a cached system prompt and history, and 5,000 output tokens. GPT-6 Luna: about $0.004. Claude Haiku 4.5: $0.039. GPT-6 Sol and Claude Sonnet 5: $0.078 each, identical to the cent because their rate cards match. Claude Opus 5.5: $0.148. Claude Fable 5.1: $0.36. GPT-6 Astra: $0.39. A voice or intake agent handling 3,000 such turns a month therefore costs about $12 on Luna, $117 on Haiku, $234 on Sol or Sonnet, $444 on Opus 5.5 and $1,170 on Astra in model fees, before telephony or tooling. Change the shape and the ranking shifts: our cached input pricing analysis shows why Anthropic's 0.05x read rate on Opus 5.5 pulls ahead on long, repetitive agent sessions, and the 272K long context explainer shows a 400K token document job costing $1.68 on Opus 5.5 and $8.30 on Astra.
The five rules that change the bill more than the list price
First, cache reads: on every current model they are 5% to 10% of the input rate, so an agent built around a stable prefix pays a fraction of the sticker price. Second, cache writes: OpenAI charges 1.25x input on GPT-5.6 and later, Anthropic 1.25x for five minutes or 2x for an hour, so a prefix written once and never reused costs more than not caching at all. Third, the long-context tier: only OpenAI has one, and it applies to the whole request. Fourth, the speed tiers: Batch halves the bill on both vendors when latency does not matter, Fast mode doubles it when it does. Fifth, residency: a 10% uplift on both vendors for in-region processing, which matters for any client whose data cannot leave a jurisdiction. The head-to-head pieces on Astra vs Opus 5.5 and Sol vs Opus 5.5 apply these rules to the two matchups agencies ask about most, and the price war brief covers how the September 22 cuts landed.
What it means for operators
Price the workload, not the model. For high-volume short turns the four cheapest options are within a few cents of each other and the decision is quality and tooling, not price. For long documents Anthropic's flat context pricing removes a risk OpenAI's tier creates. For voice, the realtime audio rates above are the model cost inside every AI phone agent on the market, which is why an AI call answering agent for a home services business is priced on minutes and turns rather than on a flat licence. If a vendor quotes you a per-seat AI price with no token line behind it, ask which of these tables it was built on; if you would rather have the arithmetic done for you, that is where an AI automation build starts.
Frequently Asked Questions
On list price GPT-6 Luna at $0.10 input and $0.50 output per million tokens, followed by Claude Haiku 4.5 at $1.00 and $5.00. GPT-6 Sol and Claude Sonnet 5 share the same $2.00 and $10.00 rate card, and Claude Opus 5.5 is $4.00 and $20.00.
$10.00 input, $1.00 cached input, $12.50 cache writes and $50.00 output on the standard tier; $20.00, $2.00, $25.00 and $75.00 once a request passes 272K input tokens. Batch and Flex halve those figures and Fast mode doubles them.
$4.00 input and $20.00 output, with five-minute cache writes at $5.00, one-hour cache writes at $8.00 and cache reads at $0.20. Fast mode is $8.00 and $40.00. The Batch API halves input and output.
No. Anthropic's pricing page says Claude 4.6 and later models include the full 1M token context window at standard pricing, with caching and batch discounts applying across the whole window. OpenAI, by contrast, doubles input and cache rates and raises output 1.5x for any request over 272K input tokens.
On both vendors batch processing is a 50% discount on input and output. OpenAI's Fast mode doubles standard rates on every model; Anthropic's Fast mode is available for Opus 5.5 ($8 and $40), Opus 5 and Opus 4.8 ($10 and $50) on the first-party API only and cannot be combined with Batch.
It reflects the OpenAI and Anthropic pricing pages as read on September 25, 2026, after the September 22 GPT-6 Sol, GPT-6 Luna and Claude Opus 5.5 releases. We revise it when either vendor changes a published rate.