AI API PRICING DIRECTORY

Calculate model costs without flattening the billing rules.

Compare 52 official-source rate cards across nine providers. Then model requests, tokens, warm cache reads, and quality pass rates against the exact model tier.

Snapshot verified 2026-08-16 · USD per one million tokens · consumer chat subscriptions excluded

PROVIDER CALCULATORS

Start with the billing surface you use.

PAIRWISE API COSTS

Put two providers on one workload.

Choose exact model tiers, reuse the same request and token inputs, then set separate quality pass rates before calling either side cheaper.

ALL RATE CARDS

Search the current snapshot.

Keep context bands, regions, cache modes, and effective dates visible. A missing cache price is never treated as free.

Showing 52 of 52 rate cards

Official provider API rate cards, USD per one million tokens
Provider / modelPrice scopeInput
USD / 1M
Cache read
USD / 1M
Output
USD / 1M
ContextSource
AnthropicClaude Fable 5claude-fable-5StandardGlobal routing$10$1$501MOfficial ↗
AnthropicClaude Haiku 4.5claude-haiku-4-5StandardGlobal routing$1$0.1$5200KOfficial ↗
AnthropicClaude Opus 5claude-opus-5StandardGlobal routing$5$0.5$251MOfficial ↗
AnthropicClaude Sonnet 4.6claude-sonnet-4-6StandardGlobal routing$3$0.3$151MOfficial ↗
AnthropicClaude Sonnet 5claude-sonnet-5StandardGlobal routing$2$0.2$101MOfficial ↗
CohereCommand Acommand-a-03-2025Production pay-as-you-goGlobal$2.50$10256KOfficial ↗
CohereCommand Rcommand-r-08-2024Production pay-as-you-goGlobal$0.15$0.6Official ↗
CohereCommand R+command-r-plus-08-2024Production pay-as-you-goGlobal$2.50$10Official ↗
CohereCommand R7Bcommand-r7b-12-2024Production pay-as-you-goGlobal$0.037$0.15Official ↗
DeepSeekDeepSeek V4 Flashdeepseek-v4-flashOff-peak · scheduled UTC windowsGlobalFrom 16 Aug 2026$0.22$0.007$0.661MOfficial ↗
DeepSeekDeepSeek V4 Flashdeepseek-v4-flashPeak · 01:00–04:00 and 06:00–10:00 UTCGlobalFrom 16 Aug 2026$0.44$0.014$1.321MOfficial ↗
DeepSeekDeepSeek V4 Flashdeepseek-v4-flashStandard · through 2026-08-16 16:00 UTCGlobalThrough 16 Aug 2026$0.14$0.0028$0.281MOfficial ↗
DeepSeekDeepSeek V4 Prodeepseek-v4-proOff-peak · scheduled UTC windowsGlobalFrom 16 Aug 2026$0.66$0.022$1.981MOfficial ↗
DeepSeekDeepSeek V4 Prodeepseek-v4-proPeak · 01:00–04:00 and 06:00–10:00 UTCGlobalFrom 16 Aug 2026$1.32$0.044$3.961MOfficial ↗
DeepSeekDeepSeek V4 Prodeepseek-v4-proStandard · through 2026-08-16 16:00 UTCGlobalThrough 16 Aug 2026$0.435$0.003625$0.871MOfficial ↗
GoogleGemini 3.1 Flash-Litegemini-3.1-flash-liteDeveloper API · Standard · text/image/videoGemini Developer API$0.25$0.025$1.50Official ↗
GoogleGemini 3.1 Pro Previewgemini-3.1-pro-previewDeveloper API · Standard · over 200K promptGemini Developer API$4$0.4$18Official ↗
GoogleGemini 3.1 Pro Previewgemini-3.1-pro-previewDeveloper API · Standard · up to 200K promptGemini Developer API$2$0.2$12Official ↗
GoogleGemini 3.5 Flashgemini-3.5-flashDeveloper API · StandardGemini Developer API$1.50$0.15$9Official ↗
GoogleGemini 3.5 Flash-Litegemini-3.5-flash-liteDeveloper API · StandardGemini Developer API$0.3$0.03$2.50Official ↗
GoogleGemini 3.6 Flashgemini-3.6-flashDeveloper API · Standard · introductoryGemini Developer APIThrough 31 Dec 2026$0.75$0.075$3.75Official ↗
GoogleGemini 3.7 Flashgemini-3.7-flashDeveloper API · Standard · introductoryGemini Developer APIThrough 31 Dec 2026$0.75$0.075$3.75Official ↗
Kimi / Moonshot AIKimi K2.6kimi-k2.6RealtimeGlobal$0.95$0.16$4262KOfficial ↗
Kimi / Moonshot AIKimi K2.7 Code Highspeedkimi-k2.7-code-highspeedRealtime · high-speed tierGlobal$1.90$0.38$8262KOfficial ↗
Kimi / Moonshot AIKimi K2.7 Codekimi-k2.7-codeRealtimeGlobal$0.95$0.19$4262KOfficial ↗
Kimi / Moonshot AIKimi K3kimi-k3RealtimeGlobal$3$0.3$151.05MOfficial ↗
Mistral AICodestralcodestral-latestStandardGlobal$0.3$0.03$0.9Official ↗
Mistral AIMinistral 3 14Bministral-14b-latestStandardGlobal$0.2$0.02$0.2Official ↗
Mistral AIMinistral 3 3Bministral-3b-latestStandardGlobal$0.1$0.01$0.1Official ↗
Mistral AIMinistral 3 8Bministral-8b-latestStandardGlobal$0.15$0.015$0.15Official ↗
Mistral AIMistral Large 3mistral-large-latestStandardGlobal$0.5$0.05$1.50Official ↗
Mistral AIMistral Medium 3.5mistral-medium-latestStandardGlobal$1.50$0.15$7.50Official ↗
Mistral AIMistral Small 4mistral-small-latestStandardGlobal$0.15$0.015$0.6Official ↗
OpenAIGPT-5.6 Lunagpt-5.6-lunaStandard · over 272K inputGlobal$0.4$0.04$1.801.05MOfficial ↗
OpenAIGPT-5.6 Lunagpt-5.6-lunaStandard · up to 272K inputGlobal$0.2$0.02$1.201.05MOfficial ↗
OpenAIGPT-5.6 Solgpt-5.6-solStandard · over 272K inputGlobal$10$1$451.05MOfficial ↗
OpenAIGPT-5.6 Solgpt-5.6-solStandard · up to 272K inputGlobal$5$0.5$301.05MOfficial ↗
OpenAIGPT-5.6 Terragpt-5.6-terraStandard · over 272K inputGlobal$4$0.4$181.05MOfficial ↗
OpenAIGPT-5.6 Terragpt-5.6-terraStandard · up to 272K inputGlobal$2$0.2$121.05MOfficial ↗
Qwen / Alibaba CloudQwen 3.7 Flashqwen3.7-flashBeijing · 256K–1M inputChina Beijing$0.165$0.033$0.661MOfficial ↗
Qwen / Alibaba CloudQwen 3.7 Flashqwen3.7-flashBeijing · 32K–256K inputChina Beijing$0.083$0.017$0.331MOfficial ↗
Qwen / Alibaba CloudQwen 3.7 Flashqwen3.7-flashBeijing · up to 32K inputChina Beijing$0.028$0.006$0.111MOfficial ↗
Qwen / Alibaba CloudQwen 3.7 Maxqwen3.7-maxGlobal scope · up to 1M inputUS Virginia · Global deployment$1.65$0.33$4.951MOfficial ↗
Qwen / Alibaba CloudQwen 3.7 Plusqwen3.7-plusGlobal scope · 256K–1M inputUS Virginia · Global deployment$0.826$0.166$3.301MOfficial ↗
Qwen / Alibaba CloudQwen 3.7 Plusqwen3.7-plusGlobal scope · up to 256K inputUS Virginia · Global deployment$0.276$0.056$1.101MOfficial ↗
xAIGrok 4.20 Reasoninggrok-4.20-0309-reasoningStandard · below 200K promptxAI direct API$1.25$0.2$2.501MOfficial ↗
xAIGrok 4.3grok-4.3Standard · 200K+ promptxAI direct API$2.50$0.4$51MOfficial ↗
xAIGrok 4.3grok-4.3Standard · below 200K promptxAI direct API$1.25$0.2$2.501MOfficial ↗
xAIGrok 4.5grok-4.5Standard · below 200K promptxAI direct API$2$0.3$6500KOfficial ↗
xAIGrok 4.6grok-4.6Standard · 200K+ promptxAI direct API$4$1$12500KOfficial ↗
xAIGrok 4.6grok-4.6Standard · below 200K promptxAI direct API$2$0.5$6500KOfficial ↗
xAIGrok Build 0.1grok-build-0.1Standard · below 200K promptxAI direct API$1$0.2$2256KOfficial ↗

Rates in USD per 1M tokens

DeepSeek

DeepSeek V4 Flash

deepseek-v4-flash

Off-peak · scheduled UTC windowsGlobal

Input
$0.22
Cache read
$0.007
Output
$0.66
Context
1M
Effective from 16 Aug 2026Open official source in a new tab
DeepSeek

DeepSeek V4 Flash

deepseek-v4-flash

Peak · 01:00–04:00 and 06:00–10:00 UTCGlobal

Input
$0.44
Cache read
$0.014
Output
$1.32
Context
1M
Effective from 16 Aug 2026Open official source in a new tab
DeepSeek

DeepSeek V4 Flash

deepseek-v4-flash

Standard · through 2026-08-16 16:00 UTCGlobal

Input
$0.14
Cache read
$0.0028
Output
$0.28
Context
1M
Effective through 16 Aug 2026Open official source in a new tab
DeepSeek

DeepSeek V4 Pro

deepseek-v4-pro

Off-peak · scheduled UTC windowsGlobal

Input
$0.66
Cache read
$0.022
Output
$1.98
Context
1M
Effective from 16 Aug 2026Open official source in a new tab
DeepSeek

DeepSeek V4 Pro

deepseek-v4-pro

Peak · 01:00–04:00 and 06:00–10:00 UTCGlobal

Input
$1.32
Cache read
$0.044
Output
$3.96
Context
1M
Effective from 16 Aug 2026Open official source in a new tab
DeepSeek

DeepSeek V4 Pro

deepseek-v4-pro

Standard · through 2026-08-16 16:00 UTCGlobal

Input
$0.435
Cache read
$0.003625
Output
$0.87
Context
1M
Effective through 16 Aug 2026Open official source in a new tab
Google

Gemini 3.1 Flash-Lite

gemini-3.1-flash-lite

Developer API · Standard · text/image/videoGemini Developer API

Input
$0.25
Cache read
$0.025
Output
$1.50
Context
Open official source in a new tab
Google

Gemini 3.1 Pro Preview

gemini-3.1-pro-preview

Developer API · Standard · over 200K promptGemini Developer API

Input
$4
Cache read
$0.4
Output
$18
Context
Open official source in a new tab
Google

Gemini 3.1 Pro Preview

gemini-3.1-pro-preview

Developer API · Standard · up to 200K promptGemini Developer API

Input
$2
Cache read
$0.2
Output
$12
Context
Open official source in a new tab
Google

Gemini 3.6 Flash

gemini-3.6-flash

Developer API · Standard · introductoryGemini Developer API

Input
$0.75
Cache read
$0.075
Output
$3.75
Context
Effective through 31 Dec 2026Open official source in a new tab
Google

Gemini 3.7 Flash

gemini-3.7-flash

Developer API · Standard · introductoryGemini Developer API

Input
$0.75
Cache read
$0.075
Output
$3.75
Context
Effective through 31 Dec 2026Open official source in a new tab
Qwen / Alibaba Cloud

Qwen 3.7 Max

qwen3.7-max

Global scope · up to 1M inputUS Virginia · Global deployment

Input
$1.65
Cache read
$0.33
Output
$4.95
Context
1M
Open official source in a new tab
Qwen / Alibaba Cloud

Qwen 3.7 Plus

qwen3.7-plus

Global scope · 256K–1M inputUS Virginia · Global deployment

Input
$0.826
Cache read
$0.166
Output
$3.30
Context
1M
Open official source in a new tab
Qwen / Alibaba Cloud

Qwen 3.7 Plus

qwen3.7-plus

Global scope · up to 256K inputUS Virginia · Global deployment

Input
$0.276
Cache read
$0.056
Output
$1.10
Context
1M
Open official source in a new tab

USD per 1M tokens. Snapshot verified 2026-08-16. Cache writes, cache storage, tools, regions, and provider-specific thresholds may be billed separately.

A missing cache rate is shown as “—”, never treated as free. Consumer chat-plan quotas are not API prices.

WORKLOAD CALCULATOR

Turn rates into a monthly scenario.

Compare raw spend with cost per accepted answer and the candidate pass rate needed to break even. These remain estimates until a quality-gated test confirms them.

Estimated monthly API spend · OpenAI
GPT-5.6 TerraBefore · Standard · up to 272K input: $2 input · $0.2 cache read · $12 output / 1MAfter · Standard · up to 272K input: $2 input · $0.2 cache read · $12 output / 1MGlobal · price bands selected automatically from input size
Before$120.00
After$80.85
Potential saving: $39.15 (32.6%)
QUALITY-ADJUSTED COST
Before / accepted answer$0.013390% pass assumption
After / accepted answer$0.009585% pass assumption

Candidate break-even: 60.6% pass rate. Your scenario assumes 85%.

Warm cache-read scenario, not general ROI. It excludes cache writes/storage, tools, regional uplifts, retries, and quality failures. A dash means no published cache-read rate.

OpenAI API pricing
NEXT STEPValidate the 28.7% accepted-answer advantage.

The estimate is not a saving until the same workload still passes its quality bar. Compare a supported recipe in the lab, or use the evidence library to design a provider-specific test.