Claude Fable 5
claude-fable-5StandardGlobal routing
- Input
- $10
- Cache read
- $1
- Output
- $50
- Context
- 1M
Run the same task at two settings. Compare the token totals and both outputs before you change a production workflow.
The same short task and shared instructions ran on GPT-5.5 at medium and low reasoning effort. Both outputs contained the requested three bullets; a human still has to judge whether either answer meets the real product rubric.
Repeat it with your own task →Demonstration, not customer evidence or an API invoice. ChatGPT-plan requests count against plan limits; provider-key requests are billed by that provider. Token reduction is not a quality verdict or guaranteed saving.
Gemini, Grok, Kimi, and Qwen are first-class—not footnotes. Provider-specific cache, region, and long-context rules remain separate instead of being flattened into misleading averages.
| Provider / model | Price scope | Input USD / 1M | Cache read USD / 1M | Output USD / 1M | Context | Source |
|---|---|---|---|---|---|---|
AnthropicClaude Fable 5claude-fable-5 | StandardGlobal routing | $10 | $1 | $50 | 1M | Official ↗ |
AnthropicClaude Haiku 4.5claude-haiku-4-5 | StandardGlobal routing | $1 | $0.1 | $5 | 200K | Official ↗ |
AnthropicClaude Opus 5claude-opus-5 | StandardGlobal routing | $5 | $0.5 | $25 | 1M | Official ↗ |
AnthropicClaude Sonnet 4.6claude-sonnet-4-6 | StandardGlobal routing | $3 | $0.3 | $15 | 1M | Official ↗ |
AnthropicClaude Sonnet 5claude-sonnet-5 | StandardGlobal routing | $2 | $0.2 | $10 | 1M | Official ↗ |
CohereCommand Acommand-a-03-2025 | Production pay-as-you-goGlobal | $2.50 | — | $10 | 256K | Official ↗ |
CohereCommand Rcommand-r-08-2024 | Production pay-as-you-goGlobal | $0.15 | — | $0.6 | — | Official ↗ |
CohereCommand R+command-r-plus-08-2024 | Production pay-as-you-goGlobal | $2.50 | — | $10 | — | Official ↗ |
CohereCommand R7Bcommand-r7b-12-2024 | Production pay-as-you-goGlobal | $0.037 | — | $0.15 | — | Official ↗ |
DeepSeekDeepSeek V4 Flashdeepseek-v4-flash | Off-peak · scheduled UTC windowsGlobalFrom 16 Aug 2026 | $0.22 | $0.007 | $0.66 | 1M | Official ↗ |
DeepSeekDeepSeek V4 Flashdeepseek-v4-flash | Peak · 01:00–04:00 and 06:00–10:00 UTCGlobalFrom 16 Aug 2026 | $0.44 | $0.014 | $1.32 | 1M | Official ↗ |
DeepSeekDeepSeek V4 Flashdeepseek-v4-flash | Standard · through 2026-08-16 16:00 UTCGlobalThrough 16 Aug 2026 | $0.14 | $0.0028 | $0.28 | 1M | Official ↗ |
DeepSeekDeepSeek V4 Prodeepseek-v4-pro | Off-peak · scheduled UTC windowsGlobalFrom 16 Aug 2026 | $0.66 | $0.022 | $1.98 | 1M | Official ↗ |
DeepSeekDeepSeek V4 Prodeepseek-v4-pro | Peak · 01:00–04:00 and 06:00–10:00 UTCGlobalFrom 16 Aug 2026 | $1.32 | $0.044 | $3.96 | 1M | Official ↗ |
DeepSeekDeepSeek V4 Prodeepseek-v4-pro | Standard · through 2026-08-16 16:00 UTCGlobalThrough 16 Aug 2026 | $0.435 | $0.003625 | $0.87 | 1M | Official ↗ |
GoogleGemini 3.1 Flash-Litegemini-3.1-flash-lite | Developer API · Standard · text/image/videoGemini Developer API | $0.25 | $0.025 | $1.50 | — | Official ↗ |
GoogleGemini 3.1 Pro Previewgemini-3.1-pro-preview | Developer API · Standard · over 200K promptGemini Developer API | $4 | $0.4 | $18 | — | Official ↗ |
GoogleGemini 3.1 Pro Previewgemini-3.1-pro-preview | Developer API · Standard · up to 200K promptGemini Developer API | $2 | $0.2 | $12 | — | Official ↗ |
GoogleGemini 3.5 Flashgemini-3.5-flash | Developer API · StandardGemini Developer API | $1.50 | $0.15 | $9 | — | Official ↗ |
GoogleGemini 3.5 Flash-Litegemini-3.5-flash-lite | Developer API · StandardGemini Developer API | $0.3 | $0.03 | $2.50 | — | Official ↗ |
GoogleGemini 3.6 Flashgemini-3.6-flash | Developer API · Standard · introductoryGemini Developer APIThrough 31 Dec 2026 | $0.75 | $0.075 | $3.75 | — | Official ↗ |
GoogleGemini 3.7 Flashgemini-3.7-flash | Developer API · Standard · introductoryGemini Developer APIThrough 31 Dec 2026 | $0.75 | $0.075 | $3.75 | — | Official ↗ |
Kimi / Moonshot AIKimi K2.6kimi-k2.6 | RealtimeGlobal | $0.95 | $0.16 | $4 | 262K | Official ↗ |
Kimi / Moonshot AIKimi K2.7 Code Highspeedkimi-k2.7-code-highspeed | Realtime · high-speed tierGlobal | $1.90 | $0.38 | $8 | 262K | Official ↗ |
Kimi / Moonshot AIKimi K2.7 Codekimi-k2.7-code | RealtimeGlobal | $0.95 | $0.19 | $4 | 262K | Official ↗ |
Kimi / Moonshot AIKimi K3kimi-k3 | RealtimeGlobal | $3 | $0.3 | $15 | 1.05M | Official ↗ |
Mistral AICodestralcodestral-latest | StandardGlobal | $0.3 | $0.03 | $0.9 | — | Official ↗ |
Mistral AIMinistral 3 14Bministral-14b-latest | StandardGlobal | $0.2 | $0.02 | $0.2 | — | Official ↗ |
Mistral AIMinistral 3 3Bministral-3b-latest | StandardGlobal | $0.1 | $0.01 | $0.1 | — | Official ↗ |
Mistral AIMinistral 3 8Bministral-8b-latest | StandardGlobal | $0.15 | $0.015 | $0.15 | — | Official ↗ |
Mistral AIMistral Large 3mistral-large-latest | StandardGlobal | $0.5 | $0.05 | $1.50 | — | Official ↗ |
Mistral AIMistral Medium 3.5mistral-medium-latest | StandardGlobal | $1.50 | $0.15 | $7.50 | — | Official ↗ |
Mistral AIMistral Small 4mistral-small-latest | StandardGlobal | $0.15 | $0.015 | $0.6 | — | Official ↗ |
OpenAIGPT-5.6 Lunagpt-5.6-luna | Standard · over 272K inputGlobal | $0.4 | $0.04 | $1.80 | 1.05M | Official ↗ |
OpenAIGPT-5.6 Lunagpt-5.6-luna | Standard · up to 272K inputGlobal | $0.2 | $0.02 | $1.20 | 1.05M | Official ↗ |
OpenAIGPT-5.6 Solgpt-5.6-sol | Standard · over 272K inputGlobal | $10 | $1 | $45 | 1.05M | Official ↗ |
OpenAIGPT-5.6 Solgpt-5.6-sol | Standard · up to 272K inputGlobal | $5 | $0.5 | $30 | 1.05M | Official ↗ |
OpenAIGPT-5.6 Terragpt-5.6-terra | Standard · over 272K inputGlobal | $4 | $0.4 | $18 | 1.05M | Official ↗ |
OpenAIGPT-5.6 Terragpt-5.6-terra | Standard · up to 272K inputGlobal | $2 | $0.2 | $12 | 1.05M | Official ↗ |
Qwen / Alibaba CloudQwen 3.7 Flashqwen3.7-flash | Beijing · 256K–1M inputChina Beijing | $0.165 | $0.033 | $0.66 | 1M | Official ↗ |
Qwen / Alibaba CloudQwen 3.7 Flashqwen3.7-flash | Beijing · 32K–256K inputChina Beijing | $0.083 | $0.017 | $0.33 | 1M | Official ↗ |
Qwen / Alibaba CloudQwen 3.7 Flashqwen3.7-flash | Beijing · up to 32K inputChina Beijing | $0.028 | $0.006 | $0.11 | 1M | Official ↗ |
Qwen / Alibaba CloudQwen 3.7 Maxqwen3.7-max | Global scope · up to 1M inputUS Virginia · Global deployment | $1.65 | $0.33 | $4.95 | 1M | Official ↗ |
Qwen / Alibaba CloudQwen 3.7 Plusqwen3.7-plus | Global scope · 256K–1M inputUS Virginia · Global deployment | $0.826 | $0.166 | $3.30 | 1M | Official ↗ |
Qwen / Alibaba CloudQwen 3.7 Plusqwen3.7-plus | Global scope · up to 256K inputUS Virginia · Global deployment | $0.276 | $0.056 | $1.10 | 1M | Official ↗ |
xAIGrok 4.20 Reasoninggrok-4.20-0309-reasoning | Standard · below 200K promptxAI direct API | $1.25 | $0.2 | $2.50 | 1M | Official ↗ |
xAIGrok 4.3grok-4.3 | Standard · 200K+ promptxAI direct API | $2.50 | $0.4 | $5 | 1M | Official ↗ |
xAIGrok 4.3grok-4.3 | Standard · below 200K promptxAI direct API | $1.25 | $0.2 | $2.50 | 1M | Official ↗ |
xAIGrok 4.5grok-4.5 | Standard · below 200K promptxAI direct API | $2 | $0.3 | $6 | 500K | Official ↗ |
xAIGrok 4.6grok-4.6 | Standard · 200K+ promptxAI direct API | $4 | $1 | $12 | 500K | Official ↗ |
xAIGrok 4.6grok-4.6 | Standard · below 200K promptxAI direct API | $2 | $0.5 | $6 | 500K | Official ↗ |
xAIGrok Build 0.1grok-build-0.1 | Standard · below 200K promptxAI direct API | $1 | $0.2 | $2 | 256K | Official ↗ |
Rates in USD per 1M tokens
claude-fable-5StandardGlobal routing
claude-haiku-4-5StandardGlobal routing
claude-opus-5StandardGlobal routing
claude-sonnet-4-6StandardGlobal routing
claude-sonnet-5StandardGlobal routing
command-a-03-2025Production pay-as-you-goGlobal
command-r-08-2024Production pay-as-you-goGlobal
command-r-plus-08-2024Production pay-as-you-goGlobal
command-r7b-12-2024Production pay-as-you-goGlobal
deepseek-v4-flashOff-peak · scheduled UTC windowsGlobal
deepseek-v4-flashPeak · 01:00–04:00 and 06:00–10:00 UTCGlobal
deepseek-v4-flashStandard · through 2026-08-16 16:00 UTCGlobal
deepseek-v4-proOff-peak · scheduled UTC windowsGlobal
deepseek-v4-proPeak · 01:00–04:00 and 06:00–10:00 UTCGlobal
deepseek-v4-proStandard · through 2026-08-16 16:00 UTCGlobal
gemini-3.1-flash-liteDeveloper API · Standard · text/image/videoGemini Developer API
gemini-3.1-pro-previewDeveloper API · Standard · over 200K promptGemini Developer API
gemini-3.1-pro-previewDeveloper API · Standard · up to 200K promptGemini Developer API
gemini-3.5-flashDeveloper API · StandardGemini Developer API
gemini-3.5-flash-liteDeveloper API · StandardGemini Developer API
gemini-3.6-flashDeveloper API · Standard · introductoryGemini Developer API
gemini-3.7-flashDeveloper API · Standard · introductoryGemini Developer API
kimi-k2.6RealtimeGlobal
kimi-k2.7-code-highspeedRealtime · high-speed tierGlobal
kimi-k2.7-codeRealtimeGlobal
kimi-k3RealtimeGlobal
codestral-latestStandardGlobal
ministral-14b-latestStandardGlobal
ministral-3b-latestStandardGlobal
ministral-8b-latestStandardGlobal
mistral-large-latestStandardGlobal
mistral-medium-latestStandardGlobal
mistral-small-latestStandardGlobal
gpt-5.6-lunaStandard · over 272K inputGlobal
gpt-5.6-lunaStandard · up to 272K inputGlobal
gpt-5.6-solStandard · over 272K inputGlobal
gpt-5.6-solStandard · up to 272K inputGlobal
gpt-5.6-terraStandard · over 272K inputGlobal
gpt-5.6-terraStandard · up to 272K inputGlobal
qwen3.7-flashBeijing · 256K–1M inputChina Beijing
qwen3.7-flashBeijing · 32K–256K inputChina Beijing
qwen3.7-flashBeijing · up to 32K inputChina Beijing
qwen3.7-maxGlobal scope · up to 1M inputUS Virginia · Global deployment
qwen3.7-plusGlobal scope · 256K–1M inputUS Virginia · Global deployment
qwen3.7-plusGlobal scope · up to 256K inputUS Virginia · Global deployment
grok-4.20-0309-reasoningStandard · below 200K promptxAI direct API
grok-4.3Standard · 200K+ promptxAI direct API
grok-4.3Standard · below 200K promptxAI direct API
grok-4.5Standard · below 200K promptxAI direct API
grok-4.6Standard · 200K+ promptxAI direct API
grok-4.6Standard · below 200K promptxAI direct API
grok-build-0.1Standard · below 200K promptxAI direct API
USD per 1M tokens. Snapshot verified 2026-08-16. Cache writes, cache storage, tools, regions, and provider-specific thresholds may be billed separately.
A missing cache rate is shown as “—”, never treated as free. Consumer chat-plan quotas are not API prices.
Choose a real provider tier, enter your workload and quality pass rates, then compare raw spend with cost per accepted answer. Every number remains an estimate until a quality-gated test confirms it.
Candidate break-even: 60.6% pass rate. Your scenario assumes 85%.
Warm cache-read scenario, not general ROI. It excludes cache writes/storage, tools, regional uplifts, retries, and quality failures. A dash means no published cache-read rate.
OpenAI API pricing ↗Official-source snapshot: 2026-08-16. API billing is separate from ChatGPT, Claude, Gemini, Grok, or Kimi consumer-plan quotas.
Reasoning effort should be matched to task complexity rather than fixed at the highest setting.
Applies to: OpenAI
Testing: Controlled lab recipe
An output-token limit bounds the most expensive side of a runaway generation.
Applies to: OpenAI
Testing: Controlled lab recipe
Rules often appear in system text, examples, tool descriptions, and the user prompt at the same time.
Applies to: OpenAI
Testing: Guided protocol
The curated cards stay distinct from 1,316 atomic candidates and 1,184 compound configurations in the server-filtered research atlas.

Set the acceptance rubric and allowed regression before looking at the outputs.
Run the same inputs through baseline and candidate, recording exact model and cache state.
Charge retries, fallbacks, tools, latency, and rejected answers to the arm that caused them.
Choose on cost per quality-passing answer, not the nicest token-reduction percentage.
Pro includes the evidence library and every encrypted bring-your-own-key adapter. Higher tiers expand dashboard and export depth instead of withholding providers.
The complete evidence library, research atlas, and every supported bring-your-own-key adapter.
No API credits included. Provider requests use your own connection and billing. Savings are not guaranteed.
Review and export a larger set of measured experiments.
No API credits included. Provider requests use your own connection and billing. Savings are not guaranteed.
The largest current dashboard and export window for repeat testing.
No API credits included. Provider requests use your own connection and billing. Savings are not guaranteed.