LLM COST INTELLIGENCE

Should you lower reasoning effort?

Run the same task at two settings. Compare the token totals and both outputs before you change a production workflow.

CONTROLLED DEMONSTRATION · 2026-08-16

One step lower used 15 fewer tokens in this run.

The same short task and shared instructions ran on GPT-5.5 at medium and low reasoning effort. Both outputs contained the requested three bullets; a human still has to judge whether either answer meets the real product rubric.

Repeat it with your own task
Medium effort
172
63 input · 109 output · 20 reasoning
Low effort
157
63 input · 94 output · 0 reasoning
Observed delta
−8.7%
15 fewer total tokens

Demonstration, not customer evidence or an API invoice. ChatGPT-plan requests count against plan limits; provider-key requests are billed by that provider. Token reduction is not a quality verdict or guaranteed saving.

OFFICIAL API RATE CARDS

One directory.
No fake equivalence.

Gemini, Grok, Kimi, and Qwen are first-class—not footnotes. Provider-specific cache, region, and long-context rules remain separate instead of being flattened into misleading averages.

Showing 52 of 52 rate cards

Official provider API rate cards, USD per one million tokens
Provider / modelPrice scopeInput
USD / 1M
Cache read
USD / 1M
Output
USD / 1M
ContextSource
AnthropicClaude Fable 5claude-fable-5StandardGlobal routing$10$1$501MOfficial ↗
AnthropicClaude Haiku 4.5claude-haiku-4-5StandardGlobal routing$1$0.1$5200KOfficial ↗
AnthropicClaude Opus 5claude-opus-5StandardGlobal routing$5$0.5$251MOfficial ↗
AnthropicClaude Sonnet 4.6claude-sonnet-4-6StandardGlobal routing$3$0.3$151MOfficial ↗
AnthropicClaude Sonnet 5claude-sonnet-5StandardGlobal routing$2$0.2$101MOfficial ↗
CohereCommand Acommand-a-03-2025Production pay-as-you-goGlobal$2.50$10256KOfficial ↗
CohereCommand Rcommand-r-08-2024Production pay-as-you-goGlobal$0.15$0.6Official ↗
CohereCommand R+command-r-plus-08-2024Production pay-as-you-goGlobal$2.50$10Official ↗
CohereCommand R7Bcommand-r7b-12-2024Production pay-as-you-goGlobal$0.037$0.15Official ↗
DeepSeekDeepSeek V4 Flashdeepseek-v4-flashOff-peak · scheduled UTC windowsGlobalFrom 16 Aug 2026$0.22$0.007$0.661MOfficial ↗
DeepSeekDeepSeek V4 Flashdeepseek-v4-flashPeak · 01:00–04:00 and 06:00–10:00 UTCGlobalFrom 16 Aug 2026$0.44$0.014$1.321MOfficial ↗
DeepSeekDeepSeek V4 Flashdeepseek-v4-flashStandard · through 2026-08-16 16:00 UTCGlobalThrough 16 Aug 2026$0.14$0.0028$0.281MOfficial ↗
DeepSeekDeepSeek V4 Prodeepseek-v4-proOff-peak · scheduled UTC windowsGlobalFrom 16 Aug 2026$0.66$0.022$1.981MOfficial ↗
DeepSeekDeepSeek V4 Prodeepseek-v4-proPeak · 01:00–04:00 and 06:00–10:00 UTCGlobalFrom 16 Aug 2026$1.32$0.044$3.961MOfficial ↗
DeepSeekDeepSeek V4 Prodeepseek-v4-proStandard · through 2026-08-16 16:00 UTCGlobalThrough 16 Aug 2026$0.435$0.003625$0.871MOfficial ↗
GoogleGemini 3.1 Flash-Litegemini-3.1-flash-liteDeveloper API · Standard · text/image/videoGemini Developer API$0.25$0.025$1.50Official ↗
GoogleGemini 3.1 Pro Previewgemini-3.1-pro-previewDeveloper API · Standard · over 200K promptGemini Developer API$4$0.4$18Official ↗
GoogleGemini 3.1 Pro Previewgemini-3.1-pro-previewDeveloper API · Standard · up to 200K promptGemini Developer API$2$0.2$12Official ↗
GoogleGemini 3.5 Flashgemini-3.5-flashDeveloper API · StandardGemini Developer API$1.50$0.15$9Official ↗
GoogleGemini 3.5 Flash-Litegemini-3.5-flash-liteDeveloper API · StandardGemini Developer API$0.3$0.03$2.50Official ↗
GoogleGemini 3.6 Flashgemini-3.6-flashDeveloper API · Standard · introductoryGemini Developer APIThrough 31 Dec 2026$0.75$0.075$3.75Official ↗
GoogleGemini 3.7 Flashgemini-3.7-flashDeveloper API · Standard · introductoryGemini Developer APIThrough 31 Dec 2026$0.75$0.075$3.75Official ↗
Kimi / Moonshot AIKimi K2.6kimi-k2.6RealtimeGlobal$0.95$0.16$4262KOfficial ↗
Kimi / Moonshot AIKimi K2.7 Code Highspeedkimi-k2.7-code-highspeedRealtime · high-speed tierGlobal$1.90$0.38$8262KOfficial ↗
Kimi / Moonshot AIKimi K2.7 Codekimi-k2.7-codeRealtimeGlobal$0.95$0.19$4262KOfficial ↗
Kimi / Moonshot AIKimi K3kimi-k3RealtimeGlobal$3$0.3$151.05MOfficial ↗
Mistral AICodestralcodestral-latestStandardGlobal$0.3$0.03$0.9Official ↗
Mistral AIMinistral 3 14Bministral-14b-latestStandardGlobal$0.2$0.02$0.2Official ↗
Mistral AIMinistral 3 3Bministral-3b-latestStandardGlobal$0.1$0.01$0.1Official ↗
Mistral AIMinistral 3 8Bministral-8b-latestStandardGlobal$0.15$0.015$0.15Official ↗
Mistral AIMistral Large 3mistral-large-latestStandardGlobal$0.5$0.05$1.50Official ↗
Mistral AIMistral Medium 3.5mistral-medium-latestStandardGlobal$1.50$0.15$7.50Official ↗
Mistral AIMistral Small 4mistral-small-latestStandardGlobal$0.15$0.015$0.6Official ↗
OpenAIGPT-5.6 Lunagpt-5.6-lunaStandard · over 272K inputGlobal$0.4$0.04$1.801.05MOfficial ↗
OpenAIGPT-5.6 Lunagpt-5.6-lunaStandard · up to 272K inputGlobal$0.2$0.02$1.201.05MOfficial ↗
OpenAIGPT-5.6 Solgpt-5.6-solStandard · over 272K inputGlobal$10$1$451.05MOfficial ↗
OpenAIGPT-5.6 Solgpt-5.6-solStandard · up to 272K inputGlobal$5$0.5$301.05MOfficial ↗
OpenAIGPT-5.6 Terragpt-5.6-terraStandard · over 272K inputGlobal$4$0.4$181.05MOfficial ↗
OpenAIGPT-5.6 Terragpt-5.6-terraStandard · up to 272K inputGlobal$2$0.2$121.05MOfficial ↗
Qwen / Alibaba CloudQwen 3.7 Flashqwen3.7-flashBeijing · 256K–1M inputChina Beijing$0.165$0.033$0.661MOfficial ↗
Qwen / Alibaba CloudQwen 3.7 Flashqwen3.7-flashBeijing · 32K–256K inputChina Beijing$0.083$0.017$0.331MOfficial ↗
Qwen / Alibaba CloudQwen 3.7 Flashqwen3.7-flashBeijing · up to 32K inputChina Beijing$0.028$0.006$0.111MOfficial ↗
Qwen / Alibaba CloudQwen 3.7 Maxqwen3.7-maxGlobal scope · up to 1M inputUS Virginia · Global deployment$1.65$0.33$4.951MOfficial ↗
Qwen / Alibaba CloudQwen 3.7 Plusqwen3.7-plusGlobal scope · 256K–1M inputUS Virginia · Global deployment$0.826$0.166$3.301MOfficial ↗
Qwen / Alibaba CloudQwen 3.7 Plusqwen3.7-plusGlobal scope · up to 256K inputUS Virginia · Global deployment$0.276$0.056$1.101MOfficial ↗
xAIGrok 4.20 Reasoninggrok-4.20-0309-reasoningStandard · below 200K promptxAI direct API$1.25$0.2$2.501MOfficial ↗
xAIGrok 4.3grok-4.3Standard · 200K+ promptxAI direct API$2.50$0.4$51MOfficial ↗
xAIGrok 4.3grok-4.3Standard · below 200K promptxAI direct API$1.25$0.2$2.501MOfficial ↗
xAIGrok 4.5grok-4.5Standard · below 200K promptxAI direct API$2$0.3$6500KOfficial ↗
xAIGrok 4.6grok-4.6Standard · 200K+ promptxAI direct API$4$1$12500KOfficial ↗
xAIGrok 4.6grok-4.6Standard · below 200K promptxAI direct API$2$0.5$6500KOfficial ↗
xAIGrok Build 0.1grok-build-0.1Standard · below 200K promptxAI direct API$1$0.2$2256KOfficial ↗

Rates in USD per 1M tokens

DeepSeek

DeepSeek V4 Flash

deepseek-v4-flash

Off-peak · scheduled UTC windowsGlobal

Input
$0.22
Cache read
$0.007
Output
$0.66
Context
1M
Effective from 16 Aug 2026Open official source in a new tab
DeepSeek

DeepSeek V4 Flash

deepseek-v4-flash

Peak · 01:00–04:00 and 06:00–10:00 UTCGlobal

Input
$0.44
Cache read
$0.014
Output
$1.32
Context
1M
Effective from 16 Aug 2026Open official source in a new tab
DeepSeek

DeepSeek V4 Flash

deepseek-v4-flash

Standard · through 2026-08-16 16:00 UTCGlobal

Input
$0.14
Cache read
$0.0028
Output
$0.28
Context
1M
Effective through 16 Aug 2026Open official source in a new tab
DeepSeek

DeepSeek V4 Pro

deepseek-v4-pro

Off-peak · scheduled UTC windowsGlobal

Input
$0.66
Cache read
$0.022
Output
$1.98
Context
1M
Effective from 16 Aug 2026Open official source in a new tab
DeepSeek

DeepSeek V4 Pro

deepseek-v4-pro

Peak · 01:00–04:00 and 06:00–10:00 UTCGlobal

Input
$1.32
Cache read
$0.044
Output
$3.96
Context
1M
Effective from 16 Aug 2026Open official source in a new tab
DeepSeek

DeepSeek V4 Pro

deepseek-v4-pro

Standard · through 2026-08-16 16:00 UTCGlobal

Input
$0.435
Cache read
$0.003625
Output
$0.87
Context
1M
Effective through 16 Aug 2026Open official source in a new tab
Google

Gemini 3.1 Flash-Lite

gemini-3.1-flash-lite

Developer API · Standard · text/image/videoGemini Developer API

Input
$0.25
Cache read
$0.025
Output
$1.50
Context
Open official source in a new tab
Google

Gemini 3.1 Pro Preview

gemini-3.1-pro-preview

Developer API · Standard · over 200K promptGemini Developer API

Input
$4
Cache read
$0.4
Output
$18
Context
Open official source in a new tab
Google

Gemini 3.1 Pro Preview

gemini-3.1-pro-preview

Developer API · Standard · up to 200K promptGemini Developer API

Input
$2
Cache read
$0.2
Output
$12
Context
Open official source in a new tab
Google

Gemini 3.6 Flash

gemini-3.6-flash

Developer API · Standard · introductoryGemini Developer API

Input
$0.75
Cache read
$0.075
Output
$3.75
Context
Effective through 31 Dec 2026Open official source in a new tab
Google

Gemini 3.7 Flash

gemini-3.7-flash

Developer API · Standard · introductoryGemini Developer API

Input
$0.75
Cache read
$0.075
Output
$3.75
Context
Effective through 31 Dec 2026Open official source in a new tab
Qwen / Alibaba Cloud

Qwen 3.7 Max

qwen3.7-max

Global scope · up to 1M inputUS Virginia · Global deployment

Input
$1.65
Cache read
$0.33
Output
$4.95
Context
1M
Open official source in a new tab
Qwen / Alibaba Cloud

Qwen 3.7 Plus

qwen3.7-plus

Global scope · 256K–1M inputUS Virginia · Global deployment

Input
$0.826
Cache read
$0.166
Output
$3.30
Context
1M
Open official source in a new tab
Qwen / Alibaba Cloud

Qwen 3.7 Plus

qwen3.7-plus

Global scope · up to 256K inputUS Virginia · Global deployment

Input
$0.276
Cache read
$0.056
Output
$1.10
Context
1M
Open official source in a new tab

USD per 1M tokens. Snapshot verified 2026-08-16. Cache writes, cache storage, tools, regions, and provider-specific thresholds may be billed separately.

A missing cache rate is shown as “—”, never treated as free. Consumer chat-plan quotas are not API prices.

FREE SCENARIO CALCULATOR

Price the workload,
then the optimization.

Choose a real provider tier, enter your workload and quality pass rates, then compare raw spend with cost per accepted answer. Every number remains an estimate until a quality-gated test confirms it.

Estimated monthly API spend · OpenAI
GPT-5.6 TerraBefore · Standard · up to 272K input: $2 input · $0.2 cache read · $12 output / 1MAfter · Standard · up to 272K input: $2 input · $0.2 cache read · $12 output / 1MGlobal · price bands selected automatically from input size
Before$120.00
After$80.85
Potential saving: $39.15 (32.6%)
QUALITY-ADJUSTED COST
Before / accepted answer$0.013390% pass assumption
After / accepted answer$0.009585% pass assumption

Candidate break-even: 60.6% pass rate. Your scenario assumes 85%.

Warm cache-read scenario, not general ROI. It excludes cache writes/storage, tools, regional uplifts, retries, and quality failures. A dash means no published cache-read rate.

OpenAI API pricing
NEXT STEPValidate the 28.7% accepted-answer advantage.

The estimate is not a saving until the same workload still passes its quality bar. Compare a supported recipe in the lab, or use the evidence library to design a provider-specific test.

Official-source snapshot: 2026-08-16. API billing is separate from ChatGPT, Claude, Gemini, Grok, or Kimi consumer-plan quotas.

FREE METHOD LIBRARY

Try the method.
Keep the evidence.

Official factDerived mathTest protocol
ReasoningOfficial fact

Lower reasoning effort one step at a time

Reasoning effort should be matched to task complexity rather than fixed at the highest setting.

Applies to: OpenAI
Testing: Controlled lab recipe

OpenAI model guidanceChecked 2026-08-15
OutputOfficial fact

Give every response an output ceiling

An output-token limit bounds the most expensive side of a runaway generation.

Applies to: OpenAI
Testing: Controlled lab recipe

OpenAI model guidanceChecked 2026-08-15
Prompt designTest protocol

Remove duplicated instructions before shortening prose

Rules often appear in system text, examples, tool descriptions, and the user prompt at the same time.

Applies to: OpenAI
Testing: Guided protocol

OpenAI model guidanceChecked 2026-08-15

108 Pro cards + 2,500 research rows

The curated cards stay distinct from 1,316 atomic candidates and 1,184 compound configurations in the server-filtered research atlas.

EXPERIMENT STANDARD

A cheaper answer only wins
when it still works.

Abstract token tiles moving through a calibrated gauge and emerging as a smaller organized set
  1. 01

    Declare quality

    Set the acceptance rubric and allowed regression before looking at the outputs.

  2. 02

    Pair the trials

    Run the same inputs through baseline and candidate, recording exact model and cache state.

  3. 03

    Count every attempt

    Charge retries, fallbacks, tools, latency, and rejected answers to the arm that caused them.

  4. 04

    Ship accepted-cost winners

    Choose on cost per quality-passing answer, not the nicest token-reduction percentage.

ONE-TIME ACCESS

Choose the workbench you will actually use.

Pro includes the evidence library and every encrypted bring-your-own-key adapter. Higher tiers expand dashboard and export depth instead of withholding providers.

  • The public rate directory and free methods remain open
  • API credits are never bundled or resold
  • Usage totals retained; prompts and outputs discarded
  • 14-day refund policy; statutory rights unaffected
Pro access£5£9 · launch signup price

The complete evidence library, research atlas, and every supported bring-your-own-key adapter.

  • All 120 evidence cards + 2,500-row research atlas
  • All nine encrypted provider connections
  • 100-row experiment dashboard

No API credits included. Provider requests use your own connection and billing. Savings are not guaranteed.

Pro+ access£15£19 · launch signup price

Review and export a larger set of measured experiments.

  • Everything in Pro
  • 250-row dashboard history
  • Experiment CSV export

No API credits included. Provider requests use your own connection and billing. Savings are not guaranteed.

Ultimate access£20£39 · launch signup price

The largest current dashboard and export window for repeat testing.

  • Everything in Pro+
  • 1,000-row dashboard history
  • 1,000-row CSV export

No API credits included. Provider requests use your own connection and billing. Savings are not guaranteed.