Qwen 3.7 Flash
qwen3.7-flashBeijing · 256K–1M inputChina Beijing
- Input
- $0.165
- Cache read
- $0.033
- Output
- $0.66
- Context
- 1M
Compare Qwen API rates while preserving region, input tier, cache mode, and context constraints.
Model a stable reusable prefix, each cache creation, and the reads that actually land inside its lifetime. The counterfactual weights every billing category at its published rate.
600 modeled requests across 100 cache episodes.
This is an input-cost counterfactual, not realized savings. It assumes the prefix stays byte-stable and every planned read hits within the selected TTL. Verify provider usage fields and accepted-answer quality.
Alibaba Cloud Qwen 3.7 Max pricing ↗Enter monthly workload, tokens, and quality pass rates. The calculator resolves eligible price bands, then shows cost per accepted answer and the candidate break-even pass rate.
Candidate break-even: 58.2% pass rate. Your scenario assumes 85%.
Warm cache-read scenario, not general ROI. It excludes cache writes/storage, tools, regional uplifts, retries, and quality failures. A dash means no published cache-read rate.
Alibaba Cloud Qwen 3.7 Max pricing ↗Qwen prices can depend on region, input size, and cache mode at the same time. TokenGauge retains those dimensions as distinct cards so a low tier is not silently applied to an ineligible request.
Rates are USD per one million tokens. Future and transitional cards remain labeled with their effective dates instead of silently replacing today’s tier.
| Provider / model | Price scope | Input USD / 1M | Cache read USD / 1M | Output USD / 1M | Context | Source |
|---|---|---|---|---|---|---|
Qwen / Alibaba CloudQwen 3.7 Flashqwen3.7-flash | Beijing · 256K–1M inputChina Beijing | $0.165 | $0.033 | $0.66 | 1M | Official ↗ |
Qwen / Alibaba CloudQwen 3.7 Flashqwen3.7-flash | Beijing · 32K–256K inputChina Beijing | $0.083 | $0.017 | $0.33 | 1M | Official ↗ |
Qwen / Alibaba CloudQwen 3.7 Flashqwen3.7-flash | Beijing · up to 32K inputChina Beijing | $0.028 | $0.006 | $0.11 | 1M | Official ↗ |
Qwen / Alibaba CloudQwen 3.7 Maxqwen3.7-max | Global scope · up to 1M inputUS Virginia · Global deployment | $1.65 | $0.33 | $4.95 | 1M | Official ↗ |
Qwen / Alibaba CloudQwen 3.7 Plusqwen3.7-plus | Global scope · 256K–1M inputUS Virginia · Global deployment | $0.826 | $0.166 | $3.30 | 1M | Official ↗ |
Qwen / Alibaba CloudQwen 3.7 Plusqwen3.7-plus | Global scope · up to 256K inputUS Virginia · Global deployment | $0.276 | $0.056 | $1.10 | 1M | Official ↗ |
Rates in USD per 1M tokens
qwen3.7-flashBeijing · 256K–1M inputChina Beijing
qwen3.7-flashBeijing · 32K–256K inputChina Beijing
qwen3.7-flashBeijing · up to 32K inputChina Beijing
qwen3.7-maxGlobal scope · up to 1M inputUS Virginia · Global deployment
qwen3.7-plusGlobal scope · 256K–1M inputUS Virginia · Global deployment
qwen3.7-plusGlobal scope · up to 256K inputUS Virginia · Global deployment
USD per 1M tokens. Snapshot verified 2026-08-16. Cache writes, cache storage, tools, regions, and provider-specific thresholds may be billed separately.
A missing cache rate is shown as “—”, never treated as free. Consumer chat-plan quotas are not API prices.
Qwen prices can depend on region, input size, and cache mode at the same time. TokenGauge retains those dimensions as distinct cards so a low tier is not silently applied to an ineligible request.
Count retries, fallbacks, tool calls, latency failures, and answers rejected by the quality rubric. A lower estimated token bill is useful only when the workload still succeeds.
Open the controlled A/B labThis page uses TokenGauge’s 2026-08-16 official-source snapshot. Follow the linked provider pages before making a production commitment because prices can change.
No. The workload calculator models input, output, and a warm cache-read share. The separate cache-episode calculator includes published cache writes and reads, but storage, tools, retries, regional uplifts, taxes, batch modes, and quality failures can still change the invoice.
No. ChatGPT, Claude, Gemini, Grok, Kimi, and other consumer-plan quotas are separate from provider API token billing.