Modeled token spend
Uses the selected input, cache-read, and output rates. It excludes tools and any billing category you did not enter.
Bring aggregate token totals and your provider bill. TokenGauge reconciles 52 dated rate cards, cache share, accepted-answer economics, and an explicitly approximate retry burden—entirely in your browser.
Rate snapshot verified 2026-08-16 · no login · no usage upload
Select the exact provider rate card, copy aggregate usage totals, and compare the token model with the reported bill. A remaining gap is a question to investigate—not an automatic overcharge claim.
Browser-local calculation. Token totals and bill values are not submitted to TokenGauge or stored.
Output is the largest modeled token-cost bucket. A published cache-read rate is applied only to the cached-input bucket.
Uniform-attempt estimate: $15.57 of modeled spend sits on non-accepted attempts. Real retry cost needs retry-specific token buckets.
The variance is a diagnostic gap, not proof of overbilling. Tools, cache writes/storage, images, audio, regional or priority uplifts, taxes, credits, rounding, and mixed price bands can explain it.
OpenAI API pricing ↗Uses the selected input, cache-read, and output rates. It excludes tools and any billing category you did not enter.
Flags the unexplained gap for investigation. Mixed tiers, media, cache writes, credits, or tax may account for it.
Stops cheap failed attempts from looking efficient. Validate changes against a declared quality bar.