Model pricing

AI Model Pricing

Every rate here was read from the provider's own pricing page on August 22, 2026. Compare input, cached and output prices across 56 models from 11 providers — then enter your own token volumes to see what a single request would cost on each one.

  • 56 models
  • 11 providers
  • Checked August 22, 2026
Price a request
Presets

Click a column heading to sort.

Input, cached input and output prices per million tokens for 56 AI models, with each model's context window, token-counting accuracy and estimated cost per request
per 1Mper 1Mper 1MCountingper request
GPT-5.6 Sol1OpenAI$4.00$0.400$20.001Mexact$0.060
GPT-5.6 Terra1OpenAI$2.00$0.200$12.001Mexact$0.032
GPT-5.6 Luna1OpenAI$0.200$0.020$1.201Mexact$0.00320
GPT-5.51OpenAI$5.00$0.500$30.001Mexact$0.080
GPT-5.41OpenAI$2.50$0.250$15.001Mexact$0.040
GPT-5.4 mini1OpenAI$0.750$0.075$4.50400Kexact$0.012
GPT-5.4 nano1OpenAI$0.200$0.020$1.25400Kexact$0.00325
GPT-5.11OpenAI$1.25$0.125$10.00400Kexact$0.022
GPT-4.11OpenAI$2.00$0.500$8.001Mexact$0.028
GPT-4o1OpenAI$2.50$1.25$10.00128Kexact$0.035
GPT-4o Mini1OpenAI$0.150$0.075$0.600128Kexact$0.00210
Claude Fable 52Anthropic$10.00$1.00$50.001Mest.$0.150
Claude Opus 52Anthropic$5.00$0.500$25.001Mest.$0.075
Claude Opus 4.82Anthropic$5.00$0.500$25.001Mest.$0.075
Claude Opus 4.72Anthropic$5.00$0.500$25.001Mest.$0.075
Claude Opus 4.62Anthropic$5.00$0.500$25.001Mest.$0.075
Claude Opus 4.52Anthropic$5.00$0.500$25.00200Kest.$0.075
Claude Sonnet 52Anthropic$2.00$0.200$10.001Mest.$0.030
Claude Sonnet 4.62Anthropic$3.00$0.300$15.001Mest.$0.045
Claude Sonnet 4.52Anthropic$3.00$0.300$15.00200Kest.$0.045
Claude Haiku 4.52Anthropic$1.00$0.100$5.00200Kest.$0.015
Gemini 3.1 Pro3Google$2.00$0.200$12.001Mest.$0.032
Gemini 3.7 FlashGoogle$0.750$0.075$3.751Mest.$0.011
Gemini 3.5 FlashGoogle$1.50$0.150$9.001Mest.$0.024
Gemini 3.5 Flash-LiteGoogle$0.300$0.030$2.501Mest.$0.00550
Gemini 2.5 Pro3Google$1.25$0.125$10.001Mest.$0.022
Grok 4.64xAI$2.00$0.500$6.00500Kest.$0.026
Grok 4.54xAI$2.00$0.300$6.00500Kest.$0.026
Grok 4.34xAI$1.25$0.200$2.501Mest.$0.015
Grok Build 0.14xAI$1.00$0.200$2.00256Kest.$0.012
DeepSeek V4 Pro5DeepSeek$1.32$0.044$3.961Mest.$0.017
DeepSeek V4 Flash5DeepSeek$0.440$0.014$1.321Mest.$0.00572
LLaMA 3.3 70B6Meta (via API)$1.04$1.04128Kest.$0.011
Qwen3.7 Max1617Alibaba$2.50$0.250$7.501Mest.$0.033
Sonar7918Perplexity$1.00$1.00est.$0.011
Sonar Pro7918Perplexity$3.00$15.00est.$0.045
Sonar Reasoning Pro918Perplexity$2.00$8.00est.$0.028
Sonar Deep Research8918Perplexity$2.00$8.00est.$0.028
Kimi K310Moonshot AI$3.00$0.300$15.001Mest.$0.045
Kimi K2.7 Code1011Moonshot AI$0.950$0.190$4.00262Kest.$0.013
Kimi K2.610Moonshot AI$0.950$0.160$4.00262Kest.$0.013
MiniMax M31213MiniMax$0.300$0.060$1.201Mest.$0.00420
MiniMax M2.714MiniMax$0.300$0.060$1.20205Kest.$0.00420
MiniMax M2.5MiniMax$0.300$0.030$1.20205Kest.$0.00420
GLM-5.315Z.ai$1.40$0.280$4.401Mest.$0.018
GLM-5.215Z.ai$1.40$0.280$4.401Mest.$0.018
GLM-515Z.ai$1.00$0.200$3.20200Kest.$0.013
GLM-5-Turbo15Z.ai$1.20$0.240$4.00200Kest.$0.016
GLM-4.715Z.ai$0.600$0.120$2.20200Kest.$0.00820
GLM-4.7-FlashX15Z.ai$0.070$0.014$0.400200Kest.$0.00110
GLM-4.5-Air15Z.ai$0.200$0.040$1.10128Kest.$0.00310
GLM-4.5-X1518Z.ai$2.20$0.440$8.90est.$0.031
Qwen3.7 Plus1617Alibaba$0.400$0.040$1.601Mest.$0.00560
Qwen3.6 Flash1617Alibaba$0.250$0.025$1.501Mest.$0.00400
Qwen3.5 Flash17Alibaba$0.100$0.010$0.4001Mest.$0.00140
Qwen3 Coder Plus1617Alibaba$1.00$0.100$5.001Mest.$0.015

Notes

  1. OpenAI bills prompts above 272K input tokens at a higher rate; the figure shown is the standard tier.
  2. The cached column is Anthropic’s cache-hit read rate. Writing to the cache costs more than ordinary input.
  3. Google bills prompts above 200K input tokens at roughly double these rates.
  4. From 200K prompt tokens, xAI bills the whole request at double these input, cached and output rates.
  5. Standard (peak) rates. DeepSeek bills 01:00–04:00 and 06:00–10:00 UTC at half of them.
  6. Meta ships weights rather than an API, so this is a hosted serverless rate (Together AI) rather than a first-party price.
  7. On top of the token rates, Perplexity charges for search per request: $5, $8 or $12 per 1,000 requests depending on search context size.
  8. Deep Research also bills $2 per 1M citation tokens, $3 per 1M reasoning tokens and $5 per 1,000 searches.
  9. Sonar chat completions have moved to Perplexity’s Agent API; the Sonar models are supported until 27 September 2026.
  10. Moonshot quotes these rates excluding tax.
  11. The -highspeed variant bills input and output at double these rates.
  12. Standard service tier at up to 512K input tokens. Above that the whole request bills at double, and the priority tier costs about 1.5×.
  13. MiniMax lists this as a standing 50% discount; the undiscounted list price is double.
  14. The -highspeed variant bills at double these rates.
  15. Z.ai’s cache-hit reads bill at one fifth of the input rate.
  16. Alibaba tiers by request size: crossing an input threshold (256K on most models, 32K on Qwen3 Coder Plus) re-prices the entire request — up to $1.20 input / $4.80 output on Qwen3.7 Plus, and $6.00 / $60.00 on Qwen3 Coder Plus.
  17. Alibaba’s Singapore (international) price list. The Beijing region is billed separately and differs.
  18. The provider does not publish a context window for this model, so the column reads “—”.

All rates are in USD per 1M tokens.

Checked against each provider’s own pricing page on August 22, 2026.

Providers change prices without notice, and several bill extra beyond the token rate — confirm against the provider’s page before you budget.

Cheapest first

Best value per million tokens

Output rates in this table run from $0.400 to $50.00 per 1M tokens — a 125× spread, which is a bigger lever than any prompt tuning.

Cheapest input

  1. 1GLM-4.7-FlashXZ.ai$0.070per 1M tokens
  2. 2Qwen3.5 FlashAlibaba$0.100per 1M tokens
  3. 3GPT-4o MiniOpenAI$0.150per 1M tokens
  4. 4GLM-4.5-AirZ.ai$0.200per 1M tokens
  5. 5GPT-5.6 LunaOpenAI$0.200per 1M tokens

Cheapest output

  1. 1GLM-4.7-FlashXZ.ai$0.400per 1M tokens
  2. 2Qwen3.5 FlashAlibaba$0.400per 1M tokens
  3. 3GPT-4o MiniOpenAI$0.600per 1M tokens
  4. 4SonarPerplexity$1.00per 1M tokens
  5. 5LLaMA 3.3 70BMeta (via API)$1.04per 1M tokens
By provider

Providers in this table

11 providers, each linked to the pricing page these figures were read from.

FAQ

Frequently asked questions

Which AI model is cheapest?

On input, GLM-4.7-FlashX at $0.070 per 1M tokens. On output, GLM-4.7-FlashX at $0.400 per 1M. Which one is cheapest for you depends on the shape of your traffic: a 10,000-token prompt with a 1,000-token answer costs $0.00110 on GLM-4.7-FlashX and $0.032 on GPT-5.6 Terra. Set the two token fields above to your own volumes and the table re-sorts around them.

How do I compare model prices fairly?

Weight input and output by how you actually use the model. Providers charge three to five times more for what the model writes than for what you send — GPT-5.6 Terra is $2.00 in and $12.00 out — so a summariser that reads a lot and writes a little ranks the models completely differently from a code generator that does the opposite. Enter both volumes above rather than comparing input rates alone, and check the notes: several providers re-price the entire request once a prompt crosses an input threshold.

What is cached input pricing?

When consecutive requests share the same prefix — a system prompt, a document, a long few-shot block — providers can serve that prefix from cache and charge less for it. The cached column is that cache-hit read rate, usually 10% to 20% of the standard input rate: Claude Sonnet 5 reads cached input at $0.200 against $2.00 standard. Writing to the cache is not always free, and a dash means the provider publishes no cached rate at all.

How current is this pricing?

Every figure was read from the provider's own pricing page on August 22, 2026, and each provider card links back to the page it came from. Nothing here is scraped live, so treat it as a snapshot: providers change rates without notice, and the ones with promotional or off-peak pricing (MiniMax, DeepSeek) move most often.

Why do some cells show a dash?

Because the provider does not publish that figure, and guessing would be worse than leaving it blank. Perplexity publishes no context window for the Sonar models, Z.ai publishes no per-variant context length for GLM-4.5-X, and Meta’s hosted LLaMA row has no cached rate. A dash always means "not published", never "zero".

Which models can this site count tokens for exactly?

11 of the 56 rows — the OpenAI models, because OpenAI publishes its tokenizer and the calculator loads the real BPE vocabulary (o200k_base) in your browser. Everything else is marked "est." and counted with a characters-per-token approximation, which lands within roughly 10–15% on English prose but drifts on code, JSON and non-Latin scripts. The prices are first-party for every row; only the token counts differ in confidence.

Do these prices include search and request fees?

No — the table shows token rates only, and a few providers bill on top of them. Sonar adds $5 to $12 per 1,000 requests for search depending on context size, Sonar Deep Research adds citation-token, reasoning-token and per-search charges, and hosted open-weight models carry whatever the host charges rather than a first-party rate. The notes column flags every row where the token rate is not the whole bill.

What does a long prompt actually cost?

Take 200,000 input tokens — a full novel, or a few hundred pages of a contract. That single prompt costs $0.400 on GPT-5.6 Terra, $0.400 on Claude Sonnet 5, $0.600 on Kimi K3 and $0.020 on Qwen3.5 Flash. Watch the tier notes at that size: on Grok 4.6 a prompt of 200K tokens bills the whole request at double the listed rates.

Count the tokens first

A price per million tokens only means something once you know how many tokens your text is. The token calculator counts as you type and prices the result across models; the file counter does the same for PDFs and Word documents.