Skip to content
LLM Abacus

LLM API Pricing Comparison

Compare 55 leading large language models in one table. Prices in USD per million tokens. Flagship models verified daily; others manually monitored.

ModelProviderInputOutputCachedContextQualityHeat
Qwen3.5 Flash🇨🇳 Alibaba Qwen$0.030$0.301M
Gemini 2.5 Flash-Lite🇺🇸 Google$0.10$0.40$0.0101M
Qwen3.5 Plus🇨🇳 Alibaba Qwen$0.12$0.711M
Doubao 1.6🇨🇳 ByteDance Doubao$0.12$1.18$0.024256K
Doubao 1.5 Pro🇨🇳 ByteDance Doubao$0.12$0.30$0.024256K
文心 ERNIE 4.5 Turbo🇨🇳 Baidu ERNIE$0.12$0.47$0.030128K
混元 TurboS🇨🇳 Tencent Hunyuan$0.12$0.30256K
DeepSeek V4 Flash🇨🇳 DeepSeek$0.15$0.30$0.0031M47#27.57T
文心 ERNIE X1 Turbo🇨🇳 Baidu ERNIE$0.15$0.59128K
混元 T1🇨🇳 Tencent Hunyuan$0.15$0.5964K
MiMo-V2.5🇨🇳 小米 MiMo$0.15$0.30$0.003262K#18.90T
混元 Hy3🇨🇳 Tencent Hunyuan$0.15$0.59$0.037262K#34.79T
Step 3.7 Flash🇨🇳 阶跃星辰$0.20$1.20131K#81.74T
GPT-5.6 Luna🇺🇸 OpenAI$0.20$1.20$0.0201.1M
Gemini 3.1 Flash-Lite🇺🇸 Google$0.25$1.50$0.0251M34
DeepSeek V3.2Retired 2026-07-24🇨🇳 DeepSeek$0.30$1.18$0.074128K
GLM-4.7🇨🇳 Zhipu AI$0.30$1.18$0.059200K
Spark X2 Flash🇨🇳 iFlytek Spark$0.30$0.30200K
Spark Ultra🇨🇳 iFlytek Spark$0.30$0.30128K
Baichuan M2🇨🇳 Baichuan AI$0.30$2.96192K
Gemini 2.5 Flash🇺🇸 Google$0.30$2.50$0.0301M
MiniMax M2.7🇨🇳 MiniMax$0.31$1.24$0.0621M50
Qwen3 Max🇨🇳 Alibaba Qwen$0.37$1.48256K
DeepSeek V4 Pro🇨🇳 DeepSeek$0.44$0.89$0.0041M52#43.59T
Spark X2🇨🇳 iFlytek Spark$0.44$0.44200K
混元 2.0 Think🇨🇳 Tencent Hunyuan$0.59$2.35128K
GLM-5🇨🇳 Zhipu AI$0.59$2.66$0.15200K
文心 ERNIE 5.1🇨🇳 Baidu ERNIE$0.59$2.66128K
MiniMax M3🇨🇳 MiniMax$0.62$2.49$0.121M55#72.02T
Baichuan M3 Plus🇨🇳 Baichuan AI$0.74$1.33192K
GLM-5.1🇨🇳 Zhipu AI$0.89$3.55$0.19200K51
Doubao Seed 2.1 Pro🇨🇳 ByteDance Doubao$0.89$4.44$0.18256K
Kimi K2.6🇨🇳 Moonshot (Kimi)$0.96$3.99$0.16262K54
Claude Haiku 4.5🇺🇸 Anthropic$1.00$5.00$0.10200K37
Grok Build 0.1🇺🇸 xAI$1.00$2.00$0.20256K
Spark Pro🇨🇳 iFlytek Spark$1.04$1.04128K
GLM-5.2🇨🇳 Zhipu AI$1.18$4.141M#53.24T
GPT-5.1🇺🇸 OpenAI$1.25$10.00$0.13400K
Gemini 2.5 Pro🇺🇸 Google$1.25$10.00$0.132M35
Grok 4.3🇺🇸 xAI$1.25$2.50$0.201M53
Gemini 3.6 Flash🇺🇸 Google$1.50$7.50$0.151M
Gemini 3.5 Flash🇺🇸 Google$1.50$9.00$0.151M55
Qwen3.7 Max🇨🇳 Alibaba Qwen$1.78$5.331M57
GPT-5.6 Terra🇺🇸 OpenAI$2.00$12.00$0.201.1M
Claude Sonnet 5🇺🇸 Anthropic$2.00$10.00$0.201M57
Gemini 3.1 Pro Preview🇺🇸 Google$2.00$12.00$0.202M57
GPT-5.4🇺🇸 OpenAI$2.50$15.00$0.25400K
Kimi K3🇨🇳 Moonshot (Kimi)$2.96$14.79$0.301.0M#91.34T
Claude Sonnet 4.6🇺🇸 Anthropic$3.00$15.00$0.301M52
GPT-5.5🇺🇸 OpenAI$5.00$30.00$0.50400K60
GPT-5.6 Sol🇺🇸 OpenAI$5.00$30.00$0.501.1M59
Claude Opus 4.8🇺🇸 Anthropic$5.00$25.00$0.501M61
Claude Opus 5🇺🇸 Anthropic$5.00$25.00$0.501M61
Claude Opus 4.7🇺🇸 Anthropic$5.00$25.00$0.501M57
Claude Fable 5🇺🇸 Anthropic$10.00$50.00$1.001M60

USD per million tokens · Chinese providers converted at 1 USD = ¥6.7598 · Green = cheapest · Heat = OpenRouter weekly rank by token usage (popularity, not quality; checked 2026-07-31) · Source: official provider pricing pages · For reference only.

Convenience pick

Tired of signing up provider-by-provider? One key for all models

Skip per-provider signups, top-ups and key juggling — AIMLAPI gives you one key to hundreds of models, pay-as-you-go, switch anytime.

Explore AIMLAPI →

Sponsored · we may earn a commission if you sign up via this link, at no extra cost to you

How to pick the most cost-effective LLM

Input vs output pricing.Almost every provider charges separately for input tokens (your prompt) and output tokens (the generation), and output is typically 4–10× more expensive. So “short prompt, long answer” tasks (writing, code generation) are dominated by output price, while “long prompt, short answer” tasks (summarization, classification) are dominated by input price.

Use prompt caching. If your requests share a large repeated prefix (fixed system prompt, RAG context), cached input can cost 10–20% of the normal rate. DeepSeek, OpenAI, Anthropic and Google all support context caching, but the discount varies a lot.

Don’t default to flagships. Mid-tier models like Claude Haiku 4.5, Gemini 2.5 Flash-Lite, Qwen3.5 Flash and DeepSeek V4 Flash are excellent value — for chat, translation, simple generation and classification they are more than enough at 1–5% of flagship cost.

Chinese models are aggressively priced.DeepSeek V4, Qwen, Doubao and Kimi often undercut Western models by an order of magnitude on output price. If latency to China matters or you’re cost-sensitive, they’re worth evaluating — see the dedicated Chinese LLM API pricing comparison in 2026 for all domestic providers in one table. Worried about data residency or training policies? We read all eight providers’ official policies in the Chinese AI API Trust Index.

Need to estimate your real bill? Try the calculators (Chinese UI).