Skip to content
LLM Abacus

LLM API Pricing Comparison

Compare 55 leading large language models in one table. Prices are shown in USD per million tokens. 7 of 16 providers were verified within the last 14 days; each model page shows its source and verification date.

US providers use their native USD list prices. Chinese-provider rows use mainland-China official CNY prices converted at 1 USD = ¥6.7421; international endpoints, taxes, regional promotions, and enterprise contracts may differ.

ModelProviderInputOutputCachedContextQualityHeat
Qwen3.5 Flash🇨🇳 Alibaba Qwen$0.030$0.301M
Gemini 2.5 Flash-Lite🇺🇸 Google$0.10$0.40$0.0101M
Qwen3.5 Plus🇨🇳 Alibaba Qwen$0.12$0.711M
Doubao 1.6🇨🇳 ByteDance Doubao$0.12$1.19$0.024256K
Doubao 1.5 Pro🇨🇳 ByteDance Doubao$0.12$0.30$0.024256K
文心 ERNIE 4.5 Turbo🇨🇳 Baidu ERNIE$0.12$0.47$0.030128K
混元 TurboS🇨🇳 Tencent Hunyuan$0.12$0.30256K
Spark Ultra🇨🇳 iFlytek Spark$0.12$0.12128K
文心 ERNIE X1 Turbo🇨🇳 Baidu ERNIE$0.15$0.59128K
混元 T1🇨🇳 Tencent Hunyuan$0.15$0.5964K
MiMo-V2.5🇨🇳 小米 MiMo$0.15$0.30$0.003262K#63.79T
混元 Hy3🇨🇳 Tencent Hunyuan$0.15$0.59$0.037262K#29.95T
GPT-5.6 Luna🇺🇸 OpenAI$0.20$1.20$0.0201.1M#35.31T
Step 3.7 Flash🇨🇳 阶跃星辰$0.20$1.20131K
Gemini 3.1 Flash-Lite🇺🇸 Google$0.25$1.50$0.0251M34
DeepSeek V3.2Retired 2026-07-24🇨🇳 DeepSeek$0.30$1.19$0.074128K
GLM-4.7🇨🇳 Zhipu AI$0.30$1.19$0.059200K
Spark X2🇨🇳 iFlytek Spark$0.30$0.30200K
Spark X2 Flash🇨🇳 iFlytek Spark$0.30$0.30200K
Baichuan M2🇨🇳 Baichuan AI$0.30$2.97192K
Gemini 2.5 Flash🇺🇸 Google$0.30$2.50$0.0301M
MiniMax M3🇨🇳 MiniMax$0.31$1.25$0.0621M55
MiniMax M2.7🇨🇳 MiniMax$0.31$1.25$0.0621M50
Qwen3 Max🇨🇳 Alibaba Qwen$0.37$1.48256K
DeepSeek V4 Flash🇨🇳 DeepSeek$0.44$1.33$0.0151M47
混元 2.0 Think🇨🇳 Tencent Hunyuan$0.59$2.36128K
GLM-5🇨🇳 Zhipu AI$0.59$2.67$0.15200K
文心 ERNIE 5.1🇨🇳 Baidu ERNIE$0.59$2.67128K
Spark Pro🇨🇳 iFlytek Spark$0.74$0.74128K
Baichuan M3 Plus🇨🇳 Baichuan AI$0.74$1.33192K
GLM-5.1🇨🇳 Zhipu AI$0.89$3.56$0.19200K51
Doubao Seed 2.1 Pro🇨🇳 ByteDance Doubao$0.89$4.45$0.18256K
Kimi K2.6🇨🇳 Moonshot (Kimi)$0.96$4.00$0.16262K54
Claude Haiku 4.5🇺🇸 Anthropic$1.00$5.00$0.10200K37
Grok Build 0.1🇺🇸 xAI$1.00$2.00$0.20256K
GLM-5.2🇨🇳 Zhipu AI$1.19$4.151M#54.33T
GPT-5.1🇺🇸 OpenAI$1.25$10.00$0.13400K
Gemini 2.5 Pro🇺🇸 Google$1.25$10.00$0.132M35
Grok 4.3🇺🇸 xAI$1.25$2.50$0.201M53
DeepSeek V4 Pro🇨🇳 DeepSeek$1.33$4.00$0.0441M52
Gemini 3.6 Flash🇺🇸 Google$1.50$7.50$0.151M#92.28T
Gemini 3.5 Flash🇺🇸 Google$1.50$9.00$0.151M55
Qwen3.7 Max🇨🇳 Alibaba Qwen$1.78$5.341M57
GPT-5.6 Terra🇺🇸 OpenAI$2.00$12.00$0.201.1M
Claude Sonnet 5🇺🇸 Anthropic$2.00$10.00$0.201M57
Gemini 3.1 Pro Preview🇺🇸 Google$2.00$12.00$0.202M57
GPT-5.4🇺🇸 OpenAI$2.50$15.00$0.25400K
Kimi K3🇨🇳 Moonshot (Kimi)$2.97$14.83$0.301.0M
Claude Sonnet 4.6🇺🇸 Anthropic$3.00$15.00$0.301M52
GPT-5.5🇺🇸 OpenAI$5.00$30.00$0.50400K60
GPT-5.6 Sol🇺🇸 OpenAI$5.00$30.00$0.501.1M59
Claude Opus 4.8🇺🇸 Anthropic$5.00$25.00$0.501M61
Claude Opus 5🇺🇸 Anthropic$5.00$25.00$0.501M61#82.67T
Claude Opus 4.7🇺🇸 Anthropic$5.00$25.00$0.501M57
Claude Fable 5🇺🇸 Anthropic$10.00$50.00$1.001M60

USD per million tokens · Chinese providers converted at 1 USD = ¥6.7421 · Green = cheapest · Heat = OpenRouter weekly rank by token usage (popularity, not quality; checked 2026-08-17) · Source: official provider pricing pages · For reference only.

Convenience pick

Tired of signing up provider-by-provider? One key for all models

Skip per-provider signups, top-ups and key juggling — AIMLAPI gives you one key to hundreds of models, pay-as-you-go, switch anytime.

Explore AIMLAPI →

Sponsored · we may earn a commission if you sign up via this link, at no extra cost to you

How to pick the most cost-effective LLM

Input vs output pricing. Almost every provider charges separately for input tokens (your prompt) and output tokens (the generation), and output is typically 4–10× more expensive. So “short prompt, long answer” tasks (writing, code generation) are dominated by output price, while “long prompt, short answer” tasks (summarization, classification) are dominated by input price.

Use prompt caching. If your requests share a large repeated prefix (fixed system prompt, RAG context), cached input can cost 10–20% of the normal rate. DeepSeek, OpenAI, Anthropic and Google all support context caching, but the discount varies a lot.

Don’t default to flagships. Mid-tier models like Claude Haiku 4.5, Gemini 2.5 Flash-Lite, Qwen3.5 Flash and DeepSeek V4 Flash are excellent value — for chat, translation, simple generation and classification they are more than enough at 1–5% of flagship cost.

Chinese models are aggressively priced. DeepSeek V4, Qwen, Doubao and Kimi often undercut Western models by an order of magnitude on output price. If latency to China matters or you’re cost-sensitive, they’re worth evaluating — see the dedicated Chinese LLM API pricing comparison in 2026 for all domestic providers in one table. Worried about data residency or training policies? We read all eight providers’ official policies in the Chinese AI API Trust Index.

Need to estimate your real bill? Try the calculators (Chinese UI).