跳到主内容
算盘

Chinese LLM API Pricing Comparison 2026

All 28 models from 10Chinese providers — DeepSeek, Qwen, Kimi, GLM, Doubao, ERNIE, Hunyuan, MiniMax, Spark and Baichuan. Prices in USD per million tokens, converted from official CNY list prices and checked against each provider’s pricing page. Flagship models are verified automatically every day; others are manually monitored.

2026 Chinese LLM API Pricing Comparison Highlights

In 2026, the cheapest Chinese LLM API by input price is Qwen3.5 Flash at $0.0300.20) per million tokens. By output price, the cheapest Chinese LLM APIs are Qwen3.5 Flash, Doubao 1.5 Pro, 混元 TurboS, DeepSeek V4 Flash, Spark X2 Flash, Spark Ultra, all costing $0.302.00) per million tokens. DeepSeek V4 Flash also offers the lowest cached input price at a rock-bottom $0.0030.02) per million tokens.

ModelProviderInput (per M)Output (per M)Cached (per M)Context
DeepSeek V4 Flash🇨🇳 DeepSeek$0.15 1.00)$0.30 2.00)$0.003 (¥0.020)1M
DeepSeek V4 Pro🇨🇳 DeepSeek$0.44 3.00)$0.89 6.00)$0.004 (¥0.025)1M
Qwen3.5 Flash🇨🇳 Alibaba Qwen$0.030 0.20)$0.30 2.00)131K
Qwen3.7 Max🇨🇳 Alibaba Qwen$1.77 12.00)$5.31 36.00)1M
Kimi K2.6🇨🇳 Moonshot (Kimi)$0.96 6.50)$3.99 27.00)$0.16 (¥1.100)262K
GLM-5.1🇨🇳 Zhipu AI (GLM)$0.89 6.00)$3.54 24.00)$0.19 (¥1.300)200K
GLM-4.7🇨🇳 Zhipu AI (GLM)$0.30 2.00)$1.18 8.00)$0.059 (¥0.400)200K
文心 ERNIE 5.1🇨🇳 Baidu ERNIE$0.59 4.00)$2.66 18.00)128K
文心 ERNIE 4.5 Turbo🇨🇳 Baidu ERNIE$0.12 0.80)$0.47 3.20)$0.030 (¥0.200)128K
Doubao 1.5 Pro🇨🇳 ByteDance Doubao$0.12 0.80)$0.30 2.00)$0.024 (¥0.160)256K
Doubao Seed 2.1 Pro🇨🇳 ByteDance Doubao$0.89 6.00)$4.43 30.00)$0.18 (¥1.200)256K
混元 2.0 Think🇨🇳 Tencent Hunyuan$0.59 3.98)$2.35 15.90)128K
混元 TurboS🇨🇳 Tencent Hunyuan$0.12 0.80)$0.30 2.00)256K

* Data Source: Daily Verification Database + models.json. All prices are officially published rates, converted to USD at 1 USD = ¥6.7746. Flagship models auto-verified daily; others manually monitored.

ModelProviderInputOutputCachedContextQualityVerified
Qwen3.5 Flash🇨🇳 Alibaba Qwen$0.030$0.30131K
2026-07-28
Auto
Qwen3.5 Plus🇨🇳 Alibaba Qwen$0.12$0.71131K
2026-07-28
Auto
Doubao 1.5 Pro🇨🇳 ByteDance Doubao$0.12$0.30$0.024256K
2026-07-03
Monitored
>14d stale
文心 ERNIE 4.5 Turbo🇨🇳 Baidu ERNIE$0.12$0.47$0.030128K
2026-06-05
Monitored
>14d stale
混元 TurboS🇨🇳 Tencent Hunyuan$0.12$0.30256K
2026-06-05
Monitored
>14d stale
DeepSeek V4 Flash🇨🇳 DeepSeek$0.15$0.30$0.0031M47
2026-07-28
Auto
文心 ERNIE X1 Turbo🇨🇳 Baidu ERNIE$0.15$0.59128K
2026-06-05
Monitored
>14d stale
混元 T1🇨🇳 Tencent Hunyuan$0.15$0.5964K
2026-06-05
Monitored
>14d stale
DeepSeek V3.2🇨🇳 DeepSeek$0.30$1.18$0.074128K
2026-07-28
Auto
GLM-4.7🇨🇳 Zhipu AI (GLM)$0.30$1.18$0.059200K
2026-07-28
Auto
Spark X2 Flash🇨🇳 iFlytek Spark$0.30$0.30200K
2026-06-05
Monitored
>14d stale
Spark Ultra🇨🇳 iFlytek Spark$0.30$0.30128K
2026-06-05
Monitored
>14d stale
Baichuan M2🇨🇳 Baichuan AI$0.30$2.95192K
2026-06-05
Monitored
>14d stale
MiniMax M2.7🇨🇳 MiniMax$0.31$1.24$0.0621M50
2026-06-05
Monitored
>14d stale
Doubao 1.6🇨🇳 ByteDance Doubao$0.35$3.54256K
2026-07-03
Monitored
>14d stale
Qwen3 Max🇨🇳 Alibaba Qwen$0.37$1.48131K
2026-07-28
Auto
DeepSeek V4 Pro🇨🇳 DeepSeek$0.44$0.89$0.0041M52
2026-07-28
Auto
Spark X2🇨🇳 iFlytek Spark$0.44$0.44200K
2026-06-05
Monitored
>14d stale
混元 2.0 Think🇨🇳 Tencent Hunyuan$0.59$2.35128K
2026-06-05
Monitored
>14d stale
GLM-5🇨🇳 Zhipu AI (GLM)$0.59$2.66$0.15200K
2026-07-28
Auto
文心 ERNIE 5.1🇨🇳 Baidu ERNIE$0.59$2.66128K
2026-06-05
Monitored
>14d stale
MiniMax M3🇨🇳 MiniMax$0.62$2.48$0.121M55
2026-06-05
Monitored
>14d stale
Baichuan M3 Plus🇨🇳 Baichuan AI$0.74$1.33192K
2026-06-05
Monitored
>14d stale
GLM-5.1🇨🇳 Zhipu AI (GLM)$0.89$3.54$0.19200K51
2026-07-28
Auto
Doubao Seed 2.1 Pro🇨🇳 ByteDance Doubao$0.89$4.43$0.18256K
2026-07-03
Monitored
>14d stale
Kimi K2.6🇨🇳 Moonshot (Kimi)$0.96$3.99$0.16262K54
2026-07-28
Auto
Spark Pro🇨🇳 iFlytek Spark$1.03$1.03128K
2026-06-05
Monitored
>14d stale
Qwen3.7 Max🇨🇳 Alibaba Qwen$1.77$5.311M57
2026-07-28
Auto

USD per million tokens · Converted at 1 USD = ¥6.7746(daily-calibrated) · Green = cheapest · Quality = Artificial Analysis intelligence index · Source: official provider pricing pages · For reference only.

The short version

Cheapest input: Qwen3.5 Flash at $0.030 per million input tokens. Cheapest output: Qwen3.5 Flash / Doubao 1.5 Pro / 混元 TurboS / DeepSeek V4 Flash / Spark X2 Flash / Spark Ultra at $0.30 per million output tokens.

Why Chinese models are worth a look. Flagship models from DeepSeek, Qwen and Kimi sit near the top of public leaderboards while pricing well below Western flagships — output-token prices are often an order of magnitude lower. For translation, summarization, classification and agent workloads where you control the evaluation, the savings are real.

What to watch.Some providers bill only in CNY or require a mainland-China identity for signup; international endpoints (DeepSeek, Alibaba Model Studio, Moonshot, Zhipu) are the practical route for overseas teams. Latency and data-residency requirements deserve the same scrutiny you’d give any provider — we compared all five providers’ legal entities, data residency and training policies in the Chinese AI API Trust Index.

Frequently asked questions

What is the cheapest Chinese LLM API in 2026?

By input price, Qwen3.5 Flash (Alibaba Qwen) is currently the cheapest at $0.030 per million input tokens. By output price, the cheapest Chinese LLM APIs are Qwen3.5 Flash, Doubao 1.5 Pro, 混元 TurboS, DeepSeek V4 Flash, Spark X2 Flash, and Spark Ultra at $0.30 (¥2.00) per million output tokens. Prices are converted from official CNY list prices at 1 USD = ¥6.7746.

How much is DeepSeek V4 Flash per million tokens in 2026?

DeepSeek V4 Flash API pricing in 2026 is ¥1.00 ($0.15) per million non-cached input tokens, ¥2.00 ($0.30) per million output tokens, and ¥0.02 ($0.003) per million cached input tokens. Its context window is 1M tokens. (Verified from official sources on 2026-07-28).

What are the pricing details of leading Chinese LLM models in 2026 (DeepSeek, Qwen, Kimi, GLM, Ernie)?

Pricing for leading Chinese LLM models in 2026 (per million tokens) includes: DeepSeek V4 Pro at ¥3.00 non-cached input / ¥6.00 output; Qwen3.7 Max at ¥12.00 input / ¥36.00 output; Kimi K2.6 at ¥6.50 input / ¥27.00 output; GLM-5.1 at ¥6.00 input / ¥24.00 output; and 文心 ERNIE 5.1 at ¥4.00 input / ¥18.00 output. DeepSeek, Qwen, Kimi and GLM are verified automatically against their official pricing pages; others are monitored manually. See each provider's specific last verification date in our comparison table.

Which is the latest Doubao, Qwen, Kimi, and Hunyuan model and what do they cost?

The latest models in 2026 include ByteDance's Doubao Seed 2.1 Pro (¥6.00 input / ¥30.00 output), Alibaba's Qwen3.7 Max (¥12.00 input / ¥36.00 output), Moonshot's Kimi K2.6 (¥6.50 input / ¥27.00 output), and Tencent's 混元 2.0 Think (¥3.975 input / ¥15.90 output) or 混元 TurboS (¥0.80 input / ¥2.00 output). Check our comparison table for their full parameters.

What is an LLM API Relay (大模型API中转) and is it safe?

An LLM API Relay (大模型API中转) is a third-party gateway service that aggregates multiple AI providers under a single API key, often billing in domestic currency. While convenient, relay services introduce risks: they may log your prompts/keys, alter system prompts, or silently swap models (e.g. serving a cheaper model instead of a flagship). We recommend using official international developer platforms (like DeepSeek Platform or Alibaba Cloud Model Studio) to ensure direct security and compliance.

Can I use Chinese LLM APIs outside China?

Mostly yes. DeepSeek, Alibaba (Qwen via Model Studio), Moonshot (Kimi) and Zhipu (GLM) all offer international endpoints with English documentation and card payment. Some providers — for example Baidu ERNIE and iFlytek Spark — primarily target the domestic market, so check signup requirements before committing.

How do Chinese LLM API prices compare with GPT-5.5 or Claude?

Chinese models are aggressively priced: flagship-level models from DeepSeek, Qwen and Kimi often cost an order of magnitude less per output token than Western flagships, and budget tiers go far lower. See our full comparison table for a side-by-side view with OpenAI, Anthropic and Google models.

How accurate are these prices?

Prices are checked against each provider's official pricing page — flagship vendors are verified automatically, and others are manually monitored. Check the 'Verified' column in our comparison table for the exact verification date of each provider. CNY prices are converted to USD at a daily-calibrated rate (currently 1 USD = ¥6.7746).

Where can I download the raw pricing dataset?

You can download the raw daily pricing snapshot in JSON format from our GitHub repository (https://github.com/szp2005/llm-prices-cn). The dataset is licensed under the Creative Commons Attribution 4.0 International License (CC-BY-4.0), allowing for free commercial use and adaptation, provided you attribute and link back to LLM Abacus.

Does LLM Abacus provide an API or Model Context Protocol (MCP) server?

Yes! We expose a remote MCP server at https://www.llmabacus.com/api/mcp/mcp that supports tools like `query_model_price` and `estimate_cost`. You can also read the raw daily pricing dataset directly from our public endpoints.

Developer API & Open-Source Dataset

To support open research and AI tool auditing, the pricing data displayed on this page is open-sourced on GitHub under the Creative Commons Attribution 4.0 International License (CC-BY-4.0). You are free to copy, adapt, and use the pricing data commercially, provided that you attribute and link back to LLM Abacus (llmabacus.com).

GitHub Dataset Repository

The repository contains the daily pricing snapshot prices.json and a synchronization script sync_prices.py.

JSON Schema of prices.json

The dataset uses a clean, developer-friendly flat JSON structure. The core schema includes:

  • last_updated: The last verified date (e.g. 2026-07-03).
  • usd_to_cny_rate: Exchange rate used for conversions (e.g. 6.7855).
  • models: Array of model objects.
    • id / name: Unique ID and display name of the model.
    • vendor_name: The model provider (e.g., Alibaba Qwen, DeepSeek).
    • billing_currency: The billing currency of the vendor (USD or CNY).
    • input_price_usd_per_m / output_price_usd_per_m: Pricing in USD per million tokens.
    • input_price_cny_per_m / output_price_cny_per_m: Pricing in CNY per million tokens.
    • context_window / max_output: Input context and output generation token limits.

Remote Model Context Protocol (MCP) Server

You can connect your AI agent (like Cursor, Claude Desktop, or Windsurf) directly to our remote MCP server to query live model pricing and estimate costs in real time:

npx -y mcp-remote https://www.llmabacus.com/api/mcp/mcp

Want the full picture with GPT-5.5, Claude and Gemini in the same table?