LLM API Latency Benchmark
Independent probes measure time to first token (TTFT) and output speed (tokens per second) for each API endpoint on a continuous schedule. For prices, see the model pricing table.
Where it is measured from: 2 locations, reported separately
🇭🇰 Hong Kong, 🇺🇸 US Central (Chicago). The two columns are measured independently and never averaged together. The same endpoint routinely differs several-fold between the two, and an average of those would be a number nobody could use. The Hong Kong probe measures international and cross-border routes, which is not the same as a connection originating inside mainland China. Calls to Chinese models from the US probe cross the Pacific. There is no mainland probe yet; if one is added it will be labelled separately.
How it is measured
TTFT uses a streaming request with max_tokens between 8 and 16, recording when the first token arrives. Output speed caps generation at 256 tokens and divides the provider’s own completion_tokens by elapsed time. Reasoning mode is explicitly disabled for both, because chain-of-thought tokens count toward completion_tokens and ignore max_tokens (one model returned 896 tokens for a 256 token request). Leaving it on would compare thinking speed against answering speed. The cost is coverage: Gemini 3.x refuses to disable reasoning across the whole family, so this board uses Gemini 2.5 Flash instead. Newer flagships increasingly force reasoning on, so this comparable set will narrow rather than widen.
What this table does not do
It does not rank which provider is best. Latency is one dimension of a choice, and samples from two locations cannot represent worldwide experience. Sample counts are in the table, so you can judge for yourself how much weight a row carries.
| Model & provider | Endpoint type | Price (per 1M) | Cost/Task🛈 | 🇭🇰hk ·· | 🇺🇸us-central ·· |
|---|---|---|---|---|---|
DeepSeek V4 Flash DeepSeek Official | Direct | In: ¥1.00 Out: ¥2.00 | ¥0.000503 | 144 ms 85.7 t/s · p95: 148ms | 643 ms 70.0 t/s · p95: 897ms |
DeepSeek V4 Pro DeepSeek Official | Direct | In: ¥3.00 Out: ¥6.00 | ¥0.002943 | 148 ms 37.0 t/s · p95: 409ms | 325 ms 32.6 t/s · p95: 1298ms |
Qwen3.5 Flash Alibaba Model Studio | Direct | In: ¥0.20 Out: ¥2.00 | ¥0.002243 | 410 ms 69.8 t/s · p95: 1178ms | 1253 ms 41.1 t/s · p95: 1598ms |
Kimi K3 OpenRouter(Relay) | Relay | In: $3.00 Out: $15.00 | $0.010608 | 413 ms 32.0 t/s · p95: 413ms | 795 ms 65.9 t/s · p95: 795ms |
Qwen3 Max Alibaba Model Studio | Direct | In: ¥2.50 Out: ¥10.00 | ¥0.002418 | 535 ms 32.5 t/s · p95: 596ms | 754 ms 29.9 t/s · p95: 1128ms |
GLM-5.2 Zhipu AI GLM Official | Direct | In: ¥8.00 Out: ¥28.00 | ¥0.023688 | 592 ms 47.4 t/s · p95: 713ms | 726 ms 47.6 t/s · p95: 916ms |
Kimi K2.6 Moonshot Kimi Official | Direct | In: ¥6.50 Out: ¥27.00 | ¥0.023199 | 630 ms 47.2 t/s · p95: 697ms | 1926 ms 33.3 t/s · p95: 2240ms |
GLM-5.2 OpenRouter(Relay) | Relay | In: $1.12 Out: $3.52 | $0.001732 | 666 ms 71.1 t/s · p95: 973ms | 358 ms 121.9 t/s · p95: 469ms |
GLM-4.7 Zhipu AI GLM Official | Direct | In: ¥2.00 Out: ¥8.00 | ¥0.005408 | 734 ms 29.4 t/s · p95: 1294ms | 1751 ms 24.5 t/s · p95: 3219ms |
Qwen3.7 Max Alibaba Model Studio | Direct | In: ¥12.00 Out: ¥36.00 | ¥0.057420 | 784 ms 40.5 t/s · p95: 996ms | 872 ms 26.5 t/s · p95: 993ms |
Qwen3.7 Max OpenRouter(Relay) | Relay | In: $1.48 Out: $4.42 | $0.006029 | 874 ms 34.8 t/s · p95: 1772ms | 1109 ms 31.5 t/s · p95: 8166ms |
DeepSeek V4 Pro OpenRouter(Relay) | Relay | In: $0.43 Out: $0.87 | $0.000449 | 1224 ms 30.1 t/s · p95: 1496ms | 1256 ms 22.2 t/s · p95: 1321ms |
Kimi K3 Moonshot Kimi Official | Direct | In: ¥20.00 Out: ¥100.00 | ¥0.062960 | 1767 ms 21.4 t/s · p95: 1767ms | 2919 ms 16.3 t/s · p95: 2919ms |
Claude Sonnet 5 OpenRouter(Relay) | Relay | In: $2.00 Out: $10.00 | $0.005766 | — | 1152 ms 66.7 t/s · p95: 1222ms |
Gemini 2.5 Flash OpenRouter(Relay) | Relay | In: $0.30 Out: $2.50 | $0.000877 | — | 422 ms 96.5 t/s · p95: 427ms |
Citing this data
You are welcome to cite these figures in articles, documentation or a GitHub README under CC BY 4.0; please keep a link back to this page. The badge below shows that endpoint’s median TTFT from the most recent hourly aggregation:
[](https://llmabacus.com/en/live)Swap deepseek-v4-pro@official-deepseek for any endpoint ID in the table. The badge is served from Cloudflare’s edge with a 600 second cache, so embedding it puts no extra load on this site.