LLM API Latency Benchmark
Independent probes measure time to first token (TTFT) and output speed (tokens per second) for each API endpoint on a continuous schedule. For prices, see the model pricing table.
Where it is measured from: 2 locations, reported separately
🇭🇰 Hong Kong, 🇺🇸 US Central (Chicago). The two columns are measured independently and never averaged together. The same endpoint routinely differs several-fold between the two, and an average of those would be a number nobody could use. The Hong Kong probe measures international and cross-border routes, which is not the same as a connection originating inside mainland China. Calls to Chinese models from the US probe cross the Pacific. There is no mainland probe yet; if one is added it will be labelled separately.
How it is measured
TTFT uses a streaming request with max_tokens between 8 and 16, recording when the first token arrives. Output speed caps generation at 256 tokens and divides the provider’s own completion_tokens by elapsed time. Reasoning mode is explicitly disabled for both, because chain-of-thought tokens count toward completion_tokens and ignore max_tokens (one model returned 896 tokens for a 256 token request). Leaving it on would compare thinking speed against answering speed. The cost is coverage: Gemini 3.x refuses to disable reasoning across the whole family, so this board uses Gemini 2.5 Flash instead. Newer flagships increasingly force reasoning on, so this comparable set will narrow rather than widen.
For rolling 7 and 30 day p50 / p95 figures with sample counts, see the latency dataset (JSON and CSV downloads included).
What this table does not do
It does not rank which provider is best. Latency is one dimension of a choice, and samples from two locations cannot represent worldwide experience. Sample counts are in the table, so you can judge for yourself how much weight a row carries.
| Model & provider | Endpoint type | Price (per 1M) | Cost/Task🛈 | 🇭🇰hk ·· | 🇺🇸us-central ·· |
|---|---|---|---|---|---|
DeepSeek V4 Flash DeepSeek Official | Direct | In: ¥1.00 Out: ¥2.00 | — | 146 ms 63.8 t/s · p95: 177ms | 549 ms 70.3 t/s · p95: 616ms |
DeepSeek V4 Pro DeepSeek Official | Direct | In: ¥3.00 Out: ¥6.00 | — | 168 ms 30.9 t/s · p95: 214ms | 326 ms 31.6 t/s · p95: 411ms |
Kimi K3 OpenRouter(Relay) | Relay | In: $3.00 Out: $15.00 | — | 470 ms TPS not measured this hour · p95: 470ms | 1385 ms TPS not measured this hour · p95: 1385ms |
GLM-5.2 OpenRouter(Relay) | Relay | In: $1.12 Out: $3.52 | — | 488 ms TPS not measured this hour · p95: 883ms | 395 ms TPS not measured this hour · p95: 550ms |
Qwen3 Max Alibaba Model Studio | Direct | In: ¥2.50 Out: ¥10.00 | — | 606 ms 31.0 t/s · p95: 638ms | 596 ms 33.6 t/s · p95: 1393ms |
GLM-5.2 Zhipu AI GLM Official | Direct | In: ¥8.00 Out: ¥28.00 | — | 670 ms TPS not measured this hour · p95: 11454ms | 786 ms TPS not measured this hour · p95: 1965ms |
Qwen3.7 Max OpenRouter(Relay) | Relay | In: $1.48 Out: $4.42 | — | 776 ms TPS not measured this hour · p95: 1166ms | 1130 ms TPS not measured this hour · p95: 1354ms |
Qwen3.5 Flash Alibaba Model Studio | Direct | In: ¥0.20 Out: ¥2.00 | — | 877 ms 53.3 t/s · p95: 1622ms | 1387 ms 37.2 t/s · p95: 2023ms |
Qwen3.7 Max Alibaba Model Studio | Direct | In: ¥12.00 Out: ¥36.00 | — | 880 ms TPS not measured this hour · p95: 1001ms | 1042 ms TPS not measured this hour · p95: 2223ms |
Kimi K2.6 Moonshot Kimi Official | Direct | In: ¥6.50 Out: ¥27.00 | — | 1280 ms TPS not measured this hour · p95: 2622ms | 1857 ms TPS not measured this hour · p95: 2159ms |
GLM-4.7 Zhipu AI GLM Official | Direct | In: ¥2.00 Out: ¥8.00 | — | 1282 ms 22.5 t/s · p95: 21765ms | 1945 ms 18.7 t/s · p95: 2017ms |
DeepSeek V4 Pro OpenRouter(Relay) | Relay | In: $0.43 Out: $0.87 | — | 1604 ms 33.3 t/s · p95: 1825ms | 1714 ms 23.2 t/s · p95: 1867ms |
Kimi K3 Moonshot Kimi Official | Direct | In: ¥20.00 Out: ¥100.00 | — | 6388 ms TPS not measured this hour · p95: 6388ms | 2145 ms TPS not measured this hour · p95: 2145ms |
Claude Sonnet 5 OpenRouter(Relay) | Relay | In: $2.00 Out: $10.00 | — | — | 1702 ms TPS not measured this hour · p95: 1702ms |
Gemini 2.5 Flash OpenRouter(Relay) | Relay | In: $0.30 Out: $2.50 | — | — | 440 ms TPS not measured this hour · p95: 440ms |
Citing this data
You are welcome to cite these figures in articles, documentation or a GitHub README under CC BY 4.0; please keep a link back to this page. The badge below shows that endpoint’s median TTFT from the most recent hourly aggregation:
[](https://llmabacus.com/en/live)Swap deepseek-v4-pro@official-deepseek for any endpoint ID in the table. The badge is served from Cloudflare’s edge with a 600 second cache, so embedding it puts no extra load on this site.