Skip to content
LLM Abacus
Two probe locations · refreshed every 10 minutes3588 samples in the last 24 hours · data through the hour beginning Sep 5, 00:00 UTC

LLM API Latency Benchmark

Independent probes measure time to first token (TTFT) and output speed (tokens per second) for each API endpoint on a continuous schedule. For prices, see the model pricing table.

Where it is measured from: 2 locations, reported separately

🇭🇰 Hong Kong, 🇺🇸 US Central (Chicago). The two columns are measured independently and never averaged together. The same endpoint routinely differs several-fold between the two, and an average of those would be a number nobody could use. The Hong Kong probe measures international and cross-border routes, which is not the same as a connection originating inside mainland China. Calls to Chinese models from the US probe cross the Pacific. There is no mainland probe yet; if one is added it will be labelled separately.

How it is measured

TTFT uses a streaming request with max_tokens between 8 and 16, recording when the first token arrives. Output speed caps generation at 256 tokens and divides the provider’s own completion_tokens by elapsed time. Reasoning mode is explicitly disabled for both, because chain-of-thought tokens count toward completion_tokens and ignore max_tokens (one model returned 896 tokens for a 256 token request). Leaving it on would compare thinking speed against answering speed. The cost is coverage: Gemini 3.x refuses to disable reasoning across the whole family, so this board uses Gemini 2.5 Flash instead. Newer flagships increasingly force reasoning on, so this comparable set will narrow rather than widen.

What this table does not do

It does not rank which provider is best. Latency is one dimension of a choice, and samples from two locations cannot represent worldwide experience. Sample counts are in the table, so you can judge for yourself how much weight a row carries.

Active Region:
Type:
Currency:
Model & providerEndpoint typePrice (per 1M)
Cost/Task🛈
🇭🇰hk
··
🇺🇸us-central
··
DeepSeek V4 Flash
DeepSeek Official
Direct
In: ¥1.00
Out: ¥2.00
¥0.000503
144 ms
85.7 t/s · p95: 148ms
643 ms
70.0 t/s · p95: 897ms
DeepSeek V4 Pro
DeepSeek Official
Direct
In: ¥3.00
Out: ¥6.00
¥0.002943
148 ms
37.0 t/s · p95: 409ms
325 ms
32.6 t/s · p95: 1298ms
Qwen3.5 Flash
Alibaba Model Studio
Direct
In: ¥0.20
Out: ¥2.00
¥0.002243
410 ms
69.8 t/s · p95: 1178ms
1253 ms
41.1 t/s · p95: 1598ms
Kimi K3
OpenRouter(Relay)
Relay
In: $3.00
Out: $15.00
$0.010608
413 ms
32.0 t/s · p95: 413ms
795 ms
65.9 t/s · p95: 795ms
Qwen3 Max
Alibaba Model Studio
Direct
In: ¥2.50
Out: ¥10.00
¥0.002418
535 ms
32.5 t/s · p95: 596ms
754 ms
29.9 t/s · p95: 1128ms
GLM-5.2
Zhipu AI GLM Official
Direct
In: ¥8.00
Out: ¥28.00
¥0.023688
592 ms
47.4 t/s · p95: 713ms
726 ms
47.6 t/s · p95: 916ms
Kimi K2.6
Moonshot Kimi Official
Direct
In: ¥6.50
Out: ¥27.00
¥0.023199
630 ms
47.2 t/s · p95: 697ms
1926 ms
33.3 t/s · p95: 2240ms
GLM-5.2
OpenRouter(Relay)
Relay
In: $1.12
Out: $3.52
$0.001732
666 ms
71.1 t/s · p95: 973ms
358 ms
121.9 t/s · p95: 469ms
GLM-4.7
Zhipu AI GLM Official
Direct
In: ¥2.00
Out: ¥8.00
¥0.005408
734 ms
29.4 t/s · p95: 1294ms
1751 ms
24.5 t/s · p95: 3219ms
Qwen3.7 Max
Alibaba Model Studio
Direct
In: ¥12.00
Out: ¥36.00
¥0.057420
784 ms
40.5 t/s · p95: 996ms
872 ms
26.5 t/s · p95: 993ms
Qwen3.7 Max
OpenRouter(Relay)
Relay
In: $1.48
Out: $4.42
$0.006029
874 ms
34.8 t/s · p95: 1772ms
1109 ms
31.5 t/s · p95: 8166ms
DeepSeek V4 Pro
OpenRouter(Relay)
Relay
In: $0.43
Out: $0.87
$0.000449
1224 ms
30.1 t/s · p95: 1496ms
1256 ms
22.2 t/s · p95: 1321ms
Kimi K3
Moonshot Kimi Official
Direct
In: ¥20.00
Out: ¥100.00
¥0.062960
1767 ms
21.4 t/s · p95: 1767ms
2919 ms
16.3 t/s · p95: 2919ms
Claude Sonnet 5
OpenRouter(Relay)
Relay
In: $2.00
Out: $10.00
$0.005766
1152 ms
66.7 t/s · p95: 1222ms
Gemini 2.5 Flash
OpenRouter(Relay)
Relay
In: $0.30
Out: $2.50
$0.000877
422 ms
96.5 t/s · p95: 427ms

Citing this data

You are welcome to cite these figures in articles, documentation or a GitHub README under CC BY 4.0; please keep a link back to this page. The badge below shows that endpoint’s median TTFT from the most recent hourly aggregation:

[![DeepSeek V4 Pro latency](https://llmabacus-probe-badge.john-jones871024.workers.dev/badge/deepseek-v4-pro@official-deepseek)](https://llmabacus.com/en/live)

Swap deepseek-v4-pro@official-deepseek for any endpoint ID in the table. The badge is served from Cloudflare’s edge with a 600 second cache, so embedding it puts no extra load on this site.