Skip to content
LLM Abacus
Two probe locations · refreshed every 10 minutes3583 samples in the last 24 hours · data through the hour beginning Sep 17, 10:00 UTC

LLM API Latency Benchmark

Independent probes measure time to first token (TTFT) and output speed (tokens per second) for each API endpoint on a continuous schedule. For prices, see the model pricing table.

Where it is measured from: 2 locations, reported separately

🇭🇰 Hong Kong, 🇺🇸 US Central (Chicago). The two columns are measured independently and never averaged together. The same endpoint routinely differs several-fold between the two, and an average of those would be a number nobody could use. The Hong Kong probe measures international and cross-border routes, which is not the same as a connection originating inside mainland China. Calls to Chinese models from the US probe cross the Pacific. There is no mainland probe yet; if one is added it will be labelled separately.

How it is measured

TTFT uses a streaming request with max_tokens between 8 and 16, recording when the first token arrives. Output speed caps generation at 256 tokens and divides the provider’s own completion_tokens by elapsed time. Reasoning mode is explicitly disabled for both, because chain-of-thought tokens count toward completion_tokens and ignore max_tokens (one model returned 896 tokens for a 256 token request). Leaving it on would compare thinking speed against answering speed. The cost is coverage: Gemini 3.x refuses to disable reasoning across the whole family, so this board uses Gemini 2.5 Flash instead. Newer flagships increasingly force reasoning on, so this comparable set will narrow rather than widen.

For rolling 7 and 30 day p50 / p95 figures with sample counts, see the latency dataset (JSON and CSV downloads included).

What this table does not do

It does not rank which provider is best. Latency is one dimension of a choice, and samples from two locations cannot represent worldwide experience. Sample counts are in the table, so you can judge for yourself how much weight a row carries.

Active Region:
Type:
Currency:
Model & providerEndpoint typePrice (per 1M)
Cost/Task🛈
🇭🇰hk
··
🇺🇸us-central
··
DeepSeek V4 Flash
DeepSeek Official
Direct
In: ¥1.00
Out: ¥2.00
146 ms
63.8 t/s · p95: 177ms
549 ms
70.3 t/s · p95: 616ms
DeepSeek V4 Pro
DeepSeek Official
Direct
In: ¥3.00
Out: ¥6.00
168 ms
30.9 t/s · p95: 214ms
326 ms
31.6 t/s · p95: 411ms
Kimi K3
OpenRouter(Relay)
Relay
In: $3.00
Out: $15.00
470 ms
TPS not measured this hour · p95: 470ms
1385 ms
TPS not measured this hour · p95: 1385ms
GLM-5.2
OpenRouter(Relay)
Relay
In: $1.12
Out: $3.52
488 ms
TPS not measured this hour · p95: 883ms
395 ms
TPS not measured this hour · p95: 550ms
Qwen3 Max
Alibaba Model Studio
Direct
In: ¥2.50
Out: ¥10.00
606 ms
31.0 t/s · p95: 638ms
596 ms
33.6 t/s · p95: 1393ms
GLM-5.2
Zhipu AI GLM Official
Direct
In: ¥8.00
Out: ¥28.00
670 ms
TPS not measured this hour · p95: 11454ms
786 ms
TPS not measured this hour · p95: 1965ms
Qwen3.7 Max
OpenRouter(Relay)
Relay
In: $1.48
Out: $4.42
776 ms
TPS not measured this hour · p95: 1166ms
1130 ms
TPS not measured this hour · p95: 1354ms
Qwen3.5 Flash
Alibaba Model Studio
Direct
In: ¥0.20
Out: ¥2.00
877 ms
53.3 t/s · p95: 1622ms
1387 ms
37.2 t/s · p95: 2023ms
Qwen3.7 Max
Alibaba Model Studio
Direct
In: ¥12.00
Out: ¥36.00
880 ms
TPS not measured this hour · p95: 1001ms
1042 ms
TPS not measured this hour · p95: 2223ms
Kimi K2.6
Moonshot Kimi Official
Direct
In: ¥6.50
Out: ¥27.00
1280 ms
TPS not measured this hour · p95: 2622ms
1857 ms
TPS not measured this hour · p95: 2159ms
GLM-4.7
Zhipu AI GLM Official
Direct
In: ¥2.00
Out: ¥8.00
1282 ms
22.5 t/s · p95: 21765ms
1945 ms
18.7 t/s · p95: 2017ms
DeepSeek V4 Pro
OpenRouter(Relay)
Relay
In: $0.43
Out: $0.87
1604 ms
33.3 t/s · p95: 1825ms
1714 ms
23.2 t/s · p95: 1867ms
Kimi K3
Moonshot Kimi Official
Direct
In: ¥20.00
Out: ¥100.00
6388 ms
TPS not measured this hour · p95: 6388ms
2145 ms
TPS not measured this hour · p95: 2145ms
Claude Sonnet 5
OpenRouter(Relay)
Relay
In: $2.00
Out: $10.00
1702 ms
TPS not measured this hour · p95: 1702ms
Gemini 2.5 Flash
OpenRouter(Relay)
Relay
In: $0.30
Out: $2.50
440 ms
TPS not measured this hour · p95: 440ms

Citing this data

You are welcome to cite these figures in articles, documentation or a GitHub README under CC BY 4.0; please keep a link back to this page. The badge below shows that endpoint’s median TTFT from the most recent hourly aggregation:

[![DeepSeek V4 Pro latency](https://llmabacus-probe-badge.john-jones871024.workers.dev/badge/deepseek-v4-pro@official-deepseek)](https://llmabacus.com/en/live)

Swap deepseek-v4-pro@official-deepseek for any endpoint ID in the table. The badge is served from Cloudflare’s edge with a 600 second cache, so embedding it puts no extra load on this site.