LLM API latency dataset
Hourly probe measurements re-aggregated over the last 7 and the last 30 days: median and 95th percentile time to first token, output speed, error rate and sample count per endpoint. For how an endpoint behaves right now, see the live latency board. This page answers the other question: how did it hold up over the week and the month. The data is free to cite with attribution.
2026-09-02T13:42:39.422Z → 2026-09-09T13:42:39.422Z
| Model / endpoint | Endpoint type | TTFT p50 ↑ | TTFT p95 typical hour / worst hour | Output speed | Error rate | Samples |
|---|---|---|---|---|---|---|
DeepSeek V4 Pro deepseek-v4-pro@official-deepseek | Official | 145ms | 163 / 1001 | 34.3tok/s | 31.3% | 1144 (786 ok) 115/168 hours coverage |
DeepSeek V4 Flash deepseek-v4-flash@official-deepseek | Official | 146ms | 159 / 5212 | 75.2tok/s | 31.4% | 1164 (799 ok) 117/168 hours coverage |
Qwen3.5 Flash qwen3-5-flash@official-aliyun | Official | 373ms | 462 / 3351 | 56.7tok/s | 0.1% | 1136 (1135 ok) 166/168 hours coverage |
Qwen3 Max qwen3-max@official-aliyun | Official | 511ms | 804 / 3449 | 35.3tok/s | 0.0% | 1142 (1142 ok) 167/168 hours coverage |
GLM-5.2 glm-5-2@openrouter | Relay | 630ms | 1273 / 12126 | 48.1tok/s | 0.0% | 1009 (1009 ok) 165/168 hours coverage |
GLM-5.2 glm-5-2@official-zhipu | Official | 634ms | 1162 / 46299 | 41.7tok/s | 0.0% | 1010 (1010 ok) 165/168 hours coverage |
Kimi K2.6 kimi-k2-6@official-moonshot | Official | 688ms | 753 / 6029 | 37.0tok/s | 0.1% | 1021 (1020 ok) 167/168 hours coverage |
Qwen3.7 Max qwen3-7-max@official-aliyun | Official | 720ms | 934 / 2126 | 38.1tok/s | 0.0% | 1013 (1013 ok) 166/168 hours coverage |
Qwen3.7 Max qwen3-7-max@openrouter | Relay | 863ms | 1073 / 2786 | 34.4tok/s | 0.2% | 1021 (1019 ok) 167/168 hours coverage |
GLM-4.7 glm-4-7@official-zhipu | Official | 1148ms | 1812 / 45540 | 34.1tok/s | 1.0% | 1137 (1126 ok) 166/168 hours coverage |
DeepSeek V4 Pro deepseek-v4-pro@openrouter | Relay | 1226ms | 1511 / 9236 | 30.1tok/s | 0.0% | 1150 (1150 ok) 168/168 hours coverage |
Kimi K3 kimi-k3@openrouter | Relay | 1278ms | 1278 / 11297 | 25.6tok/s | 0.5% | 191 (190 ok) 162/168 hours coverage |
Kimi K3 kimi-k3@official-moonshot | Official | 2086ms | 2086 / 12451 | 22.1tok/s | 0.5% | 193 (192 ok) 163/168 hours coverage |
Claude Sonnet 5No valid samples claude-sonnet-5@openrouter | Relay | — | — | — | 100.0% | 376 (0 ok) 0/168 hours coverage |
Gemini 2.5 FlashNo valid samples gemini-2-5-flash@openrouter | Relay | — | — | — | 100.0% | 369 (0 ok) 0/168 hours coverage |
How these numbers are produced
Where the probes run
🇭🇰 Hong Kong, 🇺🇸 US Central (Chicago). The two locations are listed separately and never averaged, because the same endpoint routinely differs by several times between them and the average would describe nobody. The Hong Kong node measures a cross-border international route, which is not the same as a mainland China direct connection. The US node reaching Chinese endpoints crosses the Pacific. There is no mainland node yet; if one is added it will be labelled separately.
How often it measures
TTFT is probed every 10 minutes by default and output speed every hour. A few endpoints are throttled to one round every 30 or 60 minutes to conserve provider quota (currently Kimi K3, Claude Sonnet 5 and Gemini 2.5 Flash), so rows differ widely in how many samples an hour holds. Read the sample count on the row before trusting a figure rather than assuming a uniform cadence. Results are aggregated into one row per endpoint per location per hour, and this page reads those hourly rows. A separate cost probe runs a standard task with reasoning deliberately left on; it uses a different setup and is not part of this dataset.
Reasoning mode is switched off
Reasoning mode is explicitly disabled on every probe request, so these figures describe time to first answer token, not time to finish thinking. Models that cannot disable reasoning are out of scope. Chain of thought tokens count towards completion_tokens and ignore the max_tokens cap (one model returned 896 tokens for a 256 token request), so leaving reasoning on would mix thinking speed into answering speed. The cost is coverage: Gemini 3.x will not disable reasoning, so the Gemini entry here is 2.5 Flash. Newer flagships lean towards forced reasoning, so this comparable set narrows over time.
Window percentiles are second-order aggregates
Window-level p50 and p95 are second-order aggregates of the hourly percentiles, not exact percentiles over every request in the window: raw per-request logs are retained for 14 days and are not publicly readable. ttft_p50_ms is the successful-sample-weighted median of the hourly p50 values. ttft_p95_typical_ms is the same weighted median of the hourly p95 values, i.e. the tail latency of a typical hour. ttft_p95_worst_ms is the highest hourly p95 in the window. The true window p95 lies between the typical and worst figures.
ttft_p50_ms— successful-sample-weighted median of the hourly p50 values.ttft_p95_typical_ms— the same weighted median over the hourly p95 values: the tail latency of a typical hour.ttft_p95_worst_ms— the highest hourly p95 in the window.- For endpoints measured only once or twice an hour, the typical-hour p95 collapses onto that single measurement and can equal the p50. That is a shortage of samples in the hour, not a flat tail. Divide the successful sample count by the hour coverage on the same row to see how many measurements an hour actually holds.
- tps_avg is the successful-sample-weighted mean of the hourly average tokens per second.
- error_rate is exact: the sum of error_rate x sample_count over the window divided by total samples. Hours in which every request failed carry no percentiles but still count towards samples and error_rate.
Why p95 matters, not just p50
The median tells you what a good moment feels like. The 95th percentile is the slowest of every 20 requests, and it is what decides when users start calling your app laggy, where your timeout has to sit, and how often your retry path fires. Two endpoints can both sit at 500 ms p50 while one has an 800 ms p95 and the other a 6 second p95; for anything interactive those are different products. Both the typical hour and the worst hour are listed here because "fine most of the time, awful occasionally" is the common failure shape.
How thin samples are flagged
Rows with measurements in under 50% of the window's hours, or fewer than 100 successful samples, are marked "few samples". The numbers are still shown, but they move around more than the rest. Rows where every hour failed are marked "no valid samples" and the latency columns stay blank rather than showing zero. Failed hours are not dropped from the error rate denominator: an endpoint unreachable half the time would otherwise look as fast as a stable one.
Limits
Two probe locations do not represent global experience, and they do not represent the route from your own servers. The probes send a fixed short prompt, so these figures do not describe long context, high concurrency or tool-calling workloads. Latency is one dimension of endpoint choice and this page does not rank vendors. To judge whether a number is trustworthy, read the sample count and hour coverage on the same row.
Cite and download
Published under CC BY 4.0. Free to use in articles, reports and research with attribution and a link back to this page. Both endpoints carry the same rows and refresh hourly along with the page.
curl https://www.llmabacus.com/api/latency
curl https://www.llmabacus.com/api/latency/csvSuggested citation: