Measuring Latency

Metrics

Part of: Streaming, Latency & Caching

An engineer reports "our average latency is 800 milliseconds, we are fine." Meanwhile one in twenty users waits four seconds and churns. The average hid them. To actually run a fast AI product, you measure the right numbers, and the average is rarely one of them. What it is Latency metrics are the specific numbers you track to know how responsive your system is. For streaming LLMs, three matter most: - TTFT : time to first token, the responsiveness users feel first. - Total latency : time from request to the final token. - Tokens per second : the streaming throughput, total output tokens divided by total time. And one statistical idea changes everything: you summarize these with percentiles , not averages. How it works A percentile answers "what value is X percent of requests at or below?" The p50 (median) is the typical experience; the p95 and p99 are the slow tail that frustrates real users. You sort your measurements and read off the value at that position: The average of that list is 1.24s, which describes nobody: most requests are far faster, and the worst is far slower. The p95 of 5.0s is what your unhappiest users actually live with. That is why teams set targets like "p95 T

Challenge: The Latency SLO Checker