Streaming, Latency & Caching

A user waits eight seconds staring at a spinner, then a full paragraph appears at once. The same answer, streamed, shows its first word in 400 milliseconds and feels instant. The total time barely changed, but the perceived time collapsed, and that gap is the

8 lessons, each with runnable code in the browser.

  1. Why Streaming Matters (Time To First Token)
  2. Server-Sent Events
  3. Prompt Caching
  4. Batching & Concurrency
  5. Measuring Latency
  6. Token Throughput
  7. Semantic Caching
  8. Cost vs Latency Tradeoffs

Compilearn home