Prefill vs decode
“Walk me through what happens between a prompt arriving and tokens streaming back. Where do TTFT and ITL come from?”
Answer out loud first, then reveal
- What it is
- Prefill runs the whole prompt through the model in one parallel pass and builds the KV cache. Decode then generates one token at a time.
- Why it matters
- Prefill is compute-bound and sets time to first token; decode is memory-bandwidth-bound and sets inter-token latency. They need different tuning.
- Scenario
- “The first response takes 5 seconds” → look at queueing and prefill. “The first word is quick but the answer takes forever” → look at decode throughput and ITL.
- Staff answer
- “Prefill processes the input and builds the KV cache, so it’s primarily compute-bound and affects TTFT. Decode generates tokens autoregressively and is more memory-bandwidth and KV-cache sensitive, affecting inter-token latency. That’s why production systems often optimize or scale the two phases independently.”