Track 05 · Advanced · ~25 min
The unit of serving becomes a workflow.
Modern workloads aren’t one prompt and one answer. Mixture-of-experts routes every token across GPUs, multimodal models chain encoders and LLMs, RL puts inference inside the training loop, and agents turn one request into dozens of cached turns.
Send tokens to experts.
Each token’s router picks its top-2 of 8 experts, spread over 4 GPUs. Every MoE layer waits for the slowest GPU. Skew the traffic toward code and watch one GPU become the straggler, then switch on EPLB.
Lockstep means the busiest GPU sets the pace.
What
MoE layers hold many expert FFNs, but each token activates only a few. Memory scales with total parameters and compute with active ones: DeepSeek-R1 uses 37B of 671B per token.
Small batches
Tokens per expert ≈ batch × top-k ÷ experts. DeepSeek-V3 sends each token to 8 of 256 experts, so a 256-token decode batch gives each expert about 8 tokens: a bandwidth-bound GEMM.
EPLB
vLLM’s expert-parallel load balancer records per-expert load, periodically computes a new placement with redundant copies of hot experts, and shuffles weights without a restart. Redundant experts cost HBM that could hold KV.
Staff answer
“MoE serving cuts active compute per token but adds all-to-all token routing. Because every rank must join each MoE layer’s collective, the group moves at the pace of its slowest rank, so expert load balance, DP-aware routing and prefill scheduling decide throughput. I’d measure the max/mean tokens per rank per layer before and after EPLB.”
A picture is worth hundreds of tokens.
A user uploads a screenshot of a failing pod and asks what’s wrong. Follow the image through preprocessing, the vision encoder and into the LLM, then send it again and watch the caches kick in.
Image → patches → embeddings → LLM
Off-loop preprocessing
Decoding, resizing and cropping images runs in a separate process, with a cache for repeated inputs, so the GPU loop never waits on it.
Encoder cache
Vision embeddings are kept after the encoder runs, so a long prefill can be chunked across steps without re-running the encoder.
Multimodal prefix caching
Image hashes join token IDs in block hashes, so multi-turn chats about the same image hit the cache. Client-side re-encoding breaks it: normalize media first.
Staff answer
“Multimodal serving chains heterogeneous stages: preprocessing on CPU, vision or audio encoders, then the LLM. I’d treat the encoder as cacheable, separately scaled work, via E/PD disaggregation at scale, and pick per-stage metrics, because TTFT and TPOT don’t describe a vocoder or a diffusion stage.”
Inference inside the training loop.
Reasoning models are post-trained with reinforcement learning: generate attempts, score them, update the weights, repeat. The generation half is an inference workload, and it must never starve the trainer.
Rollout → reward → train → sync weights
Synchronous: each side waits for the other. Rollout GPUs busy 72%, trainer busy 24% (hatched = idle).
Fast weight sync
New policy weights must reach rollout engines without restarts: modular weight-sync APIs, NCCL broadcast, and sleep mode to hand GPU memory between trainer and engine.
The silent bug
After a weight update, cached KV was computed with the old weights. Forget to reset the prefix cache and rollouts become subtly off-policy, with no error anywhere.
Batch invariance
Floating-point reductions change with batch shape, so a request’s logprobs can shift with its batchmates. RL importance ratios then drift for no visible reason.
Staff answer
“RL infra has two competing workloads, training and rollout generation. I treat rollouts as a scalable inference workload, overlap them with training where the algorithm allows, and make weight sync, prefix-cache invalidation and batch-invariant logprobs part of the contract, so training is never starved and never silently off-policy.”
An agent is a session, not a request.
An SRE agent investigates “why is checkout failing?”. Every turn resends the whole growing context plus one new tool result. Route turns stickily or load-balance them, add a shared KV pool, and compare what you pay.
Cached prefix vs new tokens, turn by turn
Real agent traffic
SemiAnalysis’ AgentX traces: a median of 43 turns per session, 142K input and 444 output tokens per request, and a prefix-cache hit rate above 96%. 44% of sessions spawn subagents.
What’s scarce
With 96% cached input, a turn costs its decode, a small prefill and the memory to keep its KV resident. KV capacity, not FLOPs, becomes the bottleneck.
Bitter lessons
On agent traces, session-sticky routing beat balancing by queue, tokens or KV, and pipeline parallelism lost because warm turns add too few new tokens to fill a pipeline.
Staff answer
“Agentic serving changes the unit from one model call to a multi-step session. I’d route by session ID, sticky until a replica saturates, keep KV warm in tiers between turns, cap long fresh prefills so they don’t block short turns, and measure tokens per GPU-second at a P90 interactivity floor, reporting cached, uncached and output tokens separately.”
Check yourself.
Frontier-workload questions from real staff loops.
1. DeepSeek-R1 has 671B parameters, 37B active per token. What scales with the 671B?
2. In a 32-rank wide-EP decode group, one rank starts a long prefill. What happens?
3. In a multi-turn image chat, turn 2 about the same image is as slow as turn 1. Likely cause?
4. After a weight update, RL rollouts are subtly off-policy but nothing errors. First suspect?
5. Which routing policy serves long agentic sessions best?
0 / 5 answered