Track 04 · Intermediate · ~35 min
Kubernetes, but it understands tokens.
A plain Kubernetes Service balances connections. LLM serving needs a platform that knows about prompts, queues, KV caches and GPU topology. Walk a request through llm-d + vLLM, read a real pod spec, pack GPUs, tune an autoscaler and debug a live incident.
Follow one request through llm-d.
Compare a naive Deployment + Service with llm-d: an Envoy gateway, an endpoint picker that scores pods on prefix cache, queue and KV usage, and separate prefill and decode pools linked by NIXL. Click any box to learn what it is.
llm-d: inference-aware routing on Kubernetes
Client
Your app, agent or SDK calling an OpenAI-compatible API (/v1/chat/completions). It has no idea how many GPUs sit behind the endpoint.
POST /v1/chat/completions with model: "support-bot".Gateway API Inference Extension
Standard Kubernetes APIs for inference routing: an InferencePool groups model-server pods, and an Endpoint Picker (EPP) chooses one pod per request through Envoy’s ext-proc.
llm-d
A Kubernetes-native distributed inference stack on vLLM: an inference scheduler (EPP scorers), P/D disaggregation, wide-EP “well-lit paths”, KV-cache-aware routing and autoscaling.
Two routers, two jobs
A semantic router picks which model (cheap vs strong). The EPP picks which replica of that model. They compose: Envoy → semantic router → EPP → vLLM pod.
Staff answer
“I keep the gateway generic and put LLM-specific judgment in the endpoint picker: filter by health, role and loaded adapters, then score on prefix overlap, queue depth and KV utilization. That runs across heterogeneous pods and plugs into Kubernetes rollouts and autoscaling, which one big monolithic deployment can’t do.”
Read a vLLM pod spec line by line.
Click any highlighted line of this Deployment to see what it controls and which interview topic it connects to. The diagram shows the processes running inside the pod.
Tensor parallelism = 4
This pod shards every layer across its 4 GPUs over NVLink. Must match the GPU limit below. Pick the smallest TP with enough KV headroom.
Be the scheduler: pack the GPUs.
Drag each pod onto a node, or tap a pod and then tap a node. Tensor-parallel pods need all their GPUs on one NVLink node, and some need a specific GPU type. Place everything without stranding GPUs.
What K8s knows
The NVIDIA device plugin / GPU Operator exposes nvidia.com/gpu as a countable resource, with node labels for GPU model. The default scheduler counts GPUs; it doesn’t see NVLink islands.
What you add
nodeSelector or affinity for GPU type, topology-aware placement, gang scheduling for multi-node groups (e.g. LeaderWorkerSet, Kueue, Volcano), and MIG partitions for small models.
Fragmentation
Four free GPUs spread over four nodes can’t host one TP=4 pod. Big jobs first, or a descheduler, keeps whole nodes free.
Staff answer
“GPU scheduling is placement on the right accelerator type and topology, not just finding a free GPU. For TP I need all ranks inside one NVLink domain; for multi-node groups I need gang scheduling so a half-placed group doesn’t sit on idle GPUs. Poor placement shows up directly as latency and throughput loss.”
Scale on the right signal.
A morning ramp and a lunchtime spike hit a vLLM deployment. Pick the metric your autoscaler watches, set how long a new replica takes to become ready, and see who blows the TTFT SLO.
Same traffic, different autoscaling signal
Good signals
vllm:num_requests_waiting, vllm:kv_cache_usage_perc, TTFT/ITL SLO attainment, preemption rate. With P/D: queue and TTFT for prefill, KV usage and ITL for decode.
Misleading signals
CPU barely moves. GPU utilization pegs near 100% at moderate load, so it can’t tell “busy” from “drowning”.
Cold start
A new replica loads tens to hundreds of GB of weights and captures CUDA graphs: minutes. Use warm pools, compile caches, fast weight streaming, and scale ahead of known peaks.
Staff answer
“For LLM serving I don’t autoscale on CPU or raw GPU utilization. Queue depth, KV-cache pressure, TTFT and token throughput track user-visible saturation. I also budget for cold start: new replicas have cold prefix caches, so I ramp them in and let cache-aware routing warm them.”
You’re on call. What broke?
A live dashboard for a vLLM fleet. Start an incident, watch which panels move, then pick the root cause. Each answer explains the tell-tale signals, the way you’d narrate it in an interview.
All systems normal
llm-d, Dynamo, Ray Serve, KServe, production-stack?
All of them can run vLLM underneath. They differ in the platform they assume and what they optimize. Answer three questions for a starting recommendation, then read the trade-offs.
| Stack | Built on | Strengths | Pick it when |
|---|---|---|---|
| llm-d | Kubernetes, Gateway API Inference Extension | Cache- and load-aware scheduling (EPP), P/D and wide-EP “well-lit paths” | You run Kubernetes and want upstream-aligned parts |
| NVIDIA Dynamo | Orchestration above vLLM, SGLang or TensorRT-LLM | KV-aware routing, KV Block Manager tiers, Planner for dynamic P/D | You want several engines behind one layer on NVIDIA fleets |
| Ray Serve LLM | Ray and KubeRay | P/D, data-parallel attention, prefix-affinity routing; ties into Ray Data and RL | Your platform already runs on Ray |
| vLLM production-stack | Helm charts on Kubernetes | Reference router, LMCache integration, observability | You need a simpler quick-start deployment |
| KServe | Kubernetes serving CRDs | Standard InferenceService API, vLLM runtime, multi-framework model serving | You serve many model types (not only LLMs) behind one API |
Quick chooser
A starting point, not a verdict. Keep an OpenAI-compatible API and engine-neutral metrics at the edge so you can switch later.
Check yourself.
Platform questions staff interviews love.
1. In llm-d, which component decides which vLLM pod serves a given request?
2. Your HPA targets 70% CPU and never scales, even while TTFT is 8 s. Why?
3. A TP=4 pod stays Pending although the cluster has 6 free H100s. Most likely?
4. Why mount a memory-backed volume at /dev/shm for a TP vLLM pod?
5. During an incident, NCCL bandwidth drops, ITL rises, throughput halves and GPU utilization FALLS. Likely cause?
0 / 5 answered