inference atlas

Track 02 · Intermediate · ~30 min

When one GPU is not enough.

Models outgrow a single GPU, and traffic outgrows a single replica. Learn the five ways to split a model, how a smart router sends each request to the GPU that already remembers its prompt, and why the network becomes part of the compute path.

2.1Intermediate

Will it fit?

Pick a model and a GPU, then raise the tensor-parallel degree. Each GPU holds a slice of the weights, and whatever is left becomes KV cache. That leftover is why going from 1 to 2 GPUs can more than double throughput.

Lab · sharding weights across GPUs

Llama-3.1-70B on 4 × H100 80GB

fits
Weight precision
Tensor-parallel GPUs
GPU 080 GB
weights 35.3 GB
KV free 34.7 GB
GPU 180 GB
weights 35.3 GB
KV free 34.7 GB
GPU 280 GB
weights 35.3 GB
KV free 34.7 GB
GPU 380 GB
weights 35.3 GB
KV free 34.7 GB
Weights total141 GB
Per GPU35.3 GB
KV per token per GPU80.0 KiB
Users at 8K context51
Smallest TP that fits4
Rule: pick the smallest TP inside one NVLink node that holds weights plus enough KV for your concurrency. Every extra TP rank adds all-reduce traffic on every layer, so scale out with replicas (DP) after that.
2.2Intermediatecore idea

Five ways to split a model.

Each strategy cuts along a different axis: inside a layer (TP), between layers (PP), whole copies (DP), across experts (EP), or along the sequence (CP). Pick one and step through what moves where.

Lab · parallelism explorer

Tensor parallelism · split inside each layer

One layer’s weight matrix W
Every token multiplies by W
GPU 0 · ¼ of W
XY0 partialΣ Y · full
GPU 1 · ¼ of W
XY1 partialΣ Y · full
GPU 2 · ¼ of W
XY2 partialΣ Y · full
GPU 3 · ¼ of W
XY3 partialΣ Y · full
NVLink all-reduce · every GPU sends and receives partial sums
Step 1/7
One transformer layer multiplies activations X by a big weight matrix W. For Llama-70B one layer is about 1.7 GB of weights, and there are 80 layers.
inside a layer

TP · Tensor

Split each weight matrix across GPUs.

Talks: 2 all-reduces per layer, every step.

Use when: Model won’t fit, or you need lower latency. Stay within NVLink.

between layers

PP · Pipeline

Consecutive layers on different GPUs.

Talks: Activations at stage boundaries only.

Use when: Too big for one node and inter-node links are slow. Watch bubbles.

whole copies

DP · Data

Independent replicas serve different requests.

Talks: None between replicas.

Use when: You need throughput. Pair with cache-aware routing.

across experts

EP · Expert

MoE experts live on different GPUs.

Talks: All-to-all dispatch + combine per MoE layer.

Use when: MoE models, usually with DP attention (“wide EP”).

along the sequence

CP · Context

One request’s context is sharded across GPUs.

Talks: Query gather + partial-output merge.

Use when: Very long contexts: DCP for decode, PCP for prefill.

2.3Intermediatesimulator

Route to the GPU that remembers.

Four vLLM replicas each keep a prefix cache of the prompts they’ve seen. A request whose prefix is cached skips most of its prefill. Switch routing policies, turn on a hot prompt, and watch hit rate, TTFT and queues react.

Simulator · 4 replicas × 2 decode slots · 2 prefixes cached per replica

Same traffic. Four routing brains.

Routing policy
Incoming requests
ASupport bot6K system prompt
BCode assistant8K tool defs
CHR docs RAG5K documents
DAgent sessionlong context
Router
Next pod in rotation. Ignores load and caches.
pod-1 · vLLMserved 0
KV cache
——
Running
Queue
empty
pod-2 · vLLMserved 0
KV cache
——
Running
Queue
empty
pod-3 · vLLMserved 0
KV cache
——
Running
Queue
empty
pod-4 · vLLMserved 0
KV cache
——
Running
Queue
empty
Prefix-cache hit rate—
Avg TTFT—
P95 TTFT—
Longest queue now0
Load imbalance1.00× max/mean
Round robin · t=0
Warming up: caches start cold, so the first requests are all misses (✗ = full prefill, ✓ = prefix reused).

Why round-robin fails

LLM requests are expensive and uneven, replicas hold cached state, and prefill/decode stress different hardware. Spreading requests evenly spreads cache misses evenly.

Good vs bad signals

Good: queue depth, KV-cache usage, prefix overlap, predicted latency. Bad: CPU or GPU utilization and connection counts.

In the real world

llm-d’s endpoint picker cut mean TTFT ~3× and served ~50–100% more QPS within a 2 s P95 TTFT SLO versus baseline routing. See it on Kubernetes →

Staff answer

“Inference scheduling isn’t CPU-style load balancing. I’d score replicas on prefix-cache overlap, queue depth and KV-cache pressure, then break ties by load, so we’re sticky until saturated. The cheapest-looking routing decision can otherwise create expensive recomputation.”

2.4Intermediate

GPUs talking: collectives and links.

Split models must exchange data every step. NCCL provides the collective operations; the interconnect decides how long they take. Step through each collective, then compare links.

Lab · NCCL collectives on 4 GPUs

All-reduce

GPU 0
a0
b0
c0
d0
GPU 1
a1
b1
c1
d1
GPU 2
a2
b2
c2
d2
GPU 3
a3
b3
c3
d3

Used by tensor parallelism: partial results are summed after attention and after the MLP, twice per layer.

Step 1/4
Each GPU holds a full buffer of partial results (4 chunks: a–d). Goal: every GPU ends with the element-wise sum across all GPUs.

The bandwidth ladder

Time to move 3.3 GB, the KV cache of a 10K-token prompt on Llama-3.1-70B (≈ 320 KiB/token). Log scale: every gridline step is ×10.

HBM3 · the GPU’s own memory (H100)1 ms · 3,350 GB/s
NVLink 5 · B200, per direction4 ms · 900 GB/s
NVLink 4 · H100, per direction7 ms · 450 GB/s
PCIe Gen5 x16 · GPU ↔ CPU52 ms · 64 GB/s
InfiniBand NDR · 400 Gb/s66 ms · 50 GB/s
RoCE · 200 Gb/s132 ms · 25 GB/s
Ethernet · 100 Gb/s264 ms · 12.5 GB/s

Per-direction peak rates; real transfers reach less. HBM is the GPU’s own memory, shown for scale.

Inside a node

NVLink/NVSwitch joins 8 GPUs at hundreds of GB/s. Keep tensor parallelism here: it all-reduces twice per layer, every step.

Across nodes

InfiniBand or RoCE with RDMA (GPU memory → NIC with no CPU copy). Roughly 10× slower than NVLink, so use PP, EP or DP across nodes.

Staff answer

“In multi-node inference the network is part of the compute path: TP, PP and MoE all move tensors every step. I treat bandwidth, latency, RDMA and topology as first-class constraints, and when throughput drops going from one node to two, I check NCCL collectives before blaming the model.”

2.5Quiz

Pick the right split.

Scenario questions like the ones interviewers ask. For drag-and-drop matching, visit the playground.

1. Llama-3.1-70B in BF16 (~141 GB) must serve a chat product with low latency on one 8×H100 node. Best starting layout?

2. A 405B model doesn’t fit on one node. Nodes connect over 400 Gb/s InfiniBand. How do you split it?

3. Why is plain TP a poor fit for DeepSeek-V3’s MLA attention?

4. Round robin gave a 30% prefix hit rate. Prefix-hash routing gave 90%, but P99 TTFT exploded when one system prompt went viral. The fix?

5. Which NCCL collective does MoE token dispatch use?

0 / 5 answered