inference atlas

Practice · drag and drop · ~20 min

Learn it with your hands.

Five games that turn the atlas into muscle memory. Drag a card onto a target, or tap a card and then tap a target (works on phones and with the keyboard). Every drop explains itself, right or wrong.

Game 1Beginner

Build the request path.

Put the eight hops a chat request takes in order, from the user’s app to the streamed answer. Get them all right and the request runs through your pipeline.

Drag the hops into the numbered slots, then press Check order.
Game 2Intermediate

Which layer does it live in?

Sort 20 concepts into the five layers of the atlas. Each drop is checked instantly, with a one-line reason and a link to the lab.

Foundationsone request, one GPU
Scale outmany GPUs, routing
Optimizedo less, reuse more
PlatformKubernetes & ops
FrontierMoE, multimodal, RL, agents
Drop each card into a column. Your first-try accuracy is scored.
Game 3Staff

Diagnose the bottleneck.

The staff interview framework: is it compute, memory, communication, scheduling or workload? Drag each symptom to its most likely category.

ComputeFLOPs saturated
MemoryKV capacity or bandwidth
CommunicationNCCL · network · all-to-all
Schedulingqueue · placement · locality
Workloadthe traffic changed
Drop each card into a column. Your first-try accuracy is scored.
Game 4Staff

Match the problem to the fix.

Eight production problems, eight techniques. Drop each fix onto the problem it solves best.

The same 6K-token system prompt is recomputed for every request.

A 70B model in BF16 doesn’t fit on one 80 GB GPU.

Long document prompts make everyone’s token stream stutter.

Decode is slow at low load while GPU compute sits idle.

A viral system prompt overloads the one replica that caches it.

Long contexts fill the KV cache and cause preemptions.

The HPA never scales, although TTFT is 8 s.

One MoE rank is always the straggler.

Drag a fix onto the problem it solves.
Game 5Explore

The complete map.

Every concept in the atlas and how they connect. Click a node to light up its neighbours and read what it is. Drag nodes to untangle the map your own way. Follow the link to its lab.

Concept map

Pick any node

32 concepts, 42 connections. Click a node to see how it connects; drag nodes to rearrange. Nodes you’ve opened count toward your scoreboard.

Explored: 0 / 32

LabPractice

Practice lab.

400 hands-on exercises across every track: tune, predict, calculate, fix, order, diagnose, respond to incidents, build commands, classify and recall. Raise your confidence to 90%.

Confidence0% Not startedAverage best score over all 400 exercises; untried ones count as 0. Target: 90%.
Tried0 / 400
Solved 90+0
Cards due today—

Coverage map

Every lesson in the atlas, coloured by your best scores on the exercises that practise it. Pick one to list them.

  • Not started
  • Learning
  • Solid
  • Mastered
  • No exercises yet

01Foundations50 exercises

02Scale out50 exercises

03Optimize50 exercises

04Kubernetes platform50 exercises

05Frontier serving50 exercises

06Inside vLLM50 exercises

07Kubernetes for GPU workloads50 exercises

08Interview kit50 exercises

Exercises

1–24 of 400 exercises

  1. CalculatorCore4 minA 36,000-word contract in tokens and KV01 · Tokens01 · GPU memory budgetNew
  2. ClassifyCore5 minWhich token cost does it cut?01 · Tokens01 · Prefill → KV → Decode+1New
  3. Predict then runAdvanced5 minSame characters, twice the tokens01 · Tokens01 · TTFT · ITL · E2E+1New
  4. Fix the configAdvanced6 minFix the context budget01 · Tokens01 · GPU memory budgetNew
  5. Order the stepsCore4 minOne request, from text to its last token01 · Prefill → KV → Decode01 · Tokens+1New
  6. Predict then runCore4 minSwitch the KV cache off01 · Prefill → KV → Decode01 · Attention (Q·K·V)New
  7. CalculatorStaff8 minThe fastest a decode step can be01 · Prefill → KV → Decode01 · GPU memory budget+1New
  8. Spot the bottleneckAdvanced5 minBuy GPUs with twice the FLOPs?01 · Prefill → KV → Decode01 · TTFT · ITL · E2E+1New
  9. ClassifyCore5 minCache it, or use it once?01 · Attention (Q·K·V)01 · Prefill → KV → DecodeNew
  10. Order the stepsAdvanced5 minOne attention layer, one decode step01 · Attention (Q·K·V)01 · Prefill → KV → DecodeNew
  11. Predict then runAdvanced5 min32 KV heads to 8: what changes?01 · Attention (Q·K·V)01 · GPU memory budgetNew
  12. Fix the configStaff7 minFix the attention pseudo-code01 · Attention (Q·K·V)01 · Prefill → KV → DecodeNew
  13. ClassifyCore5 minTTFT, ITL or neither?01 · TTFT · ITL · E2E01 · Prefill → KV → Decode+1New
  14. Spot the bottleneckAdvanced5 minTTFT 20 s, GPUs at 70%01 · TTFT · ITL · E2E01 · Continuous batchingNew
  15. Command builderAdvanced6 minMeasure latency at a stated load01 · TTFT · ITL · E2E01 · Continuous batchingNew
  16. IncidentStaff9 min“It feels slow” after a context-limit release01 · TTFT · ITL · E2E01 · GPU memory budget+1New
  17. Predict then runCore5 minEight requests, four slots01 · Continuous batching01 · TTFT · ITL · E2ENew
  18. CalculatorAdvanced5 minHow big is the batch, really?01 · Continuous batching01 · TTFT · ITL · E2ENew
  19. Tune to targetStaff8 minPick the batch: throughput vs per-user speed01 · Continuous batching01 · TTFT · ITL · E2E+1New
  20. CalculatorStaff8 minLlama-3.1-70B on four H100s: how many 16K conversations?01 · GPU memory budget02 · Will it fit?New
  21. Command builderCore5 minFit Llama-3.1-8B on one L401 · GPU memory budget06 · vllm serve builderNew
  22. Fix the configAdvanced6 minFix a long-context serve script01 · GPU memory budget06 · vllm serve builderNew
  23. Order the stepsAdvanced5 minHow vLLM turns GPU memory into KV capacity01 · GPU memory budget06 · Capacity mathsNew
  24. FlashcardsCore7 minTokens, prefill, decode and attention: one-sentence answers01 · Prefill → KV → Decode01 · Tokens+1New