Agentic serving
Serving multi-turn agent sessions: long, 96%-cached contexts where KV residency and sticky routing dominate cost.
See it in the lab →Reference · plain English
Short definitions for everything in the atlas, each with a “like…” analogy where it helps and a link to the lab where you can watch it working.
Serving multi-turn agent sessions: long, 96%-cached contexts where KV residency and sticky routing dominate cost.
See it in the lab →Every GPU ends with the sum of everyone’s buffers. Used by TP after attention and MLP.
See it in the lab →Every GPU sends a different chunk to every other GPU at once. Used for MoE dispatch and combine.
See it in the lab →Planning step N+1 while step N runs on the GPU, so the GPU never waits on the CPU.
See it in the lab →Each token’s query is compared with earlier tokens’ keys; the resulting weights mix their values. K and V are reused by every future token; Q isn’t.
See it in the lab →Each new token depends on all previous tokens, so decode steps can’t run in parallel for one sequence.
See it in the lab →A prompt gets identical logprobs whatever else is in its batch. Needed for reproducible RL and evals.
See it in the lab →A per-tenant value mixed into the first block hash so tenants never share KV blocks.
See it in the lab →Sending each request to the replica that already holds its prefix, balanced against load.
See it in the lab →Capacity bought ahead for one instance type in one zone, sometimes for a fixed window (EC2 Capacity Blocks). Outside that box, reserved-only pods wait.
See it in the lab →On-demand, spot or reserved. A pod or pool that allows only one type can only use offerings of that type.
See it in the lab →Splitting a long prompt’s prefill across several steps so it shares the GPU with ongoing decodes.
See it in the lab →A Kueue quota pool, often one per team. ClusterQueues in one cohort can borrow each other’s idle quota, and lenders can reclaim it.
See it in the lab →The minutes a new replica spends loading weights and capturing CUDA graphs before it can serve.
See it in the lab →Karpenter deleting or replacing nodes to cut cost. consolidateAfter sets how long a node must be stable first (default 0s; Never turns it off).
Shard one long sequence’s KV (decode CP) or prompt (prefill CP) across GPUs; partial attention merged exactly.
See it in the lab →The batch is re-formed every step: finished sequences leave, waiting ones join immediately.
A restaurant seating new guests the moment a table frees up.
See it in the lab →Cordon marks a node unschedulable; drain then evicts its pods through the Eviction API, which respects PDBs.
See it in the lab →Record a sequence of kernel launches once and replay it, removing CPU launch overhead. vLLM uses FULL_AND_PIECEWISE.
See it in the lab →Independent model replicas serving different requests. For MoE, “DP attention” gives each rank its own requests’ KV.
See it in the lab →Dual-batch overlap: splits a step into two micro-batches so one’s all-to-all overlaps the other’s compute.
See it in the lab →Generating output one token per step; each step reads all weights plus the KV cache. Memory-bandwidth-bound; sets ITL.
Writing the answer one word at a time, glancing back at your notes each time.
See it in the lab →A node agent that registers devices with the kubelet as allocatable resources such as nvidia.com/gpu. Until it runs, a Ready node advertises 0 GPUs.
See it in the lab →The karpenter.sh/do-not-disrupt pod annotation: acts like a single-pod blocking PDB, permanently or for a set duration.
A node no longer matches its NodePool or node class, for example after a new image. Karpenter replaces drifted nodes gracefully, within disruption budgets.
See it in the lab →Drafters: EAGLE-3 is a small head reading the target’s hidden states; MTP uses the model’s own multi-token prediction heads.
See it in the lab →Keep vision embeddings so images aren’t re-encoded; at scale run encoders on their own workers (encode/prefill/decode).
See it in the lab →Called by the gateway via ext-proc; runs filters (health, role, adapter) then scorers (prefix, queue, KV).
See it in the lab →Expert-parallel load balancer: replicates hot experts and reshuffles placement live so lockstep ranks stay balanced.
See it in the lab →MoE experts spread across GPUs; tokens travel all-to-all to their experts and back.
See it in the lab →Maximum node lifetime (default 720h). Expiry is forceful: draining starts without a pre-spun replacement.
See it in the lab →Enabled (a pool exists), allocatable (a node advertises GPUs), requested (pods ask for them), consumed (GPUs bound to running pods). Four numbers, four owners.
See it in the lab →Storing KV in 8 bits: half the bytes per token, roughly double the capacity. Validate long-context accuracy.
See it in the lab →Place all pods of a multi-node group together or none (LeaderWorkerSet, Kueue, Volcano), so half-placed groups don’t idle GPUs.
See it in the lab →Kubernetes APIs for inference routing: InferencePool groups model-server pods; an Endpoint Picker chooses one per request.
See it in the lab →Requests per second that meet every SLO at once. Throughput that users actually benefit from.
See it in the lab →Grouped-query attention: several query heads share one K/V head, cutting KV memory (Llama-3.1: 8 KV heads vs 32 in Llama-2-7B).
See it in the lab →Graceful (Karpenter drift and consolidation) pre-spins a replacement node and waits for it. Forceful (expiration, interruption, repair) starts draining immediately.
See it in the lab →One long request consumes every step’s budget so short requests behind it can’t start. Fixed with a per-request prefill cap.
See it in the lab →Kubernetes autoscalers. For LLMs, drive them with queue depth or KV usage, not CPU.
See it in the lab →Node-to-node networks; RDMA lets a NIC read and write GPU memory without CPU copies.
See it in the lab →The cloud can’t supply that instance type in that zone right now. Pods that allow other zones or types still land.
See it in the lab →Tokens per second per user. Benchmarks report throughput per GPU at a fixed interactivity floor (e.g. 50 tok/s/user).
See it in the lab →Inter-token latency (per-token gaps) and time per output token (per-request mean). What users feel as streaming speed.
See it in the lab →Kubernetes model-serving CRDs (InferenceService) with runtimes such as vLLM, for many model types behind one API.
See it in the lab →Job queueing and quota for Kubernetes: decides when a job is admitted, waits or is preempted. It doesn’t create nodes.
See it in the lab →The Key and Value vectors of every earlier token at every layer, kept in GPU memory so decode never recomputes them.
Your notes from reading the question, so you don’t reread it for every word.
See it in the lab →vLLM’s interface for loading and saving KV outside the GPU: P/D transfer, CPU offload, remote caches.
See it in the lab →Moving KV blocks to CPU memory or disk (DMA copies) so they can be reloaded instead of recomputed.
See it in the lab →In-flight requests = arrival rate × time in system. 50 req/s × 20 s decode ≈ 1,000 concurrent sequences.
See it in the lab →Kubernetes-native distributed inference on vLLM: inference scheduler (EPP), P/D and wide-EP “well-lit paths”, KV-aware routing.
See it in the lab →KV cache layers outside the engine: shared pools across pods and nodes for reuse and P/D transfer.
See it in the lab →Multi-Instance GPU: partitions one GPU into isolated slices, useful for small models.
See it in the lab →Many expert FFNs; each token activates a few. Memory ∝ total params, compute ∝ active (DeepSeek-R1: 37B of 671B).
See it in the lab →Multi-head latent attention (DeepSeek): stores one compressed latent per token instead of per-head K/V. Tiny KV, but TP can’t split it.
See it in the lab →vLLM’s 2026 rewrite of the model runner: GPU-native input prep, async-first, Triton sampler (+56% on a small model).
See it in the lab →Many fine-tuned adapters on shared base weights, batched together; the LoRA ID is part of every block hash.
See it in the lab →One spare node’s worth of capacity so a replacement can be Ready before the old node drains.
A spare tyre: useless until the day you need it.
See it in the lab →NVIDIA Collective Communications Library: all-reduce, all-gather, reduce-scatter, broadcast, all-to-all across GPUs.
See it in the lab →NVIDIA Inference Xfer Library: async KV transfer over RDMA/NVLink/storage; vLLM’s NixlConnector uses it for P/D.
See it in the lab →Karpenter feature (alpha) that force-replaces nodes whose health conditions stay bad, e.g. AcceleratedHardwareReady false for 10 minutes. It pauses if over 20% of a pool is unhealthy.
See it in the lab →The cloud launch settings a NodePool uses (EC2NodeClass on AWS): image, subnets, security groups, disks and capacity reservations.
See it in the lab →One concrete node Karpenter decided to launch (instance type, zone, capacity type). It’s initialized once the node is Ready, startup taints are gone and requested resources are registered.
See it in the lab →A Karpenter policy for new nodes: allowed instance types, zones and capacity types, plus taints, limits and disruption rules. Ready means valid config, not running nodes.
See it in the lab →A cap on the total CPU, memory or GPUs a NodePool may provision. Once it’s reached, Karpenter launches nothing more and pods stay Pending.
See it in the lab →Orchestration above vLLM, SGLang or TensorRT-LLM: KV-aware routing, KV block manager tiers, planner for P/D.
See it in the lab →The extended resource the NVIDIA device plugin exposes; pods request whole GPUs through it.
See it in the lab →Direct GPU-to-GPU links inside a node (hundreds of GB/s per GPU per direction).
See it in the lab →Separate prefill and decode GPU pools; KV moves between them. Better latency control; needs rate-matched pools.
A kitchen with separate prep and plating stations.
See it in the lab →KV stored in fixed 16-token blocks mapped by a per-request block table; cut KV waste from 60–80% to under 4%.
Virtual memory pages for the KV cache.
See it in the lab →Different layer ranges on different GPUs; micro-batches flow through; costs pipeline bubbles.
An assembly line.
See it in the lab →Limits how many replicas voluntary evictions may take down at once. Honoured by kubectl drain and Karpenter; can’t prevent involuntary failures.
See it in the lab →When KV blocks run out, the scheduler evicts a running request and recomputes it later.
See it in the lab →The phase that runs the whole prompt through the model in one parallel pass and writes the KV cache. Compute-bound; sets TTFT.
Reading the whole question before you start answering.
See it in the lab →Full KV blocks are hashed (chained through their parents) so later requests with the same prefix reuse them.
See it in the lab →Fewer bits per weight/activation. W4A16 = 4-bit weights, 16-bit maths; W8A8 = 8-bit both (FP8 on Hopper).
See it in the lab →Maintenance rule: the new replica is Ready before the old one is evicted. Needs two or more spread replicas, a PDB and room for a surge node.
See it in the lab →A namespace cap on requested resources, checked when pods are created. Over quota, pod creation is forbidden.
See it in the lab →Samples generated by the current policy during RL post-training. An inference workload inside the training loop.
See it in the lab →FLOPs per byte moved. Below the GPU’s ridge (~300 FLOPs/byte on H100) you’re bandwidth-bound: decode lives there.
See it in the lab →Chooses which model serves a request (cheap vs strong), with safety plugins. The EPP then picks the replica.
See it in the lab →A cheap drafter proposes k tokens; the target verifies them in one pass. Same output, more tokens per step at low load.
An intern drafts, the senior engineer approves several lines at once.
See it in the lab →Stay on the warm replica for a prefix or session until it’s overloaded, then spill to the next best.
See it in the lab →Grammar-constrained decoding: the allowed next tokens become a bitmask on the logits (XGrammar, llguidance).
See it in the lab →A taint repels pods that don’t tolerate it. GPU nodes are usually tainted nvidia.com/gpu:NoSchedule so only GPU pods land there.
A “staff only” sign that only badge holders ignore.
See it in the lab →Shard every layer’s matrices across GPUs; two all-reduces per layer per step. Keep it inside an NVLink node.
Four people each multiplying a quarter of a big table, then adding up.
See it in the lab →Upper bound on draining a node. After it, remaining pods are deleted even if a PDB or do-not-disrupt blocks them.
See it in the lab →A piece of text from the model’s fixed vocabulary, with an integer ID. Every cost in serving is counted in tokens.
Like syllables: common words are one, rare words are several.
See it in the lab →Max tokens per GPU step (--max-num-batched-tokens): decodes take 1 each, prefills take chunks of the rest.
Turns text into token IDs using merges learned from data (byte-pair encoding). English averages ~4 characters per token.
See it in the lab →The chat template renders tools into the prompt; a model-specific parser turns the model’s call into OpenAI tool_calls.
See it in the lab →Time to first token = queueing + prefill (+ KV transfer). What users feel as “thinking…”.
See it in the lab →vLLM’s 2025 architecture: an API-server process plus an EngineCore busy loop (schedule → execute → sample → update).
See it in the lab →Serving any-to-any models as a graph of stages (e.g. Thinker → Talker → vocoder), each with its own GPUs.
See it in the lab →A Kueue setting that evicts and requeues an admitted job whose pods don’t all become ready in time, so it stops holding quota.
See it in the lab →Pushing updated policy weights to rollout engines. Reset the prefix cache afterwards, or KV goes stale.
See it in the lab →DP attention + expert parallelism across many GPUs, as used for DeepSeek-class models (~2.2k tok/s per H200 decode).
See it in the lab →