# How AI Models Scale Beyond a Single GPU Across LLM Workloads

## Executive summary

Due to the massive size of modern AI models (potentially exceeding two terabytes), single GPUs are insufficient for hosting large language models (LLMs). The solution is distributed inference, which coordinates computational load across multiple GPUs and servers. The video details five primary methods—Data, Pipeline, Tensor, Expert, and Prefill/Decode disaggregation—that allow models to scale across multiple interconnected devices while managing constraints related to model memory footprint, KV cache growth, and request throughput.

## Key takeaways

- Data Parallelism: To handle high concurrent user traffic when the model fits on one GPU, the entire model is copied onto multiple GPUs (replicas). Requests are then intelligently routed to available replicas.
- Pipeline Parallelism: When the model is too large to fit, it can be split by layers (sharded). Each GPU holds a distinct set of layers, passing the input sequentially through the 'assembly line' of GPUs.
- Tensor Parallelism: This method splits individual layers horizontally. Each GPU computes a slice of the layer's math, and the partial results are combined via a collective operation. This requires high-bandwidth, low-latency interconnects.
- Expert Parallelism (MoE): For Mixture-of-Experts (MoE) models, the model's weights are split across GPUs. Instead of running all weights, each token is dispatched only to a handful of specialized subnetworks (experts), significantly reducing computation per token.
- Prefill/Decode Disaggregation: The inference process is split into two stages: Prefill (compute-bound, builds the KV cache) and Decode (memory-bandwidth-bound, generates tokens). Separating these phases onto dedicated GPU pools optimizes performance, provided the inter-pool connection is extremely fast.

## Technical details

- Model Constraints and Memory: Running LLMs requires managing three constraints: the physical model memory footprint (weights), KV cache growth (working memory for conversation context), and request throughput (handling concurrent users).
- Parallelism Techniques: Production systems use 'multi-dimensional parallelism,' combining techniques like Tensor parallelism within a server, Pipeline parallelism across servers, Data parallelism across replicas, and Expert parallelism for MoE models.
- Interconnect Requirements: Tensor parallelism and Prefill/Decode disaggregation require dedicated, low-latency networking (beyond standard Ethernet/TCP) to prevent the network link from becoming the primary bottleneck.

## Practical implications

- Designing high-scale AI serving infrastructure requires implementing multi-dimensional parallelism (combining data, pipeline, tensor, and expert methods).
- The orchestration layer is critical for routing requests, balancing load, and handling hardware failures across distributed GPU pools.
- The performance of the handoff between separated inference phases (e.g., Prefill to Decode) is often the limiting factor, necessitating specialized, low-latency networking.

## Topics

Large Language Models (LLMs), Distributed Computing, GPU Architecture, Machine Learning Operations (MLOps), Parallel Processing, IBM Technology, IBM

Source: https://www.youtube.com/watch?v=qZBibWYcKH4
