Schedule

Module 1 - Introduction

Module 2 - Systems implications of the Transformer architecture

Module 3 - Hardware infrastructure for machine learning

Mon, Sep 21
AI infrastructure I: from one GPU to many
Using the roofline model, communication as the third resource, the bandwidth hierarchy, scale-up versus scale-out, xAI’s Colossus
Optional reading: xAI’s Colossus 2: First Gigawatt Datacenter in the World (SemiAnalysis)
Wed, Sep 23
AI infrastructure II: scale-up domains and TPUs
NVIDIA DGX servers, NVLink and NVSwitch, GB200 NVL72, the TPU chip and its systolic array
Optional reading: (1) NVIDIA Hopper Architecture In-Depth (2) NVIDIA GB200 NVL72 (3) How to Scale Your Model, How to Think About TPUs
Mon, Sep 28
AI infrastructure III: TPU pods and scale-out networks
TPU generations, the ICI torus and optical circuit switches, Clos and rail-optimized fabrics
Optional reading: (1) TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning (2) A Scalable, Commodity Data Center Network Architecture (3) NVIDIA DGX SuperPOD reference architecture: network fabrics
Wed, Sep 30
AI infrastructure IV: RDMA
Why not TCP, registered memory and queue pairs, InfiniBand and RoCE, priority flow control (PFC), what breaks at 100,000 GPUs, OpenAI’s Multipath Reliable Connection (MRC)
Optional reading: (1) RDMA lecture slides (Radhika Mittal, UIUC) (2) RDMA over Ethernet for Distributed AI Training at Meta Scale (3) The Multipath Reliable Connection (MRC) Transport
Try it on your own
Train a small LLM from scratch
Not a class meeting; work through the notebook before Module 4 starts

Module 4 - Distributed training with data parallelism

Mon, Oct 5
Introduction to distributed training and data parallelism
OpenAI’s MRC in production, why the network matters for training, model FLOPs utilization (MFU), data parallelism with the AllReduce, overlapping the AllReduce with the backward pass, bucketing
Wed, Oct 7
Data parallelism with ZeRO and FSDP
The training state per parameter, ReduceScatter and AllGather, ZeRO stages 1 to 3, PyTorch FSDP
Mon, Oct 12
No Class (Fall Break)

Module 5 - Distributed training with tensor, pipeline, sequence and expert parallelism

Wed, Oct 14
Tensor parallelism
Splitting a matrix multiply by columns and by rows, tensor parallelism in the MLP and in attention (Megatron-LM), four AllReduces per layer, tensor parallelism versus ZeRO-3
Mon, Oct 19
Pipeline, sequence and context parallelism
Sequence parallelism, pipeline schedules (naive, all-forward-all-backward, 1F1B) and the pipeline bubble, context parallelism with ring attention and Ulysses
Wed, Oct 21
Expert parallelism and hybrid parallelism
Mixture-of-experts models, expert parallelism, putting the parallelisms together
Mon, Oct 26
Collective communication
All-reduce algorithms, NCCL

Module 6 - Machine learning compilers and guest lectures

Wed, Oct 28
Kernel fusion, FlashAttention and machine learning compilers
Optional reading: FlashAttention
Mon, Nov 2
Guest lecture: Abhinav Jangda (Microsoft), GPU kernels and machine learning compilers
Title and abstract to be announced
Wed, Nov 4
Guest lecture: Ben Klenk (NVIDIA), communication for machine learning
Title and abstract to be announced

Project mid-point presentations

Mon, Nov 9
Project mid-point presentations

Module 7 - Inference and post-training

Wed, Nov 11
LLM inference optimizations
KV caching, batching, speculative decoding
Mon, Nov 16
Post-training LLMs
Supervised finetuning, parameter-efficient finetuning, reinforcement learning from feedback

Module 8 - Agents, harnesses and world models

Wed, Nov 18
LLM agents
Mon, Nov 23
Agent harnesses
Tool use, sandboxes, evaluation of agentic systems
Wed, Nov 25
No Class (Thanksgiving Break)
Mon, Nov 30
Guest lecture: Vibhaalakshmi Sivaraman (World Labs), world models
Title and abstract to be announced

Module 9 - Wrap-up and final poster session

Wed, Dec 2
Course wrap-up
End-of-semester survey
Mon, Dec 7
Final project poster session
Last day of instruction