Not a class meeting; work through the notebook before Module 4 starts
Module 4 - Distributed training with data parallelism
Mon, Oct 5
Introduction to distributed training and data parallelism
OpenAI’s MRC in production, why the network matters for training, model FLOPs utilization (MFU), data parallelism with the AllReduce, overlapping the AllReduce with the backward pass, bucketing
Wed, Oct 7
Data parallelism with ZeRO and FSDP
The training state per parameter, ReduceScatter and AllGather, ZeRO stages 1 to 3, PyTorch FSDP
Mon, Oct 12
No Class (Fall Break)
Module 5 - Distributed training with tensor, pipeline, sequence and expert parallelism
Wed, Oct 14
Tensor parallelism
Splitting a matrix multiply by columns and by rows, tensor parallelism in the MLP and in attention (Megatron-LM), four AllReduces per layer, tensor parallelism versus ZeRO-3
Mon, Oct 19
Pipeline, sequence and context parallelism
Sequence parallelism, pipeline schedules (naive, all-forward-all-backward, 1F1B) and the pipeline bubble, context parallelism with ring attention and Ulysses
Wed, Oct 21
Expert parallelism and hybrid parallelism
Mixture-of-experts models, expert parallelism, putting the parallelisms together
Mon, Oct 26
Collective communication
All-reduce algorithms, NCCL
Module 6 - Machine learning compilers and guest lectures
Wed, Oct 28
Kernel fusion, FlashAttention and machine learning compilers