📄️ Transformers and attention
Build a Transformer by hand in PyTorch, block by block, from single-head attention to a working encoder-decoder that translates dates. Duration: 8h, applied projects, a 40-question exam and a verifiable certificate.
📄️ 1. Recurrent limits
Module 1 of the Transformers premium course: why sequentiality and the context vector bottleneck stopped recurrent networks from scaling, and how attention lifted both.
📄️ 2. Query, key, value
Module 2 of the Transformers premium course: compute scaled dot-product attention by hand on three tokens, then implement the first block of our red-thread Transformer in PyTorch.
📄️ 3. Multi-head attention
Module 3 of the Transformers premium course: split attention across several heads, understand what each head learns, and count the parameters correctly.
📄️ 4. Positional encoding
Module 4 of the Transformers premium course: why attention is order-blind, how sinusoidal, learned and rotary positional encodings inject position back in, and what length extrapolation costs.
📄️ 5. Residuals and LayerNorm
Module 5 of the Transformers premium course: how residuals and LayerNorm keep gradients alive, why pre-norm won over post-norm, and how the feed-forward block completes the Transformer layer.
📄️ 6. Encoder and BERT
Module 6 of the Transformers premium course: stack encoder layers into BERT, understand masked language modelling, and fine-tune for classification and extraction.
📄️ 7. Decoder and GPT
Module 7 of the Transformers premium course: build the decoder with a causal mask, understand greedy versus sampling decoding, and use a KV cache to make inference tractable.
📄️ 8. Encoder-decoder and T5
Module 8 of the Transformers premium course: connect an encoder and a decoder with cross-attention, understand T5's text-to-text framing, and choose the right family for a task.
📄️ 9. Quadratic cost
Module 9 of the Transformers premium course: measure the O(n squared) memory of attention, then survey FlashAttention, sparse and windowed attention, and understand what long context really costs.
📄️ 10. Implementing the block
Module 10 of the Transformers premium course: assemble every part from modules 2 to 9 into a working PyTorch Transformer, train it on date translation, and read its attention maps.
📄️ Recap and exam
Complete recap of the Transformers premium course: recurrent limits, attention step by step, multi-head, positional encoding, residuals, encoder, decoder, encoder-decoder, quadratic cost, then the 40-question exam.