Transformers & Attention
Self-attention, transformer architecture, BERT, GPT foundations.
Course Duration: 12 hours
What You'll Learn
- Attention mechanisms
- Self-attention
- Multi-head attention
- Transformer architecture
- Positional encoding
- BERT and GPT foundations
Prerequisites
- RNN & LSTM
- Strong Python skills
Course Modules
- Attention Mechanism
- Self-Attention
- Scaled Dot-Product Attention
- Multi-Head Attention
- Positional Encoding
- Transformer Encoder
- Transformer Decoder
- "Attention Is All You Need" Paper
- BERT Architecture
- GPT Architecture
- Vision Transformers (ViT)
- Implementing Transformer from Scratch
Key Concepts
- Query, Key, Value
- Attention weights
- Encoder-decoder architecture
- Pre-training and fine-tuning