Skip to main content

Transformers & Attention

Self-attention, transformer architecture, BERT, GPT foundations.

Course Duration: 12 hours

What You'll Learn

  • Attention mechanisms
  • Self-attention
  • Multi-head attention
  • Transformer architecture
  • Positional encoding
  • BERT and GPT foundations

Prerequisites

  • RNN & LSTM
  • Strong Python skills

Course Modules

  1. Attention Mechanism
  2. Self-Attention
  3. Scaled Dot-Product Attention
  4. Multi-Head Attention
  5. Positional Encoding
  6. Transformer Encoder
  7. Transformer Decoder
  8. "Attention Is All You Need" Paper
  9. BERT Architecture
  10. GPT Architecture
  11. Vision Transformers (ViT)
  12. Implementing Transformer from Scratch

Key Concepts

  • Query, Key, Value
  • Attention weights
  • Encoder-decoder architecture
  • Pre-training and fine-tuning