📄️ Reinforcement learning
How an agent learns to act by trial and reward, from tabular Q-learning on FrozenLake to PPO on LunarLander, with the pitfalls that break real deployments.
📄️ 1. Agent and environment
Module 1 of the Reinforcement Learning premium course: the interaction loop, discounted return, what separates RL from supervised learning, and the pitfalls of reward design.
📄️ 2. Markov decision processes
Module 2 of the Reinforcement Learning premium course: states, actions, transitions, policies, the Markov assumption and where it fails, with FrozenLake formalized end to end.
📄️ 3. Value functions and Bellman
Module 3 of the Reinforcement Learning premium course: V and Q, expectation and optimality equations, worked by hand on a 3x3 grid then formalized in code.
📄️ 4. Dynamic programming, Monte Carlo
Module 4 of the Reinforcement Learning premium course: value iteration and policy iteration when the model is known, Monte Carlo estimation when it is not, and the variance of returns.
📄️ 5. Temporal difference, Q-learning
Module 5 of the Reinforcement Learning premium course: TD update, SARSA versus Q-learning (on versus off-policy), tabular Q-learning on FrozenLake, learning rate and discount.
📄️ 6. Exploration
Module 6 of the Reinforcement Learning premium course: the exploration-exploitation dilemma, epsilon decay, softmax and UCB, with under-exploration shown as a failure mode.
📄️ 7. Deep Q-networks, replay
Module 7 of the Reinforcement Learning premium course: Q-approximation by a neural network, replay buffer, target network, instabilities, Double DQN, CartPole solved.
📄️ 8. Policy gradient
Module 8 of the Reinforcement Learning premium course: stochastic policy, REINFORCE, variance and baseline, continuous actions, CartPole solved with REINFORCE.
📄️ 9. Actor-critic, A2C, PPO
Module 9 of the Reinforcement Learning premium course: advantage, A2C, PPO clipped objective, hyperparameters that matter, homemade PPO compared with Stable-Baselines3 on LunarLander.
📄️ 10. Project: Gymnasium agent
Module 10 of the Reinforcement Learning premium course: LunarLander end to end, seeds and run-to-run variability, curves with intervals, video recording, what does not transfer to the real world.
📄️ Recap and exam
Complete recap of the Reinforcement Learning premium course: MDPs, Bellman, Q-learning, DQN, policy gradient, PPO, then the 40-question exam.