AI-2 5DV181 Course | Prepares for Assignment 2: MAB and Pong Q-Learning
Date/Time:
11 - December - 2025
The main goal is to transition from theoretical understanding (lectures) to practical implementation (assignment code) by experimenting with core RL algorithms in Python.
Ensure you have reviewed the necessary lecture material *before* the session begins.
Lecture Notes (Required):
Download: Making Complex Decisions (PDF)Follow these steps sequentially. Work in pairs and discuss the conceptual questions in the notebooks before running the code.
Review the definition of the Bellman Optimality Equation (Value Iteration) and the TD Update Rule (Q-Learning). Confirm that you understand the terms γ (discount factor) and α (learning rate).
Focus on the Exploration vs. Exploitation trade-off:
Focus on the TD update and Off-Policy nature of Q-Learning:
Ensure you have the required files for Assignment 2 checked out locally from GitLab and that your Python environment (virtual environment recommended) is correctly installed.
Download Assignment Code (GitLab)There is no submission required for this lab session itself. The material practiced here directly prepares you for the submission of Assignment 2 (MyBandit.py and Agent.py) on January 9, 2026.