Reinforcement Learning Lab Session

AI-2 5DV181 Course | Prepares for Assignment 2: MAB and Pong Q-Learning

RL Icon

Logistics

Date/Time:

11 - December - 2025

Location:

Hörsal UB.A.240 - Lindellhallen 4

View Map

Lab Objectives

The main goal is to transition from theoretical understanding (lectures) to practical implementation (assignment code) by experimenting with core RL algorithms in Python.

Required Pre-Reading

Ensure you have reviewed the necessary lecture material *before* the session begins.

Lecture Notes (Required):

Download: Making Complex Decisions (PDF)

Laboratory Procedure

Follow these steps sequentially. Work in pairs and discuss the conceptual questions in the notebooks before running the code.

Step 1: Check Theoretical Foundation (15 min)

Review the definition of the Bellman Optimality Equation (Value Iteration) and the TD Update Rule (Q-Learning). Confirm that you understand the terms γ (discount factor) and α (learning rate).

Step 2: Multi-Armed Bandits (MAB) Practice (45 min)

Focus on the Exploration vs. Exploitation trade-off:

  • Simple Base: Run the Car Engine Selection to see ε-Greedy (Exploration) vs. Purely Greedy (Exploitation) performance.
  • Advanced: Run the Headline Optimization (UCB) and Ad Optimization (Gradient Bandit) notebooks. Compare how UCB's "bonus term" differs from the Gradient Bandit's "Softmax probability" approach to exploration.

Step 3: Q-Learning and Off-Policy Control (45 min)

Focus on the TD update and Off-Policy nature of Q-Learning:

  • Simple: Trace the Q-table updates in the Grid-world example. Pay attention to how a large penalty (-10) in the Danger Zone propagates backward through the states.
  • Control: Review the Tennis Game code. Identify where the behavior policy (ε-Greedy) chooses the action, and where the target policy ($\max$ over $Q$) is used in the update rule (demonstrating off-policy learning).

Step 4: Final Check & Setup

Ensure you have the required files for Assignment 2 checked out locally from GitLab and that your Python environment (virtual environment recommended) is correctly installed.

Download Assignment Code (GitLab)

Deliverable

There is no submission required for this lab session itself. The material practiced here directly prepares you for the submission of Assignment 2 (MyBandit.py and Agent.py) on January 9, 2026.