04Learning
Q-Learning
Reinforcement Learning
Bootstrap value estimates from delayed reward.
Applications
- Robotic control
- Game-playing agents
- Adaptive systems
Reasoning Trace
Q-Learning on 4×4 grid. Goal ★ = +1, trap × = −1. α=0.4, γ=0.9, ε=0.6. Agent learns by trial-and-error.
Principle in Pseudocode
- 1Q(s,a) ← 0
- 2for each episode:
- 3 s ← start
- 4 while not terminal:
- 5 a ← ε-greedy(Q, s)
- 6 observe r, s'
- 7 Q(s,a) ← Q(s,a) + α[r + γ·max Q(s',·) − Q(s,a)]
- 8 s ← s'
Reference implementation · pseudo
// Tabular Q-Learning
// See pseudocode →