Q-learning on a five-tile corridor
The agent starts on tile 1 knowing nothing. Right pays nothing until tile 5 (+10); Left jumps straight to tile 1 for +2. Q-learning learns, by trial and error, which action is preferable in each state.
On tile 5, Right hits the wall; on tile 1, Left stays put. Both count as re-entry (+10, +2) unless unticked in Settings.
Q(s,a) ← Q(s,a) + α [ r + γ maxa′ Q(s′,a′) − Q(s,a) ]
Keyboard: Space for next stage, S to play one full step, P to play or pause.
Q-table
Current tileLast updated Q(s,a)
This step
- Choose an action
- Move
- Collect the reward
- Update the Q-value
Settings
Changing α, γ, ε or the wall rule takes effect from the next step. Changing the seed or episode length resets the run.