Q-learning on a five-tile corridor

The agent starts on tile 1 knowing nothing. Right pays nothing until tile 5 (+10); Left jumps straight to tile 1 for +2. Q-learning learns, by trial and error, which action is preferable in each state.

On tile 5, Right hits the wall; on tile 1, Left stays put. Both count as re-entry (+10, +2) unless unticked in Settings.

Q(s,a) ← Q(s,a) + α [ r + γ maxa′ Q(s′,a′) − Q(s,a) ]

Keyboard: Space for next stage, S to play one full step, P to play or pause.

Q-table

Current tileLast updated Q(s,a)

This step
  1. Choose an action
  2. Move
  3. Collect the reward
  4. Update the Q-value
Settings

Changing α, γ, ε or the wall rule takes effect from the next step. Changing the seed or episode length resets the run.