DQN v1.0: experience replay and fixed targets
ℒ(wi) = 1BΣi ∈ batch[ ri + γ maxa Q̂(si+1, a; w−) − Q̂(si, ai; wi) ]2
wi ← wi − η ∇wi ℒ w− ← wi every k steps
ReadyPress “Next ▶” to follow one transition through DQN with experience replay and a target network.
Stage – of 9
Keyboard: Space or → for next stage, ← for previous, S to play one full step, P to play or pause.