DQN v1.0: experience replay and fixed targets

ℒ(wi) = 1BΣi ∈ batch[ ri + γ maxa Q̂(si+1, a; w−) − Q̂(si, ai; wi) ]2
wi ← wi − η ∇wi ℒ    w− ← wi  every k steps
DQN v1.0 ENV State (st) Action (at) Reward (rt) Next State (st+1) π action (at) ℒ mini- batch si​ si+1​ ri​, ai​ update w wi​ w−​ Copy every k timesteps max Q̂(si+1​, a; w−) ② Q̂(si​, ai​; wi​) ① at​ = ε-greedy argmaxa Q̂(st​, a; wi​)
ℒ = 1BΣi[ ri + γ max Q̂(si+1, a; w−) ② − Q̂(si, ai; wi) ① ]2
ReadyPress “Next ▶” to follow one transition through DQN with experience replay and a target network.
Stage – of 9

Keyboard: Space or → for next stage, ← for previous, S to play one full step, P to play or pause.

Settings