Tag: DQN
-
Off-Policy RL Replay Buffer Memory Leak: Fix 2M Step Crash
Fix off-policy RL replay buffer memory leak causing 2M step crashes. Learn circular buffer implementation and memory-efficient sampling patterns.
-
DQN vs Rainbow: 4.8x Score Gain From 6 Extensions
Compare DQN and Rainbow's 6 RL extensions that achieve 4.8x higher Atari scores. See how prioritized replay and dueling nets stack up.
-
DQN vs Double DQN vs Dueling DQN: Atari Breakout Benchmark
Compare DQN, Double DQN, and Dueling DQN performance on Atari Breakout. Which architecture solves overestimation and learns faster? Benchmark results inside.
-
PPO vs DQN: Discrete Action Spaces Beat Continuous 3x
Compare PPO and DQN performance on discrete vs continuous control tasks. Surprising speed differences revealed through benchmark experiments.
-
DQN Overestimation Bias: 3 Double-Q Fixes That Work
DQN agents plateau at 60% optimal? Overestimation bias is why. Compare 3 Double-Q fixes on LunarLander with real training curves.
-
DQN vs PPO vs SAC: MuJoCo Training Speed Benchmarks
DQN fails on continuous control. SAC beats PPO 2-3x in sample efficiency but costs 20% more wall-clock time. Real benchmarks on HalfCheetah, Hopper, Ant.
-
SimpleRL: Building DQN to PPO from Scratch in 500 Lines
Build DQN and PPO reinforcement learning algorithms from scratch in under 500 lines. Step-by-step implementation guide with minimal dependencies.
-
From Q-Learning to DQN: Your First RL Algorithms
Implement Q-Learning from scratch, then scale up to Deep Q-Networks with experience replay and target networks. Complete PyTorch code included.