Tag: Continuous Control
-
DDPG from Scratch: 400-Line PyTorch Implementation
Build a DDPG agent from scratch in 400 lines of PyTorch. Learn continuous action RL with policy gradients, replay buffers, and target networks.
-
PPO vs SAC Sparse Rewards: 3x Sample Efficiency Gap
PPO vs SAC on sparse rewards: which RL algorithm learns faster? Benchmark shows 3x sample efficiency gap. Compare training curves and understand why.
-
PPO vs SAC: Real Robot Benchmark on 3 Manipulation Tasks
Compare PPO vs SAC on real robot manipulation tasks. Performance metrics, training stability, and practical deployment insights revealed.