Tag: MuJoCo
-
High Discount Factor ฮณ=0.99 Causes Divergence: Fix Guide
Fix high discount factor ฮณ=0.99 divergence in RL with gradient clipping, reward scaling, and target networks. Practical debugging solutions.
-
CleanRL vs Stable Baselines3: PPO Training 2.3x Faster
Compare CleanRL vs Stable Baselines3 PPO implementations and discover why CleanRL achieves 2.3x faster training with cleaner, hackable code.
-
SAC Entropy Tuning: Auto-Alpha Cuts Failures by 80%
Learn SAC entropy tuning with auto-alpha to slash RL training failures by 80%. Discover the temperature coefficient trick that stabilizes learning.
-
PPO vs SAC vs TD3: MuJoCo Humanoid Training in 5M Steps
Compare PPO, SAC, and TD3 on MuJoCo Humanoid: sample efficiency, stability, and final performance revealed in a 5M-step RL benchmark showdown.
-
PPO Training Diverges After 1M Steps: Clipping & LR Fixes
PPO training collapse after 1M steps? Learn how gradient clipping and learning rate schedules prevent policy divergence in deep RL implementations.
-
DQN vs PPO vs SAC: MuJoCo Training Speed Benchmarks
DQN fails on continuous control. SAC beats PPO 2-3x in sample efficiency but costs 20% more wall-clock time. Real benchmarks on HalfCheetah, Hopper, Ant.