Tag: Reinforcement Learning
-
High Discount Factor ฮณ=0.99 Causes Divergence: Fix Guide
Fix high discount factor ฮณ=0.99 divergence in RL with gradient clipping, reward scaling, and target networks. Practical debugging solutions.
-
Gymnasium render_mode=’human’ Crashes Training: 3 Fixes
Gymnasium render_mode='human' crashes your RL training? Discover 3 proven fixes for headless servers and stable visualization workflows.
-
DDPG from Scratch: 400-Line PyTorch Implementation
Build a DDPG agent from scratch in 400 lines of PyTorch. Learn continuous action RL with policy gradients, replay buffers, and target networks.
-
PPO vs SAC: 1-GPU Memory & Compute Cost Benchmark
SAC uses 40% more VRAM than PPO on the same taskโbut reaches target rewards 34% faster. Real benchmarks on RTX 3090 with memory and compute trade-offs.
-
Stable Baselines3 VecEnv Reset Bug: 100K Step Desync Fix
Fix the Stable Baselines3 VecEnv reset bug causing 100K step desyncs. Learn why auto_reset breaks training and how to solve it properly.
-
Gymnasium Custom Env Step() Returns Invalid Shape: 5 Fixes
Fix Gymnasium custom env step() shape errors with 5 proven solutions. Learn proper observation space handling and avoid common pitfalls.
-
PPO vs SAC Sparse Rewards: 3x Sample Efficiency Gap
PPO vs SAC on sparse rewards: which RL algorithm learns faster? Benchmark shows 3x sample efficiency gap. Compare training curves and understand why.
-
DQN vs Rainbow: 4.8x Score Gain From 6 Extensions
Compare DQN and Rainbow's 6 RL extensions that achieve 4.8x higher Atari scores. See how prioritized replay and dueling nets stack up.