Tag: GPU Memory
-
PPO vs SAC: 1-GPU Memory & Compute Cost Benchmark
SAC uses 40% more VRAM than PPO on the same taskโbut reaches target rewards 34% faster. Real benchmarks on RTX 3090 with memory and compute trade-offs.
-
LLM Memory Calculator: Online Estimators Miss 40% Usage
Calculate LLM memory needs accurately. Why online tools fail at KV cache estimation and how to fix it with real GPU profiling methods.
-
Gradient Accumulation vs Large Batch: Memory & Cost Test
Compare gradient accumulation vs large batch training in real GPU memory testsโdiscover which method saves more VRAM and when to use each approach
-
Speculative Decoding: Why 2x Faster Inference Fails
Speculative decoding promises 2x faster LLM inference, but real-world gains often disappoint. Debug the hidden bottlenecks killing your speedup.