Tag: training divergence
-
Transformer NaN Loss: 7 Fixes That Actually Work
Transformer NaN loss? Don't lower learning rate first. Fix attention overflow, fp16 underflow, and LayerNorm edge cases with these 7 patches.
Transformer NaN loss? Don't lower learning rate first. Fix attention overflow, fp16 underflow, and LayerNorm edge cases with these 7 patches.