- FSDP params are DTensors sharded across ranks - torch.nn.utils.clip_grad_norm_ computes LOCAL norm only - Each rank would clip by a different factor, causing gradient divergence - Fix: compute local norm, all-reduce squared sum, sqrt for global norm - Default reshard_after_forward=False (forward then backward makes reshard redundant) - Reduces per-step time by ~19% (1033ms to 839ms on 2xL20)