- Add SVD effective rank (ER@99%) and condition number analysis across
three checkpoints (kami-15bt, norm-15bt, muon-25bt) in appendix
- Add post-training weight distribution analysis as §5.4, verifying
residual scaling persists throughout training
- Add per-component weight std table in appendix showing Muon constrains
early-stage drift vs Normal init
- Add Muon optimizer training dynamics figure (muon_pt.png)
- Add SFT training metrics figure (sft_metrics.png) on mixed CN-EN data
- Restructure: move effective rank to appendix, weight distribution to
§5.4, delete standalone §6, tighten abstract to prose-only
- Update conclusion to reference new analyses