- Add SVD effective rank (ER@99%) and condition number analysis across three checkpoints (kami-15bt, norm-15bt, muon-25bt) in appendix - Add post-training weight distribution analysis as §5.4, verifying residual scaling persists throughout training - Add per-component weight std table in appendix showing Muon constrains early-stage drift vs Normal init - Add Muon optimizer training dynamics figure (muon_pt.png) - Add SFT training metrics figure (sft_metrics.png) on mixed CN-EN data - Restructure: move effective rank to appendix, weight distribution to §5.4, delete standalone §6, tighten abstract to prose-only - Update conclusion to reference new analyses
176 KiB
2700x750px
176 KiB
2700x750px