fix: align SFT/DPO figures and text with actual training data
- SFT: 1,000 steps/WSD → ~3,800 steps/cosine; loss ~2.5→1.6 → ~2.1→1.5 - DPO: fix preference-loss narrative to match training-loss curve; correct initial grad-norm (~50 → ~200) - Table 2: peak LR 1.5e-4 → 2.0e-4 (matches pt_metric.png) - Clarify scheduling: pretraining=WSD, SFT+DPO=cosine - Add disclaimer to ckpt_weight_density_per_run caption for legacy iter labels
This commit is contained in:
Binary file not shown.
|
Before Width: | Height: | Size: 92 KiB After Width: | Height: | Size: 86 KiB |
Binary file not shown.
|
Before Width: | Height: | Size: 85 KiB After Width: | Height: | Size: 84 KiB |
Reference in New Issue
Block a user