Fix training budget (19k steps / ~20B tokens), add Muon optimizer, correct variance scaling math, move figures to §Training Config, tighten abstract
This commit is contained in:
Binary file not shown.
|
Before Width: | Height: | Size: 273 KiB After Width: | Height: | Size: 267 KiB |
Reference in New Issue
Block a user