-
e6be33aa53
doc: add SFT length-filter rationale — high PPL + high variance from A.3 IFD figure
master
ViperEkura
2026-07-24 06:37:35 +08:00
-
0a1d0573ae
doc: add SFT length filter (15-token floor) with IFD bias rationale
ViperEkura
2026-07-24 06:27:10 +08:00
-
0c2bc916f2
fix: loss_compare caption token count (20B -> 5B)
ViperEkura
2026-07-24 06:21:52 +08:00
-
4e70e827ff
remove: drop ckpt_weight_density_per_run.png figure
ViperEkura
2026-07-24 06:20:08 +08:00
-
a7bbc7b29f
fix: align SFT/DPO figures and text with actual training data
ViperEkura
2026-07-24 06:15:42 +08:00
-
4b37f289c0
fix: attribute muon-25bt larger std to Muon characteristic, not training duration alone
ViperEkura
2026-07-20 11:59:09 +08:00
-
a83555f326
fix: align conclusion, abstract, introduction with body data and full paper scope
ViperEkura
2026-07-20 11:51:48 +08:00
-
d93ff48320
docs: fix storage backends (2->3) and per-field length filter
ViperEkura
2026-07-20 11:27:39 +08:00
-
da0f536526
Update DPO iteration count: 1,500 -> 3,000
ViperEkura
2026-07-20 03:59:55 +08:00
-
c775a2b3e0
Update DPO metrics figure
ViperEkura
2026-07-20 03:59:04 +08:00
-
3f0ff911a8
Fix lingering cosine->WSD in Conclusion; update AGENTS.md with checkpoint labels and scheduler rules
ViperEkura
2026-07-19 15:30:49 +08:00
-
e149997200
Add DPO data generation & training; fix checkpoint labels; switch to WSD scheduler; update optimizer ablations in abstract/conclusion
ViperEkura
2026-07-19 15:25:42 +08:00
-
6c2e04a86f
Add DPO training subsection with dpo_metrics figure and Rafailov et al. citation
ViperEkura
2026-07-19 15:01:23 +08:00
-
411354eeb1
Update sft_metrics.png
ViperEkura
2026-07-18 00:52:52 +08:00
-
a0c39601a0
Add effective rank analysis, weight distribution evolution, Muon/SFT training figures
ViperEkura
2026-07-17 23:05:21 +08:00
-
bfc8ff6098
Fix training budget (19k steps / ~20B tokens), add Muon optimizer, correct variance scaling math, move figures to §Training Config, tighten abstract
ViperEkura
2026-07-08 22:49:43 +08:00
-
4a143b056d
Tighten abstract, add Muon optimizer mention, update ckpt_comparison figure
ViperEkura
2026-07-08 22:27:33 +08:00
-
5726bf6c84
Add optimizer/initialization comparison figure (ckpt_comparison)
ViperEkura
2026-07-08 21:59:56 +08:00
-
ff13a45d95
Restructure data pipeline: replace IFD analysis with length bias appendix
ViperEkura
2026-07-05 18:59:29 +08:00
-
2d52a00149
Update ifd_loss_ratio_density.png
ViperEkura
2026-07-05 12:38:43 +08:00
-
f2c2b159ba
Add IFD loss ratio density analysis, weight distribution appendix, and rewrite abstract
ViperEkura
2026-07-05 12:31:40 +08:00
-
459523f0e4
Fix IFD citation: correct author list and venue (NAACL 2024, not NeurIPS)
ViperEkura
2026-07-05 00:24:43 +08:00
-
4a7bf7ac3f
Add GQA/SwiGLU/RoPE formula blocks, Alpaca-GPT4 citation, and improve IFD-Loss Ratio interpretation
ViperEkura
2026-07-05 00:17:09 +08:00
-
5c915514b3
Add correlation analysis interpretation to IFD bias section
ViperEkura
2026-07-04 23:45:55 +08:00
-
bc8d4066c1
Add IFD analysis: density/length-bias figures and appendix
ViperEkura
2026-07-04 23:33:12 +08:00
-
9534316bac
Fix duplicate paragraph and misleading table caption
ViperEkura
2026-07-04 22:39:16 +08:00
-
b40fcc86c7
Refine abstract, switch to Times font, remove date, fix overfull hboxes with microtype
ViperEkura
2026-07-04 22:31:45 +08:00
-
dd412b3c9c
Add IFD analysis, appendix, and improve abstract
ViperEkura
2026-07-04 22:25:18 +08:00
-
73cb02f0aa
fix training hyperparams to match pretraining script: lr 1.5e-4, wd 0.1, warmup 0.02
ViperEkura
2026-06-28 16:57:30 +08:00
-
adc284920a
add Alembic SFT data cleaning section with MinHash algorithm and fix bibliography order
ViperEkura
2026-06-28 16:52:11 +08:00
-
13089b002a
fix training config: grad_accum 32, remove unused validation & activation ckpt
ViperEkura
2026-06-28 15:35:15 +08:00
-
d3fbe0ecc6
refactor: numerical stability analysis with residual scaling comparison
ViperEkura
2026-06-28 14:36:12 +08:00
-
3db1096c78
fix: add caption package, fix overfull hbox, correct RoFormer author names
ViperEkura
2026-06-25 16:53:26 +08:00
-
6f9278b35b
Initial commit: BF16 weight lock-in paper for 1.2B Transformer training with AstrAI
ViperEkura
2026-06-25 15:51:05 +08:00