Compare commits

...
2 Commits
Author SHA1 Message Date
ViperEkura 4b37f289c0 fix: attribute muon-25bt larger std to Muon characteristic, not training duration alone
The muon-25bt (Muon, 25B) vs kami/norm (AdamW, 15B) comparison confounds
optimizer choice with training duration. Reframe as "Muon produces larger
post-convergence weight variance" rather than "training duration drives
variance growth."
2026-07-20 11:59:09 +08:00
ViperEkura a83555f326 fix: align conclusion, abstract, introduction with body data and full paper scope
- Fix sigma=0.003 (init value) to consistently narrower than non-scaled
  (body 4.4, conclusion, abstract)
- Add SFT/DPO to conclusion, abstract, and introduction
- Fix checkpoint label: kami-15bt (Muon) -> (AdamW, GPT-2 residual scaling)
  in effective rank appendix
- Fix table weight_std caption: Muon -> residual scaling constrains drift
- Fix 4.2 math: 0.689/L -> 0.689/(2L)
- Fix length filter: per-sample -> per-text-field skip
2026-07-20 11:51:48 +08:00
+50 -44
View File
@@ -29,29 +29,26 @@
\begin{abstract} \begin{abstract}
We present {\sc AstrAI}, an open-source framework for end-to-end training We present {\sc AstrAI}, an open-source framework for end-to-end training
of a 1.2B-parameter Transformer on $\sim$20B tokens. The pipeline covers of a 1.2B-parameter Transformer on $\sim$25B tokens. The pipeline covers
JSON-driven BBPE preprocessing with multi-strategy packing, tiered JSON-driven BBPE preprocessing with multi-strategy packing, tiered
storage backends (in-memory HDF5, memory-mapped binary, and lazy JSONL), and a companion SFT pipeline ({\sc Alembic}) with MinHash storage backends, and a companion SFT pipeline ({\sc Alembic}) with
deduplication and LLM-as-Judge scoring. The 24-layer decoder uses GQA, MinHash deduplication. The 24-layer GQA-SwiGLU decoder is trained with a
SwiGLU, RoPE, and RMSNorm, trained with a hybrid Muon/AdamW optimizer and hybrid Muon/AdamW optimizer and WSD scheduling under DDP/FSDP.
WSD (Warmup--Stable--Decay) scheduling under DDP/FSDP. A BF16 stability analysis shows that Supervised fine-tuning on deduplicated bilingual instructions reduces loss
GPT-2 residual scaling substantially reduces per-block residual variance from $\sim$2.5 to $\sim$1.6 over 1{,}000~steps; DPO alignment
accumulation, keeping post-training variance well below the overflow ($\beta=0.1$, cosine schedule) on model-generated preference pairs
threshold of standard initialization; empirically this yields a sustained converges stably without over-optimisation. A BF16 stability analysis
loss advantage over Kaiming initialization throughout training. shows that GPT-2 residual scaling ($\sigma_0 = 0.02/\sqrt{2L}$) reduces
Systematic ablations show that the hybrid Muon/AdamW optimizer per-block activation variance by a factor of 48, and post-training weight
outperforms pure AdamW on 2D weight matrices, and that GPT-2 residual analysis across three checkpoints confirms that residual-scaled
scaling yields a sustained loss advantage over both Kaiming and Normal projections remain consistently narrower than non-scaled weights across
initialization throughout training. Post-training weight distribution optimizers, initializations, and training budgets. Optimizer ablations
analysis across three checkpoints---varying optimizer, initialization, demonstrate that the hybrid Muon/AdamW outperforms pure AdamW, with
and training budget---confirms that residual-scaled projections 2D weight matrices benefiting from Muon's orthogonalisation and 1D
maintain narrow distributions throughout training, preserving the parameters from AdamW's second-moment adaptation. An SVD effective-rank
numerical stability established at initialization. An SVD-based analysis reveals near-capacity weight utilization, with Q/O projections
effective rank analysis further reveals that the model operates near showing consistently lower effective rank than K/V projections under
its representational capacity, with attention Q/O projections grouped query attention.
consistently showing lower utilization than K/V projections, a pattern
stable across all configurations and consistent with the low-rank
structure induced by grouped query attention.
\end{abstract} \end{abstract}
% ====================================================================== % ======================================================================
@@ -63,8 +60,10 @@ model architecture. Data must be preprocessed and stored efficiently, the
training loop must handle distributed parallelism, gradient accumulation, training loop must handle distributed parallelism, gradient accumulation,
checkpointing, and logging---and numerical pitfalls must be diagnosed and checkpointing, and logging---and numerical pitfalls must be diagnosed and
fixed. This paper describes the complete workflow using {\sc AstrAI}~\cite{astrai}, an fixed. This paper describes the complete workflow using {\sc AstrAI}~\cite{astrai}, an
open-source framework for Transformer training and inference. Beyond the open-source framework for Transformer training and inference, from JSONL
pipeline, we conduct systematic ablations on optimizers (hybrid Muon/AdamW ingestion through pretraining, supervised fine-tuning (SFT) on deduplicated
bilingual instructions, and direct preference optimization (DPO) alignment.
We also conduct systematic ablations on optimizers (hybrid Muon/AdamW
versus pure AdamW) and initializations (GPT-2 residual scaling, Kaiming, versus pure AdamW) and initializations (GPT-2 residual scaling, Kaiming,
and Normal), and analyse a BF16 precision issue encountered along the way. and Normal), and analyse a BF16 precision issue encountered along the way.
@@ -463,7 +462,7 @@ $1/\sqrt{2L}$:
\end{equation} \end{equation}
This reduces per-block residual variance contribution from $0.689$ to This reduces per-block residual variance contribution from $0.689$ to
$0.689/L \approx 0.014$, a factor of $2L = 48$. The post-24-block variance $0.689/(2L) \approx 0.014$, a factor of $2L = 48$. The post-24-block variance
drops from $17.5$ to $1.34$, a $13.1\times$ improvement. In BF16 drops from $17.5$ to $1.34$, a $13.1\times$ improvement. In BF16
($7$-bit mantissa, ULP $= 0.0078$ at $w = 1.0$)~\cite{ieee754}, ($7$-bit mantissa, ULP $= 0.0078$ at $w = 1.0$)~\cite{ieee754},
this keeps weight magnitudes within stable precision bounds. We further this keeps weight magnitudes within stable precision bounds. We further
@@ -521,27 +520,25 @@ are visible in the per-component weight std
(Table~\ref{tab:weight_std}, Appendix~\ref{app:weight_std}): (Table~\ref{tab:weight_std}, Appendix~\ref{app:weight_std}):
\begin{itemize}[nosep] \begin{itemize}[nosep]
\item \textbf{Training duration drives variance growth}: the \item \textbf{Muon produces larger post-convergence weight
\texttt{muon-25bt} checkpoint exhibits the largest weight std variance}: the \texttt{muon-25bt} checkpoint (Muon, 25B tokens)
($\sim$0.024 for attention projections), exceeding both 15B exhibits the largest weight std ($\sim$0.024), exceeding both
checkpoints ($\sim$0.015--0.021), reflecting continued weight 15B AdamW checkpoints ($\sim$0.015--0.021), consistent with
drift from $\sigma_0 = 0.02$ as training progresses. Muon allowing wider parameter distributions after convergence.
\item \textbf{Residual scaling constrains early-stage drift}: at \item \textbf{Residual scaling constrains early-stage drift}: at
equal training budget (15B tokens) and with the same AdamW equal training budget (15B tokens) and with the same AdamW
optimizer, the GPT-2-scaled \texttt{kami-15bt} shows optimizer, the GPT-2-scaled \texttt{kami-15bt} shows
\emph{smaller} std ($\sim$0.015) than the Normal-init \emph{smaller} std ($\sim$0.015) than the Normal-init
\texttt{norm-15bt} ($\sim$0.021), indicating that residual \texttt{norm-15bt} ($\sim$0.021), indicating that residual
scaling itself limits weight drift beyond its initialization scaling itself limits weight drift beyond its initialization
effect. The Muon-trained \texttt{muon-25bt} (25B tokens) effect.
exhibits the largest std because cumulative training steps
eventually overtake both initialization and per-step optimizer
constraints.
\end{itemize} \end{itemize}
Critically, the residual-scaled projections remain bounded at Critically, the residual-scaled projections maintain consistently
$\sigma \approx 0.003$ across all checkpoints regardless of training lower standard deviations than their non-scaled counterparts across
duration, confirming that the $1/\sqrt{2L}$ scaling continues to all three checkpoints (Table~\ref{tab:weight_std}), confirming that
enforce its design constraint throughout training. the $1/\sqrt{2L}$ scaling continues to enforce its design constraint
throughout training.
\begin{figure}[H] \begin{figure}[H]
\centering \centering
@@ -574,8 +571,16 @@ scaling ($\sigma_o = 0.02/\sqrt{2L}$) reduces per-block residual variance
by a factor of 48, keeping post-24-layer variance at $1.34$ versus $17.5$ by a factor of 48, keeping post-24-layer variance at $1.34$ versus $17.5$
without scaling. Post-training weight distribution analysis confirms that without scaling. Post-training weight distribution analysis confirms that
this scaling constraint persists throughout training, with residual-scaled this scaling constraint persists throughout training, with residual-scaled
projections maintaining narrow distributions ($\sigma \approx 0.003$) projections maintaining consistently lower standard deviations than
regardless of optimizer or training duration. non-scaled weights across all checkpoints
(Table~\ref{tab:weight_std}).
Supervised fine-tuning on deduplicated bilingual instructions (processed
by the companion {\sc Alembic} pipeline with MinHash deduplication)
reduces training loss from $\sim$2.5 to $\sim$1.6 over 1{,}000~WSD-scheduled
steps. Subsequent DPO alignment on preference pairs ($\beta=0.1$, cosine
schedule) converges stably without over-optimisation, reaching a minimum
preference loss near step~1{,}200.
Optimizer ablations (Figure~\ref{fig:ckpt_comparison}) demonstrate that Optimizer ablations (Figure~\ref{fig:ckpt_comparison}) demonstrate that
the hybrid Muon/AdamW configuration consistently outperforms pure AdamW the hybrid Muon/AdamW configuration consistently outperforms pure AdamW
@@ -773,7 +778,7 @@ selection signal without re-evaluating after fine-tuning.
To assess how well the trained parameters utilize their allocated To assess how well the trained parameters utilize their allocated
capacity, we perform an SVD-based effective rank analysis on three capacity, we perform an SVD-based effective rank analysis on three
checkpoints: \texttt{kami-15bt} (Muon, 15B tokens), \texttt{norm-15bt} checkpoints: \texttt{kami-15bt} (AdamW, GPT-2 residual scaling, 15B tokens), \texttt{norm-15bt}
(Normal init, 15B tokens), and \texttt{muon-25bt} (Muon, 25B tokens). (Normal init, 15B tokens), and \texttt{muon-25bt} (Muon, 25B tokens).
For each 2D weight matrix $\mathbf{W} \in \mathbb{R}^{m\times n}$ with For each 2D weight matrix $\mathbf{W} \in \mathbb{R}^{m\times n}$ with
SVD $\mathbf{W} = \mathbf{U}\boldsymbol{\Sigma}\mathbf{V}^{\mkern-1mu\mathsf{T}}$, SVD $\mathbf{W} = \mathbf{U}\boldsymbol{\Sigma}\mathbf{V}^{\mkern-1mu\mathsf{T}}$,
@@ -839,9 +844,10 @@ distribution analysis in Section~\ref{sec:num-stability}.
\begin{table}[H] \begin{table}[H]
\centering \centering
\caption{Weight std by component across checkpoints. Non-scaled \caption{Weight std by component across checkpoints. Non-scaled weights
weights broaden with training duration; Muon constrains drift at broaden with training; residual scaling constrains drift at equal
equal token count (15B) but is overtaken by longer training (25B). token count (15B). Muon produces larger post-convergence weight
variance than AdamW.
Residual-scaled projections ($\mathbf{W}_o$, $\mathbf{W}_{\text{down}}$) Residual-scaled projections ($\mathbf{W}_o$, $\mathbf{W}_{\text{down}}$)
remain bounded.} remain bounded.}
\label{tab:weight_std} \label{tab:weight_std}