fix: align conclusion, abstract, introduction with body data and full paper scope
- Fix sigma=0.003 (init value) to consistently narrower than non-scaled (body 4.4, conclusion, abstract) - Add SFT/DPO to conclusion, abstract, and introduction - Fix checkpoint label: kami-15bt (Muon) -> (AdamW, GPT-2 residual scaling) in effective rank appendix - Fix table weight_std caption: Muon -> residual scaling constrains drift - Fix 4.2 math: 0.689/L -> 0.689/(2L) - Fix length filter: per-sample -> per-text-field skip
This commit is contained in:
@@ -29,29 +29,26 @@
|
|||||||
|
|
||||||
\begin{abstract}
|
\begin{abstract}
|
||||||
We present {\sc AstrAI}, an open-source framework for end-to-end training
|
We present {\sc AstrAI}, an open-source framework for end-to-end training
|
||||||
of a 1.2B-parameter Transformer on $\sim$20B tokens. The pipeline covers
|
of a 1.2B-parameter Transformer on $\sim$25B tokens. The pipeline covers
|
||||||
JSON-driven BBPE preprocessing with multi-strategy packing, tiered
|
JSON-driven BBPE preprocessing with multi-strategy packing, tiered
|
||||||
storage backends (in-memory HDF5, memory-mapped binary, and lazy JSONL), and a companion SFT pipeline ({\sc Alembic}) with MinHash
|
storage backends, and a companion SFT pipeline ({\sc Alembic}) with
|
||||||
deduplication and LLM-as-Judge scoring. The 24-layer decoder uses GQA,
|
MinHash deduplication. The 24-layer GQA-SwiGLU decoder is trained with a
|
||||||
SwiGLU, RoPE, and RMSNorm, trained with a hybrid Muon/AdamW optimizer and
|
hybrid Muon/AdamW optimizer and WSD scheduling under DDP/FSDP.
|
||||||
WSD (Warmup--Stable--Decay) scheduling under DDP/FSDP. A BF16 stability analysis shows that
|
Supervised fine-tuning on deduplicated bilingual instructions reduces loss
|
||||||
GPT-2 residual scaling substantially reduces per-block residual variance
|
from $\sim$2.5 to $\sim$1.6 over 1{,}000~steps; DPO alignment
|
||||||
accumulation, keeping post-training variance well below the overflow
|
($\beta=0.1$, cosine schedule) on model-generated preference pairs
|
||||||
threshold of standard initialization; empirically this yields a sustained
|
converges stably without over-optimisation. A BF16 stability analysis
|
||||||
loss advantage over Kaiming initialization throughout training.
|
shows that GPT-2 residual scaling ($\sigma_0 = 0.02/\sqrt{2L}$) reduces
|
||||||
Systematic ablations show that the hybrid Muon/AdamW optimizer
|
per-block activation variance by a factor of 48, and post-training weight
|
||||||
outperforms pure AdamW on 2D weight matrices, and that GPT-2 residual
|
analysis across three checkpoints confirms that residual-scaled
|
||||||
scaling yields a sustained loss advantage over both Kaiming and Normal
|
projections remain consistently narrower than non-scaled weights across
|
||||||
initialization throughout training. Post-training weight distribution
|
optimizers, initializations, and training budgets. Optimizer ablations
|
||||||
analysis across three checkpoints---varying optimizer, initialization,
|
demonstrate that the hybrid Muon/AdamW outperforms pure AdamW, with
|
||||||
and training budget---confirms that residual-scaled projections
|
2D weight matrices benefiting from Muon's orthogonalisation and 1D
|
||||||
maintain narrow distributions throughout training, preserving the
|
parameters from AdamW's second-moment adaptation. An SVD effective-rank
|
||||||
numerical stability established at initialization. An SVD-based
|
analysis reveals near-capacity weight utilization, with Q/O projections
|
||||||
effective rank analysis further reveals that the model operates near
|
showing consistently lower effective rank than K/V projections under
|
||||||
its representational capacity, with attention Q/O projections
|
grouped query attention.
|
||||||
consistently showing lower utilization than K/V projections, a pattern
|
|
||||||
stable across all configurations and consistent with the low-rank
|
|
||||||
structure induced by grouped query attention.
|
|
||||||
\end{abstract}
|
\end{abstract}
|
||||||
|
|
||||||
% ======================================================================
|
% ======================================================================
|
||||||
@@ -63,8 +60,10 @@ model architecture. Data must be preprocessed and stored efficiently, the
|
|||||||
training loop must handle distributed parallelism, gradient accumulation,
|
training loop must handle distributed parallelism, gradient accumulation,
|
||||||
checkpointing, and logging---and numerical pitfalls must be diagnosed and
|
checkpointing, and logging---and numerical pitfalls must be diagnosed and
|
||||||
fixed. This paper describes the complete workflow using {\sc AstrAI}~\cite{astrai}, an
|
fixed. This paper describes the complete workflow using {\sc AstrAI}~\cite{astrai}, an
|
||||||
open-source framework for Transformer training and inference. Beyond the
|
open-source framework for Transformer training and inference, from JSONL
|
||||||
pipeline, we conduct systematic ablations on optimizers (hybrid Muon/AdamW
|
ingestion through pretraining, supervised fine-tuning (SFT) on deduplicated
|
||||||
|
bilingual instructions, and direct preference optimization (DPO) alignment.
|
||||||
|
We also conduct systematic ablations on optimizers (hybrid Muon/AdamW
|
||||||
versus pure AdamW) and initializations (GPT-2 residual scaling, Kaiming,
|
versus pure AdamW) and initializations (GPT-2 residual scaling, Kaiming,
|
||||||
and Normal), and analyse a BF16 precision issue encountered along the way.
|
and Normal), and analyse a BF16 precision issue encountered along the way.
|
||||||
|
|
||||||
@@ -463,7 +462,7 @@ $1/\sqrt{2L}$:
|
|||||||
\end{equation}
|
\end{equation}
|
||||||
|
|
||||||
This reduces per-block residual variance contribution from $0.689$ to
|
This reduces per-block residual variance contribution from $0.689$ to
|
||||||
$0.689/L \approx 0.014$, a factor of $2L = 48$. The post-24-block variance
|
$0.689/(2L) \approx 0.014$, a factor of $2L = 48$. The post-24-block variance
|
||||||
drops from $17.5$ to $1.34$, a $13.1\times$ improvement. In BF16
|
drops from $17.5$ to $1.34$, a $13.1\times$ improvement. In BF16
|
||||||
($7$-bit mantissa, ULP $= 0.0078$ at $w = 1.0$)~\cite{ieee754},
|
($7$-bit mantissa, ULP $= 0.0078$ at $w = 1.0$)~\cite{ieee754},
|
||||||
this keeps weight magnitudes within stable precision bounds. We further
|
this keeps weight magnitudes within stable precision bounds. We further
|
||||||
@@ -538,10 +537,11 @@ are visible in the per-component weight std
|
|||||||
constraints.
|
constraints.
|
||||||
\end{itemize}
|
\end{itemize}
|
||||||
|
|
||||||
Critically, the residual-scaled projections remain bounded at
|
Critically, the residual-scaled projections maintain consistently
|
||||||
$\sigma \approx 0.003$ across all checkpoints regardless of training
|
lower standard deviations than their non-scaled counterparts across
|
||||||
duration, confirming that the $1/\sqrt{2L}$ scaling continues to
|
all three checkpoints (Table~\ref{tab:weight_std}), confirming that
|
||||||
enforce its design constraint throughout training.
|
the $1/\sqrt{2L}$ scaling continues to enforce its design constraint
|
||||||
|
throughout training.
|
||||||
|
|
||||||
\begin{figure}[H]
|
\begin{figure}[H]
|
||||||
\centering
|
\centering
|
||||||
@@ -574,8 +574,16 @@ scaling ($\sigma_o = 0.02/\sqrt{2L}$) reduces per-block residual variance
|
|||||||
by a factor of 48, keeping post-24-layer variance at $1.34$ versus $17.5$
|
by a factor of 48, keeping post-24-layer variance at $1.34$ versus $17.5$
|
||||||
without scaling. Post-training weight distribution analysis confirms that
|
without scaling. Post-training weight distribution analysis confirms that
|
||||||
this scaling constraint persists throughout training, with residual-scaled
|
this scaling constraint persists throughout training, with residual-scaled
|
||||||
projections maintaining narrow distributions ($\sigma \approx 0.003$)
|
projections maintaining consistently lower standard deviations than
|
||||||
regardless of optimizer or training duration.
|
non-scaled weights across all checkpoints
|
||||||
|
(Table~\ref{tab:weight_std}).
|
||||||
|
|
||||||
|
Supervised fine-tuning on deduplicated bilingual instructions (processed
|
||||||
|
by the companion {\sc Alembic} pipeline with MinHash deduplication)
|
||||||
|
reduces training loss from $\sim$2.5 to $\sim$1.6 over 1{,}000~WSD-scheduled
|
||||||
|
steps. Subsequent DPO alignment on preference pairs ($\beta=0.1$, cosine
|
||||||
|
schedule) converges stably without over-optimisation, reaching a minimum
|
||||||
|
preference loss near step~1{,}200.
|
||||||
|
|
||||||
Optimizer ablations (Figure~\ref{fig:ckpt_comparison}) demonstrate that
|
Optimizer ablations (Figure~\ref{fig:ckpt_comparison}) demonstrate that
|
||||||
the hybrid Muon/AdamW configuration consistently outperforms pure AdamW
|
the hybrid Muon/AdamW configuration consistently outperforms pure AdamW
|
||||||
@@ -773,7 +781,7 @@ selection signal without re-evaluating after fine-tuning.
|
|||||||
|
|
||||||
To assess how well the trained parameters utilize their allocated
|
To assess how well the trained parameters utilize their allocated
|
||||||
capacity, we perform an SVD-based effective rank analysis on three
|
capacity, we perform an SVD-based effective rank analysis on three
|
||||||
checkpoints: \texttt{kami-15bt} (Muon, 15B tokens), \texttt{norm-15bt}
|
checkpoints: \texttt{kami-15bt} (AdamW, GPT-2 residual scaling, 15B tokens), \texttt{norm-15bt}
|
||||||
(Normal init, 15B tokens), and \texttt{muon-25bt} (Muon, 25B tokens).
|
(Normal init, 15B tokens), and \texttt{muon-25bt} (Muon, 25B tokens).
|
||||||
For each 2D weight matrix $\mathbf{W} \in \mathbb{R}^{m\times n}$ with
|
For each 2D weight matrix $\mathbf{W} \in \mathbb{R}^{m\times n}$ with
|
||||||
SVD $\mathbf{W} = \mathbf{U}\boldsymbol{\Sigma}\mathbf{V}^{\mkern-1mu\mathsf{T}}$,
|
SVD $\mathbf{W} = \mathbf{U}\boldsymbol{\Sigma}\mathbf{V}^{\mkern-1mu\mathsf{T}}$,
|
||||||
@@ -839,9 +847,9 @@ distribution analysis in Section~\ref{sec:num-stability}.
|
|||||||
|
|
||||||
\begin{table}[H]
|
\begin{table}[H]
|
||||||
\centering
|
\centering
|
||||||
\caption{Weight std by component across checkpoints. Non-scaled
|
\caption{Weight std by component across checkpoints. Non-scaled weights broaden with training duration; residual scaling
|
||||||
weights broaden with training duration; Muon constrains drift at
|
constrains drift at equal token count (15B) but is overtaken by
|
||||||
equal token count (15B) but is overtaken by longer training (25B).
|
longer training (25B).
|
||||||
Residual-scaled projections ($\mathbf{W}_o$, $\mathbf{W}_{\text{down}}$)
|
Residual-scaled projections ($\mathbf{W}_o$, $\mathbf{W}_{\text{down}}$)
|
||||||
remain bounded.}
|
remain bounded.}
|
||||||
\label{tab:weight_std}
|
\label{tab:weight_std}
|
||||||
|
|||||||
Reference in New Issue
Block a user