Add DPO training subsection with dpo_metrics figure and Rafailov et al. citation

This commit is contained in:
2026-07-19 15:01:23 +08:00
parent 411354eeb1
commit 6c2e04a86f
2 changed files with 48 additions and 0 deletions
Binary file not shown.

After

Width:  |  Height:  |  Size: 343 KiB

+48
View File
@@ -313,6 +313,49 @@ Chinese--English instruction dataset decreases from $\sim$2.5 to
$\sim$1.6 over 1{,}000 steps, with the gradient norm converging to a
stable range after the warmup phase.
\subsection{Direct Preference Optimization}
\label{sec:dpo}
Following supervised fine-tuning, we align the model with pairwise
human preferences via Direct Preference Optimization
(DPO)~\cite{rafailov2023dpo}. DPO avoids explicit reward-model training
by optimizing the policy $\pi$ directly against a frozen reference
policy $\pi_{\text{ref}}$ (the SFT checkpoint). For a prompt $x$ with
preferred response $y_w$ and dispreferred response $y_l$, the loss is:
\begin{equation}
\mathcal{L}_{\text{DPO}} = -\log \sigma\!\left(
\beta \Bigl[
\log\tfrac{\pi(y_w\mid x)}{\pi_{\text{ref}}(y_w\mid x)}
-\log\tfrac{\pi(y_l\mid x)}{\pi_{\text{ref}}(y_l\mid x)}
\Bigr]\right),
\end{equation}
where $\beta$ controls the KL-divergence penalty against the reference.
We use $\beta=0.1$, a batch of 64 preference pairs per step, and the
same hybrid Muon/AdamW optimizer with cosine scheduling.
\begin{figure}[H]
\centering
\includegraphics[width=0.95\linewidth]{data/dpo_metrics.png}
\caption{DPO training metrics over $\sim$1{,}500 alignment steps:
preference loss (raw and 100-step moving average), learning rate
schedule, and gradient norm.}
\label{fig:dpo_metrics}
\end{figure}
Figure~\ref{fig:dpo_metrics} summarises the DPO training dynamics.
The raw preference loss (left panel) starts near $0.69$ and is visibly
noisy, a hallmark of pairwise preference sampling. The 100-step moving
average reveals a steady downward trend that reaches a minimum of
$\sim$0.51 near step 1{,}200, after which it gently rebounds to
$\sim$0.52 and oscillates, suggesting the policy has converged to a
stable preference boundary rather than over-optimising away from the
reference distribution. The learning-rate schedule (centre panel) peaks
at $5\times10^{-6}$ after a short linear warmup and then follows cosine
decay to a floor of $\sim$0.5\,$\times\,$10$^{-6}$. Gradient norms (right panel) start
near 50 with occasional spikes above 55, then gradually decline and
stabilise in the 35--40 range after step 400, indicating consistent
gradient magnitudes throughout alignment.
% ======================================================================
\section{Numerical Stability via Residual Scaling}
\label{sec:num-stability}
@@ -808,6 +851,11 @@ A.~Radford, J.~Wu, R.~Child, D.~Luan, D.~Amodei, I.~Sutskever.
Language models are unsupervised multitask learners.
\textit{OpenAI Blog}, 2019.
\bibitem{rafailov2023dpo}
R.~Rafailov, A.~Sharma, E.~Mitchell, C.~D.~Manning, S.~Ermon, C.~Finn.
Direct Preference Optimization: Your language model is secretly a reward model.
\textit{NeurIPS}, 2023.
\bibitem{shazeer2020glu}
N.~Shazeer. GLU variants improve Transformer.
\textit{arXiv:2002.05202}, 2020.