Compare commits

...
2 Commits
Author SHA1 Message Date
ViperEkura e6be33aa53 doc: add SFT length-filter rationale — high PPL + high variance from A.3 IFD figure
- Add 15-token length floor to SFT samples, with explicit reference
  to Appendix A.3 / Figure 5 (ifd_length_grid)
- Short replies (<10 tokens) show both high per-token perplexity
  (L_uncond ~6-8, PPL ~400-3000 vs long replies ~2-3, PPL ~7-20)
  and wide variance in L_cond (span 0-17.5), which distorts
  downstream IFD-based difficulty estimates.
2026-07-24 06:37:35 +08:00
ViperEkura 0a1d0573ae doc: add SFT length filter (15-token floor) with IFD bias rationale
Per-field length filtering in SFT drops instruction--response pairs
with responses shorter than 15 tokens. Short replies exhibit high
per-token variance in both conditional and unconditional loss
(Appendix A.3 / Figure 5), which would distort downstream IFD-based
difficulty estimates.
2026-07-24 06:27:10 +08:00
+10 -2
View File
@@ -154,8 +154,16 @@ pipeline proceeds as follows:
previously kept sample $\mathbf{s}'$.
\end{enumerate}
An optional LLM-as-Judge scoring module provides multi-dimensional
quality scores that can be used to filter low-quality samples.
A length filter is applied to SFT samples based on the IFD
length-bias analysis in Appendix~\ref{sec:ifd_bias}
(Figure~\ref{fig:length_bias}): instruction--response pairs
whose response contains fewer than 15 tokens are discarded,
because short replies exhibit both high per-token perplexity
($L_{\text{uncond}} \approx 6\text{--}8$, PPL~$\approx 400\text{--}3000$)
and wide variance in both $L_{\text{cond}}$ and $L_{\text{uncond}}$,
which would distort downstream IFD-based difficulty estimates.
The threshold is applied per-field, analogous to the pretraining
filter described above.
\subsection{DPO Data Generation}