feat: add online ppo with value-model critic and gae advantages
- register online_ppo train type backed by PPOStrategy: token-level clipped surrogate over GAE advantages plus masked value regression against rollout-pinned returns, with explained-variance metrics - fold the reference-KL penalty (k3 estimator) into per-token rewards before GAE and pin advantages/returns on RolloutResult so replayed gradient steps optimize fixed targets - add self-contained ValueModel critic with a zero-initialized value head and backbone warm-started from policy weights; AutoRegressiveLM stays untouched and trunk parity is pinned by tests - step the critic's own optimizer outside the policy-version lock with the same max_grad_norm clipping as the policy - persist critic state as value_model.pt/value_optimizer.pt checkpoint extras; resume restores it, fails loudly when missing, and the train.sh completeness check requires the extras for online_ppo configs - extract shared rollout sequence/logprob helpers from GRPO (behavior unchanged) and add ppo_gamma/ppo_gae_lambda/ppo_vf_coef CLI options
This commit is contained in:
@@ -22,13 +22,26 @@ validate_job_name() {
|
||||
die "Invalid TRAIN_JOB_NAME '$1'; use letters, numbers, dot, underscore, or dash"
|
||||
}
|
||||
|
||||
checkpoint_extra_files() {
|
||||
# Additional files a complete checkpoint must contain for the strategy
|
||||
# configured in the given training YAML. PPO persists critic state as
|
||||
# checkpoint extras (value_model.pt / value_optimizer.pt); a resume
|
||||
# without them must not look complete.
|
||||
local config="$1"
|
||||
|
||||
[[ -n "${config}" && -f "${config}" ]] || return 0
|
||||
if grep -Eq '^[[:space:]]*train_type:[[:space:]]*["'\'']?online_ppo' "${config}"; then
|
||||
printf 'value_model.pt value_optimizer.pt'
|
||||
fi
|
||||
}
|
||||
|
||||
checkpoint_is_complete() {
|
||||
local checkpoint="$1"
|
||||
local file
|
||||
|
||||
[[ -d "${checkpoint}" ]] || return 1
|
||||
|
||||
for file in meta.json config.json model.safetensors optimizer.pt scheduler.pt; do
|
||||
for file in meta.json config.json model.safetensors optimizer.pt scheduler.pt ${CHECKPOINT_EXTRA_FILES:-}; do
|
||||
[[ -s "${checkpoint}/${file}" ]] || return 1
|
||||
done
|
||||
|
||||
|
||||
@@ -34,6 +34,7 @@ fi
|
||||
if [[ -n "${TRAIN_CONFIG}" ]]; then
|
||||
[[ -f "${TRAIN_CONFIG}" ]] || die "Training config not found: ${TRAIN_CONFIG}"
|
||||
fi
|
||||
export CHECKPOINT_EXTRA_FILES="$(checkpoint_extra_files "${TRAIN_CONFIG}")"
|
||||
[[ -r /data ]] || die "Training data directory is not readable: /data"
|
||||
|
||||
mkdir -p "${CHECKPOINT_DIR}"
|
||||
|
||||
Reference in New Issue
Block a user