fix: resolve audited training, import, and serving bugs
- shard the Muon Newton-Schulz orthogonalization over the FSDP mesh instead of partial local slices - import HF checkpoints faithfully: per-head RoPE permutation for q/k projections and qk-norm, qwen3, shared experts, and qk-norm before RoPE (changes numerics for existing use_qk_norm checkpoints) - make preprocessing and resume self-contained: backfill realigned bucket keys by semantics (masks ones, rest zeros) and snapshot tokenizer files into every checkpoint - keep RL consistent: sync the offline GRPO old_model each optimizer step and validate online strategies through a public one-off-rollout hook that leaves the replay cache untouched - fix streaming serving: withhold partial tool-call prefixes with a stream-end flush, stream tool-call arguments from the raw source span, and terminate SSE frames with a blank line - fix sampling semantics: capture logprobs before top-k/top-p mutate logits in place and detect greedy pipelines polymorphically instead of isinstance bookkeeping
This commit is contained in:
@@ -551,3 +551,16 @@ class RolloutRunner:
|
||||
# cache publication. Reward scoring itself intentionally remains
|
||||
# outside the policy lock because it may call an external service.
|
||||
return self.generator.with_policy_snapshot(commit)
|
||||
|
||||
def evaluate(self, batch: Dict) -> RolloutResult:
|
||||
"""One-off rollout + scoring that leaves the replay cache untouched.
|
||||
|
||||
Used by validation on online strategies: the training cache, its
|
||||
cadence counter, and the cache key stay intact, so evaluation
|
||||
prompts never disturb the rollout replay schedule.
|
||||
"""
|
||||
raw = self.generator.generate(batch)
|
||||
self._validate_policy_version(raw)
|
||||
scored = self._score(raw)
|
||||
self._validate_policy_version(scored)
|
||||
return scored
|
||||
|
||||
Reference in New Issue
Block a user