CLI Parameter Reference
Contents
Training Parameters
Basic Parameters
| Parameter |
Description |
Default |
--train_type |
Training type (seq, sft, dpo, grpo, online_grpo, online_dpo) |
required |
--data_root_path |
Dataset root directory |
required |
--param_path |
Model parameters or checkpoint path |
required |
--n_epoch |
Total training epochs |
1 |
--batch_per_device |
Batch size per device |
1 |
--grad_accum_steps |
Gradient accumulation steps between optimizer steps |
1 |
Learning Rate Scheduling
| Parameter |
Description |
Default |
--warmup_ratio |
Fraction of total steps used for LR warmup |
0.05 |
--max_lr |
NAdamW learning rate; schedulers scale every optimizer group proportionally |
3e-4 |
--max_grad_norm |
Maximum gradient norm for clipping (None disables) |
1.0 |
Optimizer
The default nora_nadamw optimizer sends internal Linear.weight matrices to
Nora and embeddings, the LM head, norms, biases, LoRA factors, and fallback
parameters to NAdamW. Parameters are classified by module role and identity,
so tied embedding/head weights occur in exactly one group. Nora requires complete
rows under DTensor sharding and rejects layouts sharded along the last dimension.
| Parameter |
Description |
Default |
--optimizer |
Built-in optimizer (nora_nadamw, muon_adamw) |
nora_nadamw |
--weight_decay |
NAdamW decay for eligible fallback parameters; known embeddings, heads, norms, biases, and LoRA factors use 0 |
0.1 |
--nora_lr |
Nora learning rate |
5e-3 |
--nora_beta |
Nora momentum-buffer EMA factor |
0.95 |
--nora_momentum |
Nora Nesterov interpolation factor |
0.95 |
--nora_weight_decay |
Nora matrix weight decay |
0.0 |
muon_adamw preserves the previous MuonMix behavior and the following options:
| Parameter |
Description |
Default |
--muon_momentum |
Muon momentum factor |
0.95 |
--muon_nesterov |
Enable Nesterov momentum for Muon |
True |
--muon_ns_steps |
Newton-Schulz iteration steps for Muon |
5 |
--muon_adjust_lr |
Muon LR adjustment strategy (original, match_rms_adamw) |
match_rms_adamw |
Optimizer identity and hyperparameters are saved in checkpoint metadata. Optimizer
states are intentionally not interchangeable: resume older MuonMix checkpoints
with --optimizer=muon_adamw.
Data Loading
| Parameter |
Description |
Default |
--window_size |
Max input sequence length |
model config max_position_embeddings |
--stride |
Stride for sliding window over sequences |
None |
--random_seed |
Random seed for reproducibility |
3407 |
--num_workers |
DataLoader worker processes |
4 |
--no_pin_memory |
Disable pin_memory (enabled by default) |
(flag) |
Checkpoint & Resume
| Parameter |
Description |
Default |
--ckpt_interval |
Iterations between checkpoints |
5000 |
--ckpt_dir |
Checkpoint save directory |
checkpoint |
--start_epoch |
Resume from epoch (0 = from scratch) |
0 |
--start_samples |
Resume from sample count per rank |
0 |
Validation
| Parameter |
Description |
Default |
--val_split |
Ratio to split from training dataset for validation (e.g. 0.05) |
None |
--val_step |
Number of optimizer steps between validation runs |
1000 |
Logging
| Parameter |
Description |
Default |
--log_dir |
Directory for metric logs |
checkpoint/logs |
--metrics |
Metrics to log (e.g. --metrics loss lr val_loss) |
["loss", "lr", "grad_norm"] |
Gradient Checkpointing
| Parameter |
Description |
Default |
--gradient_checkpointing |
Enable activation checkpointing for DecoderBlock modules |
False |
Distributed Training
| Parameter |
Description |
Default |
--nprocs |
Number of GPUs / processes |
1 |
--parallel_mode |
Parallel strategy (none, ddp, fsdp) |
fsdp |
--device_type |
Device type |
cuda |
--start_method |
Multiprocessing start method (spawn, fork, forkserver) |
spawn |
--backend |
Distributed training backend |
nccl |
--master_addr |
Master node address |
localhost |
--master_port |
Master node port |
29500 |
Strategy-specific
| Parameter |
Description |
Default |
Used by |
--dpo_beta |
DPO beta value |
0.1 |
dpo |
--label_smoothing |
Label smoothing for cross-entropy loss |
0.0 |
seq, sft |
--group_size |
GRPO group size |
4 |
grpo |
--grpo_clip_eps |
GRPO clipping epsilon |
0.2 |
grpo |
--grpo_kl_coef |
GRPO KL penalty coefficient |
0.01 |
grpo |
--neftune_alpha |
NEFTune noise alpha (0=disabled, typical: 5.0) |
0.0 |
sft |
Online Rollout
These options apply to online_grpo and online_dpo. Online strategies require
a BaseRewardModel factory in TrainConfig; train.py does not currently
provide a command-line option for configuring one.
| Parameter |
Description |
Default |
--rollout_interval |
Optimizer steps between rollout refreshes |
512 |
--rollout_temperature |
Rollout sampling temperature |
0.7 |
--rollout_top_k |
Rollout top-k filtering (0 disables) |
0 |
--rollout_top_p |
Rollout nucleus sampling threshold |
0.9 |
--rollout_max_tokens |
Maximum generated tokens per response |
1024 |
Scheduler
| Parameter |
Description |
Default |
--schedule_type |
LR scheduler type (cosine, sgdr, wsd) |
cosine |
--min_rate |
Minimum LR as fraction of base LR |
None (scheduler default: 0.05 for cosine/SGDR, 0.0 for WSD) |
--cycle_length |
SGDR first cycle length in steps |
None (total_steps - warmup_steps) |
--t_mult |
SGDR cycle length multiplier per restart |
2 |
--stable_steps |
WSD stable plateau steps |
None (80% of post-warmup steps) |
--decay_steps |
WSD decay steps |
None (total_steps - warmup_steps - stable_steps) |
Usage Example
Inference Server (server.py)
| Parameter |
Type |
Default |
Description |
--host |
str |
0.0.0.0 |
Host address |
--port |
int |
8000 |
Port number |
--param_path |
path |
project_root/params |
Path to model parameters |
--device |
str |
cuda |
Device to load model on |
--dtype |
str |
bfloat16 |
Model weights dtype (bfloat16, float16, float32) |
--max_batch_size |
int |
16 |
Maximum batch size for continuous batching |
--max_seq_len |
int |
model config max_position_embeddings |
Maximum sequence length (KV cache size + prompt truncation) |
--reload |
flag |
False |
Enable auto-reload for development |
Usage:
See Inference Guide for HTTP API documentation.
Preprocess
See Preprocessing Guide for config file format and examples.
Generate (generate.py)
| Parameter |
Type |
Default |
Description |
--param_path |
str |
required |
Path to the model directory |
--input_json_file |
str |
required |
Path to the input JSONL file |
--output_json_file |
str |
required |
Path to the output JSONL file |
--question_key |
str |
question |
Key for the question in input JSON |
--response_key |
str |
response |
Key for the response in output JSON |
--temperature |
float |
0.60 |
Sampling temperature |
--top_k |
int |
30 |
Top-k filtering |
--top_p |
float |
0.95 |
Nucleus sampling threshold |
--batch_size |
int |
1 |
Batch size for generation |
--num_samples |
int |
1 |
Responses per prompt |
--max_tokens |
int |
model config max_position_embeddings |
Maximum tokens to generate |
--cache_len |
int |
2048 |
KV cache length |
--frequency_penalty |
float |
0.0 |
Frequency penalty |
--rep_window |
int |
64 |
Window size for frequency penalty |
Usage:
Preprocess (preprocess.py)
| Parameter |
Type |
Default |
Description |
input_files |
path(s) |
required |
Input JSONL file(s), supports glob (data/*.jsonl) |
--output_dir, -o |
path |
required |
Output directory for processed data |
--config, -c |
path |
required |
Preprocessing pipeline config (JSON) |
--tokenizer_path |
str |
params |
Path to tokenizer directory |
Usage:
See Preprocessing Guide for config file format and examples.
Document Update Time: 2026-07-20