refactor: 重构打包模块,新增 BFD/FFD/Greedy 三种 bin-packing 算法,默认 BFD
- 将 pipeline/packing.py 拆分为 packing/ 子包 (base/stream/binpack) - 新增 BfdPacker(默认)/FfDPacker/GreedyPacker,移除 StreamingPacker - 超长序列直接截断至 pack_size - group_size 语义改为"每 N 个 chunk 合并为一块",默认 1000 - 新增 AutoTokenizer.token_to_id(),修复 ChatML 中 hacky 的 nl_id 获取 - pad_value 默认改为 2(pad_token_id),position_ids pad=0, loss_mask pad=False - 新增 position_ids 打包后归零一致性测试 - scripts/cache_h5.py 新增 --pack-algo 参数
This commit is contained in:
+12
-2
@@ -28,7 +28,13 @@ Usage::
|
||||
from pipeline.pipeline import Pipeline, PipelineConfig, Stage, TransformStage
|
||||
from pipeline.tokenize import AutoTokenizer, ChatTemplate, train_bpe_tokenizer
|
||||
from pipeline.text import TextNormalizer
|
||||
from pipeline.packing import SequencePacker
|
||||
from pipeline.packing import (
|
||||
GreedyPacker,
|
||||
FfDPacker,
|
||||
BfdPacker,
|
||||
BasePacker,
|
||||
pack_tensors,
|
||||
)
|
||||
|
||||
# I/O module
|
||||
from pipeline.io import FileScanner, HDF5Handler, export_dataset, cache_jsonl
|
||||
@@ -70,7 +76,11 @@ __all__ = [
|
||||
"train_bpe_tokenizer",
|
||||
# Text processing
|
||||
"TextNormalizer",
|
||||
"SequencePacker",
|
||||
"GreedyPacker",
|
||||
"FfDPacker",
|
||||
"BfdPacker",
|
||||
"BasePacker",
|
||||
"pack_tensors",
|
||||
# I/O
|
||||
"FileScanner",
|
||||
"HDF5Handler",
|
||||
|
||||
Reference in New Issue
Block a user