refactor : 基于声明式 JSON 配置的预处理管线重构
- 用工厂注册的 MaskBuilder(chat/instruction/text)替换硬编码的 _transform_* 方法 - mask 规则以 role-to-action 映射声明在配置中,与 chat_template 完全解耦 - 单次编码 + role-span 追踪替代两次编码 + 长度差计算 mask 的方式 - 支持多轮对话训练:所有 assistant 轮次参与训练,而非仅最后一轮 - 新建 astrai.preprocessing 包(builder.py + pipeline.py),删除 astrai/preprocess.py - CLI 精简为 --config 参数,所有参数通过 PipelineConfig JSON 配置 - 新增 PipelineConfig、InputConfig、ProcessingConfig、OutputConfig dataclass - 文档:assets/docs/preprocessing.md - 27 个测试覆盖 mask builder、pipeline、配置序列化、工厂注册
This commit is contained in:
@@ -0,0 +1,227 @@
|
||||
# Preprocessing Pipeline
|
||||
|
||||
Declarative JSON-driven data preprocessing. No code needed -- describe your input format and mask rules in a config file, the engine does the rest.
|
||||
|
||||
## Philosophy
|
||||
|
||||
| Component | Responsibility |
|
||||
|-----------|---------------|
|
||||
| `tokenizer_config.json` (`chat_template`) | Formatting -- how roles become tokens |
|
||||
| `pipeline.json` (`mask`) | Masking -- which roles participate in training |
|
||||
|
||||
The two are fully decoupled. A single config file captures the entire pipeline, reusable and version-controllable. Extension is via factory registration (`@MaskBuilderFactory.register`) -- no need to touch existing code.
|
||||
|
||||
## Quick Start
|
||||
|
||||
### SFT Chat
|
||||
|
||||
```json
|
||||
{
|
||||
"version": 1,
|
||||
"input": {
|
||||
"type": "chat",
|
||||
"messages_key": "messages"
|
||||
},
|
||||
"mask": {
|
||||
"system": "mask",
|
||||
"user": "mask",
|
||||
"assistant": "train"
|
||||
},
|
||||
"mask_default": "mask",
|
||||
"preprocessing": {
|
||||
"max_seq_len": 2048,
|
||||
"deduplicate": true
|
||||
},
|
||||
"output": {
|
||||
"domain_key": "source",
|
||||
"storage_format": "bin",
|
||||
"max_tokens_per_shard": 100000000
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Three lines of mask rules cover the most common SFT case: train on assistant turns, mask everything else.
|
||||
|
||||
### Instruction Tuning
|
||||
|
||||
```json
|
||||
{
|
||||
"version": 1,
|
||||
"input": {
|
||||
"type": "instruction",
|
||||
"prompt_key": "instruction",
|
||||
"response_key": "output"
|
||||
},
|
||||
"mask": {
|
||||
"prompt": "mask",
|
||||
"response": "train"
|
||||
},
|
||||
"mask_default": "mask",
|
||||
"preprocessing": {
|
||||
"max_seq_len": 2048
|
||||
},
|
||||
"output": {
|
||||
"storage_format": "bin"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Mask splits at the prompt/response field boundary.
|
||||
|
||||
### Pretraining
|
||||
|
||||
```json
|
||||
{
|
||||
"version": 1,
|
||||
"input": {
|
||||
"type": "text",
|
||||
"text_key": "content"
|
||||
},
|
||||
"mask": {},
|
||||
"preprocessing": {
|
||||
"max_seq_len": 2048,
|
||||
"min_chars": 50
|
||||
},
|
||||
"output": {
|
||||
"storage_format": "bin"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
No mask -- train on all tokens.
|
||||
|
||||
### Run
|
||||
|
||||
```bash
|
||||
python scripts/tools/preprocess.py data/*.jsonl -o output/ -c sft.json
|
||||
```
|
||||
|
||||
## Configuration Reference
|
||||
|
||||
### `input`
|
||||
|
||||
| Field | Type | Required | Default | Description |
|
||||
|-------|------|----------|---------|-------------|
|
||||
| `type` | string | yes | `"chat"` | Format: `"chat"`, `"instruction"`, or `"text"` |
|
||||
| `messages_key` | string | no | `"messages"` | JSON key for messages array (chat) |
|
||||
| `prompt_key` | string | no | `"prompt"` | JSON key for prompt field (instruction) |
|
||||
| `response_key` | string | no | `"response"` | JSON key for response field (instruction) |
|
||||
| `text_key` | string | no | `"text"` | JSON key for text field |
|
||||
|
||||
### `mask`
|
||||
|
||||
A map of `{role_or_field: "mask" | "train"}`. The engine uses this to build `loss_mask`:
|
||||
|
||||
- `"mask"` -- tokens in this span are ignored during training (`loss_mask=0`)
|
||||
- `"train"` -- tokens in this span contribute to the loss (`loss_mask=1`)
|
||||
|
||||
For chat mode, keys are role names (`system`, `user`, `assistant`, ...).
|
||||
For instruction mode, keys are `"prompt"` and `"response"`.
|
||||
|
||||
| Field | Type | Default | Description |
|
||||
|-------|------|---------|-------------|
|
||||
| `mask` | dict | `{}` | Role/field to action mapping |
|
||||
| `mask_default` | string | `"mask"` | Default action for unlisted roles |
|
||||
|
||||
### `preprocessing`
|
||||
|
||||
| Field | Type | Default | Description |
|
||||
|-------|------|---------|-------------|
|
||||
| `max_seq_len` | int | `2048` | Maximum token length; truncated if exceeded |
|
||||
| `min_chars` | int | `50` | Minimum character length; dropped if shorter (text mode only) |
|
||||
| `max_chars` | int | `2000000` | Maximum character length; dropped if longer (text mode only) |
|
||||
| `deduplicate` | bool | `true` | Remove exact duplicates via MD5 of first 200 chars |
|
||||
| `max_items` | int or null | `null` | Maximum items to process; `null` = unlimited |
|
||||
|
||||
### `output`
|
||||
|
||||
| Field | Type | Default | Description |
|
||||
|-------|------|---------|-------------|
|
||||
| `domain_key` | string or null | `null` | JSON key for domain grouping; `null` = all output to `__default__` |
|
||||
| `storage_format` | string | `"bin"` | `"bin"` (mmap, zero-copy) or `"h5"` (HDF5) |
|
||||
| `max_tokens_per_shard` | int | `100000000` | Max tokens per output shard |
|
||||
|
||||
## Mask Algorithm
|
||||
|
||||
### Chat Mode (role-span tracking)
|
||||
|
||||
For each message in the `messages` array:
|
||||
|
||||
1. Render through the chat template for that single message
|
||||
2. Encode the rendered text, record token span `(start, end, role)`
|
||||
3. Concatenate all spans -- special tokens from the chat template naturally prevent BPE merging across message boundaries
|
||||
4. Fill `loss_mask` from the mask rules
|
||||
|
||||
**Multi-turn example**:
|
||||
|
||||
```
|
||||
Data:
|
||||
[system: "You are helpful."]
|
||||
[user: "What is 2+2?"]
|
||||
[assistant: "4"]
|
||||
[user: "What is 3+3?"]
|
||||
[assistant: "6"]
|
||||
|
||||
Config:
|
||||
"mask": {"system": "mask", "user": "mask", "assistant": "train"}
|
||||
|
||||
Result:
|
||||
tokens: <bos> [system span] [user span] [assistant:4 span] [user span] [assistant:6 span]
|
||||
mask: 0 0 0 1 0 1
|
||||
```
|
||||
|
||||
Both assistant turns are trained. All system and user tokens are masked.
|
||||
|
||||
### Instruction Mode (field boundary)
|
||||
|
||||
Encode the prompt and response fields independently, then split the mask at the field boundary.
|
||||
|
||||
- `"prompt": "mask", "response": "train"` -- mask the left half, train the right half
|
||||
- `"prompt": "train", "response": "mask"` -- the reverse
|
||||
|
||||
### Text Mode (no mask)
|
||||
|
||||
Pure tokenization. No `loss_mask` is produced. Used for pretraining.
|
||||
|
||||
## Output Layout
|
||||
|
||||
```
|
||||
output_dir/
|
||||
__default__/ # when domain_key is null
|
||||
meta.json # {"sequence": {"shape": [N], "dtype": "int64"}, ...}
|
||||
sequence.bin # int64 raw bytes, mmap-able for zero-copy reads
|
||||
loss_mask.bin # int64 raw bytes
|
||||
wiki/ # when domain_key="source" and item["source"]="wiki"
|
||||
meta.json
|
||||
sequence.bin
|
||||
loss_mask.bin
|
||||
```
|
||||
|
||||
## Extension
|
||||
|
||||
Register a custom builder for new formats:
|
||||
|
||||
```python
|
||||
from astrai.preprocessing.builder import BaseMaskBuilder, MaskBuilderFactory
|
||||
|
||||
@MaskBuilderFactory.register("my_format")
|
||||
class MyFormatBuilder(BaseMaskBuilder):
|
||||
def build(self, item: dict, config, tokenizer) -> dict | None:
|
||||
# Return {"ids": [...], "loss_mask": [...], "domain": "..."}
|
||||
# Return None to skip this item
|
||||
...
|
||||
```
|
||||
|
||||
Then set `"input": {"type": "my_format"}` in your config.
|
||||
|
||||
## Compared to Old Pipeline
|
||||
|
||||
| Old (`astrai.preprocess.Pipeline`) | New (`astrai.preprocessing.pipeline.Pipeline`) |
|
||||
|---|---|
|
||||
| Configured via constructor arguments | Configured via JSON file |
|
||||
| Hardcoded `_transform_chat` / `_transform_text` | Factory-registered `Builder` with declarative mask rules |
|
||||
| Auto-detects format via magic key lists | Explicit `input.type` declaration |
|
||||
| Double-encodes (full + prompt), uses length diff for mask | Single-encode with role-span tracking |
|
||||
| Only trains the last assistant turn | Configurable: multi-turn, single-turn, or no mask |
|
||||
|
||||
> Document Update Time: 2026-05-30
|
||||
@@ -1,38 +1,5 @@
|
||||
# Training
|
||||
|
||||
## Model Architecture
|
||||
|
||||
The model uses a decoder-only Transformer with **GQA** (Grouped Query Attention) and optional **MLA** (Multi-head Latent Attention). 1.0 billion parameters, Chinese–English bilingual.
|
||||
|
||||
```mermaid
|
||||
flowchart TB
|
||||
subgraph Layers["Transformer Layers"]
|
||||
direction TB
|
||||
A[Input Embedding] --> B[Transformer Block\nLayer 1]
|
||||
B --> C[Transformer Block\nLayer ...]
|
||||
C --> D[Transformer Block\nLayer ...]
|
||||
D --> E[RMSNorm]
|
||||
E --> F[Linear]
|
||||
F --> G[SoftMax]
|
||||
end
|
||||
|
||||
subgraph TransformerBlock["Transformer Block"]
|
||||
direction TB
|
||||
H[x] --> I[RMSNorm]
|
||||
I --> J[Linear → Q/K/V]
|
||||
J --> K[Q]; J --> L[K]; J --> M[V]
|
||||
K --> N[RoPE]; L --> O[RoPE]
|
||||
N --> P["Q @ K^T / sqrt(d)"]; O --> P
|
||||
P --> Q[Masked SoftMax]; Q --> R[S @ V]; M --> R
|
||||
R --> S[Linear]; S --> T[+]; H --> T
|
||||
T --> U[RMSNorm]
|
||||
U --> V["Linear (gate)"]; U --> W["Linear (up)"]
|
||||
V --> X[SiLU]; X --> Y[×]; W --> Y
|
||||
Y --> Z["Linear (down)"]; Z --> AA[+]; T --> AA
|
||||
AA --> BB[x']
|
||||
end
|
||||
```
|
||||
|
||||
### Autoregression
|
||||
|
||||
Given a token sequence, the model predicts the probability of the next token. Each generated token is appended to the input and fed back, repeating until an end-of-sequence token or max length.
|
||||
|
||||
Reference in New Issue
Block a user