37 Commits
Author SHA1 Message Date
ViperEkura 663ef900fc refactor: move sample-id indexing from dataset to store
- Store owns window_size/stride and __getitem__/__len__/sample_window
- Dataset classes become thin delegators binding a Store to a train-type key mapping
- Drop BaseDataset.get_index and the RecordDataset中间类 (window死代码)
- DatasetFactory forces window_size=0 for record datasets so record semantics never get window-tainted
- token_count/num_records split the legacy len() semantics (raw stream length vs record count)
- Update tests to the new .store/.token_count API and window/record mode switching
2026-07-19 16:02:50 +08:00
ViperEkura 7d478a54db docs: update HF org from ViperEk to ViperEkura
- Replace 4 HF links in README.md and README-zh-CN.md to point to ViperEkura
- Update download.py default repo to AstrAI-V1-instruct under ViperEkura
2026-07-19 14:49:58 +08:00
ViperEkura f3eaaef842 refactor: remove redundant strategy/executor code
- Drop BaseStrategy.model_fn (stored but never read)
- Drop model_fn= passed to StrategyFactory.create in train_context
- Simplify FSDPExecutor.clip_grad_norm None branch to delegate to super()
- Remove DDPExecutor._gather_state_dict override (identical to base)
2026-07-19 12:45:58 +08:00
ViperEkura d655b65027 docs: sync architecture/dataflow/training/params with code
- dataflow.md: update DatasetFactory.load signature, stream vs record access, Store._offsets
- architecture.md: add tokenizer to Pipeline, TokenizeTransform class, RecordDataset, Streamable/Recordable mixins, fix GRPOStrategy (old_model/sync_old_model)
- training.md: DPO reduction="sum", GRPO rho_t uses pi_old, gradient_clipping always registered
- params.md: --max_grad_norm default None
2026-07-19 12:33:35 +08:00
ViperEkura 31c22dc043 refactor: deduplicate preprocessing kernel and BFD packing
- Extract shared core (mask building, primary-id extraction, tensorisation, position-id generation) to astrai/preprocessing/core.py; Pipeline and TokenizeTransform both consume it, eliminating ~60% duplicated logic
- Promote BFD _plan to module-level plan_bfd(lengths, max_len) returning pure index bins; BFDPacking.apply and evaluate_ifd._pack_bins both call it, removing the second BFD implementation
- Split Pipeline._flush (49 lines) into _inject_doc_reset_position_ids + _inject_continuous_position_ids + _to_tensors; split Pipeline.run by delegating record iteration to core.iter_raw_records
- Remove dead no-op pop/塞回 in Pipeline.run (L110-111)
2026-07-19 12:27:56 +08:00
ViperEkura 17127f8b3c fix: make tokenizer picklable for spawn multiprocessing
- ChatTemplate: defer Jinja2 compilation to cached_property, exclude compiled template from __getstate__ (its dynamic root function has __module__=None and falls back to __main__, breaking pickle)
- AutoTokenizer: bypass __getattr__ for underscore-prefixed attrs to prevent infinite recursion during unpickle when __dict__ is empty
2026-07-19 11:59:55 +08:00
ViperEkura d7695b40e3 feat: make max_grad_norm optional (None disables clipping)
- TrainConfig.max_grad_norm defaults to None
- executor.clip_grad_norm returns grad norm without clipping when None
- train.py --max_grad_norm defaults to None
2026-07-19 00:08:18 +08:00
ViperEkura fc62890e70 fix: apply chat template in DPO tokenization
- dpo_tokenize now uses tokenizer.apply_chat_template to match SFT format
- Prompt rendered with add_generation_prompt=True
- Chosen/rejected appended as assistant turn
- Remove leftover dead code from _extract_text
- Update tests to mock apply_chat_template
2026-07-19 00:00:51 +08:00
ViperEkura f433672140 fix: use sum reduction for DPO sequence logprob
- DPO requires sequence-level sum of token logprobs, not per-token mean
- mean reduction made beta*ratio_diff ~0.03 (near-zero gradient)
- loss stalled at 0.6931 because logsigmoid(0.03) has vanishing grad
- sum gives beta*ratio_diff ~10 with meaningful gradients
2026-07-18 23:48:35 +08:00
ViperEkura 7e1e5b6e6a refactor: DatasetFactory.load accepts pre-built store instance
- load(store=...) binds directly, skipping format detection/processor
- load_path now optional when store is given
- Remove redundant from_store (merged into load)
- Caller can fully control Store construction + processor setup
2026-07-18 23:23:51 +08:00
ViperEkura 553a42702d refactor: replace diamond inheritance with mixin composition
- StreamStore/RecordStore → Streamable/Recordable (stateless mixins)
- Store is sole base class, no MRO ambiguity
- H5Store/MmapStore/JsonlStore mix in both traits explicitly
- segments_are_records declared per-subclass (H5/Jsonl=True, bin=False)
- Add tests for dpo_tokenize, lazy jsonl, dual-mode H5, stream-only bin
- Remove unused _to_tensor helper
2026-07-18 23:20:41 +08:00
ViperEkura b133fc9c07 refactor: split Store into StreamStore and RecordStore
- StreamStore: fetch(begin, end, key) for stream access (SEQ/SFT)
- RecordStore: mixin with fetch_record(i, key) for record access
- H5Store/MmapStore/JsonlStore now dual-inherit both (C3 MRO)
- JsonlStore supports lazy mode via processor= (no TokenizeTransform)
- RecordDataset base class holds processor, DPO/GRPO simplified
- dpo_tokenize pure function for on-the-fly JSONL tokenisation
- DatasetFactory builds processor for jsonl+record datasets
- train.py passes tokenizer_path=param_path uniformly
- progress: len(dataset) returns sample count (stream=windows, record=records)
- json no longer auto-detected as jsonl format
2026-07-18 23:04:31 +08:00
ViperEkura b33250dc28 refactor: decouple tokenizer from Store into Transform layer
- Extract tokenization/mask/position logic from JsonlStore into TokenizeTransform
- JsonlStore now pure reader: reads JSON records, delegates to transform
- Store no longer imports tokenizer or preprocessing components
- Replace per_record param with segments_are_records class attribute
- Store subclasses declare segment semantics as format-level property
2026-07-18 21:37:31 +08:00
ViperEkura a74e5b91a3 feat: add record-mode to Store for DPO/GRPO
- Store gains fetch_record/num_records alongside stream fetch/__len__
- save_bin/load_bin support per-record offsets via record_keys param
- H5Store/MmapStore/JsonlStore all support dual stream+record access
- DPODataset/GRPODataset use fetch_record, no cross-record concat
- dpo_collate_fn + collate_fn wired through TrainConfig
- fixes attention context leakage in DPO from windowed concatenation
2026-07-18 21:02:29 +08:00
ViperEkura 28886e4241 fix: make system prompt optional across scripts
- stream_chat: default empty system_prompt, single-turn mode
- generate_batch: drop hardcoded system role
- generate.py: preserve original fields in messages branch
  and use response_key for the output column name
2026-07-18 14:10:37 +08:00
ViperEkura 9d3ccfdffc fix: incremental decode to avoid U+FFFD in streaming
- StreamDecoder buffers incomplete multi-byte sequences
- Task.decode_new_token replaces per-token decode in scheduler
- flush_remaining emits final buffered text on task finish
2026-07-18 13:05:36 +08:00
ViperEkura a24a7b4da5 perf: merge decode batch for 10x throughput
- merge all active decode tasks into single forward pass (was grouped by next_pos)
- add per-task write_positions to ContiguousCacheView for correct KV writes
- override ContiguousCache.task_cached (base returned 0, caused prefill loops)
- add --cache_len/--frequency_penalty/--rep_window to generate.py
- chunked batch processing with tqdm progress

bench (1.2B model, 128 prompts, 64 tok, batch=128):
  before: 77.2s, ~111 tok/s
  after:   7.1s, ~1210 tok/s (10.9x)
2026-07-18 08:50:46 +08:00
ViperEkura f7df02f9a3 feat: add --num_samples to batch generation script 2026-07-18 01:13:38 +08:00
ViperEkura ee450686f3 fix: add option permutation to MMLU eval
- Few-shot examples now include subject preamble (consistent format)
- Add --seed flag for option permutation (default 0, -1 to disable)
- Shuffles A/B/C/D positions per-question to neutralise positional bias
2026-07-18 00:14:34 +08:00
ViperEkura 2565755e45 refactor: switch eval datasets to HuggingFace source
- Replace GitHub/berkeley direct downloads with HF datasets API
- MMLU: cais/mmlu (all config), map val->validation split, write per-subject CSV
- HumanEval: openai/openai_humaneval
- IFEval: google/IFEval
- Enables HF_ENDPOINT mirror for faster downloads in CN
2026-07-18 00:09:31 +08:00
ViperEkura d08a92c7bd feat: add frequency penalty to inference sampling pipeline
- Add FrequencyPenaltyStrategy (logit -= penalty * count)
- Per-task rep_window for penalty history lookup
- Wire through engine, task, executor, API layer
- Add --frequency_penalty and --rep_window to stream_chat.py
- 9 unit tests for frequency penalty strategy
2026-07-17 21:28:31 +08:00
ViperEkura a1ea26d367 fix: rewrite GRPO data pipeline for offline record-level access
- process_list_field returns List[List[int]] preserving per-response boundaries
- GRPODataset rewritten to record-level __getitem__ (no windowing/stride)
- grpo_collate_fn pads variable-length responses into [B, G, R] tensors
- JsonlStore detects nested List[List[int]] and stores List[Tensor] per record
- Store._normalize skips nested-list keys from cumsum bookkeeping
- Pipeline._flush handles nested lists without cross-record flattening
- Export grpo_collate_fn from astrai.dataset
- 6 new GRPO tests + 2 updated builder tests, 114 total pass
2026-07-17 14:34:41 +08:00
ViperEkura c17aa0dc54 fix: eval script bugs and add missing features
- evaluate_mmlu: fix double few-shot injection (build_prompt no longer
  adds few-shot, apply_chat handles it once)
- evaluate_humaneval: fix pass@k k-filtering to be per-problem instead
  of using first problem's n globally; reuse ProcessPoolExecutor across
  problems; fix closure UnboundLocalError in test_one; handle None in
  report when k > n
- evaluate_ifd: remove dead code (score_plain/score_messages); add
  multi-file/directory input support with --input_path/--output_dir;
  add summary.json aggregation and --max_samples; add --dtype flag
- evaluate_ppl: add --device and --dtype flags (was hardcoded to cuda)
- evaluate_ifeval: fix docstring path (scripts/tools -> scripts/eval)
- analyze_weights: add --output JSON export; fix dead code filter
  ("_norm" not in r was always True)
2026-07-17 14:02:58 +08:00
ViperEkura b12b24eadc feat: rewrite evaluate_ppl with token-level loss and multi-file support
- Support multiple input files, glob patterns, and directory input
- Add --token_level flag: per-record token_ids + log_probs JSONL output
- Add --max_samples for random subsampling per file
- LossAccumulator: streaming mode (histogram-based percentiles, low memory) vs exact mode (full token list)
- Token type analysis (ascii/cjk/non_ascii/special) when token_level=True
- Fix token_ids/log_probs alignment (shift offset)
- Cache frozenset(stop_ids) outside loop for performance
- Aggregate stats: mean/median/ppl/p50/p90/p95/p99
- Summary JSON with all datasets in one file
2026-07-17 13:14:15 +08:00
ViperEkura cd14d53707 feat: implement bfd_split packing strategy
- BFDSplitPacking splits over-length sequences into chunks before BFD
- All keys (loss_mask, position_ids, ...) split in lockstep for alignment
- No tokens lost vs bfd which truncates over-length sequences
- Tests: token preservation, chunk alignment, short unchanged, vs bfd
2026-07-17 12:38:56 +08:00
ViperEkura e220413035 feat: support raw JSON files in dataset pipeline and JsonlStore
- detect_format now recognizes .json directories as jsonl store
- JsonlStore loads .json arrays and dicts alongside .jsonl
- tokenizer_path defaults to dataset dir when omitted
- Pipeline._iter_items handles .json files (arrays/single dict)
- Tests: detect_format, seq load, self-contained dataset dir
2026-07-17 12:20:03 +08:00
ViperEkura 84ed2327f5 feat: add --resume flag to decouple weight loading from training resumption
- Add --resume bool flag to train.py CLI
- --param_path always loads weights only by default
- --resume restores epoch, consumed_samples, optimizer & scheduler
- Checkpoint.load() now preserves full meta dict
- Update test_early_stopping to use new param_path/resume API
2026-07-16 14:23:23 +08:00
ViperEkura b14f301730 fix: init last_ckpt_step and last_log_flush_step from context.optimizer_step 2026-07-15 22:15:55 +08:00
ViperEkura 0654b4b916 refactor: template combine kernel, fix mask bug, unify dispatch
- Template combine kernel, share macros, extract entry_utils helpers
- Fix mask indexing (pass stride not pre-multiplied base)
- Remove !p.use_mask — MMA handles mask
2026-07-15 21:44:17 +08:00
ViperEkura 1f0be382ad refactor: extract load_q_mma_frags template, unify comment style
- Add load_q_mma_frags<KD>() shared template in attn_mma_utils.cuh
- Replace ~15 duplicated Q-load lines in 3 MMA kernels
- Unify section header comment style to // ---- Section ----
- Remove duplicate separator line in attn_mma_utils.cuh
2026-07-15 19:07:18 +08:00
ViperEkura bb175fda91 fix: resume optimizer LR, step display, and consumed_samples alignment 2026-07-15 08:59:52 +08:00
ViperEkura 13998da15a fix: uninitialized strides in decode test and wrong stride helper in paged test
- decode test main() missing set_default_strides → illegal memory access
- paged test used set_default_strides on PagedAttentionParams → compile error
2026-07-14 23:58:30 +08:00
ViperEkura 57729fd92d refactor: stride-based attn interface with layout and causal mask
- Replace is_causal + causal_offset with unified causal_offset (-1 = off, >=0 = first Q pos)
- Causal and mask can now coexist (was mutually exclusive)
- Add stride-based addressing for Q/KV/O (layout-agnostic, zero-copy)
- Add layout param ("bhld"/"blhd") parsed in Python, passed as int to C++
- Support 2D [batch, kv_len] and 3D [batch, q_len, kv_len] mask
- Vectorize paged KV gather in Python fallback (was per-token Python loop)
- Extract shared helpers: compute_num_splits, alloc_split_partials, dispatch_head_dim
- Unify paged_decode entry via attn_pack_paged_params
- Update mma_softmax_tile for 3D mask with per-row qrow indexing
2026-07-14 21:34:42 +08:00
ViperEkura 2c7a71a9c0 refactor: separate old policy and ref model in GRPO strategy
- Split single ref_model into old_model (importance sampling ratio) and ref_model (frozen KL regularizer)
- Move ref_model/old_model creation from strategy __init__ to TrainContextBuilder, pass as explicit parameters
- Remove periodic sync_ref_model + sync_interval; add sync_old_model for external rollout loop to call
- DPOStrategy also receives ref_model from builder
- Fix std to use unbiased=False (population std per GRPO paper)
- Remove redundant tests (test_grpo_kl_zero_at_init, test_grpo_no_sync_interval_param)
- Remove --grpo_sync_interval CLI arg
2026-07-14 20:03:45 +08:00
ViperEkura 3e0007fc91 docs : fix factory lists and MaskBuilderFactory docs
- Add MaskBuilderFactory, StoreWriterFactory, PackingStrategyFactory, PositionIdStrategyFactory to architecture design patterns
- Clarify MaskBuilderFactory three registered names (single, multi, sectioned) in preprocessing docs
2026-07-13 15:18:06 +08:00
ViperEkura b092316385 feat : add distributed checkpoint via executor checkpoint_context
- Add checkpoint_context context manager to BaseExecutor with entry/exit barrier
- Add _gather_state_dict hook overridden per executor (template method)
- DDPExecutor skips unwrap on non-rank-0 to avoid redundant state_dict gather
- FSDPExecutor uses rank0_only=True to reduce memory on non-writers
- Remove redundant rank-0 guard from Checkpoint.save and manual barrier from Callback
2026-07-13 12:27:09 +08:00
ViperEkura 9bcd696580 fix: token-level ratio and prompt masking in GRPO strategy
- Mask prompt tokens to 0 so their logprobs excluded from ratio/KL
- Switch to token-level ratio + PPO clipping via reduction='none'
- Slice response token logprobs from full sequence output
- Replace k3 KL estimator with non-negative k1 estimator
- Fix epsilon from finifo.eps (~1e-38) to 1e-8
- Remove unused 'reduction' param from GRPOStrategy.__init__
- Clarify offline batch semantics in docstring
- Add 11 unit tests for masking, advantage, KL, sync, clipping
- Sync training.md and architecture.md docs
2026-07-12 21:24:09 +08:00
74 changed files with 4639 additions and 1428 deletions
+2 -2
View File
@@ -20,7 +20,7 @@
<a href="assets/docs/README-zh-CN.md">中文</a> • <a href="assets/docs/README-zh-CN.md">中文</a> •
<a href="https://github.com/ViperEkura/AstrAI/issues">Issue Tracker</a> • <a href="https://github.com/ViperEkura/AstrAI/issues">Issue Tracker</a> •
<a href="https://github.com/ViperEkura/AstrAI/discussions">Discussions</a> • <a href="https://github.com/ViperEkura/AstrAI/discussions">Discussions</a> •
<a href="https://huggingface.co/ViperEk/">HuggingFace</a> <a href="https://huggingface.co/ViperEkura">HuggingFace</a>
</div> </div>
<br> <br>
@@ -241,7 +241,7 @@ For major changes, please open an issue first to discuss what you would like to
- **GitHub Issues**: [Issue Tracker](https://github.com/ViperEkura/AstrAI/issues) - **GitHub Issues**: [Issue Tracker](https://github.com/ViperEkura/AstrAI/issues)
- **Discussions**: [GitHub Discussions](https://github.com/ViperEkura/AstrAI/discussions) - **Discussions**: [GitHub Discussions](https://github.com/ViperEkura/AstrAI/discussions)
- **HuggingFace**: [Model Hub](https://huggingface.co/ViperEk) - **HuggingFace**: [Model Hub](https://huggingface.co/ViperEkura)
### License ### License
+2 -2
View File
@@ -27,7 +27,7 @@
<a href="#chinese">中文</a> • <a href="#chinese">中文</a> •
<a href="https://github.com/ViperEkura/AstrAI/issues">问题追踪</a> • <a href="https://github.com/ViperEkura/AstrAI/issues">问题追踪</a> •
<a href="https://github.com/ViperEkura/AstrAI/discussions">讨论区</a> • <a href="https://github.com/ViperEkura/AstrAI/discussions">讨论区</a> •
<a href="https://huggingface.co/ViperEk">HuggingFace</a> <a href="https://huggingface.co/ViperEkura">HuggingFace</a>
</div> </div>
<br> <br>
@@ -247,7 +247,7 @@ SSE 流式格式、错误码和统计端点详见[推理文档](./inference.md)
- **GitHub Issues**: [问题追踪](https://github.com/ViperEkura/AstrAI/issues) - **GitHub Issues**: [问题追踪](https://github.com/ViperEkura/AstrAI/issues)
- **Discussions**: [GitHub 讨论区](https://github.com/ViperEkura/AstrAI/discussions) - **Discussions**: [GitHub 讨论区](https://github.com/ViperEkura/AstrAI/discussions)
- **HuggingFace**: [模型中心](https://huggingface.co/ViperEk) - **HuggingFace**: [模型中心](https://huggingface.co/ViperEkura)
### 许可证 ### 许可证
+65 -16
View File
@@ -117,7 +117,7 @@ classDiagram
+int n_epoch +int n_epoch
+int batch_per_device +int batch_per_device
+int grad_accum_steps +int grad_accum_steps
+float max_grad_norm +Optional[float] max_grad_norm
+list gradient_checkpointing_modules +list gradient_checkpointing_modules
+int start_epoch +int start_epoch
+int start_samples +int start_samples
@@ -166,6 +166,13 @@ classDiagram
+__getitem__(index) Dict +__getitem__(index) Dict
} }
class RecordDataset {
+Optional[Callable] processor
+load(load_path, storage_type)
+__getitem__(index)
+__len__()
}
class DPODataset { class DPODataset {
+__getitem__(index) Dict +__getitem__(index) Dict
} }
@@ -177,13 +184,26 @@ classDiagram
class Store { class Store {
+Dict[str, List[Tensor]] _data +Dict[str, List[Tensor]] _data
+Dict[str, List[int]] _cum +Dict[str, List[int]] _cum
+Dict[str, List[int]] _offsets
+int _length +int _length
+int _num_records
+keys (property) +keys (property)
+load(path) +load(path)
+fetch(begin, end, keys)
+__len__() +__len__()
-_fetch_key(key, begin, end) Tensor -_normalize(raw, offsets)
-_normalize(raw) }
class Streamable {
<<mixin>>
+fetch(begin, end, keys)
-_fetch_stream_key(key, begin, end) Tensor
}
class Recordable {
<<mixin>>
+num_records (property)
+fetch_record(index, keys)
-_fetch_record_key(key, index) Tensor
} }
class H5Store { class H5Store {
@@ -195,6 +215,13 @@ classDiagram
+load(path) +load(path)
} }
class JsonlStore {
+JsonlSource _source
+Callable _processor
+load(path, transform, processor)
+fetch_record(index, keys)
}
class ResumableDistributedSampler { class ResumableDistributedSampler {
+int epoch +int epoch
+int iter +int iter
@@ -210,7 +237,7 @@ classDiagram
+Dict _entries +Dict _entries
+register(name) decorator +register(name) decorator
+create(train_type, window_size, stride) BaseDataset +create(train_type, window_size, stride) BaseDataset
+load(train_type, load_path, window_size, stride, storage_type) BaseDataset +load(train_type, load_path, window_size, stride, storage_type, tokenizer_path, max_len, store) BaseDataset
} }
} }
@@ -378,6 +405,7 @@ classDiagram
+List[str] paths +List[str] paths
+str output_dir +str output_dir
+str tokenizer_path +str tokenizer_path
+AutoTokenizer tokenizer
+BaseMaskBuilder mask_builder +BaseMaskBuilder mask_builder
+PackingStrategy _packer +PackingStrategy _packer
+PositionIdStrategy _position_id +PositionIdStrategy _position_id
@@ -385,6 +413,18 @@ classDiagram
+transform(item) Optional[dict] +transform(item) Optional[dict]
+run() +run()
+_flush(domains, shard_idx) +_flush(domains, shard_idx)
+_inject_doc_reset_position_ids(keys, mode, seqs) Dict
+_inject_continuous_position_ids(tensors, mode, seqs) Dict
+_to_tensors(keys) Dict
}
class TokenizeTransform {
+PipelineConfig config
+AutoTokenizer tokenizer
+BaseMaskBuilder mask_builder
+PositionIdStrategy position_strategy
+from_config_file(path) TokenizeTransform
+apply(records) Dict[str, list]
} }
} }
@@ -495,14 +535,13 @@ classDiagram
} }
class GRPOStrategy { class GRPOStrategy {
+nn.Module old_model
+nn.Module ref_model +nn.Module ref_model
+float clip_eps +float clip_eps
+float kl_coef +float kl_coef
+int group_size +int group_size
+str reduction
+int sync_interval
+compute_loss(batch) Tensor +compute_loss(batch) Tensor
+sync_ref_model() +sync_old_model()
} }
class BaseScheduler { class BaseScheduler {
@@ -552,7 +591,7 @@ classDiagram
} }
class GradientClippingCallback { class GradientClippingCallback {
+float max_grad_norm +Optional[float] max_grad_norm
+on_optimizer_step(context) +on_optimizer_step(context)
} }
@@ -1065,11 +1104,18 @@ classDiagram
TrainCallback <|-- MetricCallback TrainCallback <|-- MetricCallback
BaseDataset <|-- SEQDataset BaseDataset <|-- SEQDataset
BaseDataset <|-- SFTDataset BaseDataset <|-- SFTDataset
BaseDataset <|-- DPODataset BaseDataset <|-- RecordDataset
BaseDataset <|-- GRPODataset RecordDataset <|-- DPODataset
RecordDataset <|-- GRPODataset
Store <|-- H5Store Store <|-- H5Store
Store <|-- MmapStore Store <|-- MmapStore
Store <|-- JsonlStore Store <|-- JsonlStore
H5Store --|> Streamable
H5Store --|> Recordable
MmapStore --|> Streamable
MmapStore --|> Recordable
JsonlStore --|> Streamable
JsonlStore --|> Recordable
BaseSamplingStrategy <|-- TemperatureStrategy BaseSamplingStrategy <|-- TemperatureStrategy
BaseSamplingStrategy <|-- TopKStrategy BaseSamplingStrategy <|-- TopKStrategy
BaseSamplingStrategy <|-- TopPStrategy BaseSamplingStrategy <|-- TopPStrategy
@@ -1144,6 +1190,9 @@ classDiagram
BaseDataset o-- Store BaseDataset o-- Store
Pipeline o-- PipelineConfig Pipeline o-- PipelineConfig
Pipeline o-- BaseMaskBuilder Pipeline o-- BaseMaskBuilder
Pipeline o-- AutoTokenizer
TokenizeTransform o-- AutoTokenizer
TokenizeTransform o-- BaseMaskBuilder
%% --- Dependency (uses temporarily) --- %% --- Dependency (uses temporarily) ---
TrainConfig ..> BaseStrategy : selects TrainConfig ..> BaseStrategy : selects
@@ -1187,7 +1236,7 @@ classDiagram
%% --- Association (general usage) --- %% --- Association (general usage) ---
Trainer --> TrainConfig Trainer --> TrainConfig
DPOStrategy --> AutoModel DPOStrategy --> AutoModel
GRPOStrategy --> AutoModel GRPOStrategy --> AutoModel : policy/old/ref
InferenceScheduler --> Task InferenceScheduler --> Task
InferenceScheduler --> TaskStatus InferenceScheduler --> TaskStatus
Task --> TaskStatus Task --> TaskStatus
@@ -1204,8 +1253,8 @@ classDiagram
| Module | Components | Description | | Module | Components | Description |
|--------|------------|-------------| |--------|------------|-------------|
| **astrai.config** | BaseConfig, BaseModelConfig, AutoRegressiveLMConfig, EncoderConfig, ConfigFactory, TrainConfig, PipelineConfig, InputConfig, ProcessingConfig, OutputConfig | Configuration management (to_dict/from_dict, to_file/from_file) | | **astrai.config** | BaseConfig, BaseModelConfig, AutoRegressiveLMConfig, EncoderConfig, ConfigFactory, TrainConfig, PipelineConfig, InputConfig, ProcessingConfig, OutputConfig | Configuration management (to_dict/from_dict, to_file/from_file) |
| **astrai.preprocessing** | BaseMaskBuilder, MaskBuilderFactory, SectionedMaskBuilder, Pipeline, filter_by_length, PackingStrategy, PackingStrategyFactory, PositionIdStrategy, PositionIdStrategyFactory, StoreWriter, StoreWriterFactory | Declarative JSON-driven data preprocessing | | **astrai.preprocessing** | BaseMaskBuilder, MaskBuilderFactory, SectionedMaskBuilder, SingleOutputMaskBuilder, MultiOutputMaskBuilder, Pipeline, TokenizeTransform, filter_by_length, PackingStrategy, PackingStrategyFactory, plan_bfd, PositionIdStrategy, PositionIdStrategyFactory, StoreWriter, StoreWriterFactory, core (shared helpers) | Declarative JSON-driven data preprocessing |
| **astrai.dataset** | BaseDatasetGRPODataset, StoreJsonlStore/MmapStore/H5Store, StoreFactory, ResumableDistributedSampler, DatasetFactory | Dataset loading and management | | **astrai.dataset** | BaseDatasetRecordDatasetDPO/GRPODataset, SEQDataset, SFTDataset, Store, Streamable, Recordable, H5Store, MmapStore, JsonlStore, StoreFactory, ResumableDistributedSampler, DatasetFactory | Dataset loading and management |
| **astrai.serialization** | Checkpoint | Model serialization | | **astrai.serialization** | Checkpoint | Model serialization |
| **astrai.model** | AutoModel, AutoRegressiveLM, EmbeddingEncoder, DecoderBlock, GQA, MLA, MLP, DeepSeekMoE, AttnFactory, FFNFactory, RMSNorm, Linear, RotaryEmbedding, Embedding | Neural network model | | **astrai.model** | AutoModel, AutoRegressiveLM, EmbeddingEncoder, DecoderBlock, GQA, MLA, MLP, DeepSeekMoE, AttnFactory, FFNFactory, RMSNorm, Linear, RotaryEmbedding, Embedding | Neural network model |
| **astrai.tokenize** | AutoTokenizer, ChatTemplate | Tokenizer and chat template | | **astrai.tokenize** | AutoTokenizer, ChatTemplate | Tokenizer and chat template |
@@ -1219,7 +1268,7 @@ classDiagram
| Pattern | Classes | Purpose | | Pattern | Classes | Purpose |
|---------|---------|---------| |---------|---------|---------|
| **Factory** | `AttnFactory`, `FFNFactory`, `StrategyFactory`, `DatasetFactory`, `SchedulerFactory`, `CallbackFactory`, `StoreFactory`, `ConfigFactory`, `ExecutorFactory` | Decorator-based component creation | | **Factory** | `AttnFactory`, `FFNFactory`, `StrategyFactory`, `DatasetFactory`, `SchedulerFactory`, `CallbackFactory`, `StoreFactory`, `ConfigFactory`, `ExecutorFactory`, `MaskBuilderFactory`, `StoreWriterFactory`, `PackingStrategyFactory`, `PositionIdStrategyFactory` | Decorator-based component creation |
| **Registry** | `BaseFactory` | Component registration | | **Registry** | `BaseFactory` | Component registration |
| **Strategy** | `SEQStrategy`, `SFTStrategy`, `DPOStrategy`, `GRPOStrategy` | Training strategy switching | | **Strategy** | `SEQStrategy`, `SFTStrategy`, `DPOStrategy`, `GRPOStrategy` | Training strategy switching |
| **Strategy (Sampling)** | `TemperatureStrategy`, `TopKStrategy`, `TopPStrategy`, `SamplingPipeline` | Composable logit transformations | | **Strategy (Sampling)** | `TemperatureStrategy`, `TopKStrategy`, `TopPStrategy`, `SamplingPipeline` | Composable logit transformations |
@@ -1247,4 +1296,4 @@ classDiagram
10. **AutoModel**: `from_pretrained()` loads `config.json` + `model.safetensors`, `_disable_random_init` replaces `nn.init.*` with no-ops 10. **AutoModel**: `from_pretrained()` loads `config.json` + `model.safetensors`, `_disable_random_init` replaces `nn.init.*` with no-ops
11. **Protocols**: `OptimizerProtocol` / `SchedulerProtocol` — structural subtyping for `AccumOptimizer` / `AccumScheduler` wrappers 11. **Protocols**: `OptimizerProtocol` / `SchedulerProtocol` — structural subtyping for `AccumOptimizer` / `AccumScheduler` wrappers
> Document Update Time: 2026-07-09 > Document Update Time: 2026-07-19
+36 -18
View File
@@ -61,41 +61,59 @@ StoreFactory.create("bin") → MmapStore
StoreFactory.create("jsonl") → JsonlStore StoreFactory.create("jsonl") → JsonlStore
``` ```
**H5Store**: Reads HDF5 files. Tensors are loaded into host memory and normalized into segmented storage. All three inherit `Store` (base, owns `_data`/`_cum`/`_offsets`/`_normalize`) plus the `Streamable` and `Recordable` mixins, so every backend supports both `fetch(begin, end, keys)` (stream) and `fetch_record(index, keys)` (record) APIs.
**MmapStore**: Memory-maps `.bin` files. OS page cache sharing is native — no explicit `share_memory_()` needed. Uses `torch.from_numpy(np.memmap(...))`. **H5Store**: Reads HDF5 files. Tensors are loaded into host memory and normalized into segmented storage. `segments_are_records=True` — each `data_i` dataset is one record.
**JsonlStore**: On-the-fly tokenization of raw JSONL files at load time. Requires a `dataset_config.json` alongside the `.jsonl` files following the same `PipelineConfig` schema with an additional `tokenizer_path` field. **MmapStore**: Memory-maps `.bin` files. OS page cache sharing is native — no explicit `share_memory_()` needed. Uses `torch.from_numpy(np.memmap(...))`. `segments_are_records=False` — bin segments are contiguous streams; record access is driven by `_offsets` (written when `save_bin(..., record_keys=...)` was used at preprocessing time).
All backends normalise tensors into `Store._data[Dict[str, List[Tensor]]]` + `Store._cum[Dict[str, List[int]]]` (cumulative lengths for bisect-based indexing). **JsonlStore**: On-the-fly tokenization of raw JSONL files at load time. Requires a `dataset_config.json` alongside the `.jsonl` files following the same `PipelineConfig` schema with an additional `tokenizer_path` field. Two modes: eager (default, applies `TokenizeTransform` to all records at load) and lazy (`processor=fn` given, defers tokenisation to `fetch_record` — used by DPO/GRPO).
All backends normalise tensors into `Store._data[Dict[str, List[Tensor]]]` + `Store._cum[Dict[str, List[int]]]` (cumulative lengths for bisect-based stream indexing) + `Store._offsets[Dict[str, List[int]]]` (per-record offsets for record-mode indexing). Nested keys (GRPO `responses`/`masks` as `List[List[Tensor]]`) are stored as-is and excluded from both bookkeepings — they are only accessed record-by-record.
## Data Keys by Training Type ## Data Keys by Training Type
| Type | Storage Keys | | Type | Storage Keys | Access Mode |
|------|-------------| |------|-------------|-------------|
| `seq` | `sequence` (→ input_ids, target_ids via offset-by-1) | | `seq` | `sequence` (→ input_ids, target_ids via offset-by-1) | stream (`fetch`) |
| `sft` | `sequence`, `loss_mask`, `position_ids` | | `sft` | `sequence`, `loss_mask`, `position_ids` | stream (`fetch`) |
| `dpo` | `chosen`, `rejected`, `chosen_mask`, `rejected_mask` | | `dpo` | `chosen`, `rejected`, `chosen_mask`, `rejected_mask` | record (`fetch_record`) |
| `grpo` | `prompts`, `responses`, `masks`, `rewards` | | `grpo` | `prompts`, `responses`, `masks`, `rewards` | record (`fetch_record`) |
## Dataset Architecture ## Dataset Architecture
``` ```
DatasetFactory.load(train_type, load_path, window_size, stride=None, storage_type=None) DatasetFactory.load(train_type, load_path, window_size, stride=None,
storage_type=None, tokenizer_path=None,
max_len=2048, store=None)
→ BaseDataset.load(load_path, storage_type=None) → BaseDataset.load(load_path, storage_type=None)
→ detect_format(load_path) → detect_format(load_path)
→ StoreFactory.create(storage_type) → StoreFactory.create(storage_type)
→ Store.load(load_path) → Store.load(load_path)
→ _normalize(raw) # base Store, shared by both backends → _normalize(raw) # base Store, shared by both backends
→ Store._data[Dict[str, List[Tensor]]] + _cum[Dict[str, List[int]]] → Store._data[Dict[str, List[Tensor]]]
→ BaseDataset.__getitem__(idx) + _cum[Dict[str, List[int]]] (stream mode)
→ get_index(idx) → [begin, end) + _offsets[Dict[str, List[int]]] (record mode)
→ Store.fetch(begin, end, keys) → Tensor / Dict[str, Tensor]
Stream datasets (SEQ/SFT):
BaseDataset.__getitem__(idx)
→ get_index(idx) → [begin, end)
→ Store.fetch(begin, end, keys) → Tensor / Dict[str, Tensor]
Record datasets (DPO/GRPO via RecordDataset):
RecordDataset.__getitem__(idx)
→ Store.fetch_record(idx, keys) → Tensor / Dict[str, Tensor]
``` ```
`window_size` = max input length, `stride` = step between consecutive samples (defaults to `window_size`, optional). `storage_type` defaults to `None` (auto-detect via `detect_format`). Class hierarchy: `BaseDataset``SEQDataset` / `SFTDataset` (stream); `BaseDataset``RecordDataset``DPODataset` / `GRPODataset` (record).
`Store.fetch(begin, end, keys)` accepts a single key (`str`) returning a `Tensor`, or a list of keys returning `Dict[str, Tensor]`. Internally uses `bisect` across multi-segment tensors. Raises `RuntimeError("Store not loaded")` if called before `load()`. `window_size` = max input length, `stride` = step between consecutive samples (defaults to `window_size`, optional). Only meaningful for stream datasets — record datasets ignore both. `storage_type` defaults to `None` (auto-detect via `detect_format`).
`tokenizer_path` triggers lazy on-the-fly tokenisation for record datasets on raw JSONL (DPO builds a `dpo_processor`; SEQ/SFT/pre-tokenised backends ignore it). `store` (pre-built `Store`) bypasses `load_path`/`storage_type`/`tokenizer_path` entirely — the caller controls Store construction.
`Store.fetch(begin, end, keys)` (stream mode, on `Streamable`): accepts a single key (`str`) returning a `Tensor`, or a list of keys returning `Dict[str, Tensor]`. Internally uses `bisect` across multi-segment tensors. Raises `RuntimeError("Store not loaded")` if called before `load()`.
`Store.fetch_record(index, keys)` (record mode, on `Recordable`): same key API. Uses `_offsets[key]` when present (bin layout with per-record offsets), otherwise indexes `_data[key]` directly (H5/JSONL where each segment is one record).
## Sampler ## Sampler
@@ -109,4 +127,4 @@ DatasetFactory.load(train_type, load_path, window_size, stride=None, storage_typ
Standard PyTorch `DataLoader` with configurable `batch_size`, `num_workers`, `pin_memory`, `prefetch_factor`. Sampler produces indices; dataloader fetches tensor batches via `__getitem__`. Standard PyTorch `DataLoader` with configurable `batch_size`, `num_workers`, `pin_memory`, `prefetch_factor`. Sampler produces indices; dataloader fetches tensor batches via `__getitem__`.
> Document Update Time: 2026-07-09 > Document Update Time: 2026-07-19
+2 -2
View File
@@ -26,7 +26,7 @@
|-----------|-------------|---------| |-----------|-------------|---------|
| `--warmup_ratio` | Fraction of total steps used for LR warmup | 0.05 | | `--warmup_ratio` | Fraction of total steps used for LR warmup | 0.05 |
| `--max_lr` | Maximum learning rate (cosine decay after warmup) | 3e-4 | | `--max_lr` | Maximum learning rate (cosine decay after warmup) | 3e-4 |
| `--max_grad_norm` | Maximum gradient norm for clipping | 1.0 | | `--max_grad_norm` | Maximum gradient norm for clipping (None disables) | None |
### Optimizer (MuonMix) ### Optimizer (MuonMix)
@@ -201,4 +201,4 @@ See [Preprocessing Guide](preprocessing.md) for config file format and examples.
--- ---
> Document Update Time: 2026-07-09 > Document Update Time: 2026-07-19
+1 -1
View File
@@ -1,6 +1,6 @@
# Preprocessing Pipeline # Preprocessing Pipeline
Declarative JSON-driven data preprocessing. One `SectionedMaskBuilder` handles all formats via `input.sections` (single-output) or `input.sources` (multi-output). Declarative JSON-driven data preprocessing. `MaskBuilderFactory` supports three registered builders: `"single"` (single-output via `input.sections`), `"multi"` (multi-output via `input.sources`), and `"sectioned"` (façade dispatching to `single` or `multi` based on config).
## Contents ## Contents
+16 -6
View File
@@ -86,7 +86,7 @@ on_train_end
| `on_error` | On exception during training | `CheckpointCallback`, `MetricCallback` | | `on_error` | On exception during training | `CheckpointCallback`, `MetricCallback` |
| `on_train_end` | Training ends (always via finally) | `CheckpointCallback`, `MetricCallback`, `GradientCheckpointingCallback` | | `on_train_end` | Training ends (always via finally) | `CheckpointCallback`, `MetricCallback`, `GradientCheckpointingCallback` |
Default callbacks (in order): `gradient_checkpointing` (activation checkpointing, optional), `checkpoint` (safetensors, rank-0), `metric` (JSONL + validation, rank-0), `progress_bar` (tqdm), `gradient_clipping`. Default callbacks (in order): `gradient_checkpointing` (activation checkpointing, optional), `checkpoint` (safetensors, rank-0), `metric` (JSONL + validation, rank-0), `progress_bar` (tqdm), `gradient_clipping` (always registered; computes grad norm, clips only when `max_grad_norm` is not `None`).
## Strategies ## Strategies
@@ -118,21 +118,31 @@ $$
L_{\text{DPO}} = -\mathbb{E}\left[\log\sigma\left(\beta\log\frac{\pi_\theta(y_w\mid x)}{\pi_{\text{ref}}(y_w\mid x)} - \beta\log\frac{\pi_\theta(y_l\mid x)}{\pi_{\text{ref}}(y_l\mid x)}\right)\right] L_{\text{DPO}} = -\mathbb{E}\left[\log\sigma\left(\beta\log\frac{\pi_\theta(y_w\mid x)}{\pi_{\text{ref}}(y_w\mid x)} - \beta\log\frac{\pi_\theta(y_l\mid x)}{\pi_{\text{ref}}(y_l\mid x)}\right)\right]
$$ $$
Parameters: `beta=0.1`, `reduction="mean"`. Keys: `chosen`, `rejected`, `chosen_mask`, `rejected_mask`. Parameters: `beta=0.1`, `reduction="sum"`. Keys: `chosen`, `rejected`, `chosen_mask`, `rejected_mask`.
### GRPO (Group Relative Policy Optimization) ### GRPO (Group Relative Policy Optimization)
On-policy PPO with group-normalized advantages: Token-level PPO with group-normalized advantages. Advantages are derived from
scalar per-response rewards, group-normalized, and broadcast across all response
tokens. Only response tokens contribute to the loss (prompt tokens are masked
out):
$$ $$
\text{Advantage}_i = \frac{r_i - \mu}{\sigma + \epsilon} \text{Advantage}_i = \frac{r_i - \mu}{\sigma + \epsilon}
$$ $$
$$ $$
L_{\text{GRPO}} = -\mathbb{E}\left[\min\left(\frac{\pi_\theta}{\pi_{\text{ref}}}A,\; \text{clip}\left(\frac{\pi_\theta}{\pi_{\text{ref}}}, 1-\epsilon, 1+\epsilon\right)A\right)\right] + \lambda \cdot \mathbb{E}\left[(\log\pi_\theta - \log\pi_{\text{ref}})^2\right] L_{\text{GRPO}} = -\mathbb{E}_t\left[\min\left(\rho_t A,\; \text{clip}\left(\rho_t, 1-\epsilon, 1+\epsilon\right)A\right)\right] + \lambda \cdot \mathbb{E}_t\left[\frac{\pi_{\text{ref}}}{\pi_\theta} - \log\frac{\pi_{\text{ref}}}{\pi_\theta} - 1\right]
$$ $$
Parameters: `group_size=4`, `clip_eps=0.2`, `kl_coef=0.01`, `sync_interval=200`, `reduction="mean"`. where $\rho_t = \pi_\theta(a_t|s_t) / \pi_{\text{old}}(a_t|s_t)$ is the
per-token importance sampling ratio against the behaviour policy
(`old_model`, synced externally between data-generation rounds) and the
expectations are over valid response tokens. The KL term regularises
$\pi_\theta$ towards a frozen reference model (`ref_model`, typically
the SFT checkpoint).
Parameters: `group_size=4`, `clip_eps=0.2`, `kl_coef=0.01`. External sync of `old_model` weights via `sync_old_model()` between data-generation rounds.
Keys: `prompts`, `responses`, `masks`, `rewards`. Keys: `prompts`, `responses`, `masks`, `rewards`.
@@ -212,4 +222,4 @@ nohup python scripts/tools/train.py \
Full parameter reference at [params.md](params.md). Full parameter reference at [params.md](params.md).
> Document Update Time: 2026-07-09 > Document Update Time: 2026-07-19
+2 -2
View File
@@ -12,7 +12,7 @@ from astrai.config import (
from astrai.dataset import ( from astrai.dataset import (
BaseDataset, BaseDataset,
DatasetFactory, DatasetFactory,
ResumableDistributedSampler, RDSampler,
Store, Store,
StoreFactory, StoreFactory,
) )
@@ -77,7 +77,7 @@ __all__ = [
"Pipeline", "Pipeline",
"PipelineConfig", "PipelineConfig",
"ProtocolHandler", "ProtocolHandler",
"ResumableDistributedSampler", "RDSampler",
"SamplingPipeline", "SamplingPipeline",
"SchedulerFactory", "SchedulerFactory",
"Store", "Store",
+7 -2
View File
@@ -37,8 +37,9 @@ class TrainConfig(BaseConfig):
grad_accum_steps: int = field( grad_accum_steps: int = field(
default=1, metadata={"help": "Number of iterations between steps."} default=1, metadata={"help": "Number of iterations between steps."}
) )
max_grad_norm: float = field( max_grad_norm: Optional[float] = field(
default=1.0, metadata={"help": "Maximum gradient norm."} default=None,
metadata={"help": "Maximum gradient norm. None disables clipping."},
) )
gradient_checkpointing_modules: List[str] = field( gradient_checkpointing_modules: List[str] = field(
default_factory=list, default_factory=list,
@@ -87,6 +88,10 @@ class TrainConfig(BaseConfig):
pin_memory: bool = field( pin_memory: bool = field(
default=False, metadata={"help": "Pin memory for dataloader."} default=False, metadata={"help": "Pin memory for dataloader."}
) )
collate_fn: Optional[Callable[[List[Any]], Any]] = field(
default=None,
metadata={"help": "Collate function for dataloader (e.g. dpo_collate_fn)."},
)
# distributed training # distributed training
nprocs: int = field( nprocs: int = field(
+10 -2
View File
@@ -1,14 +1,18 @@
from astrai.dataset.dataset import ( from astrai.dataset.dataset import (
BaseDataset, BaseDataset,
DatasetFactory, DatasetFactory,
dpo_collate_fn,
grpo_collate_fn,
) )
from astrai.dataset.sampler import ResumableDistributedSampler from astrai.dataset.sampler import RDSampler
from astrai.dataset.storage import ( from astrai.dataset.storage import (
H5Store, H5Store,
JsonlStore, JsonlStore,
MmapStore, MmapStore,
Recordable,
Store, Store,
StoreFactory, StoreFactory,
Streamable,
detect_format, detect_format,
) )
from astrai.serialization import ( from astrai.serialization import (
@@ -21,7 +25,11 @@ from astrai.serialization import (
__all__ = [ __all__ = [
"BaseDataset", "BaseDataset",
"DatasetFactory", "DatasetFactory",
"dpo_collate_fn",
"grpo_collate_fn",
"Store", "Store",
"Streamable",
"Recordable",
"StoreFactory", "StoreFactory",
"H5Store", "H5Store",
"MmapStore", "MmapStore",
@@ -31,5 +39,5 @@ __all__ = [
"load_h5", "load_h5",
"save_bin", "save_bin",
"load_bin", "load_bin",
"ResumableDistributedSampler", "RDSampler",
] ]
+400 -183
View File
@@ -1,7 +1,31 @@
"""Dataset implementations with factory pattern for training.""" """Dataset implementations for training.
Composition over inheritance — every dataset is a thin wrapper that
binds a :class:`Store` to a particular train-type's key mapping. All
sample-id → token/record indexing lives on the Store; datasets never
know about window/stride math or segment layouts.
Class hierarchy:
BaseDataset (ABC) — holds a Store, exposes __len__/keys,
overrides __getitem__
├── SEQDataset — next-token prediction (stream)
├── SFTDataset — loss-mask + position_ids (stream)
├── DPODataset — chosen/rejected pairs (record)
└── GRPODataset — prompt + response group (record)
``DatasetFactory.load(train_type, load_path, window_size, stride, …)``
builds the Store (auto-detecting format) before constructing the
matching dataset. Passing ``store=`` skips Store construction.
When a record dataset (DPO) reads from raw JSONL, a *processor*
function (pure ``record -> Dict[str, Tensor]``) is forwarded to
:class:`JsonlStore` so tokenisation happens on the fly.
"""
from abc import ABC, abstractmethod from abc import ABC, abstractmethod
from typing import Dict, List, Optional from functools import partial
from typing import Callable, Dict, List, Optional
import torch import torch
from torch import Tensor from torch import Tensor
@@ -13,202 +37,389 @@ from astrai.dataset.storage import (
detect_format, detect_format,
) )
from astrai.factory import BaseFactory from astrai.factory import BaseFactory
from astrai.tokenize import AutoTokenizer
def dpo_tokenize(
record: dict,
tokenizer,
max_len: int = 2048,
) -> Optional[dict]:
"""Tokenize one DPO record into chosen/rejected + masks.
Applies the tokenizer's chat template so token sequences match the
SFT checkpoint's format. Prompt is rendered with
``add_generation_prompt=True``; chosen/rejected are appended as a
single assistant turn.
Accepts:
- Flat: ``{"prompt": str, "chosen": str, "rejected": str}``
- Conv: ``{"prompt": [{role, content}, ...], "chosen": [...], ...}``
- Legacy: ``{"input": str, "chosen": str, "rejected": str}``
No packing, no ``position_ids`` — DPO sequences are independent.
"""
prompt = record.get("prompt") or record.get("input")
chosen = record.get("chosen")
rejected = record.get("rejected")
if prompt is None or chosen is None or rejected is None:
return None
prompt_messages = _to_messages(prompt)
chosen_text = _extract_text(chosen)
rejected_text = _extract_text(rejected)
if chosen_text is None or rejected_text is None:
return None
chosen_messages = prompt_messages + [{"role": "assistant", "content": chosen_text}]
rejected_messages = prompt_messages + [
{"role": "assistant", "content": rejected_text}
]
prompt_ids = tokenizer.apply_chat_template(
prompt_messages, tokenize=True, add_generation_prompt=True
)
ch_ids = tokenizer.apply_chat_template(
chosen_messages, tokenize=True, add_generation_prompt=False
)
re_ids = tokenizer.apply_chat_template(
rejected_messages, tokenize=True, add_generation_prompt=False
)
full_ch = ch_ids[:max_len]
full_re = re_ids[:max_len]
prompt_len = min(len(prompt_ids), max_len)
ch_mask = [0] * prompt_len + [1] * max(0, len(full_ch) - prompt_len)
ch_mask = ch_mask[:max_len]
re_mask = [0] * prompt_len + [1] * max(0, len(full_re) - prompt_len)
re_mask = re_mask[:max_len]
return {
"chosen": full_ch,
"rejected": full_re,
"chosen_mask": ch_mask,
"rejected_mask": re_mask,
}
def _to_messages(value) -> list:
"""Accept str or conversation list; return message list."""
if isinstance(value, str):
return [{"role": "user", "content": value}]
if isinstance(value, list):
return value
return [{"role": "user", "content": str(value)}]
def _extract_text(value) -> Optional[str]:
"""Accept str or conversation list; return plain text."""
if value is None:
return None
if isinstance(value, str):
return value
if isinstance(value, list):
return "".join(m.get("content", "") for m in value if isinstance(m, dict))
return None
def dpo_processor(
record: dict,
tokenizer,
max_len: int = 2048,
) -> Dict[str, Tensor]:
"""DPO processor: wraps :func:`dpo_tokenize` and returns tensors."""
result = dpo_tokenize(record, tokenizer, max_len=max_len)
if result is None:
raise ValueError(f"Malformed DPO record: {list(record.keys())}")
return {
"chosen": torch.tensor(result["chosen"], dtype=torch.int32),
"rejected": torch.tensor(result["rejected"], dtype=torch.int32),
"chosen_mask": torch.tensor(result["chosen_mask"], dtype=torch.bool),
"rejected_mask": torch.tensor(result["rejected_mask"], dtype=torch.bool),
}
def dpo_collate_fn(batch: List[Dict[str, Tensor]]) -> Dict[str, Tensor]:
"""Collate variable-length DPO samples into padded 2-D tensors.
Input: list of dicts, each with:
- chosen: [C_i]
- rejected: [R_i]
- chosen_mask: [C_i]
- rejected_mask: [R_i]
Output (padded to the max length across chosen/rejected within the batch):
- chosen: [B, S_max]
- rejected: [B, S_max]
- chosen_mask: [B, S_max]
- rejected_mask: [B, S_max]
"""
B = len(batch)
S_max = max(b["chosen"].size(0) for b in batch)
S_max = max(S_max, max(b["rejected"].size(0) for b in batch))
chosen = torch.zeros(B, S_max, dtype=torch.long)
rejected = torch.zeros(B, S_max, dtype=torch.long)
chosen_mask = torch.zeros(B, S_max, dtype=torch.bool)
rejected_mask = torch.zeros(B, S_max, dtype=torch.bool)
for i, b in enumerate(batch):
c_len = b["chosen"].size(0)
r_len = b["rejected"].size(0)
chosen[i, :c_len] = b["chosen"]
rejected[i, :r_len] = b["rejected"]
chosen_mask[i, :c_len] = b["chosen_mask"]
rejected_mask[i, :r_len] = b["rejected_mask"]
return {
"chosen": chosen,
"rejected": rejected,
"chosen_mask": chosen_mask,
"rejected_mask": rejected_mask,
}
def grpo_collate_fn(batch: List[Dict[str, Tensor]]) -> Dict[str, Tensor]:
"""Collate variable-length GRPO samples into padded 3-D tensors.
Input: list of dicts, each with:
- prompts: [P_i]
- responses: list of G tensors, each [R_ij]
- masks: list of G tensors, each [R_ij]
- rewards: [G]
Output:
- prompts: [B, P_max]
- responses: [B, G, R_max]
- masks: [B, G, R_max]
- rewards: [B, G]
"""
B = len(batch)
G = len(batch[0]["responses"])
P_max = max(b["prompts"].size(0) for b in batch)
R_max = max(r.size(0) for b in batch for r in b["responses"])
prompts = torch.zeros(B, P_max, dtype=torch.long)
responses = torch.zeros(B, G, R_max, dtype=torch.long)
masks = torch.zeros(B, G, R_max, dtype=torch.bool)
rewards = torch.zeros(B, G, dtype=torch.float32)
for i, b in enumerate(batch):
p_len = b["prompts"].size(0)
prompts[i, :p_len] = b["prompts"]
rewards[i, : b["rewards"].size(0)] = b["rewards"]
for g in range(min(G, len(b["responses"]))):
r_len = b["responses"][g].size(0)
responses[i, g, :r_len] = b["responses"][g]
if g < len(b["masks"]):
masks[i, g, :r_len] = b["masks"][g]
return {
"prompts": prompts,
"responses": responses,
"masks": masks,
"rewards": rewards,
}
def validate_keys(store: Store, required: List[str]) -> None:
"""Raise ``KeyError`` if *store* is missing any *required* key."""
if not required:
return
actual = set(store.keys)
missing = [k for k in required if k not in actual]
if missing:
raise KeyError(
f"Store at {getattr(store, '_load_path', '?')} is missing required "
f"keys {missing}; available keys are {sorted(actual)}."
)
class BaseDataset(Dataset, ABC): class BaseDataset(Dataset, ABC):
"""Abstract base class for all dataset types. """Abstract base class for dataset types.
Implements common functionality for window-based data fetching. Holds a :class:`Store`. All sample-id indexing is delegated to the
Uses a storage abstraction for format-agnostic data loading. store — this class exposes ``__len__`` as ``len(store)`` and the
``keys`` property as ``store.keys``. Subclasses implement
``__getitem__`` with the train-type-specific key mapping and any
training-only index arithmetic (e.g. the next-token ``+1`` shift).
""" """
def __init__(self, window_size: int, stride: int): required_keys: List[str] = []
def __init__(self, store: Store):
super().__init__() super().__init__()
self.window_size = window_size self.store: Store = store
self.stride = stride validate_keys(store, self.required_keys)
self.storage: Optional[Store] = None
@property def __len__(self) -> int:
def required_keys(self) -> List[str]: return len(self.store)
"""Return required storage keys for this dataset type.
Subclasses should override to specify expected keys.
"""
return []
def _validate_keys(self):
if not self.required_keys:
return
actual_keys = set(self.storage.keys)
missing = [k for k in self.required_keys if k not in actual_keys]
if missing:
raise KeyError(
f"Dataset {type(self).__name__} requires keys {self.required_keys}, "
f"but storage at {self._load_path} only has {sorted(actual_keys)}. "
f"Missing: {missing}"
)
def load(self, load_path: str, storage_type: Optional[str] = None, **kwargs):
"""Load dataset from the given path.
Auto-detects the storage format if not specified.
Args:
load_path: Path to the data directory or file
storage_type: Force a specific storage type ("h5", "bin", "jsonl"),
or None for auto-detection
**kwargs: Extra arguments forwarded to the store constructor and
to ``store.load()``.
Raises:
KeyError: If the loaded storage is missing required keys.
"""
if storage_type is None:
storage_type = detect_format(load_path)
self.storage = StoreFactory.create(storage_type, **kwargs)
self._load_path = load_path
self.storage.load(load_path, **kwargs)
self._validate_keys()
@property
def count(self) -> int:
"""Return the total number of raw elements (tokens) in the dataset."""
if self.storage is None:
return 0
return len(self.storage)
@property @property
def keys(self) -> List[str]: def keys(self) -> List[str]:
"""Return the available data keys.""" return self.store.keys
if self.storage is None:
return []
return self.storage.keys
def get_index(self, index: int) -> tuple: @property
"""Calculate begin and end indices for a sample. def token_count(self) -> int:
return self.store.token_count
Args:
index: Sample index
Returns:
Tuple of (begin_idx, end_idx)
"""
if self.storage is None:
raise RuntimeError("Dataset not loaded, call load() first")
total = len(self.storage)
if total <= self.window_size:
raise ValueError(
f"Data too short: {total} tokens <= window_size {self.window_size}"
)
begin_idx = min(index * self.stride, total - 1 - self.window_size)
end_idx = min(begin_idx + self.window_size, total - 1)
return begin_idx, end_idx
@abstractmethod @abstractmethod
def __getitem__(self, index: int) -> Dict[str, Tensor]: def __getitem__(self, index: int) -> Dict[str, Tensor]:
"""Get a single sample by index.
Must be implemented by subclasses.
"""
raise NotImplementedError raise NotImplementedError
def __len__(self) -> int:
if self.storage is None:
return 0
total = len(self.storage)
if total <= self.window_size:
return 0
return (total - 1 - self.window_size) // self.stride + 1
class DatasetFactory(BaseFactory["BaseDataset"]): class DatasetFactory(BaseFactory["BaseDataset"]):
"""Factory class for creating dataset instances. """Factory for creating dataset instances by train-type.
Supports decorator-based registration for extensible dataset types. Use :meth:`DatasetFactory.register("custom")` to register new
All default dataset types (seq, sft, dpo, grpo) are registered automatically dataset classes; they must inherit from :class:`BaseDataset`.
when their classes are defined with the decorator.
Example usage:
@DatasetFactory.register("custom")
class CustomDataset(BaseDataset):
...
dataset = DatasetFactory.create("custom", window_size, stride)
""" """
@classmethod @classmethod
def load( def load(
cls, cls,
train_type: str, train_type: str,
load_path: str, load_path: Optional[str] = None,
window_size: int, window_size: int = 0,
stride: Optional[int] = None, stride: Optional[int] = None,
storage_type: Optional[str] = None, storage_type: Optional[str] = None,
tokenizer_path: Optional[str] = None,
max_len: int = 2048,
store: Optional[Store] = None,
**kwargs, **kwargs,
) -> "BaseDataset": ) -> "BaseDataset":
"""Create and load a dataset in one step. """Create and load a dataset in one step.
Two entry points:
- **store given**: bind it directly — the caller fully controls
Store construction and processor setup. *load_path*,
*storage_type*, *tokenizer_path*, *window_size*, *stride* are
ignored.
- **store is None**: build a Store from *load_path*, auto-detecting
format and constructing a processor when *tokenizer_path* is
given for a record dataset on JSONL.
Args: Args:
train_type: Type of training dataset train_type: Registered dataset name ("seq", "sft", "dpo",
load_path: Path to the data file "grpo", …).
window_size: Window size for data sampling load_path: Path to the data file or directory (ignored if
stride: Stride between consecutive samples (default: same as window_size) *store* is given).
storage_type: Storage type ("h5", "bin", "jsonl") or None for auto-detection window_size: Stream window length — only meaningful for
**kwargs: Extra arguments forwarded to ``dataset.load()``. stream datasets (SEQ/SFT). Record datasets ignore it.
stride: Stride between consecutive stream samples
(default: same as *window_size*).
storage_type: Storage backend ("h5", "bin", "jsonl") or
None for auto-detection.
tokenizer_path: Path to tokenizer for lazy JSONL
tokenisation (record datasets only).
max_len: Max sequence length forwarded to processors.
store: Pre-built, already-loaded Store instance.
**kwargs: Extra arguments forwarded to ``store.load()``.
Returns: Returns:
Loaded dataset instance Loaded dataset instance.
""" """
if store is not None:
return cls.create(train_type, store=store)
if load_path is None:
raise ValueError("Either load_path or store must be provided")
if storage_type is None:
storage_type = detect_format(load_path)
if stride is None: if stride is None:
stride = window_size stride = window_size
dataset = cls.create(train_type, window_size, stride) processor = cls._maybe_build_processor(
dataset.load(load_path, storage_type=storage_type, **kwargs) train_type, storage_type, tokenizer_path, max_len
)
return dataset store_window = cls._store_window_for(train_type, window_size)
store = StoreFactory.create(
storage_type,
window_size=store_window,
stride=stride if stride else store_window,
)
if processor is not None:
store.load(load_path, processor=processor, **kwargs)
else:
store.load(load_path, **kwargs)
return cls.create(train_type, store=store)
@staticmethod
def _store_window_for(train_type: str, window_size: int) -> int:
"""Stream datasets consume ``window_size``; record datasets ignore it.
Record datasets (dpo/grpo) treat each record as an independent
training unit and never window, so the store is built with
``window_size=0`` and ``len(store)`` returns the record count.
"""
if train_type in ("seq", "sft"):
return window_size
return 0
@staticmethod
def _maybe_build_processor(
train_type: str,
storage_type: str,
tokenizer_path: Optional[str],
max_len: int,
) -> Optional[Callable[[dict], Dict[str, Tensor]]]:
"""Build an on-the-fly tokenisation processor if applicable.
Only raw JSONL + record datasets (DPO/GRPO) need a processor;
pre-tokenised backends (H5/bin) and stream datasets (SEQ/SFT)
return ``None`` so no tokenizer is loaded.
"""
if tokenizer_path is None or storage_type != "jsonl":
return None
if train_type == "dpo":
tokenizer = AutoTokenizer.from_pretrained(tokenizer_path)
return partial(dpo_processor, tokenizer=tokenizer, max_len=max_len)
return None
@DatasetFactory.register("seq") @DatasetFactory.register("seq")
class SEQDataset(BaseDataset): class SEQDataset(BaseDataset):
"""Dataset for sequential next-token prediction training.""" """Dataset for sequential next-token prediction training.
@property Stream mode: ``store.fetch(begin, end, "sequence")`` returns the
def required_keys(self) -> List[str]: input window; the +1 shifted call returns the next-token target.
return ["sequence"] """
def _fetch_data(self, begin_idx: int, end_idx: int) -> Tensor: required_keys = ["sequence"]
return self.storage.fetch(begin_idx, end_idx, "sequence")
def __getitem__(self, index): def __getitem__(self, index: int):
begin_idx, end_idx = self.get_index(index) begin, end = self.store.sample_window(index)
x = self.store.fetch(begin, end, "sequence")
x = self._fetch_data(begin_idx, end_idx).to(dtype=torch.long) y = self.store.fetch(begin + 1, end + 1, "sequence")
y = self._fetch_data(begin_idx + 1, end_idx + 1).to(dtype=torch.long) return {
"input_ids": x.to(dtype=torch.long),
return {"input_ids": x, "target_ids": y} "target_ids": y.to(dtype=torch.long),
}
@DatasetFactory.register("sft") @DatasetFactory.register("sft")
class SFTDataset(BaseDataset): class SFTDataset(BaseDataset):
"""Dataset for supervised fine-tuning with loss masking.""" """Dataset for supervised fine-tuning with loss masking.
@property Stream mode: ``sequence``/``loss_mask``/``position_ids`` are sliced
def required_keys(self) -> List[str]: to the window. ``loss_mask`` and ``target_ids`` use the +1 shifted
return ["sequence", "loss_mask", "position_ids"] slice so they align with the predicted positions.
"""
def _fetch_data(self, begin_idx: int, end_idx: int, key: str) -> Tensor: required_keys = ["sequence", "loss_mask", "position_ids"]
return self.storage.fetch(begin_idx, end_idx, key)
def __getitem__(self, index):
begin_idx, end_idx = self.get_index(index)
x = self._fetch_data(begin_idx, end_idx, "sequence")
y = self._fetch_data(begin_idx + 1, end_idx + 1, "sequence")
position_ids = self._fetch_data(begin_idx, end_idx, "position_ids")
loss_mask = self._fetch_data(begin_idx + 1, end_idx + 1, "loss_mask")
def __getitem__(self, index: int):
begin, end = self.store.sample_window(index)
x = self.store.fetch(begin, end, "sequence")
y = self.store.fetch(begin + 1, end + 1, "sequence")
position_ids = self.store.fetch(begin, end, "position_ids")
loss_mask = self.store.fetch(begin + 1, end + 1, "loss_mask")
return { return {
"input_ids": x.to(dtype=torch.long), "input_ids": x.to(dtype=torch.long),
"target_ids": y.to(dtype=torch.long), "target_ids": y.to(dtype=torch.long),
@@ -219,59 +430,65 @@ class SFTDataset(BaseDataset):
@DatasetFactory.register("dpo") @DatasetFactory.register("dpo")
class DPODataset(BaseDataset): class DPODataset(BaseDataset):
"""Dataset for Direct Preference Optimization training.""" """Record-structured dataset for Direct Preference Optimization.
@property Each sample is one preference pair (chosen + rejected) and is an
def required_keys(self) -> List[str]: independent training unit — no windowing, stride, or cross-record
return ["chosen", "rejected", "chosen_mask", "rejected_mask"] concatenation. This keeps each sequence self-contained so attention
never leaks across preference pairs.
def _fetch_data(self, begin_idx: int, end_idx: int, key: str) -> Tensor: Two loading paths (handled by :class:`DatasetFactory`):
return self.storage.fetch(begin_idx, end_idx, key)
def __getitem__(self, index: int): - **Pre-tokenized** (H5/bin): ``store.load(path)`` reads per-record
begin_idx, end_idx = self.get_index(index) tensors; ``__getitem__`` returns them directly.
- **Raw JSONL** (``tokenizer_path=...``): builds a lazy processor
via :func:`dpo_processor` that tokenises on the fly — no packing,
no ``position_ids``.
"""
chosen = self._fetch_data(begin_idx, end_idx, "chosen").to(dtype=torch.long) required_keys = ["chosen", "rejected", "chosen_mask", "rejected_mask"]
rejected = self._fetch_data(begin_idx, end_idx, "rejected").to(dtype=torch.long)
chosen_mask = self._fetch_data(begin_idx, end_idx, "chosen_mask").to(
dtype=torch.bool
)
rejected_mask = self._fetch_data(begin_idx, end_idx, "rejected_mask").to(
dtype=torch.bool
)
def make_processor(self, tokenizer, max_len: int):
return partial(dpo_processor, tokenizer=tokenizer, max_len=max_len)
def __getitem__(self, index: int) -> Dict[str, Tensor]:
return { return {
"chosen": chosen, "chosen": self.store.fetch_record(index, "chosen").to(dtype=torch.long),
"rejected": rejected, "rejected": self.store.fetch_record(index, "rejected").to(dtype=torch.long),
"chosen_mask": chosen_mask, "chosen_mask": self.store.fetch_record(index, "chosen_mask").to(
"rejected_mask": rejected_mask, dtype=torch.bool
),
"rejected_mask": self.store.fetch_record(index, "rejected_mask").to(
dtype=torch.bool
),
} }
@DatasetFactory.register("grpo") @DatasetFactory.register("grpo")
class GRPODataset(BaseDataset): class GRPODataset(BaseDataset):
"""Dataset for Group Relative Policy Optimization training.""" """Dataset for offline Group Relative Policy Optimization.
@property Each sample is one prompt with its group of responses and scalar
def required_keys(self) -> List[str]: rewards — an independent training unit with no windowing or stride.
return ["prompts", "responses", "masks", "rewards"]
def _fetch_data(self, begin_idx: int, end_idx: int, key: str) -> Tensor: Expected storage layout (produced by JsonlStore or pre-tokenized):
return self.storage.fetch(begin_idx, end_idx, key)
- ``prompts``: List[Tensor] — one 1-D token tensor per record
- ``responses``: List[List[Tensor]] — G response tensors per record
- ``masks``: List[List[Tensor]] — G mask tensors per record
- ``rewards``: List[Tensor] — one 1-D float tensor (len G) per record
"""
required_keys = ["prompts", "responses", "masks", "rewards"]
def __getitem__(self, index: int) -> Dict[str, Tensor]: def __getitem__(self, index: int) -> Dict[str, Tensor]:
begin_idx, end_idx = self.get_index(index) prompts = self.store.fetch_record(index, "prompts")
responses = self.store.fetch_record(index, "responses")
prompts = self._fetch_data(begin_idx, end_idx, "prompts").to(dtype=torch.long) masks = self.store.fetch_record(index, "masks")
responses = self._fetch_data(begin_idx, end_idx, "responses").to( rewards = self.store.fetch_record(index, "rewards")
dtype=torch.long
)
masks = self._fetch_data(begin_idx, end_idx, "masks").to(dtype=torch.bool)
rewards = self._fetch_data(begin_idx, end_idx, "rewards")
return { return {
"prompts": prompts, "prompts": prompts.to(dtype=torch.long),
"responses": responses, "responses": [r.to(dtype=torch.long) for r in responses],
"masks": masks, "masks": [m.to(dtype=torch.bool) for m in masks],
"rewards": rewards, "rewards": rewards.to(dtype=torch.float32),
} }
+9 -1
View File
@@ -5,7 +5,15 @@ import torch.distributed as dist
from torch.utils.data import Dataset, Sampler from torch.utils.data import Dataset, Sampler
class ResumableDistributedSampler(Sampler[int]): class RDSampler(Sampler[int]):
"""Resumable Distributed Sampler.
A distributed sampler that supports checkpoint-based resume: iteration
state (epoch, position) is tracked so training can continue from the
exact sample after a restart. Shards the dataset across
``dist.world_size`` replicas with optional shuffling.
"""
def __init__( def __init__(
self, self,
data_source: Dataset, data_source: Dataset,
+498 -174
View File
@@ -1,20 +1,48 @@
"""Storage backends for different data formats. """Storage backends for different data formats.
Layers: Architecture (composition over inheritance):
- I/O layer: save_* / load_* functions, read/write raw files (HDF5/bin)
return Dict[str, List[Tensor]] — format-specific, no state
- Store (ABC): central abstraction, normalizes multi-segment into
Dict[str, List[Tensor]] per key via _normalize(),
fetch() uses bisect across segments — no forced concat
- Dataset layer: BaseDataset owns a Store, only calls store.fetch(begin, end, key)
Key properties: Store (ABC) — owns _data/_cum/_offsets bookkeeping
- Multi-segment: segments kept as-is, no forced concatenation — safe for + window_size/stride for sample-id
datasets larger than RAM indexing. __getitem__/__len__ produce
- Explicit length: _length = min(total elements across keys), set at load, the smallest iterable unit so Dataset
__len__ returns O(1) classes are pure delegators.
- Zero-copy mmap: MmapStore wraps np.memmap(mode="r"), all DataLoader Streamable (mixin) — raw token slice fetch(begin, end, keys)
workers share OS page-cache pages Recordable (mixin) — raw record slice fetch_record(idx, keys)
H5Store(Store, Streamable, Recordable)
MmapStore(Store, Streamable, Recordable)
JsonlStore(Store, Streamable, Recordable)
Each mixin is a stateless trait that relies on ``self._data`` etc.
provided by :class:`Store`. Concrete stores mix in whichever access
primitives they support — ``Store`` is the sole base class, so there is
no diamond inheritance or MRO ambiguity.
Sample-id indexing lives on :class:`Store`, not on the dataset:
- **Stream mode** (``window_size > 0``): ``len(store)`` returns the number
of ``(window_size, stride)`` windows that fit in the token river;
``store[i]`` returns the *i*-th window as a dict of per-key tensors;
``store.sample_window(i)`` exposes the underlying ``(begin, end)``
token slice for callers (e.g. next-token trainers) that need a +1
shifted companion window.
- **Record mode** (``num_records > 0``): ``len(store)`` returns the
record count; ``store[i]`` returns the *i*-th record dict.
Raw token/record access via :meth:`fetch` / :meth:`fetch_record`
remains available for low-level callers that want explicit index
control. ``store.token_count`` is the total stream token count (what
``len(store)`` used to mean in the legacy stream-only API).
``segments_are_records`` (class attribute on each Store subclass)
tells ``_normalize`` whether segments are inherently per-record (H5/
JSONL) or opaque shards (bin). Record access for bin relies on
``_offsets`` instead.
:class:`JsonlStore` supports a lazy mode (``processor=fn``) that keeps
raw records and defers tokenisation to ``fetch_record`` — used by DPO
to train directly from a ``.jsonl`` file without a pre-tokenised copy.
""" """
import bisect import bisect
@@ -23,20 +51,18 @@ import json
import logging import logging
from abc import ABC, abstractmethod from abc import ABC, abstractmethod
from pathlib import Path from pathlib import Path
from typing import Dict, List, Union from typing import Callable, Dict, List, Optional, Tuple, Union
import torch import torch
from torch import Tensor from torch import Tensor
from astrai.config.preprocess_config import PipelineConfig
from astrai.factory import BaseFactory from astrai.factory import BaseFactory
from astrai.preprocessing.builder import MaskBuilderFactory from astrai.preprocessing.transform import TokenizeTransform
from astrai.preprocessing.position_id import PositionIdStrategyFactory
from astrai.serialization import ( from astrai.serialization import (
load_bin, load_bin,
load_bin_offsets,
load_h5, load_h5,
) )
from astrai.tokenize import AutoTokenizer
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
@@ -48,7 +74,7 @@ def detect_format(load_path: str) -> str:
load_path: Directory or file path load_path: Directory or file path
Returns: Returns:
Format string ("h5", "bin", or "jsonl") Format string ("h5", "bin", "jsonl", or "processed")
Raises: Raises:
FileNotFoundError: If no supported data files are found FileNotFoundError: If no supported data files are found
@@ -85,228 +111,526 @@ def detect_format(load_path: str) -> str:
class Store(ABC): class Store(ABC):
"""String keys -> segmented tensors with ``fetch(begin, end, keys)``. """Common base for all storage backends.
Each key maps to one or more tensor segments (no forced concatenation). A Store owns both its data layout AND its sample-id → token/record
``len(store)`` returns ``self._length`` (explicit, O(1)), the minimum index translation. Datasets are thin wrappers that bind a Store
total element count across all keys. to a particular train-type's key mapping; they never know about
window/stride math.
Subclasses fill ``self._data`` and ``self._cum`` during ``load()`` Two iteration modes:
via ``_normalize()``.
- **Stream** (``window_size > 0``): data is treated as one long
token river. ``len(store)`` returns the number of windows;
``store[i]`` slices every stream-compatible key to window ``i``;
``store.sample_window(i)`` returns the ``(begin, end)`` token
slice for callers needing a +1 shifted companion window.
- **Record** (``num_records > 0``): data is per-record.
``len(store)`` returns ``num_records``; ``store[i]`` returns
the *i*-th record as a dict.
Raw token slicing is still available via :meth:`fetch` (mixed in
by :class:`Streamable`) when a store has stream support configured.
Raw record slicing via :meth:`fetch_record` (mixed in by
:class:`Recordable`) when a store has record support.
``token_count`` exposes the raw total stream length — this is what
``len(store)`` returned in the legacy stream-only API and what
stream-bound ``fetch`` uses for its bounds check.
""" """
def __init__(self): segments_are_records: bool = False
def __init__(
self,
window_size: int = 0,
stride: Optional[int] = None,
):
self._data: Dict[str, List[Tensor]] = {} self._data: Dict[str, List[Tensor]] = {}
self._cum: Dict[str, List[int]] = {} self._cum: Dict[str, List[int]] = {}
self._offsets: Dict[str, List[int]] = {}
self._length: int = 0 self._length: int = 0
self._num_records: int = 0
self._window_size: int = int(window_size)
self._stride: int = int(stride) if stride is not None else int(window_size)
@abstractmethod @abstractmethod
def load(self, path: str) -> None: def load(self, path: str, **kwargs) -> None:
raise NotImplementedError raise NotImplementedError
@property @property
def keys(self) -> List[str]: def keys(self) -> List[str]:
return list(self._data.keys()) return list(self._data.keys())
def __len__(self) -> int: @property
def window_size(self) -> int:
return self._window_size
@property
def stride(self) -> int:
return self._stride
@property
def token_count(self) -> int:
"""Total tokens across all stream segments.
Useful for the bounds-checked raw :meth:`fetch` and as the
legacy ``len(store)`` value.
"""
return self._length return self._length
@property
def num_records(self) -> int:
"""Number of records available via :meth:`fetch_record`.
Non-zero only when the backing layout provides per-record
indexing (H5/JSONL segments or bin ``_offsets``).
"""
return self._num_records
@property
def num_samples(self) -> int:
"""Number of items produced by ``__getitem__``.
Stream-mode wins when ``window_size > 0`` and there are tokens
to slice; otherwise falls back to ``num_records``.
"""
if self._window_size > 0 and self._length > 0:
total = self._length
w = self._window_size
if total <= w:
return 0
return (total - 1 - w) // self._stride + 1
return self._num_records
def __len__(self) -> int:
return self.num_samples
def __getitem__(self, index: int) -> Dict[str, Tensor]:
if index < 0:
index += self.num_samples
if not 0 <= index < self.num_samples:
raise IndexError(
f"Store index out of range: {index}, num_samples={self.num_samples}"
)
if self._window_size > 0 and self._length > 0:
begin, end = self.sample_window(index)
keys = self._stream_keys()
return {k: self.fetch(begin, end, k) for k in keys}
return self.fetch_record(index, self._record_keys())
def sample_window(self, index: int) -> Tuple[int, int]:
"""Return ``(begin, end)`` token positions for stream sample *index*.
The clipped tail keeps the last reachable window inside the
token river instead of overshooting. Caller is responsible
for staying within :attr:`num_samples`: an out-of-range index
raises ``IndexError``.
"""
if self._window_size <= 0:
raise RuntimeError("sample_window() requires window_size > 0 (stream mode)")
if self._window_size <= 0 or self._length <= self._window_size:
raise IndexError(
f"Data too short for window: token_count={self._length}, "
f"window_size={self._window_size}"
)
if not 0 <= index < self.num_samples:
raise IndexError(
f"Sample index out of range: {index}, num_samples={self.num_samples}"
)
total = self._length
begin = min(index * self._stride, total - 1 - self._window_size)
end = min(begin + self._window_size, total - 1)
return begin, end
def _stream_keys(self) -> List[str]:
out: List[str] = []
for k, tensors in self._data.items():
if tensors and isinstance(tensors[0], list):
continue
out.append(k)
return out
def _record_keys(self) -> List[str]:
return list(self._data.keys())
def _normalize(
self,
raw: Dict[str, list],
offsets: Optional[Dict[str, List[int]]] = None,
):
"""Register segments and pre-compute indices for both access modes.
Stream mode: ``_cum[key]`` accumulates per-segment lengths so
``Streamable._fetch_stream_key`` can bisect across segments
without concatenation.
Record mode: if *offsets* is provided (bin layout),
``_offsets[key]`` stores cumulative per-record offsets into the
single concatenated segment. Otherwise, when
``segments_are_records`` is True (H5/JSONL), ``_data[key]`` is
a per-record list and ``fetch_record`` indexes it directly.
Nested keys (GRPO ``responses``/``masks`` as
``List[List[Tensor]]``) are stored as-is and excluded from both
cumulative bookkeepings — they are only accessed record-by-record.
"""
flat_lengths = []
for key, tensors in raw.items():
self._data[key] = tensors
if not tensors:
self._cum[key] = []
flat_lengths.append(0)
continue
if isinstance(tensors[0], list):
self._cum[key] = []
continue
cum = []
total = 0
for t in tensors:
total += t.shape[0]
cum.append(total)
self._cum[key] = cum
flat_lengths.append(cum[-1] if cum else 0)
self._length = min(flat_lengths) if flat_lengths else 0
valid_offsets: Dict[str, List[int]] = {}
if offsets:
for key, off in offsets.items():
segs = self._data.get(key, [])
if len(segs) == 1 and len(off) > 1:
valid_offsets[key] = off
elif len(segs) > 1:
logger.warning(
"Key '%s' has %d segments with offsets — record mode "
"disabled for this key (multi-shard bin+offsets not "
"supported). Merge shards or use H5/JSONL.",
key,
len(segs),
)
self._offsets = valid_offsets
if valid_offsets:
record_counts = [len(v) - 1 for v in valid_offsets.values()]
self._num_records = min(record_counts) if record_counts else 0
elif self.segments_are_records:
per_record_counts = []
for key, tensors in self._data.items():
if tensors and isinstance(tensors[0], list):
continue
per_record_counts.append(len(tensors))
self._num_records = min(per_record_counts) if per_record_counts else 0
else:
self._num_records = 0
class Streamable:
"""Mixin granting raw token-stream access via :meth:`fetch`.
Stateless trait relying on ``self._data``, ``self._cum``,
``self._length`` maintained by :class:`Store`. Stream mode is
active when the owning store has ``window_size > 0``; for stores
that can also serve record access (H5/JSONL/bin+offsets), the
``fetch_record`` API from :class:`Recordable` is used instead.
"""
def fetch( def fetch(
self, self,
begin: int, begin: int,
end: int, end: int,
keys: Union[str, List[str]], keys: Union[str, List[str]],
): ):
if not self._data: return _stream_fetch(self, begin, end, keys)
raise RuntimeError("Store not loaded")
if not (0 <= begin < self._length and 0 <= end <= self._length):
raise ValueError(
f"Index out of bounds: begin={begin}, end={end}, length={self._length}"
)
if isinstance(keys, str):
return self._fetch_key(keys, begin, end)
return {k: self._fetch_key(k, begin, end) for k in keys}
def _fetch_key(self, key: str, begin: int, end: int) -> Tensor:
"""Fetch slice [begin, end) across potentially multiple segments."""
segments = self._data[key]
cum = self._cum[key]
seg_start = bisect.bisect_right(cum, begin)
seg_end = bisect.bisect_left(cum, end)
results = [] def _stream_fetch(self, begin: int, end: int, keys: Union[str, List[str]]):
for i in range(seg_start, seg_end + 1): if not getattr(self, "_data", None):
prev = cum[i - 1] if i > 0 else 0 raise RuntimeError("Store not loaded")
s = max(begin - prev, 0) if not (0 <= begin < self._length and 0 <= end <= self._length):
e = min(end - prev, segments[i].shape[0]) raise ValueError(
results.append(segments[i][s:e]) f"Index out of bounds: begin={begin}, end={end}, length={self._length}"
return results[0] if len(results) == 1 else torch.cat(results, dim=0)
def _normalize(self, raw: Dict[str, List[Tensor]]):
"""Register segments and pre-compute cumulative lengths.
Does NOT concatenate — segments are kept as-is to avoid OOM on
large datasets. Sets ``self._length`` to the minimum total
element count across all keys.
"""
for key, tensors in raw.items():
self._data[key] = tensors
cum = []
total = 0
for t in tensors:
total += t.shape[0]
cum.append(total)
self._cum[key] = cum
self._length = (
min((cum[-1] if cum else 0) for cum in self._cum.values())
if self._cum
else 0
) )
if isinstance(keys, str):
return _fetch_stream_key(self, keys, begin, end)
return {k: _fetch_stream_key(self, k, begin, end) for k in keys}
def _fetch_stream_key(self, key: str, begin: int, end: int) -> Tensor:
segments = self._data[key]
cum = self._cum[key]
seg_start = bisect.bisect_right(cum, begin)
seg_end = bisect.bisect_left(cum, end)
results = []
for i in range(seg_start, seg_end + 1):
prev = cum[i - 1] if i > 0 else 0
s = max(begin - prev, 0)
e = min(end - prev, segments[i].shape[0])
results.append(segments[i][s:e])
return results[0] if len(results) == 1 else torch.cat(results, dim=0)
class Recordable:
"""Mixin granting raw record access via :meth:`fetch_record`.
Stateless trait relying on ``self._data``, ``self._offsets``,
``self._num_records`` maintained by :class:`Store`.
"""
def fetch_record(
self,
index: int,
keys: Union[str, List[str]],
):
return _record_fetch(self, index, keys)
def _record_fetch(self, index: int, keys: Union[str, List[str]]):
if not getattr(self, "_data", None) and self._num_records == 0:
raise RuntimeError("Store not loaded")
if not 0 <= index < self._num_records:
raise ValueError(
f"Record index out of bounds: {index}, num_records={self._num_records}"
)
if isinstance(keys, str):
return _fetch_record_key(self, keys, index)
return {k: _fetch_record_key(self, k, index) for k in keys}
def _fetch_record_key(self, key: str, index: int):
offsets = self._offsets.get(key)
if offsets:
start = offsets[index]
end = (
offsets[index + 1]
if index + 1 < len(offsets)
else self._data[key][0].shape[0]
)
return self._data[key][0][start:end]
return self._data[key][index]
class StoreFactory(BaseFactory["Store"]): class StoreFactory(BaseFactory["Store"]):
"""Factory for creating Store instances by type name. """Factory for creating Store instances by type name."""
Example::
@StoreFactory.register("custom")
class CustomStore(Store):
...
"""
@StoreFactory.register("h5") @StoreFactory.register("h5")
class H5Store(Store): class H5Store(Store, Streamable, Recordable):
"""HDF5-based storage backend (pre-tokenized data).""" """HDF5-based storage backend (pre-tokenized data).
def load(self, path: str): Each key is stored as a group of per-record datasets (``data_0``,
``data_1``, …). Supports both access modes:
- **Stream**: ``fetch(begin, end, key)`` and ``store[i]`` slice
across concatenated records via ``_cum`` — used by SEQ/SFT.
- **Record**: ``fetch_record(i, key)`` and ``store[i]`` (when
``window_size == 0``) index ``_data[key]`` directly — used by
DPO/GRPO.
"""
segments_are_records = True
def __init__(
self,
window_size: int = 0,
stride: Optional[int] = None,
):
super().__init__(window_size=window_size, stride=stride)
def load(self, path: str, **kwargs):
self._normalize(load_h5(path)) self._normalize(load_h5(path))
@StoreFactory.register("bin") @StoreFactory.register("bin")
class MmapStore(Store): class MmapStore(Store, Streamable, Recordable):
"""Memory-mapped binary storage backend. """Memory-mapped binary storage backend.
Each key is a single .bin file backed by ``np.memmap(mode="r")``. Each key is a single .bin file backed by ``np.memmap(mode="r")``.
No per-process memory duplication — all DataLoader workers share the No per-process memory duplication — all DataLoader workers share the
same OS page-cache pages. same OS page-cache pages.
Format on disk:: Supports both access modes:
data_root/ - **Stream**: always available via :meth:`fetch`.
meta.json # {key: {shape, dtype}, ...} - **Record** (``fetch_record(i, key)``): only when ``meta.json``
<key>.bin # raw numpy array, one per key contains per-record ``offsets`` (written via
``save_bin(..., record_keys=...)``). Legacy bin files without
offsets have ``num_records == 0`` and ``len(store)`` reflects the
windowed sample count when ``window_size > 0``.
``segments_are_records`` is ``False`` here (bin segments are
contiguous streams, not per-record) — record access is driven
purely by ``_offsets``.
""" """
def load(self, path: str): segments_are_records = False
def __init__(
self,
window_size: int = 0,
stride: Optional[int] = None,
):
super().__init__(window_size=window_size, stride=stride)
self._mmap_refs: List[Tensor] = []
def load(self, path: str, **kwargs):
self._mmap_refs = [] self._mmap_refs = []
root = Path(path) root = Path(path)
all_raw: Dict[str, List[Tensor]] = {} all_raw: Dict[str, List[Tensor]] = {}
all_offsets: Dict[str, List[int]] = {}
meta_paths = [ meta_paths = [
Path(p) for p in glob.glob(str(root / "**" / "meta.json"), recursive=True) Path(p) for p in glob.glob(str(root / "**" / "meta.json"), recursive=True)
] ]
for meta_path in meta_paths: for meta_path in meta_paths:
raw = load_bin(str(meta_path.parent)) raw = load_bin(str(meta_path.parent))
off = load_bin_offsets(str(meta_path.parent))
for key, tensors in raw.items(): for key, tensors in raw.items():
if key not in all_raw: if key not in all_raw:
all_raw[key] = [] all_raw[key] = []
all_raw[key].extend(tensors) all_raw[key].extend(tensors)
for key, o in off.items():
if key not in all_offsets:
all_offsets[key] = []
all_offsets[key].extend(o)
if not meta_paths: if not meta_paths:
raise FileNotFoundError(f"No meta.json found under {path}") raise FileNotFoundError(f"No meta.json found under {path}")
self._normalize(all_raw) self._normalize(all_raw, offsets=all_offsets or None)
for tensors in self._data.values(): for tensors in self._data.values():
self._mmap_refs.extend(tensors) self._mmap_refs.extend(tensors)
@StoreFactory.register("jsonl") class JsonlSource:
class JsonlStore(Store): """Read raw JSON records from a ``.jsonl`` file or directory.
"""On-the-fly tokenization store for raw JSONL files.
A JSONL dataset directory contains ``*.jsonl`` files plus a A thin reader used by :class:`JsonlStore` in processor mode — holds
``dataset_config.json`` file that follows the same schema as no tokenizer, performs no tokenisation, just yields dicts.
:class:`PipelineConfig` with an additional ``tokenizer_path`` field. """
Records are tokenized when the store is loaded and concatenated into
segmented tensors matching the key layout expected by the dataset def __init__(self, path: str):
classes (``sequence``, ``loss_mask``, ``position_ids``, ...). self.path = Path(path)
self._records: Optional[List[dict]] = None
def load(self) -> List[dict]:
if self._records is None:
self._records = self._read(self.path)
return self._records
@staticmethod
def _read(root: Path) -> List[dict]:
if root.is_file():
return JsonlSource._read_file(root)
return JsonlSource._read_dir(root)
@staticmethod
def _read_file(path: Path) -> List[dict]:
records: List[dict] = []
with open(path, "r", encoding="utf-8") as f:
for line in f:
line = line.strip()
if not line:
continue
try:
records.append(json.loads(line))
except json.JSONDecodeError:
logger.warning("Failed to parse JSON line in %s, skipping", path)
return records
@staticmethod
def _read_dir(root: Path) -> List[dict]:
records: List[dict] = []
for jsonl_path in sorted(root.glob("*.jsonl")):
records.extend(JsonlSource._read_file(jsonl_path))
return records
@StoreFactory.register("jsonl")
class JsonlStore(Store, Streamable, Recordable):
"""JSONL reader with two tokenisation modes.
A JSONL dataset is a ``.jsonl`` file or a directory of ``*.jsonl``
files plus (optionally) a ``dataset_config.json`` describing the
tokenization pipeline.
Two modes, selected at :meth:`load` time:
- **Eager** (default): applies a :class:`TokenizeTransform` to every
record at load time and registers per-key tensors via
``_normalize``. Both ``fetch`` (stream) and ``fetch_record``
(record) work.
- **Lazy** (``processor=fn`` passed): keeps raw records and defers
tokenisation to ``fetch_record``. Only record access works —
``len(store)`` returns ``num_records``; stream primitives raise.
""" """
CONFIG_NAME = "dataset_config.json" CONFIG_NAME = "dataset_config.json"
segments_are_records = True
def load(self, path: str): def __init__(
root = Path(path) self,
config_path = root / self.CONFIG_NAME window_size: int = 0,
if not config_path.exists(): stride: Optional[int] = None,
raise FileNotFoundError( ):
f"JSONL dataset config not found: {config_path}. " super().__init__(window_size=window_size, stride=stride)
f"Expected {self.CONFIG_NAME} alongside *.jsonl files." self._source: Optional[JsonlSource] = None
self._processor: Optional[Callable[[dict], Dict[str, Tensor]]] = None
self._keys_cache: Optional[List[str]] = None
def load(self, path: str, transform=None, processor=None, **kwargs):
self._source = JsonlSource(path)
records = self._source.load()
if processor is not None:
self._processor = processor
self._num_records = len(records)
return
if transform is None:
root = Path(path)
config_path = root / self.CONFIG_NAME if root.is_dir() else None
if config_path is None or not config_path.exists():
raise FileNotFoundError(
f"JSONL dataset config not found. Expected "
f"{self.CONFIG_NAME} alongside *.jsonl files, pass an "
f"explicit transform, or pass processor= for lazy "
f"on-the-fly tokenisation."
)
transform = TokenizeTransform.from_config_file(str(config_path))
transformed = transform.apply(records)
self._normalize(transformed)
@property
def keys(self) -> List[str]:
if self._processor is not None:
if self._keys_cache is None and self._num_records > 0:
sample = self._processor(self._source.load()[0])
self._keys_cache = list(sample.keys())
return self._keys_cache or []
return list(self._data.keys())
def fetch_record(self, index: int, keys: Union[str, List[str]]):
if self._processor is not None:
if not 0 <= index < self._num_records:
raise ValueError(
f"Record index out of bounds: {index}, "
f"num_records={self._num_records}"
)
record = self._source.load()[index]
data = self._processor(record)
if isinstance(keys, str):
return data[keys]
return {k: data[k] for k in keys}
return _record_fetch(self, index, keys)
def fetch(self, begin: int, end: int, keys: Union[str, List[str]]):
if self._processor is not None:
raise RuntimeError(
"JsonlStore in lazy (processor) mode does not support "
"stream fetch(); use fetch_record() instead."
) )
return _stream_fetch(self, begin, end, keys)
with open(config_path, "r", encoding="utf-8") as f: def __getitem__(self, index: int) -> Dict[str, Tensor]:
raw_config = json.load(f) if self._processor is not None:
return self.fetch_record(index, self._record_keys())
tokenizer_path = raw_config.pop("tokenizer_path", None) return super().__getitem__(index)
if tokenizer_path is None:
raise ValueError(
f"JSONL dataset config must specify 'tokenizer_path': {config_path}"
)
self.config = PipelineConfig.from_dict(raw_config)
tokenizer = AutoTokenizer.from_pretrained(tokenizer_path)
mask_builder = MaskBuilderFactory.create("sectioned")
position_strategy = PositionIdStrategyFactory.create(
self.config.output.position_ids_mode
)
raw: Dict[str, List[Tensor]] = {}
doc_sequences: List[List[int]] = []
for jsonl_path in sorted(root.glob("*.jsonl")):
with open(jsonl_path, "r", encoding="utf-8") as f:
for line in f:
line = line.strip()
if not line:
continue
try:
item = json.loads(line)
except json.JSONDecodeError:
logger.warning(
"Failed to parse JSON line in %s, skipping", jsonl_path
)
continue
result = mask_builder.build(item, self.config, tokenizer)
if result is None:
continue
result.pop("domain", None)
primary_ids = self._primary_ids(result)
if not primary_ids:
continue
doc_sequences.append(primary_ids)
for key, ids in result.items():
if key not in raw:
raw[key] = []
raw[key].append(torch.tensor(ids, dtype=self._infer_dtype(ids)))
pos_ids = position_strategy.generate(doc_sequences)
if pos_ids:
raw["position_ids"] = [torch.tensor(pos_ids, dtype=torch.int32)]
self._normalize(raw)
@staticmethod
def _primary_ids(result: dict) -> List[int]:
"""Return the first integer list in *result* as the primary id sequence."""
for val in result.values():
if isinstance(val, list) and val and isinstance(val[0], int):
return val
return []
@staticmethod
def _infer_dtype(ids: List) -> torch.dtype:
"""Infer tensor dtype from the first element of a token/value list."""
if ids and isinstance(ids[0], float):
return torch.float32
return torch.int32
+8
View File
@@ -5,6 +5,14 @@ Public API:
- ``attn_prefill`` — multi-query prefill attention - ``attn_prefill`` — multi-query prefill attention
- ``attn_paged_decode`` — paged decode attention (direct page-table access) - ``attn_paged_decode`` — paged decode attention (direct page-table access)
Interface (shared by all wrappers):
causal_offset: -1 = non-causal; >=0 = absolute position of first Q token
mask: 2D [batch, kv_len] or 3D [batch, q_len, kv_len] (bool, True = keep)
scale: 0.0 = auto (1/sqrt(head_dim)); >0 = explicit
layout: "bhld" (default) or "blhd"
Causal and mask can coexist — both are applied simultaneously.
Each wrapper dispatches to its compiled CUDA kernel (``astrai.extension.attn_*``) Each wrapper dispatches to its compiled CUDA kernel (``astrai.extension.attn_*``)
when available, otherwise falls back to ``torch.nn.functional.scaled_dot_product_attention``. when available, otherwise falls back to ``torch.nn.functional.scaled_dot_product_attention``.
""" """
+129 -32
View File
@@ -3,15 +3,43 @@
Each wrapper dispatches to its CUDA kernel (loaded in ``loader.py``) when Each wrapper dispatches to its CUDA kernel (loaded in ``loader.py``) when
available, otherwise falls back to ``torch`` SDPA. available, otherwise falls back to ``torch`` SDPA.
Interface (all functions):
causal_offset: -1 = non-causal; >=0 = absolute position of first Q token
mask: 2D [batch, kv_len] or 3D [batch, q_len, kv_len] (bool)
scale: 0.0 = auto (1/sqrt(head_dim)); >0 = explicit
layout: "bhld" (default) or "blhd"
Add new kernel wrappers here; split into per-variant files only if this file Add new kernel wrappers here; split into per-variant files only if this file
grows large. grows large.
""" """
import math
import torch import torch
import torch.nn.functional as F import torch.nn.functional as F
from astrai.extension.loader import _available, _modules from astrai.extension.loader import _available, _modules
_LAYOUT_CODES: dict[str, int] = {"bhld": 0, "blhd": 1}
def _parse_layout(layout: str | int) -> int:
if isinstance(layout, int):
return layout
code = _LAYOUT_CODES.get(layout.lower())
if code is None:
raise ValueError(
f"unknown layout '{layout}', expected one of {list(_LAYOUT_CODES)}"
)
return code
def _to_bhld(t: torch.Tensor, layout: int) -> torch.Tensor:
"""Normalize to b h l d view. Zero-copy transpose if layout==1 (b l h d)."""
if layout == 1:
return t.transpose(1, 2)
return t
def _expand_kv_heads( def _expand_kv_heads(
k: torch.Tensor, v: torch.Tensor, q_head: int k: torch.Tensor, v: torch.Tensor, q_head: int
@@ -26,20 +54,84 @@ def _expand_kv_heads(
return k, v return k, v
def _build_attn_mask(
q: torch.Tensor,
k: torch.Tensor,
mask: torch.Tensor | None,
causal_offset: int,
scale: float,
) -> tuple[torch.Tensor | None, float]:
"""Build SDPA-compatible attn_mask + resolved scale.
q and k must already be in b h l d layout.
Causal and mask can coexist: causal sets -inf above the diagonal, mask
sets -inf for padded positions. Both are OR'd into a single bool mask.
"""
q_len = q.size(2)
kv_len = k.size(2)
head_dim = q.size(3)
resolved_scale = scale if scale and scale > 0 else 1.0 / math.sqrt(head_dim)
attn_mask = None
if mask is not None:
if mask.dim() == 2:
# [batch, kv_len] → [batch, 1, 1, kv_len]
attn_mask = mask[:, None, None, :]
elif mask.dim() == 3:
# [batch, q_len, kv_len] → [batch, 1, q_len, kv_len]
attn_mask = mask[:, None, :, :]
else:
raise ValueError(f"mask must be 2D or 3D, got {mask.dim()}D")
if causal_offset >= 0:
batch = q.size(0)
# q row i attends to kv cols 0..(causal_offset + i)
q_idx = torch.arange(q_len, device=q.device).unsqueeze(1) # [q_len, 1]
kv_idx = torch.arange(kv_len, device=q.device).unsqueeze(0) # [1, kv_len]
causal_bool = kv_idx > (causal_offset + q_idx) # True = masked out
causal_mask = causal_bool.unsqueeze(0).expand(
batch, -1, -1
) # [batch, q_len, kv_len]
causal_mask = causal_mask[:, None, :, :] # [batch, 1, q_len, kv_len]
if attn_mask is not None:
attn_mask = attn_mask | causal_mask
else:
attn_mask = causal_mask
return attn_mask, resolved_scale
def _torch_fallback( def _torch_fallback(
q: torch.Tensor, q: torch.Tensor,
k: torch.Tensor, k: torch.Tensor,
v: torch.Tensor, v: torch.Tensor,
mask: torch.Tensor | None, mask: torch.Tensor | None,
is_causal: bool, causal_offset: int,
scale: float | None, scale: float,
q_layout: int,
kv_layout: int | None = None,
) -> torch.Tensor: ) -> torch.Tensor:
"""Reference attention via ``scaled_dot_product_attention``.""" """Reference attention via ``scaled_dot_product_attention``.
q_layout / kv_layout: 0 = b h l d, 1 = b l h d.
If kv_layout is None, uses q_layout (Q and K/V share the same layout).
"""
if kv_layout is None:
kv_layout = q_layout
q = _to_bhld(q, q_layout)
k = _to_bhld(k, kv_layout)
v = _to_bhld(v, kv_layout)
k, v = _expand_kv_heads(k, v, q.size(1)) k, v = _expand_kv_heads(k, v, q.size(1))
attn_mask = mask[:, None, None, :] if mask is not None else None attn_mask, resolved_scale = _build_attn_mask(q, k, mask, causal_offset, scale)
return F.scaled_dot_product_attention( out = F.scaled_dot_product_attention(
q, k, v, attn_mask=attn_mask, is_causal=is_causal and mask is None, scale=scale q, k, v, attn_mask=attn_mask, is_causal=False, scale=resolved_scale
) )
# Restore Q's original layout
if q_layout == 1:
out = out.transpose(1, 2)
return out
def _gather_kv_from_pages( def _gather_kv_from_pages(
@@ -56,23 +148,22 @@ def _gather_kv_from_pages(
k_cache : [n_pages, page_size, n_kv_heads, head_dim] k_cache : [n_pages, page_size, n_kv_heads, head_dim]
v_cache : same as k_cache v_cache : same as k_cache
Returns: Returns:
k, v : [batch, n_kv_heads, kv_len, head_dim] k, v : [batch, kv_len, n_kv_heads, head_dim] (b l h d)
""" """
batch, max_pages = page_table.shape batch, max_pages = page_table.shape
n_pages, ps, n_kv_heads, head_dim = k_cache.shape _, ps, n_kv_heads, head_dim = k_cache.shape
if ps != page_size: if ps != page_size:
raise ValueError(f"k_cache page_size mismatch: {ps} vs {page_size}") raise ValueError(f"k_cache page_size mismatch: {ps} vs {page_size}")
k = k_cache.new_empty(batch, n_kv_heads, kv_len, head_dim) # Vectorized gather: build physical page + offset indices, then advanced-index
v = v_cache.new_empty(batch, n_kv_heads, kv_len, head_dim) positions = torch.arange(kv_len, device=page_table.device)
logical_pages = positions // page_size # [kv_len]
page_offsets = positions % page_size # [kv_len]
for b in range(batch): phys_pages = page_table[:, logical_pages] # [batch, kv_len]
for pos in range(kv_len): # k_cache[phys_pages, page_offsets] → [batch, kv_len, n_kv_heads, head_dim] (b l h d)
log_pg = pos // page_size k = k_cache[phys_pages, page_offsets]
pg_off = pos % page_size v = v_cache[phys_pages, page_offsets]
phys = int(page_table[b, log_pg].item())
k[b, :, pos, :] = k_cache[phys, pg_off, :, :]
v[b, :, pos, :] = v_cache[phys, pg_off, :, :]
return k, v return k, v
@@ -81,21 +172,22 @@ def attn_decode(
k: torch.Tensor, k: torch.Tensor,
v: torch.Tensor, v: torch.Tensor,
mask: torch.Tensor | None = None, mask: torch.Tensor | None = None,
is_causal: bool = False, causal_offset: int = -1,
causal_offset: int = 0, scale: float = 0.0,
scale: float | None = None, layout: str = "bhld",
) -> torch.Tensor: ) -> torch.Tensor:
li = _parse_layout(layout)
if _available["attn_decode"]: if _available["attn_decode"]:
return _modules["attn_decode"].attn_decode( return _modules["attn_decode"].attn_decode(
q, q,
k, k,
v, v,
mask=mask, mask=mask,
is_causal=is_causal,
causal_offset=causal_offset, causal_offset=causal_offset,
scale=scale, scale=scale,
layout=li,
) )
return _torch_fallback(q, k, v, mask, is_causal, scale) return _torch_fallback(q, k, v, mask, causal_offset, scale, q_layout=li)
def attn_prefill( def attn_prefill(
@@ -103,21 +195,22 @@ def attn_prefill(
k: torch.Tensor, k: torch.Tensor,
v: torch.Tensor, v: torch.Tensor,
mask: torch.Tensor | None = None, mask: torch.Tensor | None = None,
is_causal: bool = False, causal_offset: int = -1,
causal_offset: int = 0, scale: float = 0.0,
scale: float | None = None, layout: str = "bhld",
) -> torch.Tensor: ) -> torch.Tensor:
li = _parse_layout(layout)
if _available["attn_prefill"]: if _available["attn_prefill"]:
return _modules["attn_prefill"].attn_prefill( return _modules["attn_prefill"].attn_prefill(
q, q,
k, k,
v, v,
mask=mask, mask=mask,
is_causal=is_causal,
causal_offset=causal_offset, causal_offset=causal_offset,
scale=scale, scale=scale,
layout=li,
) )
return _torch_fallback(q, k, v, mask, is_causal, scale) return _torch_fallback(q, k, v, mask, causal_offset, scale, q_layout=li)
def attn_paged_decode( def attn_paged_decode(
@@ -128,10 +221,11 @@ def attn_paged_decode(
page_size: int, page_size: int,
kv_len: int, kv_len: int,
mask: torch.Tensor | None = None, mask: torch.Tensor | None = None,
is_causal: bool = False, causal_offset: int = -1,
causal_offset: int = 0, scale: float = 0.0,
scale: float | None = None, layout: str = "bhld",
) -> torch.Tensor: ) -> torch.Tensor:
li = _parse_layout(layout)
if _available["attn_paged_decode"]: if _available["attn_paged_decode"]:
return _modules["attn_paged_decode"].attn_paged_decode( return _modules["attn_paged_decode"].attn_paged_decode(
q, q,
@@ -141,9 +235,12 @@ def attn_paged_decode(
page_size, page_size,
kv_len, kv_len,
mask=mask, mask=mask,
is_causal=is_causal,
causal_offset=causal_offset, causal_offset=causal_offset,
scale=scale, scale=scale,
layout=li,
) )
# Gathered K/V are always b l h d
k, v = _gather_kv_from_pages(page_table, k_cache, v_cache, page_size, kv_len) k, v = _gather_kv_from_pages(page_table, k_cache, v_cache, page_size, kv_len)
return _torch_fallback(q, k, v, mask, is_causal, scale) return _torch_fallback(
q, k, v, mask, causal_offset, scale, q_layout=li, kv_layout=1
)
+3 -1
View File
@@ -6,7 +6,7 @@ Layers:
- protocols/: Response builders (OpenAI, Anthropic) - protocols/: Response builders (OpenAI, Anthropic)
- transport/: SSE transport utilities - transport/: SSE transport utilities
- engine.py: Facade (InferenceEngine), Value Object (GenerationRequest) - engine.py: Facade (InferenceEngine), Value Object (GenerationRequest)
- sample.py: Strategy pattern (TemperatureStrategy, TopKStrategy, TopPStrategy) - sample.py: Strategy pattern (TemperatureStrategy, TopKStrategy, TopPStrategy, FrequencyPenaltyStrategy)
""" """
from astrai.inference.api import ( from astrai.inference.api import (
@@ -50,6 +50,7 @@ from astrai.inference.core import (
from astrai.inference.engine import GenerationRequest, InferenceEngine from astrai.inference.engine import GenerationRequest, InferenceEngine
from astrai.inference.sample import ( from astrai.inference.sample import (
BaseSamplingStrategy, BaseSamplingStrategy,
FrequencyPenaltyStrategy,
SamplingPipeline, SamplingPipeline,
TemperatureStrategy, TemperatureStrategy,
TopKStrategy, TopKStrategy,
@@ -83,6 +84,7 @@ __all__ = [
"TemperatureStrategy", "TemperatureStrategy",
"TopKStrategy", "TopKStrategy",
"TopPStrategy", "TopPStrategy",
"FrequencyPenaltyStrategy",
"SamplingPipeline", "SamplingPipeline",
"ProtocolHandler", "ProtocolHandler",
"StopChecker", "StopChecker",
-1
View File
@@ -21,7 +21,6 @@ logger = logging.getLogger(__name__)
_UNSUPPORTED_PARAMS = ( _UNSUPPORTED_PARAMS = (
"n", "n",
"presence_penalty", "presence_penalty",
"frequency_penalty",
"logit_bias", "logit_bias",
"user", "user",
) )
+1
View File
@@ -125,6 +125,7 @@ class ProtocolHandler:
temperature=self.request.temperature, temperature=self.request.temperature,
top_p=self.request.top_p, top_p=self.request.top_p,
top_k=self.request.top_k, top_k=self.request.top_k,
frequency_penalty=getattr(self.request, "frequency_penalty", 0.0),
) )
if self.request.stream: if self.request.stream:
+47 -13
View File
@@ -300,7 +300,11 @@ class KVCache(ABC):
@abstractmethod @abstractmethod
def bind_tasks( def bind_tasks(
self, task_ids: List[str], total_len: int, device: torch.device self,
task_ids: List[str],
total_len: int,
device: torch.device,
write_positions: Optional[Tensor] = None,
) -> CacheView: ... ) -> CacheView: ...
def task_cached(self, task_id: str) -> int: def task_cached(self, task_id: str) -> int:
@@ -399,7 +403,11 @@ class PageCache(KVCache):
self._pool.record(page_table[i], prompt_ids, i) self._pool.record(page_table[i], prompt_ids, i)
def bind_tasks( def bind_tasks(
self, task_ids: List[str], total_len: int, device: torch.device self,
task_ids: List[str],
total_len: int,
device: torch.device,
write_positions: Optional[Tensor] = None,
) -> PageCacheView: ) -> PageCacheView:
page_table = self._table.table_tensor(task_ids, device) page_table = self._table.table_tensor(task_ids, device)
return PageCacheView(self._storage, page_table, total_len) return PageCacheView(self._storage, page_table, total_len)
@@ -409,23 +417,37 @@ class ContiguousCacheView(CacheView):
"""Contiguous KV-cache view for attention layers.""" """Contiguous KV-cache view for attention layers."""
def __init__( def __init__(
self, cache: "ContiguousCache", batch_indices: Tensor, total_len: int = 0 self,
cache: "ContiguousCache",
batch_indices: Tensor,
total_len: int = 0,
write_positions: Optional[Tensor] = None,
): ):
self._cache = cache self._cache = cache
self._batch_indices = batch_indices self._batch_indices = batch_indices
self._total_len = total_len self._total_len = total_len
self._write_positions = write_positions
def write(self, layer_id: int, k: Tensor, v: Tensor): def write(self, layer_id: int, k: Tensor, v: Tensor):
seq_len = k.size(1) seq_len = k.size(1)
start_pos = self._total_len - seq_len
indices = self._batch_indices indices = self._batch_indices
self._cache.k[layer_id, indices, start_pos : start_pos + seq_len] = k if self._write_positions is not None and seq_len == 1:
self._cache.v[layer_id, indices, start_pos : start_pos + seq_len] = v pos = self._write_positions
new_len = start_pos + seq_len self._cache.k[layer_id, indices, pos] = k.squeeze(1)
for s in indices.tolist(): self._cache.v[layer_id, indices, pos] = v.squeeze(1)
cur = self._cache._slot_len.get(s, 0) for s, p in zip(indices.tolist(), pos.tolist()):
if new_len > cur: cur = self._cache._slot_len.get(s, 0)
self._cache._slot_len[s] = new_len if p + 1 > cur:
self._cache._slot_len[s] = p + 1
else:
start_pos = self._total_len - seq_len
self._cache.k[layer_id, indices, start_pos : start_pos + seq_len] = k
self._cache.v[layer_id, indices, start_pos : start_pos + seq_len] = v
new_len = start_pos + seq_len
for s in indices.tolist():
cur = self._cache._slot_len.get(s, 0)
if new_len > cur:
self._cache._slot_len[s] = new_len
def gather(self, layer_id: int) -> Tuple[Tensor, Tensor]: def gather(self, layer_id: int) -> Tuple[Tensor, Tensor]:
max_len = max( max_len = max(
@@ -491,9 +513,21 @@ class ContiguousCache(KVCache):
def task_extend(self, task_id: str, pos: int) -> bool: def task_extend(self, task_id: str, pos: int) -> bool:
return pos < self.max_seq_len return pos < self.max_seq_len
def task_cached(self, task_id: str) -> int:
slot = self._task_slot.get(task_id)
if slot is None:
return 0
return self._slot_len.get(slot, 0)
def bind_tasks( def bind_tasks(
self, task_ids: List[str], total_len: int, device: torch.device self,
task_ids: List[str],
total_len: int,
device: torch.device,
write_positions: Optional[Tensor] = None,
) -> ContiguousCacheView: ) -> ContiguousCacheView:
slots = [self._task_slot[tid] for tid in task_ids] slots = [self._task_slot[tid] for tid in task_ids]
batch_indices = torch.tensor(slots, dtype=torch.long, device=device) batch_indices = torch.tensor(slots, dtype=torch.long, device=device)
return ContiguousCacheView(self, batch_indices, total_len) return ContiguousCacheView(
self, batch_indices, total_len, write_positions=write_positions
)
+36 -1
View File
@@ -75,11 +75,43 @@ class Executor:
temperatures = torch.tensor([t.temperature for t in tasks], device=self.device) temperatures = torch.tensor([t.temperature for t in tasks], device=self.device)
top_ks = torch.tensor([t.top_k for t in tasks], device=self.device) top_ks = torch.tensor([t.top_k for t in tasks], device=self.device)
top_ps = torch.tensor([t.top_p for t in tasks], device=self.device) top_ps = torch.tensor([t.top_p for t in tasks], device=self.device)
freq_penalties = torch.tensor(
[t.frequency_penalty for t in tasks], device=self.device
)
history_lists = []
mask_lists = []
for t in tasks:
window = t.rep_window
prompt_part = t.prompt_ids[-window:]
ids = prompt_part + t.output_ids
history_lists.append(ids)
mask_lists.append([True] * len(ids))
max_len = max(len(h) for h in history_lists)
padded_ids = torch.zeros(
len(tasks), max_len, dtype=torch.long, device=self.device
)
padded_mask = torch.zeros(
len(tasks), max_len, dtype=torch.bool, device=self.device
)
for i, (h, m) in enumerate(zip(history_lists, mask_lists)):
padded_ids[i, : len(h)] = torch.tensor(
h, dtype=torch.long, device=self.device
)
padded_mask[i, : len(m)] = torch.tensor(
m, dtype=torch.bool, device=self.device
)
with torch.inference_mode(): with torch.inference_mode():
outputs = self.model( outputs = self.model(
input_ids.unsqueeze(1), input_ids.unsqueeze(1),
paged_cache=self.kv_cache.bind_tasks(task_ids, total_len, self.device), paged_cache=self.kv_cache.bind_tasks(
task_ids,
total_len,
self.device,
write_positions=position_ids,
),
position_ids=position_ids.unsqueeze(1), position_ids=position_ids.unsqueeze(1),
) )
logits = outputs["logits"][:, -1, :] logits = outputs["logits"][:, -1, :]
@@ -89,4 +121,7 @@ class Executor:
temperature=temperatures, temperature=temperatures,
top_k=top_ks, top_k=top_ks,
top_p=top_ps, top_p=top_ps,
frequency_penalty=freq_penalties,
input_ids=padded_ids,
input_mask=padded_mask,
).tolist() ).tolist()
+23 -26
View File
@@ -138,36 +138,33 @@ class InferenceScheduler:
t.task_id, t.prompt_ids, start_logical_page t.task_id, t.prompt_ids, start_logical_page
) )
pos_groups: Dict[int, List[Task]] = {} decode_tasks = self._task_mgr.get_active_tasks()
for t in self._task_mgr.get_active_tasks():
pos_groups.setdefault(t.next_pos, []).append(t)
for next_pos in sorted(pos_groups.keys()): valid: List[Task] = []
group = sorted(pos_groups[next_pos], key=lambda t: t.task_id) for t in sorted(decode_tasks, key=lambda t: t.task_id):
if cache.task_extend(t.task_id, t.next_pos):
valid.append(t)
else:
t.status = TaskStatus.ABORTED
self._task_mgr.invoke_callback(t.task_id, STOP)
valid: List[Task] = [] if valid:
for t in group: next_tokens = self._executor.execute_decode(valid)
if cache.task_extend(t.task_id, t.next_pos):
valid.append(t) for t, ntok in zip(valid, next_tokens):
else: t.output_ids.append(ntok)
t.status = TaskStatus.ABORTED t.output_tokens += 1
new_text = t.decode_new_token(self._task_mgr.tokenizer)
if new_text:
self._task_mgr.invoke_callback(t.task_id, new_text)
for t in valid:
if t.is_finished(stop_ids):
remaining = t.flush_remaining(self._task_mgr.tokenizer)
if remaining:
self._task_mgr.invoke_callback(t.task_id, remaining)
self._task_mgr.invoke_callback(t.task_id, STOP) self._task_mgr.invoke_callback(t.task_id, STOP)
if valid:
next_tokens = self._executor.execute_decode(valid)
for t, ntok in zip(valid, next_tokens):
t.output_ids.append(ntok)
t.output_tokens += 1
self._task_mgr.invoke_callback(
t.task_id,
self._task_mgr.tokenizer.decode([ntok]),
)
for t in valid:
if t.is_finished(stop_ids):
self._task_mgr.invoke_callback(t.task_id, STOP)
except Exception as e: except Exception as e:
self._stop_event.set() self._stop_event.set()
logger.error(f"Scheduler loop crashed: {e}", exc_info=True) logger.error(f"Scheduler loop crashed: {e}", exc_info=True)
+70
View File
@@ -13,6 +13,40 @@ logger = logging.getLogger(__name__)
STOP = object() STOP = object()
class StreamDecoder:
"""Incremental decoder for byte-level BPE streaming.
Byte-level BPE may split a single Unicode character (e.g. em-dash,
smart quotes) across multiple tokens. Decoding such a token in
isolation produces U+FFFD (replacement char). This decoder
accumulates token IDs and only emits text once the trailing
characters are complete, buffering incomplete multi-byte sequences
until the next token arrives.
"""
__slots__ = ("_tokenizer", "_ids", "_emitted")
def __init__(self, tokenizer: AutoTokenizer):
self._tokenizer = tokenizer
self._ids: List[int] = []
self._emitted: str = ""
def push(self, token_id: int) -> str:
"""Append a token ID and return newly completed text.
Returns "" while a multi-byte character is still incomplete.
"""
self._ids.append(token_id)
full = self._tokenizer.decode(self._ids, skip_special_tokens=True)
if full.endswith("\ufffd"):
return ""
if len(full) > len(self._emitted):
diff = full[len(self._emitted) :]
self._emitted = full
return diff
return ""
class TaskStatus(Enum): class TaskStatus(Enum):
"""Task lifecycle states.""" """Task lifecycle states."""
@@ -33,6 +67,8 @@ class Task:
temperature: float = 1.0, temperature: float = 1.0,
top_p: float = 1.0, top_p: float = 1.0,
top_k: int = 50, top_k: int = 50,
frequency_penalty: float = 0.0,
rep_window: int = 64,
): ):
self.task_id = task_id self.task_id = task_id
self.prompt_ids = prompt_ids self.prompt_ids = prompt_ids
@@ -40,6 +76,8 @@ class Task:
self.temperature = temperature self.temperature = temperature
self.top_p = top_p self.top_p = top_p
self.top_k = top_k self.top_k = top_k
self.frequency_penalty = frequency_penalty
self.rep_window = rep_window
self.status = TaskStatus.PENDING self.status = TaskStatus.PENDING
self.output_ids: List[int] = [] self.output_ids: List[int] = []
@@ -47,6 +85,34 @@ class Task:
self.output_tokens: int = 0 self.output_tokens: int = 0
self.arrival_time = time.time() self.arrival_time = time.time()
self.finish_time: Optional[float] = None self.finish_time: Optional[float] = None
self._decoder: Optional[StreamDecoder] = None
def decode_new_token(self, tokenizer: AutoTokenizer) -> str:
"""Decode the last appended output token, buffering incomplete
multi-byte sequences across calls.
Lazily creates a :class:`StreamDecoder` on first use.
"""
if self._decoder is None:
self._decoder = StreamDecoder(tokenizer)
return self._decoder.push(self.output_ids[-1])
def flush_remaining(self, tokenizer: AutoTokenizer) -> str:
"""Emit any text still buffered in the decoder.
Called when generation terminates (max_tokens reached, stop
sequence, or external removal) to avoid dropping a final
incomplete-looking fragment that is actually complete when
adjacent to the stop token.
"""
if self._decoder is None or not self.output_ids:
return ""
full = tokenizer.decode(self.output_ids, skip_special_tokens=True)
if len(full) > len(self._decoder._emitted):
diff = full[len(self._decoder._emitted) :]
self._decoder._emitted = full
return diff
return ""
@property @property
def next_pos(self) -> int: def next_pos(self) -> int:
@@ -92,6 +158,8 @@ class TaskManager:
temperature: float = 1.0, temperature: float = 1.0,
top_p: float = 1.0, top_p: float = 1.0,
top_k: int = 50, top_k: int = 50,
frequency_penalty: float = 0.0,
rep_window: int = 64,
stream_callback: Optional[Callable[[str], None]] = None, stream_callback: Optional[Callable[[str], None]] = None,
) -> str: ) -> str:
task_id = f"task_{int(time.time())}_{uuid.uuid4().hex[:8]}" task_id = f"task_{int(time.time())}_{uuid.uuid4().hex[:8]}"
@@ -116,6 +184,8 @@ class TaskManager:
temperature=temperature, temperature=temperature,
top_p=top_p, top_p=top_p,
top_k=top_k, top_k=top_k,
frequency_penalty=frequency_penalty,
rep_window=rep_window,
) )
with self._lock: with self._lock:
+63 -5
View File
@@ -74,6 +74,8 @@ class GenerationRequest:
top_p: float = 1.0, top_p: float = 1.0,
temperature: float = 1.0, temperature: float = 1.0,
max_tokens: Optional[int] = None, max_tokens: Optional[int] = None,
frequency_penalty: float = 0.0,
rep_window: int = 64,
stream: bool = False, stream: bool = False,
): ):
if not (isinstance(top_k, int) and top_k >= 0): if not (isinstance(top_k, int) and top_k >= 0):
@@ -82,12 +84,21 @@ class GenerationRequest:
raise ValueError("top_p must be a float between 0.0 and 1.0") raise ValueError("top_p must be a float between 0.0 and 1.0")
if not (isinstance(temperature, (int, float)) and temperature > 0): if not (isinstance(temperature, (int, float)) and temperature > 0):
raise ValueError("temperature must be a positive number") raise ValueError("temperature must be a positive number")
if not (
isinstance(frequency_penalty, (int, float))
and -2.0 <= frequency_penalty <= 2.0
):
raise ValueError("frequency_penalty must be between -2.0 and 2.0")
if not (isinstance(rep_window, int) and rep_window > 0):
raise ValueError("rep_window must be a positive integer")
self.messages = messages self.messages = messages
self.top_k = top_k self.top_k = top_k
self.top_p = top_p self.top_p = top_p
self.temperature = temperature self.temperature = temperature
self.max_tokens = max_tokens self.max_tokens = max_tokens
self.frequency_penalty = frequency_penalty
self.rep_window = rep_window
self.stream = stream self.stream = stream
@@ -132,17 +143,33 @@ class InferenceEngine:
temperature: float = 1.0, temperature: float = 1.0,
top_p: float = 1.0, top_p: float = 1.0,
top_k: int = 50, top_k: int = 50,
frequency_penalty: float = 0.0,
rep_window: int = 64,
) -> Union[Generator, str, List[str]]: ) -> Union[Generator, str, List[str]]:
is_batch = isinstance(prompt, list) is_batch = isinstance(prompt, list)
prompts = prompt if is_batch else [prompt] prompts = prompt if is_batch else [prompt]
if stream: if stream:
return self._generate_streaming( return self._generate_streaming(
prompts, is_batch, max_tokens, temperature, top_p, top_k prompts,
is_batch,
max_tokens,
temperature,
top_p,
top_k,
frequency_penalty,
rep_window,
) )
else: else:
return self._generate_non_streaming( return self._generate_non_streaming(
prompts, is_batch, max_tokens, temperature, top_p, top_k prompts,
is_batch,
max_tokens,
temperature,
top_p,
top_k,
frequency_penalty,
rep_window,
) )
def generate_async( def generate_async(
@@ -152,9 +179,18 @@ class InferenceEngine:
temperature: float = 1.0, temperature: float = 1.0,
top_p: float = 1.0, top_p: float = 1.0,
top_k: int = 50, top_k: int = 50,
frequency_penalty: float = 0.0,
rep_window: int = 64,
) -> AsyncGenerator[str, None]: ) -> AsyncGenerator[str, None]:
sync_gen = self._generate_streaming( sync_gen = self._generate_streaming(
[prompt], False, max_tokens, temperature, top_p, top_k [prompt],
False,
max_tokens,
temperature,
top_p,
top_k,
frequency_penalty,
rep_window,
) )
async def _agen(): async def _agen():
@@ -185,6 +221,8 @@ class InferenceEngine:
temperature=request.temperature, temperature=request.temperature,
top_p=request.top_p, top_p=request.top_p,
top_k=request.top_k, top_k=request.top_k,
frequency_penalty=request.frequency_penalty,
rep_window=request.rep_window,
) )
def _submit_tasks( def _submit_tasks(
@@ -194,6 +232,8 @@ class InferenceEngine:
temperature: float, temperature: float,
top_p: float, top_p: float,
top_k: int, top_k: int,
frequency_penalty: float,
rep_window: int,
) -> Tuple[GenerateResult, List[str]]: ) -> Tuple[GenerateResult, List[str]]:
n = len(prompts) n = len(prompts)
result = GenerateResult(count=n) result = GenerateResult(count=n)
@@ -206,6 +246,8 @@ class InferenceEngine:
temperature=temperature, temperature=temperature,
top_p=top_p, top_p=top_p,
top_k=top_k, top_k=top_k,
frequency_penalty=frequency_penalty,
rep_window=rep_window,
stream_callback=cb, stream_callback=cb,
) )
task_ids.append(task_id) task_ids.append(task_id)
@@ -226,9 +268,17 @@ class InferenceEngine:
temperature: float, temperature: float,
top_p: float, top_p: float,
top_k: int, top_k: int,
frequency_penalty: float,
rep_window: int,
) -> Generator: ) -> Generator:
result, task_ids = self._submit_tasks( result, task_ids = self._submit_tasks(
prompts, max_tokens, temperature, top_p, top_k prompts,
max_tokens,
temperature,
top_p,
top_k,
frequency_penalty,
rep_window,
) )
n = len(prompts) n = len(prompts)
remaining = n remaining = n
@@ -262,9 +312,17 @@ class InferenceEngine:
temperature: float, temperature: float,
top_p: float, top_p: float,
top_k: int, top_k: int,
frequency_penalty: float,
rep_window: int,
) -> Union[str, List[str]]: ) -> Union[str, List[str]]:
result, task_ids = self._submit_tasks( result, task_ids = self._submit_tasks(
prompts, max_tokens, temperature, top_p, top_k prompts,
max_tokens,
temperature,
top_p,
top_k,
frequency_penalty,
rep_window,
) )
try: try:
+144 -13
View File
@@ -1,15 +1,15 @@
"""Composable sampling strategies for logit transformation. """Composable sampling strategies for logit transformation.
Implements the Strategy pattern: each sampling technique Implements the Strategy pattern: each sampling technique
(temperature, top-k, top-p) is a pluggable strategy that (temperature, top-k, top-p, frequency penalty) is a pluggable
can be composed into a pipeline. strategy that can be composed into a pipeline.
All strategies accept both scalar and per-sample tensor All strategies accept both scalar and per-sample tensor
parameters, so a single pipeline works for any batch size. parameters, so a single pipeline works for any batch size.
""" """
from abc import ABC, abstractmethod from abc import ABC, abstractmethod
from typing import List, Union from typing import List, Optional, Union
import torch import torch
from torch import Tensor from torch import Tensor
@@ -19,12 +19,23 @@ class BaseSamplingStrategy(ABC):
"""Abstract base for a logit transformation strategy.""" """Abstract base for a logit transformation strategy."""
@abstractmethod @abstractmethod
def apply(self, logits: Tensor, filter_value: float = -float("inf")) -> Tensor: def apply(
self,
logits: Tensor,
filter_value: float = -float("inf"),
input_ids: Optional[Tensor] = None,
input_mask: Optional[Tensor] = None,
) -> Tensor:
"""Applies the strategy to logits. """Applies the strategy to logits.
Args: Args:
logits: Raw logits tensor (batch, vocab_size). logits: Raw logits tensor (batch, vocab_size).
filter_value: Value assigned to filtered-out positions. filter_value: Value assigned to filtered-out positions.
input_ids: Previously generated token IDs ``[batch, seq_len]``,
padded with 0. Used by frequency penalty.
input_mask: Boolean mask ``[batch, seq_len]``, True for real
tokens, False for padding. Used to exclude padding from
penalty computation.
Returns: Returns:
Transformed logits tensor. Transformed logits tensor.
@@ -42,7 +53,13 @@ class TemperatureStrategy(BaseSamplingStrategy):
def __init__(self, temperature: Union[float, Tensor] = 1.0): def __init__(self, temperature: Union[float, Tensor] = 1.0):
self.temperature = temperature self.temperature = temperature
def apply(self, logits: Tensor, filter_value: float = -float("inf")) -> Tensor: def apply(
self,
logits: Tensor,
filter_value: float = -float("inf"),
input_ids: Optional[Tensor] = None,
input_mask: Optional[Tensor] = None,
) -> Tensor:
t = self.temperature t = self.temperature
if isinstance(t, Tensor): if isinstance(t, Tensor):
t = t.to(logits.device, non_blocking=True).view(-1, 1) t = t.to(logits.device, non_blocking=True).view(-1, 1)
@@ -64,7 +81,13 @@ class TopKStrategy(BaseSamplingStrategy):
def __init__(self, top_k: Union[int, Tensor] = 0): def __init__(self, top_k: Union[int, Tensor] = 0):
self.top_k = top_k self.top_k = top_k
def apply(self, logits: Tensor, filter_value: float = -float("inf")) -> Tensor: def apply(
self,
logits: Tensor,
filter_value: float = -float("inf"),
input_ids: Optional[Tensor] = None,
input_mask: Optional[Tensor] = None,
) -> Tensor:
tk = self.top_k tk = self.top_k
if isinstance(tk, Tensor): if isinstance(tk, Tensor):
tk = tk.to(logits.device, non_blocking=True).long().clamp(min=0) tk = tk.to(logits.device, non_blocking=True).long().clamp(min=0)
@@ -114,7 +137,13 @@ class TopPStrategy(BaseSamplingStrategy):
logits[mask] = filter_value logits[mask] = filter_value
return logits return logits
def apply(self, logits: Tensor, filter_value: float = -float("inf")) -> Tensor: def apply(
self,
logits: Tensor,
filter_value: float = -float("inf"),
input_ids: Optional[Tensor] = None,
input_mask: Optional[Tensor] = None,
) -> Tensor:
tp = self.top_p tp = self.top_p
if isinstance(tp, Tensor): if isinstance(tp, Tensor):
tp = tp.to(logits.device, non_blocking=True) tp = tp.to(logits.device, non_blocking=True)
@@ -125,6 +154,84 @@ class TopPStrategy(BaseSamplingStrategy):
return logits return logits
class FrequencyPenaltyStrategy(BaseSamplingStrategy):
"""Penalizes tokens based on how many times they appeared in history.
Subtracts ``penalty * count(token)`` from each token's logit, where
``count(token)`` is the number of occurrences in the generation history
(prompt + output). A penalty of ``0.0`` disables the strategy.
Unlike repetition penalty (which only checks *presence*), frequency
penalty scales linearly with occurrence count: the first use is
penalized once, the third use three times. This allows natural
repetition of common words while suppressing degenerate loops.
Reference: OpenAI API ``frequency_penalty`` parameter.
Args:
penalty: Scalar or ``[batch]`` tensor (0.0 disables, range -2.0~2.0).
"""
def __init__(self, penalty: Union[float, Tensor] = 0.0):
self.penalty = penalty
def apply(
self,
logits: Tensor,
filter_value: float = -float("inf"),
input_ids: Optional[Tensor] = None,
input_mask: Optional[Tensor] = None,
) -> Tensor:
if input_ids is None:
return logits
p = self.penalty
if isinstance(p, Tensor):
p = p.to(logits.device, non_blocking=True).view(-1, 1)
if (p == 0.0).all():
return logits
elif p == 0.0:
return logits
input_ids = input_ids.to(logits.device, non_blocking=True)
if input_mask is not None:
input_mask = input_mask.to(logits.device, non_blocking=True)
masked_ids = input_ids.clone()
masked_ids[~input_mask] = -1
else:
masked_ids = input_ids
batch_sz, seq_len = masked_ids.shape
vocab_size = logits.size(-1)
if isinstance(p, Tensor):
penalty_per_row = p.expand(batch_sz, 1)
else:
penalty_per_row = torch.full(
(batch_sz, 1), float(p), device=logits.device, dtype=logits.dtype
)
counts = torch.zeros(
batch_sz, vocab_size, device=logits.device, dtype=logits.dtype
)
valid_mask = masked_ids >= 0
if valid_mask.any():
valid_ids = masked_ids[valid_mask]
row_indices = (
torch.arange(batch_sz, device=logits.device)
.unsqueeze(1)
.expand_as(masked_ids)[valid_mask]
)
counts.index_put_(
(row_indices, valid_ids),
torch.ones_like(valid_ids, dtype=logits.dtype),
accumulate=True,
)
return logits - penalty_per_row * counts
class SamplingPipeline(BaseSamplingStrategy): class SamplingPipeline(BaseSamplingStrategy):
"""Composes multiple sampling strategies into a single transformation. """Composes multiple sampling strategies into a single transformation.
@@ -145,23 +252,39 @@ class SamplingPipeline(BaseSamplingStrategy):
def __init__(self, strategies: List[BaseSamplingStrategy]): def __init__(self, strategies: List[BaseSamplingStrategy]):
self.strategies = strategies self.strategies = strategies
def apply(self, logits: Tensor, filter_value: float = -float("inf")) -> Tensor: def apply(
self,
logits: Tensor,
filter_value: float = -float("inf"),
input_ids: Optional[Tensor] = None,
input_mask: Optional[Tensor] = None,
) -> Tensor:
for strategy in self.strategies: for strategy in self.strategies:
logits = strategy.apply(logits, filter_value) logits = strategy.apply(logits, filter_value, input_ids, input_mask)
return logits return logits
@torch.no_grad() @torch.inference_mode()
def sample(self, logits: Tensor, filter_value: float = -float("inf")) -> Tensor: def sample(
self,
logits: Tensor,
filter_value: float = -float("inf"),
input_ids: Optional[Tensor] = None,
input_mask: Optional[Tensor] = None,
) -> Tensor:
"""Apply strategies then sample (softmax + multinomial). """Apply strategies then sample (softmax + multinomial).
Args: Args:
logits: Raw logits ``[batch, vocab_size]``. logits: Raw logits ``[batch, vocab_size]``.
input_ids: Previously generated token IDs ``[batch, seq_len]``.
input_mask: Boolean mask for ``input_ids`` padding.
Returns: Returns:
Sampled token IDs ``[batch]``. Sampled token IDs ``[batch]``.
""" """
return torch.multinomial( return torch.multinomial(
torch.softmax(self.apply(logits, filter_value), dim=-1), torch.softmax(
self.apply(logits, filter_value, input_ids, input_mask), dim=-1
),
num_samples=1, num_samples=1,
).squeeze(-1) ).squeeze(-1)
@@ -172,6 +295,9 @@ def sample(
temperature: Union[float, Tensor] = 1.0, temperature: Union[float, Tensor] = 1.0,
top_k: Union[int, Tensor] = 0, top_k: Union[int, Tensor] = 0,
top_p: Union[float, Tensor] = 1.0, top_p: Union[float, Tensor] = 1.0,
frequency_penalty: Union[float, Tensor] = 0.0,
input_ids: Optional[Tensor] = None,
input_mask: Optional[Tensor] = None,
filter_value: float = -float("inf"), filter_value: float = -float("inf"),
) -> Tensor: ) -> Tensor:
"""Apply sampling strategies then sample (softmax + multinomial). """Apply sampling strategies then sample (softmax + multinomial).
@@ -180,6 +306,10 @@ def sample(
Args: Args:
logits: Raw logits ``[batch, vocab_size]``. logits: Raw logits ``[batch, vocab_size]``.
frequency_penalty: Penalty per occurrence for repeated tokens
(0.0 disables, range -2.0~2.0).
input_ids: Previously generated token IDs ``[batch, seq_len]``.
input_mask: Boolean mask for ``input_ids`` padding.
Returns: Returns:
Sampled token IDs ``[batch]``. Sampled token IDs ``[batch]``.
@@ -189,5 +319,6 @@ def sample(
TemperatureStrategy(temperature), TemperatureStrategy(temperature),
TopKStrategy(top_k), TopKStrategy(top_k),
TopPStrategy(top_p), TopPStrategy(top_p),
FrequencyPenaltyStrategy(frequency_penalty),
] ]
).sample(logits, filter_value) ).sample(logits, filter_value, input_ids, input_mask)
+28 -3
View File
@@ -7,6 +7,7 @@ from contextlib import contextmanager
from typing import Optional, Tuple from typing import Optional, Tuple
import torch import torch
import torch.distributed as dist
import torch.nn as nn import torch.nn as nn
from torch.distributed.fsdp import FullStateDictConfig, StateDictType from torch.distributed.fsdp import FullStateDictConfig, StateDictType
from torch.distributed.fsdp import FullyShardedDataParallel as FSDP from torch.distributed.fsdp import FullyShardedDataParallel as FSDP
@@ -120,6 +121,21 @@ class BaseExecutor:
def unwrap_model(self, model: nn.Module): def unwrap_model(self, model: nn.Module):
return model.state_dict() return model.state_dict()
@contextmanager
def checkpoint_context(self, model: nn.Module):
if self.use_distributed:
dist.barrier()
state_dict = self._gather_state_dict(model)
yield state_dict
if self.use_distributed:
dist.barrier()
def _gather_state_dict(self, model: nn.Module):
state_dict = self.unwrap_model(model)
if self.use_distributed and get_rank() != 0:
return None
return state_dict
@property @property
def use_distributed(self) -> bool: def use_distributed(self) -> bool:
return get_world_size() > 1 return get_world_size() > 1
@@ -132,7 +148,14 @@ class BaseExecutor:
def grad_accum_steps(self) -> int: def grad_accum_steps(self) -> int:
return self.gradient_state.num_steps return self.gradient_state.num_steps
def clip_grad_norm(self, model: nn.Module, max_norm: float) -> float: def clip_grad_norm(self, model: nn.Module, max_norm: Optional[float]) -> float:
if max_norm is None:
total_norm = torch.norm(
torch.stack(
[p.grad.norm(2) for p in model.parameters() if p.grad is not None]
)
)
return total_norm.item()
total_norm = torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm) total_norm = torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm)
if isinstance(total_norm, torch.Tensor): if isinstance(total_norm, torch.Tensor):
return total_norm.item() return total_norm.item()
@@ -266,7 +289,9 @@ class FSDPExecutor(BaseExecutor):
return model.no_sync() return model.no_sync()
return contextlib.nullcontext() return contextlib.nullcontext()
def clip_grad_norm(self, model: nn.Module, max_norm: float) -> float: def clip_grad_norm(self, model: nn.Module, max_norm: Optional[float]) -> float:
if max_norm is None:
return super().clip_grad_norm(model, max_norm)
if isinstance(model, FSDP) and self.use_distributed: if isinstance(model, FSDP) and self.use_distributed:
total_norm = model.clip_grad_norm_(max_norm) total_norm = model.clip_grad_norm_(max_norm)
if isinstance(total_norm, torch.Tensor): if isinstance(total_norm, torch.Tensor):
@@ -279,7 +304,7 @@ class FSDPExecutor(BaseExecutor):
with FSDP.state_dict_type( with FSDP.state_dict_type(
model, model,
StateDictType.FULL_STATE_DICT, StateDictType.FULL_STATE_DICT,
FullStateDictConfig(offload_to_cpu=True, rank0_only=False), FullStateDictConfig(offload_to_cpu=True, rank0_only=True),
): ):
return model.state_dict() return model.state_dict()
+4
View File
@@ -8,12 +8,14 @@ from astrai.preprocessing.builder import (
from astrai.preprocessing.packing import ( from astrai.preprocessing.packing import (
PackingStrategy, PackingStrategy,
PackingStrategyFactory, PackingStrategyFactory,
plan_bfd,
) )
from astrai.preprocessing.pipeline import Pipeline, filter_by_length from astrai.preprocessing.pipeline import Pipeline, filter_by_length
from astrai.preprocessing.position_id import ( from astrai.preprocessing.position_id import (
PositionIdStrategy, PositionIdStrategy,
PositionIdStrategyFactory, PositionIdStrategyFactory,
) )
from astrai.preprocessing.transform import TokenizeTransform
from astrai.preprocessing.writer import ( from astrai.preprocessing.writer import (
StoreWriter, StoreWriter,
StoreWriterFactory, StoreWriterFactory,
@@ -32,5 +34,7 @@ __all__ = [
"SingleOutputMaskBuilder", "SingleOutputMaskBuilder",
"StoreWriter", "StoreWriter",
"StoreWriterFactory", "StoreWriterFactory",
"TokenizeTransform",
"filter_by_length", "filter_by_length",
"plan_bfd",
] ]
+34 -21
View File
@@ -95,8 +95,15 @@ class SectionRenderer:
return all_ids, loss_mask return all_ids, loss_mask
def process_list_field(self, item: dict, sections: list, config, tokenizer): def process_list_field(self, item: dict, sections: list, config, tokenizer):
all_ids: list[int] = [] """Tokenize a list-valued field, preserving per-element boundaries.
loss_mask: list[int] = []
Returns ``(list_of_id_lists, list_of_mask_lists)`` where each
inner list corresponds to one element of the source list. This
is critical for GRPO where each response must stay a separate
sequence so the strategy can form a ``[G, R]`` tensor.
"""
per_item_ids: list[list[int]] = []
per_item_masks: list[list[int]] = []
for sec in sections: for sec in sections:
field = sec["field"] field = sec["field"]
@@ -108,17 +115,13 @@ class SectionRenderer:
continue continue
for val in values: for val in values:
ids: list[int] = []
mask: list[int] = []
if use_template: if use_template:
if isinstance(val, list): if isinstance(val, list):
wrapper = {field: val} wrapper = {field: val}
self._append_template( self._append_template(
wrapper, wrapper, field, action, tokenizer, config, ids, mask
field,
action,
tokenizer,
config,
all_ids,
loss_mask,
) )
else: else:
wrapper = {field: str(val)} wrapper = {field: str(val)}
@@ -130,17 +133,19 @@ class SectionRenderer:
False, False,
False, False,
config, config,
all_ids, ids,
loss_mask, mask,
) )
if ids:
max_len = config.preprocessing.max_seq_len
ids = ids[:max_len]
mask = mask[: len(ids)]
per_item_ids.append(ids)
per_item_masks.append(mask)
max_len = config.preprocessing.max_seq_len if not per_item_ids:
all_ids = all_ids[:max_len]
loss_mask = loss_mask[: len(all_ids)]
if not all_ids:
return None, None return None, None
return all_ids, loss_mask return per_item_ids, per_item_masks
@staticmethod @staticmethod
def is_value_section(sections: list) -> bool: def is_value_section(sections: list) -> bool:
@@ -282,10 +287,18 @@ class MultiOutputMaskBuilder(BaseMaskBuilder):
ids, mask = self.renderer.process_list_field( ids, mask = self.renderer.process_list_field(
item, sections, config, tokenizer item, sections, config, tokenizer
) )
else: if ids is None:
ids, mask = self.renderer.process_sections( continue
item, sections, config, tokenizer, is_top_level=True # ids is List[List[int]] — preserve per-response structure
) result[output_key] = ids
if mask is not None:
result[mask_key] = mask
any_output = True
continue
ids, mask = self.renderer.process_sections(
item, sections, config, tokenizer, is_top_level=True
)
if ids is None: if ids is None:
continue continue
+124
View File
@@ -0,0 +1,124 @@
"""Shared preprocessing kernel used by both :class:`Pipeline` and
:class:`TokenizeTransform`.
The two entry points previously duplicated ~60 % of their logic:
record iteration, mask-builder invocation, primary-id extraction,
per-key accumulation, dtype inference and position-id generation.
This module factors out the common core as pure functions so that
the online (``TokenizeTransform``) and offline (``Pipeline``) paths
stay in lockstep.
"""
from itertools import chain
from typing import Dict, Iterator, List, Optional
import torch
from astrai.config.preprocess_config import PipelineConfig
from astrai.preprocessing.builder import MaskBuilderFactory
from astrai.preprocessing.position_id import PositionIdStrategyFactory
from astrai.tokenize import AutoTokenizer
def build_preprocessing_components(config: PipelineConfig, tokenizer_path: str):
"""Load tokenizer, mask builder and position-id strategy together.
Both ``Pipeline`` and ``TokenizeTransform`` need the same triple;
centralising the construction avoids drift (e.g. one path forgetting
to create the position-id strategy).
"""
tokenizer = AutoTokenizer.from_pretrained(tokenizer_path)
mask_builder = MaskBuilderFactory.create("sectioned")
position_strategy = PositionIdStrategyFactory.create(
config.output.position_ids_mode
)
return tokenizer, mask_builder, position_strategy
def primary_ids(result: dict) -> List[int]:
"""Return the first flat int-list value in *result*.
Used for token counting and position-id generation when the
primary key name is not known (DPO uses ``chosen``, GRPO uses
``prompts``, SFT uses ``sequence``).
"""
for val in result.values():
if isinstance(val, list) and val and isinstance(val[0], int):
return val
return []
def infer_dtype(ids: List) -> torch.dtype:
"""Float values become float32, everything else int32."""
if ids and isinstance(ids[0], float):
return torch.float32
return torch.int32
def iter_raw_records(
records: List[dict],
mask_builder,
config: PipelineConfig,
tokenizer,
) -> Iterator[dict]:
"""Yield mask-builder output dicts for each record, skipping failures.
Drops ``domain`` from the result (callers that need it should read
it before calling this). Each yielded dict maps a key
(``sequence``, ``chosen``, ``responses``…) to either a flat
``List[int]`` or a nested ``List[List[int]]`` (GRPO responses/masks).
"""
for item in records:
result = mask_builder.build(item, config, tokenizer)
if result is None:
continue
result.pop("domain", None)
if not primary_ids(result):
continue
yield result
def to_per_record_tensors(
raw: Dict[str, list],
) -> Dict[str, List[torch.Tensor]]:
"""Convert an accumulated ``{key: [per-record ids]}`` dict to tensors.
Handles three shapes transparently:
- ``List[int]`` per record (``sequence``, ``chosen``…) → one tensor per record.
- ``List[List[int]]`` per record (GRPO ``responses``/``masks``) → one
``List[Tensor]`` per record (nested), preserving the per-response
boundary so downstream code can index responses individually.
- ``List[int]`` for the whole shard (pre-packed keys) → single tensor.
The detection mirrors the previous inline logic in
``Pipeline._flush`` and ``TokenizeTransform.apply``.
"""
tensors: Dict[str, List[torch.Tensor]] = {}
for key, ids_list in raw.items():
if ids_list and isinstance(ids_list[0], list):
tensors[key] = [
[torch.tensor(sub, dtype=infer_dtype(sub)) for sub in ids]
if ids and isinstance(ids[0], list)
else torch.tensor(ids, dtype=infer_dtype(ids))
for ids in ids_list
]
else:
tensors[key] = [
torch.tensor(list(chain.from_iterable(ids_list)), dtype=torch.int32)
]
return tensors
def build_position_ids(
sequences: List[List[int]],
strategy,
) -> Optional[List[int]]:
"""Generate position ids for *sequences* using *strategy*.
Returns ``None`` when the strategy produces no ids (e.g. ``none``
mode), so callers can skip attaching the key instead of storing
an empty list.
"""
pos_ids = strategy.generate(sequences)
return pos_ids or None
+83 -28
View File
@@ -19,6 +19,43 @@ def _truncate(seq: List[int], max_len: int, mode: str) -> List[int]:
return seq[:max_len] return seq[:max_len]
def plan_bfd(
sequences: List[List[int]], max_packed_len: int, truncation_mode: str = "keep_start"
) -> List[List[int]]:
"""Best-Fit Decreasing bin packing of *sequences* into bins.
Returns a list of bins, each bin a list of original indices into
*sequences*. Bin capacities are respected on the *truncated*
length of each sequence (so a sequence longer than
*max_packed_len* counts at *max_packed_len*).
Pure index-based so callers can apply the same plan to any
aligned key (``loss_mask``, ``position_ids``…).
"""
n = len(sequences)
order = sorted(range(n), key=lambda i: len(sequences[i]), reverse=True)
bins: List[List[int]] = []
bin_lengths: List[int] = []
for orig_idx in order:
seq_len = len(_truncate(sequences[orig_idx], max_packed_len, truncation_mode))
best_bin = None
best_remain = max_packed_len + 1
for i, bl in enumerate(bin_lengths):
remain = max_packed_len - bl
if seq_len <= remain < best_remain:
best_remain = remain
best_bin = i
if best_bin is not None:
bins[best_bin].append(orig_idx)
bin_lengths[best_bin] += seq_len
else:
bins.append([orig_idx])
bin_lengths.append(seq_len)
return bins
class PackingStrategy(ABC): class PackingStrategy(ABC):
"""Reorder and truncate sequences within a shard.""" """Reorder and truncate sequences within a shard."""
@@ -70,7 +107,7 @@ class BFDPacking(PackingStrategy):
sequences = keys.get("sequence", []) sequences = keys.get("sequence", [])
if not sequences: if not sequences:
return keys return keys
bins = self._plan(sequences, max_packed_len, truncation_mode) bins = plan_bfd(sequences, max_packed_len, truncation_mode)
packed: Dict[str, List[List[int]]] = {} packed: Dict[str, List[List[int]]] = {}
for k, vals in keys.items(): for k, vals in keys.items():
@@ -91,31 +128,49 @@ class BFDPacking(PackingStrategy):
result.extend(vals[i]) result.extend(vals[i])
return result return result
@PackingStrategyFactory.register("bfd_split")
class BFDSplitPacking(BFDPacking):
"""BFD packing with over-length sequences split into chunks.
Sequences longer than *max_packed_len* are split into consecutive
chunks of at most *max_packed_len* tokens instead of being
truncated. Each chunk becomes an independent sequence that enters
BFD planning. All keys (``loss_mask``, ``position_ids``, …) are
split in lockstep so per-token alignment is preserved.
Note: because each chunk is treated as a separate document, the
second chunk of a split sequence loses the preceding context.
"""
def apply(
self,
keys: Dict[str, List[List[int]]],
max_packed_len: int,
truncation_mode: str,
) -> Dict[str, List[List[int]]]:
sequences = keys.get("sequence", [])
if not sequences:
return keys
if max_packed_len <= 0:
return super().apply(keys, max_packed_len, truncation_mode)
split_keys = self._split_all(keys, max_packed_len)
return super().apply(split_keys, max_packed_len, truncation_mode)
@staticmethod @staticmethod
def _plan( def _split_all(
sequences: List[List[int]], max_packed_len: int, truncation_mode: str keys: Dict[str, List[List[int]]], max_packed_len: int
) -> List[List[int]]: ) -> Dict[str, List[List[int]]]:
n = len(sequences) """Split every sequence exceeding *max_packed_len* into chunks,
order = sorted(range(n), key=lambda i: len(sequences[i]), reverse=True) applying the same chunk boundaries to all keys."""
bins: List[List[int]] = [] sequences = keys["sequence"]
bin_lengths: List[int] = [] chunk_bounds = [list(range(0, len(s), max_packed_len)) for s in sequences]
result: Dict[str, List[List[int]]] = {}
for orig_idx in order: for key, vals in keys.items():
seq_len = len( split_vals: List[List[int]] = []
_truncate(sequences[orig_idx], max_packed_len, truncation_mode) for val, starts in zip(vals, chunk_bounds):
) for start in starts:
best_bin = None split_vals.append(val[start : start + max_packed_len])
best_remain = max_packed_len + 1 result[key] = split_vals
for i, bl in enumerate(bin_lengths): return result
remain = max_packed_len - bl
if seq_len <= remain < best_remain:
best_remain = remain
best_bin = i
if best_bin is not None:
bins[best_bin].append(orig_idx)
bin_lengths[best_bin] += seq_len
else:
bins.append([orig_idx])
bin_lengths.append(seq_len)
return bins
+104 -48
View File
@@ -4,6 +4,10 @@ Composes a :class:`BaseMaskBuilder` (selected by ``input.type``) with
sharding and flush to ``.h5`` / ``.bin`` storage. Packing, position-id sharding and flush to ``.h5`` / ``.bin`` storage. Packing, position-id
generation and storage writing are each delegated to pluggable strategies, generation and storage writing are each delegated to pluggable strategies,
dispatched by configuration keys. dispatched by configuration keys.
Record iteration, mask building, primary-id extraction and per-key
accumulation are shared with :class:`TokenizeTransform` via the
:mod:`astrai.preprocessing.core` helpers.
""" """
import json import json
@@ -17,11 +21,13 @@ import torch
import tqdm import tqdm
from astrai.config.preprocess_config import PipelineConfig from astrai.config.preprocess_config import PipelineConfig
from astrai.preprocessing.builder import MaskBuilderFactory from astrai.preprocessing.core import (
build_preprocessing_components,
iter_raw_records,
primary_ids,
)
from astrai.preprocessing.packing import PackingStrategyFactory from astrai.preprocessing.packing import PackingStrategyFactory
from astrai.preprocessing.position_id import PositionIdStrategyFactory
from astrai.preprocessing.writer import StoreWriterFactory from astrai.preprocessing.writer import StoreWriterFactory
from astrai.tokenize import AutoTokenizer
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
@@ -64,20 +70,18 @@ class Pipeline:
self.output_dir = output_dir self.output_dir = output_dir
self.tokenizer_path = tokenizer_path self.tokenizer_path = tokenizer_path
self.mask_builder = MaskBuilderFactory.create("sectioned") self.tokenizer, self.mask_builder, self._position_id = (
build_preprocessing_components(config, tokenizer_path)
)
self._packer = PackingStrategyFactory.create( self._packer = PackingStrategyFactory.create(
config.preprocessing.packing_strategy config.preprocessing.packing_strategy
) )
self._position_id = PositionIdStrategyFactory.create(
config.output.position_ids_mode
)
self._writer = StoreWriterFactory.create(config.output.storage_format) self._writer = StoreWriterFactory.create(config.output.storage_format)
def transform(self, item: dict) -> Optional[dict]: def transform(self, item: dict) -> Optional[dict]:
return self.mask_builder.build(item, self.config, self._tokenizer) return self.mask_builder.build(item, self.config, self.tokenizer)
def run(self): def run(self):
self._tokenizer = AutoTokenizer.from_pretrained(self.tokenizer_path)
domains: dict = defaultdict(lambda: defaultdict(list)) domains: dict = defaultdict(lambda: defaultdict(list))
total_tokens = 0 total_tokens = 0
shard_idx: dict[str, int] = defaultdict(int) shard_idx: dict[str, int] = defaultdict(int)
@@ -102,14 +106,7 @@ class Pipeline:
continue continue
domain = result.pop("domain", "__default__") domain = result.pop("domain", "__default__")
ids = primary_ids(result)
is_multi = bool(getattr(self.config.input, "sources", None))
if is_multi:
ids = self._primary_ids(result)
else:
ids = result.pop("sequence")
result["sequence"] = ids
if not ids: if not ids:
continue continue
@@ -129,15 +126,6 @@ class Pipeline:
if total_tokens > 0: if total_tokens > 0:
self._flush(domains, shard_idx) self._flush(domains, shard_idx)
@staticmethod
def _primary_ids(result: dict) -> list:
"""Return the first list-valued entry in *result* as the primary id
sequence for token counting."""
for val in result.values():
if isinstance(val, list) and val and isinstance(val[0], int):
return val
return []
@staticmethod @staticmethod
def _align_bucket(bucket: dict, result: dict, ids: list): def _align_bucket(bucket: dict, result: dict, ids: list):
"""Pad previously-accumulated keys that are missing from *result*.""" """Pad previously-accumulated keys that are missing from *result*."""
@@ -149,11 +137,18 @@ class Pipeline:
def _iter_items(self): def _iter_items(self):
for path in self.paths: for path in self.paths:
with open(path, "r", encoding="utf-8") as f: with open(path, "r", encoding="utf-8") as f:
for line in f: if path.endswith(".json"):
line = line.strip() data = json.load(f)
if not line: if isinstance(data, dict):
continue yield data
yield json.loads(line) elif isinstance(data, list):
yield from data
else:
for line in f:
line = line.strip()
if not line:
continue
yield json.loads(line)
def _flush(self, domains, shard_idx): def _flush(self, domains, shard_idx):
for domain, keys in domains.items(): for domain, keys in domains.items():
@@ -163,24 +158,12 @@ class Pipeline:
original_sequences = keys.get("sequence", []) original_sequences = keys.get("sequence", [])
mode = self.config.output.position_ids_mode mode = self.config.output.position_ids_mode
if mode == "doc_reset" and original_sequences: keys = self._inject_doc_reset_position_ids(keys, mode, original_sequences)
keys["position_ids"] = [list(range(len(s))) for s in original_sequences]
keys = self._packer.apply(dict(keys), pp.max_packed_len, pp.truncation_mode) keys = self._packer.apply(dict(keys), pp.max_packed_len, pp.truncation_mode)
tensors = self._to_tensors(keys)
tensors: Dict[str, List[torch.Tensor]] = {} tensors = self._inject_continuous_position_ids(
for key, ids_list in keys.items(): tensors, mode, keys.get("sequence", [])
dt = _STR_TO_DTYPE.get( )
self.config.output.dtype.get(key, "int32"), torch.int32
)
tensors[key] = [
torch.tensor(list(chain.from_iterable(ids_list)), dtype=dt)
]
if mode == "continuous" and original_sequences:
pos_ids = self._position_id.generate(keys.get("sequence", []))
if pos_ids:
tensors["position_ids"] = [torch.tensor(pos_ids, dtype=torch.int32)]
self._writer.save(self.output_dir, domain, idx, tensors) self._writer.save(self.output_dir, domain, idx, tensors)
shard_idx[domain] = idx + 1 shard_idx[domain] = idx + 1
@@ -190,3 +173,76 @@ class Pipeline:
f" saved {domain}/shard_{idx:04d} " f" saved {domain}/shard_{idx:04d} "
f"({tensors[first_key][0].numel():,} tokens)" f"({tensors[first_key][0].numel():,} tokens)"
) )
def _inject_doc_reset_position_ids(
self,
keys: Dict[str, list],
mode: str,
original_sequences: List[List[int]],
) -> Dict[str, list]:
"""Attach per-document position_ids before packing (``doc_reset``).
``doc_reset`` position ids must enter the packer so that each
packed bin concatenates the per-doc ranges in bin order. The
per-record structure ``[range(len(s)) for s in seqs]`` is required
by the packer (it concatenates per-record lists per bin); the
``PositionIdStrategy.generate`` flattens, so it cannot be used
directly here — it is only consulted for the ``continuous``
post-packing path.
"""
if mode != "doc_reset" or not original_sequences:
return keys
keys["position_ids"] = [list(range(len(s))) for s in original_sequences]
return keys
def _inject_continuous_position_ids(
self,
tensors: Dict[str, List[torch.Tensor]],
mode: str,
packed_sequences: List[List[int]],
) -> Dict[str, List[torch.Tensor]]:
"""Attach a single continuous position_ids tensor after packing.
``continuous`` mode spans the whole shard (post-packing), so it
cannot participate in bin packing — it is computed from the
packed sequences and appended directly to the tensor dict.
"""
if mode != "continuous" or not packed_sequences:
return tensors
pos_ids = self._position_id.generate(packed_sequences)
if pos_ids:
tensors["position_ids"] = [torch.tensor(pos_ids, dtype=torch.int32)]
return tensors
def _to_tensors(self, keys: Dict[str, list]) -> Dict[str, List[torch.Tensor]]:
"""Convert packed per-key id lists to tensors.
Honours ``config.output.dtype`` overrides per key; falls back to
``int32``. Handles three shapes (see
:func:`astrai.preprocessing.core.to_per_record_tensors` for the
equivalent online-path helper):
- ``List[int]`` per record → one tensor per record.
- ``List[List[int]]`` per record (GRPO responses/masks) → one tensor
per record, inner lists flattened.
- ``List[int]`` for the whole shard (pre-packed keys) → single tensor.
"""
tensors: Dict[str, List[torch.Tensor]] = {}
for key, ids_list in keys.items():
dt = _STR_TO_DTYPE.get(
self.config.output.dtype.get(key, "int32"), torch.int32
)
if ids_list and isinstance(ids_list[0], list):
tensors[key] = [
torch.tensor(
list(chain.from_iterable(ids))
if ids and isinstance(ids[0], list)
else ids,
dtype=dt,
)
for ids in ids_list
]
else:
tensors[key] = [
torch.tensor(list(chain.from_iterable(ids_list)), dtype=dt)
]
return tensors
+92
View File
@@ -0,0 +1,92 @@
"""Tokenization transform for JSONL record streams.
Bridges the Reader layer (``JsonlStore`` reads raw JSON records) and the
Dataset layer (expects per-record tensors). Holds the tokenizer,
mask-builder and position-id strategy together so that I/O code stays
free of model dependencies.
The record-processing core (mask building, primary-id extraction,
per-key tensorisation, position-id generation) is shared with
:class:`astrai.preprocessing.pipeline.Pipeline` via the
:mod:`astrai.preprocessing.core` helpers.
"""
import json
from pathlib import Path
from typing import Dict, List
import torch
from astrai.config.preprocess_config import PipelineConfig
from astrai.preprocessing.core import (
build_position_ids,
build_preprocessing_components,
iter_raw_records,
to_per_record_tensors,
)
class TokenizeTransform:
"""Tokenize raw JSONL record dicts into per-key tensor lists.
Owns the three preprocessing concerns that were previously inlined in
``JsonlStore``: tokenization, loss-mask construction and position-id
generation. Constructing it loads the tokenizer, so it is intentionally
cheap to pass around once built.
Args:
config: Pipeline config describing sections / masks / position mode.
tokenizer_path: Path passed to ``AutoTokenizer.from_pretrained``.
"""
def __init__(self, config: PipelineConfig, tokenizer_path: str):
self.config = config
self.tokenizer, self.mask_builder, self.position_strategy = (
build_preprocessing_components(config, tokenizer_path)
)
@classmethod
def from_config_file(cls, config_path: str) -> "TokenizeTransform":
"""Build from a ``dataset_config.json`` file path.
The config file follows :class:`PipelineConfig` schema with an
extra ``tokenizer_path`` field. When omitted, the config's
parent directory is used as the tokenizer path.
"""
root = Path(config_path).parent
with open(config_path, "r", encoding="utf-8") as f:
raw_config = json.load(f)
tokenizer_path = raw_config.pop("tokenizer_path", None) or str(root)
config = PipelineConfig.from_dict(raw_config)
return cls(config, tokenizer_path)
def apply(self, records: List[dict]) -> Dict[str, list]:
"""Tokenize a list of raw record dicts.
Returns a dict mapping key (``sequence``, ``chosen``, ``responses``,
…) to a list of per-record tensors (or nested tensor lists for
multi-response keys such as GRPO ``responses``).
"""
raw: Dict[str, list] = {}
doc_sequences: List[List[int]] = []
for result in iter_raw_records(
records, self.mask_builder, self.config, self.tokenizer
):
primary = None
for val in result.values():
if isinstance(val, list) and val and isinstance(val[0], int):
primary = val
break
if primary is not None:
doc_sequences.append(primary)
for key, ids in result.items():
raw.setdefault(key, []).append(ids)
tensors = to_per_record_tensors(raw)
pos_ids = build_position_ids(doc_sequences, self.position_strategy)
if pos_ids is not None:
tensors["position_ids"] = [torch.tensor(pos_ids, dtype=torch.int32)]
return tensors
+2
View File
@@ -19,6 +19,7 @@ from astrai.serialization.checkpoint import (
) )
from astrai.serialization.dataset import ( from astrai.serialization.dataset import (
load_bin, load_bin,
load_bin_offsets,
load_h5, load_h5,
save_bin, save_bin,
save_h5, save_h5,
@@ -37,6 +38,7 @@ __all__ = [
"save_safetensors", "save_safetensors",
"save_torch", "save_torch",
"load_bin", "load_bin",
"load_bin_offsets",
"load_h5", "load_h5",
"save_bin", "save_bin",
"save_h5", "save_h5",
+1 -4
View File
@@ -2,7 +2,6 @@
import io import io
import json import json
import os
import time import time
from dataclasses import dataclass, field from dataclasses import dataclass, field
from pathlib import Path from pathlib import Path
@@ -148,9 +147,6 @@ class Checkpoint:
save_path = Path(save_dir) save_path = Path(save_dir)
save_path.mkdir(parents=True, exist_ok=True) save_path.mkdir(parents=True, exist_ok=True)
if get_rank() != 0:
return
meta = { meta = {
"epoch": self.epoch, "epoch": self.epoch,
"consumed_samples": self.consumed_samples, "consumed_samples": self.consumed_samples,
@@ -181,6 +177,7 @@ class Checkpoint:
epoch=meta.get("epoch", 0), epoch=meta.get("epoch", 0),
consumed_samples=meta.get("consumed_samples", 0), consumed_samples=meta.get("consumed_samples", 0),
extra=extra, extra=extra,
meta=meta,
config=config, config=config,
) )
+50 -3
View File
@@ -3,7 +3,7 @@
import json import json
import os import os
from pathlib import Path from pathlib import Path
from typing import Dict, List from typing import Any, Dict, List, Optional
import h5py import h5py
import numpy as np import numpy as np
@@ -50,12 +50,43 @@ def load_h5(file_path: str, share_memory=True) -> Dict[str, List[Tensor]]:
return tensor_group return tensor_group
def save_bin(file_path: str, tensor_group: Dict[str, List[Tensor]]): def save_bin(
file_path: str,
tensor_group: Dict[str, List[Tensor]],
record_keys: Optional[List[str]] = None,
):
"""Save tensors as memory-mapped binary files.
When *record_keys* is provided, those keys are written with per-record
cumulative offsets in ``meta.json`` so that ``MmapStore.fetch_record``
can slice individual records from the concatenated binary without
cross-record concatenation. Keys not in *record_keys* (e.g. SEQ
``sequence``) are written as a single contiguous stream without
offsets, preserving backward compatibility.
Nested keys (``List[List[Tensor]]`` such as GRPO ``responses``) are
not supported in bin format — use H5 for those.
"""
os.makedirs(file_path, exist_ok=True) os.makedirs(file_path, exist_ok=True)
record_keys = set(record_keys or [])
meta = {} meta = {}
for key, tensors in tensor_group.items(): for key, tensors in tensor_group.items():
if tensors and isinstance(tensors[0], list):
raise ValueError(
f"Nested key '{key}' (List[List[Tensor]]) is not supported "
f"in bin format. Use H5 or JSONL storage instead."
)
cat = torch.cat(tensors, dim=0) cat = torch.cat(tensors, dim=0)
meta[key] = {"shape": list(cat.shape), "dtype": str(cat.dtype).split(".")[-1]} entry: Dict[str, Any] = {
"shape": list(cat.shape),
"dtype": str(cat.dtype).split(".")[-1],
}
if key in record_keys:
offsets = [0]
for t in tensors:
offsets.append(offsets[-1] + t.shape[0])
entry["offsets"] = offsets
meta[key] = entry
np.asarray(cat.cpu().numpy()).tofile(os.path.join(file_path, f"{key}.bin")) np.asarray(cat.cpu().numpy()).tofile(os.path.join(file_path, f"{key}.bin"))
with open(os.path.join(file_path, "meta.json"), "w") as f: with open(os.path.join(file_path, "meta.json"), "w") as f:
json.dump(meta, f) json.dump(meta, f)
@@ -74,3 +105,19 @@ def load_bin(file_path: str) -> Dict[str, List[Tensor]]:
) )
segments[key] = [torch.from_numpy(arr)] segments[key] = [torch.from_numpy(arr)]
return segments return segments
def load_bin_offsets(file_path: str) -> Dict[str, List[int]]:
"""Read per-record cumulative offsets from ``meta.json``.
Returns an empty dict when no key has offsets (legacy bin files),
in which case record-mode access falls back to per-record segment
indexing (H5/JSONL layout).
"""
with open(os.path.join(file_path, "meta.json"), "r") as f:
meta = json.load(f)
offsets: Dict[str, List[int]] = {}
for key, info in meta.items():
if "offsets" in info:
offsets[key] = info["offsets"]
return offsets
+14 -1
View File
@@ -1,3 +1,4 @@
from functools import cached_property
from typing import Any, Dict, List, Optional from typing import Any, Dict, List, Optional
from jinja2 import Template from jinja2 import Template
@@ -29,7 +30,19 @@ class ChatTemplate:
self.description = description self.description = description
self.default_variables = default_variables or {} self.default_variables = default_variables or {}
self.special_tokens = special_tokens or {} self.special_tokens = special_tokens or {}
self._compiled: Template = Template(template_str)
@cached_property
def _compiled(self) -> Template:
"""Lazy-compiled Jinja2 template, cached on first access.
The compiled :class:`~jinja2.Template` holds a dynamically-generated
``root`` render function whose ``__module__`` is ``None``; under
``pickle`` it falls back to ``__main__`` and breaks ``spawn``-based
multiprocessing. By deferring compilation to first access, the
default pickle protocol serialises only ``template_str``; each
worker rebuilds the cache on first render.
"""
return Template(self.template_str)
@classmethod @classmethod
def from_string( def from_string(
+7
View File
@@ -164,7 +164,14 @@ class AutoTokenizer:
- tokenizer.bos_token → returns string - tokenizer.bos_token → returns string
- tokenizer.bos_token_id → returns corresponding integer ID - tokenizer.bos_token_id → returns corresponding integer ID
- tokenizer.stop_ids → returns list of corresponding integer IDs for all special tokens - tokenizer.stop_ids → returns list of corresponding integer IDs for all special tokens
Internal/private attrs are not intercepted: during unpickle
``__dict__`` is empty, so probing ``self._special_token_map``
would recurse infinitely.
""" """
if key.startswith("_"):
raise AttributeError(key)
# Handle stop_ids - return IDs for all special tokens # Handle stop_ids - return IDs for all special tokens
if key == "stop_ids": if key == "stop_ids":
stop_ids = [] stop_ids = []
+66 -38
View File
@@ -98,7 +98,6 @@ class BaseStrategy(ABC):
self.model = model self.model = model
self.device = device self.device = device
self.executor = kwargs.pop("executor", None) self.executor = kwargs.pop("executor", None)
self.model_fn = kwargs.pop("model_fn", None)
self.extra_kwargs = kwargs self.extra_kwargs = kwargs
@abstractmethod @abstractmethod
@@ -223,14 +222,13 @@ class DPOStrategy(BaseStrategy):
self, self,
model: nn.Module, model: nn.Module,
device: str, device: str,
ref_model: nn.Module,
beta: float = 0.1, beta: float = 0.1,
reduction: str = "mean", reduction: str = "sum",
**kwargs, **kwargs,
): ):
super().__init__(model, device, **kwargs) super().__init__(model, device, **kwargs)
self.ref_model = create_ref_model( self.ref_model = ref_model
self.model_fn, self.executor.unwrap_model(model)
).to(device=self.device)
self.beta = beta self.beta = beta
self.reduction = reduction self.reduction = reduction
@@ -267,42 +265,45 @@ class DPOStrategy(BaseStrategy):
class GRPOStrategy(BaseStrategy): class GRPOStrategy(BaseStrategy):
"""Group Relative Policy Optimization strategy. """Group Relative Policy Optimization strategy.
On-policy GRPO following DeepSeek-R1: the policy model is updated while Implements GRPO following DeepSeek-R1 with token-level PPO clipping.
a frozen ref_model stores the old-policy log-probs. ratio = exp(logπ_θ - logπ_ref), Advantages are group-normalized from scalar per-response rewards and
clipped PPO objective. Call ``sync_ref_model()`` after each data-generation round. broadcast across all response tokens. The loss is computed **only on
response tokens** — prompt tokens are masked out.
Three model roles are distinguished:
* **Policy** ``self.model`` — the model being trained.
* **Old policy** ``self.old_model`` — the behaviour policy that generated
the responses. Used for the importance sampling ratio
``ρ = π_θ / π_old``. Synced externally after each data-generation round.
* **Reference model** ``self.ref_model`` — a frozen copy of the initial
policy (typically the SFT checkpoint) used **only** for the KL
regularisation term. It is never updated during training.
""" """
def __init__( def __init__(
self, self,
model: nn.Module, model: nn.Module,
device: str, device: str,
old_model: nn.Module,
ref_model: nn.Module,
clip_eps: float = 0.2, clip_eps: float = 0.2,
kl_coef: float = 0.01, kl_coef: float = 0.01,
group_size: int = 4, group_size: int = 4,
reduction: str = "mean",
sync_interval: int = 200,
**kwargs, **kwargs,
): ):
super().__init__(model, device, **kwargs) super().__init__(model, device, **kwargs)
self.ref_model = create_ref_model( self.old_model = old_model
self.model_fn, self.executor.unwrap_model(model) self.ref_model = ref_model
).to(device=self.device)
self.clip_eps = clip_eps self.clip_eps = clip_eps
self.kl_coef = kl_coef self.kl_coef = kl_coef
self.group_size = group_size self.group_size = group_size
self.reduction = reduction
self.sync_interval = sync_interval
self._step = 0
def sync_ref_model(self): def sync_old_model(self):
"""Copy current model weights to ref model.""" """Copy current policy weights to old model."""
self.ref_model.load_state_dict(self.executor.unwrap_model(self.model)) self.old_model.load_state_dict(self.executor.unwrap_model(self.model))
def compute_loss(self, batch: Dict[str, Tensor]) -> Tensor: def compute_loss(self, batch: Dict[str, Tensor]) -> Tensor:
self._step += 1
if self._step % self.sync_interval == 0:
self.sync_ref_model()
batch = move_to_device(batch, self.device) batch = move_to_device(batch, self.device)
prompts = batch["prompts"] prompts = batch["prompts"]
responses = batch["responses"] responses = batch["responses"]
@@ -313,33 +314,60 @@ class GRPOStrategy(BaseStrategy):
responses_flat = responses.view(-1, response_len) responses_flat = responses.view(-1, response_len)
masks_flat = masks.view(-1, response_len) masks_flat = masks.view(-1, response_len)
prompt_expanded = prompts.unsqueeze(1).repeat(1, group_size, 1).flatten(0, 1) prompt_expanded = prompts.unsqueeze(1).repeat(1, group_size, 1).flatten(0, 1)
prompt_len = prompt_expanded.size(1)
full_sequences = torch.cat([prompt_expanded, responses_flat], dim=-1) full_sequences = torch.cat([prompt_expanded, responses_flat], dim=-1)
full_masks = torch.cat([torch.ones_like(prompt_expanded), masks_flat], dim=-1) # Prompt tokens are masked out (0) so logprobs are computed only for
# response tokens. get_logprobs shifts the mask by one position, so
log_probs_policy = get_logprobs( # the first response token's logprob (predicted from the last prompt
self.model, full_sequences, full_masks, self.reduction # token) is correctly included.
) full_masks = torch.cat([torch.zeros_like(prompt_expanded), masks_flat], dim=-1)
log_probs_policy = log_probs_policy.view(batch_size, group_size)
# get_logprobs returns [B*G, S-1] (S = prompt_len + response_len).
# Response token logprobs occupy the last ``response_len`` positions
# (the first response token is predicted from the last prompt token).
token_log_probs_policy = get_logprobs(
self.model, full_sequences, full_masks, "none"
)[:, prompt_len - 1 :]
with torch.no_grad(): with torch.no_grad():
log_probs_ref = get_logprobs( token_log_probs_old = get_logprobs(
self.ref_model, full_sequences, full_masks, self.reduction self.old_model, full_sequences, full_masks, "none"
) )[:, prompt_len - 1 :]
log_probs_ref = log_probs_ref.view(batch_size, group_size) token_log_probs_ref = get_logprobs(
self.ref_model, full_sequences, full_masks, "none"
)[:, prompt_len - 1 :]
eps = torch.finfo(log_probs_policy.dtype).eps # Reshape to [B, G, response_len]
token_log_probs_policy = token_log_probs_policy.view(batch_size, group_size, -1)
token_log_probs_old = token_log_probs_old.view(batch_size, group_size, -1)
token_log_probs_ref = token_log_probs_ref.view(batch_size, group_size, -1)
token_masks = masks_flat.view(batch_size, group_size, -1).float()
# Group-normalized advantages from scalar per-response rewards.
eps = 1e-8
mean = rewards.mean(dim=-1, keepdim=True) mean = rewards.mean(dim=-1, keepdim=True)
std = rewards.std(dim=-1, keepdim=True) std = rewards.std(dim=-1, keepdim=True, unbiased=False)
advantages = (rewards - mean) / (std + eps) advantages = (rewards - mean) / (std + eps)
# Broadcast scalar advantage to every response token: [B, G, 1]
advantages = advantages.unsqueeze(-1)
ratio = torch.exp(log_probs_policy - log_probs_ref) # Token-level ratio (π_θ / π_old) and PPO clipping.
log_ratio = token_log_probs_policy - token_log_probs_old
ratio = torch.exp(log_ratio)
surr1 = ratio * advantages surr1 = ratio * advantages
surr2 = torch.clamp(ratio, 1 - self.clip_eps, 1 + self.clip_eps) * advantages surr2 = torch.clamp(ratio, 1 - self.clip_eps, 1 + self.clip_eps) * advantages
per_token_policy_loss = -torch.min(surr1, surr2)
token_count = token_masks.sum().clamp(min=1.0)
policy_loss = (per_token_policy_loss * token_masks).sum() / token_count
# KL penalty to frozen reference model with k1 estimator (non-negative):
# k1 = π_ref / π_θ - log(π_ref / π_θ) - 1, where π_ref / π_θ = exp(log_ref - log_policy).
log_ref_ratio = token_log_probs_ref - token_log_probs_policy
r = torch.exp(log_ref_ratio)
kl_per_token = r - torch.log(r + eps) - 1.0
kl_penalty = self.kl_coef * (kl_per_token * token_masks).sum() / token_count
policy_loss = -torch.min(surr1, surr2).mean()
kl_penalty = self.kl_coef * (log_probs_policy - log_probs_ref).square().mean()
total_loss = policy_loss + kl_penalty total_loss = policy_loss + kl_penalty
return total_loss return total_loss
+33 -23
View File
@@ -14,7 +14,7 @@ from tqdm import tqdm
from astrai.factory import BaseFactory from astrai.factory import BaseFactory
from astrai.parallel import only_on_rank from astrai.parallel import only_on_rank
from astrai.parallel.setup import get_current_device, get_rank from astrai.parallel.setup import get_current_device
from astrai.serialization import Checkpoint from astrai.serialization import Checkpoint
from astrai.trainer.metric_util import ( from astrai.trainer.metric_util import (
ctx_get_grad_norm, ctx_get_grad_norm,
@@ -139,28 +139,31 @@ class CheckpointCallback(TrainCallback):
self.interval = interval self.interval = interval
self.weight_only = weight_only self.weight_only = weight_only
self.save_extra_fn = save_extra_fn or CheckpointCallback.save_extra self.save_extra_fn = save_extra_fn or CheckpointCallback.save_extra
self.last_ckpt_step = 0 self.last_ckpt_step = None
def _save_checkpoint(self, context: TrainContext): def on_train_begin(self, context: TrainContext):
state_dict = context.executor.unwrap_model(context.model)
self.last_ckpt_step = context.optimizer_step self.last_ckpt_step = context.optimizer_step
if get_rank() == 0: def _save_checkpoint(self, context: TrainContext):
save_path = os.path.join( self.last_ckpt_step = context.optimizer_step
self.save_dir,
f"epoch_{context.epoch}_step_{context.optimizer_step}", with context.executor.checkpoint_context(context.model) as state_dict:
) if state_dict is not None:
extra = self.save_extra_fn(context) save_path = os.path.join(
meta = context.config.to_dict() self.save_dir,
context.checkpoint = Checkpoint( f"epoch_{context.epoch}_step_{context.optimizer_step}",
state_dict=state_dict, )
epoch=context.epoch, extra = self.save_extra_fn(context)
consumed_samples=context.consumed_samples, meta = context.config.to_dict()
config=context.model_config, context.checkpoint = Checkpoint(
extra=extra, state_dict=state_dict,
meta=meta, epoch=context.epoch,
) consumed_samples=context.consumed_samples,
context.checkpoint.save(save_path) config=context.model_config,
extra=extra,
meta=meta,
)
context.checkpoint.save(save_path)
def on_batch_end(self, context: TrainContext): def on_batch_end(self, context: TrainContext):
if context.optimizer_step - self.last_ckpt_step >= self.interval: if context.optimizer_step - self.last_ckpt_step >= self.interval:
@@ -210,7 +213,7 @@ class ProgressBarCallback(TrainCallback):
@only_on_rank(0) @only_on_rank(0)
def on_optimizer_step(self, context: TrainContext): def on_optimizer_step(self, context: TrainContext):
postfix = { postfix = {
"step": context.optimizer_step, "step": f"{context.optimizer_step:d}",
"loss": f"{context.loss:.4f}", "loss": f"{context.loss:.4f}",
"lr": f"{context.optimizer.param_groups[-1]['lr']:.2e}", "lr": f"{context.optimizer.param_groups[-1]['lr']:.2e}",
} }
@@ -237,7 +240,7 @@ class MetricCallback(TrainCallback):
metrics: List[str] = None, metrics: List[str] = None,
val_step: int = 0, val_step: int = 0,
): ):
self.last_log_flush_step = 0 self.last_log_flush_step = None
self.save_interval = save_interval self.save_interval = save_interval
self.metrics = metrics or ["loss", "lr"] self.metrics = metrics or ["loss", "lr"]
self.val_step = val_step self.val_step = val_step
@@ -298,6 +301,9 @@ class MetricCallback(TrainCallback):
context.model.train() context.model.train()
return avg_loss return avg_loss
def on_train_begin(self, context: TrainContext):
self.last_log_flush_step = context.optimizer_step
@only_on_rank(0) @only_on_rank(0)
def _flush(self, epoch, step): def _flush(self, epoch, step):
log_file = self.log_dir / f"epoch_{epoch}_step_{step}_metric.jsonl" log_file = self.log_dir / f"epoch_{epoch}_step_{step}_metric.jsonl"
@@ -327,8 +333,12 @@ class MetricCallback(TrainCallback):
self._append("epoch", context) self._append("epoch", context)
def on_train_end(self, context): def on_train_end(self, context):
if context.optimizer_step != self.last_log_flush_step: if (
self.last_log_flush_step is None
or context.optimizer_step != self.last_log_flush_step
):
self._flush(context.epoch, context.optimizer_step) self._flush(context.epoch, context.optimizer_step)
self.last_log_flush_step = context.optimizer_step
def on_error(self, context): def on_error(self, context):
self._flush(context.epoch, context.optimizer_step) self._flush(context.epoch, context.optimizer_step)
+47 -19
View File
@@ -7,13 +7,13 @@ import torch.nn as nn
from torch.utils.data import DataLoader, random_split from torch.utils.data import DataLoader, random_split
from astrai.config.train_config import TrainConfig from astrai.config.train_config import TrainConfig
from astrai.dataset import ResumableDistributedSampler from astrai.dataset import RDSampler
from astrai.model.components.lora import inject_lora from astrai.model.components.lora import inject_lora
from astrai.parallel.executor import BaseExecutor, ExecutorFactory from astrai.parallel.executor import BaseExecutor, ExecutorFactory
from astrai.parallel.setup import get_current_device, get_rank, get_world_size from astrai.parallel.setup import get_current_device, get_rank, get_world_size
from astrai.protocols import OptimizerProtocol, SchedulerProtocol from astrai.protocols import OptimizerProtocol, SchedulerProtocol
from astrai.serialization import Checkpoint, load_json from astrai.serialization import Checkpoint, load_json
from astrai.trainer.strategy import BaseStrategy, StrategyFactory from astrai.trainer.strategy import BaseStrategy, StrategyFactory, create_ref_model
@dataclass @dataclass
@@ -54,10 +54,12 @@ class TrainContextBuilder:
config: TrainConfig, config: TrainConfig,
): ):
self.config = config self.config = config
self._resume_dir: Optional[str] = None self._param_path: Optional[str] = None
self._resume: bool = False
def with_resume_dir(self, resume_dir: Optional[str]) -> Self: def with_param_path(self, param_path: Optional[str], resume: bool = False) -> Self:
self._resume_dir = resume_dir self._param_path = param_path
self._resume = resume
return self return self
def build(self) -> TrainContext: def build(self) -> TrainContext:
@@ -74,8 +76,8 @@ class TrainContextBuilder:
model = model.to(device=device) model = model.to(device=device)
model_config = {} model_config = {}
if self._resume_dir: if self._param_path:
config_path = Path(self._resume_dir) / "config.json" config_path = Path(self._param_path) / "config.json"
if config_path.exists(): if config_path.exists():
model_config = load_json(config_path) model_config = load_json(config_path)
@@ -91,18 +93,29 @@ class TrainContextBuilder:
executor=executor, executor=executor,
) )
if self._resume_dir: if self._param_path:
checkpoint = Checkpoint.load_any(self._resume_dir) checkpoint = Checkpoint.load_any(self._param_path)
if checkpoint is not None: if checkpoint is not None:
model.load_state_dict(checkpoint.state_dict, strict=False) model.load_state_dict(checkpoint.state_dict, strict=False)
if checkpoint.config: if checkpoint.config:
context.model_config = checkpoint.config context.model_config = checkpoint.config
context.epoch = checkpoint.epoch or cfg.start_epoch
if checkpoint.consumed_samples > 0: if self._resume:
context.consumed_samples = checkpoint.consumed_samples context.epoch = checkpoint.epoch or cfg.start_epoch
else: if checkpoint.consumed_samples > 0:
context.consumed_samples = cfg.start_samples * context.world_size per_step = (
context.checkpoint = checkpoint cfg.batch_per_device
* context.world_size
* cfg.grad_accum_steps
)
context.consumed_samples = (
checkpoint.consumed_samples // per_step
) * per_step
else:
context.consumed_samples = (
cfg.start_samples * context.world_size
)
context.checkpoint = checkpoint
if cfg.lora is not None: if cfg.lora is not None:
inject_lora( inject_lora(
@@ -128,7 +141,7 @@ class TrainContextBuilder:
) )
sampler_offset = context.consumed_samples // context.world_size sampler_offset = context.consumed_samples // context.world_size
sampler = ResumableDistributedSampler( sampler = RDSampler(
data_source=train_dataset, data_source=train_dataset,
start_epoch=context.epoch, start_epoch=context.epoch,
start_iter=sampler_offset, start_iter=sampler_offset,
@@ -141,10 +154,11 @@ class TrainContextBuilder:
num_workers=cfg.num_workers, num_workers=cfg.num_workers,
pin_memory=cfg.pin_memory, pin_memory=cfg.pin_memory,
prefetch_factor=cfg.prefetch_factor, prefetch_factor=cfg.prefetch_factor,
collate_fn=cfg.collate_fn,
) )
if val_dataset is not None: if val_dataset is not None:
val_sampler = ResumableDistributedSampler( val_sampler = RDSampler(
data_source=val_dataset, data_source=val_dataset,
start_epoch=0, start_epoch=0,
start_iter=0, start_iter=0,
@@ -158,6 +172,7 @@ class TrainContextBuilder:
num_workers=cfg.num_workers, num_workers=cfg.num_workers,
pin_memory=cfg.pin_memory, pin_memory=cfg.pin_memory,
prefetch_factor=cfg.prefetch_factor, prefetch_factor=cfg.prefetch_factor,
collate_fn=cfg.collate_fn,
) )
context.model, context.optimizer, context.dataloader, context.scheduler = ( context.model, context.optimizer, context.dataloader, context.scheduler = (
@@ -177,13 +192,26 @@ class TrainContextBuilder:
if obj is not None: if obj is not None:
obj.load_state_dict(extra[name]) obj.load_state_dict(extra[name])
strategy_kwargs = dict(cfg.extra_kwargs)
if cfg.strategy in ("dpo", "grpo"):
ref_model = create_ref_model(
cfg.model_fn, executor.unwrap_model(context.model)
).to(device=device)
strategy_kwargs["ref_model"] = ref_model
if cfg.strategy == "grpo":
old_model = create_ref_model(
cfg.model_fn, executor.unwrap_model(context.model)
).to(device=device)
strategy_kwargs["old_model"] = old_model
context.strategy = StrategyFactory.create( context.strategy = StrategyFactory.create(
cfg.strategy, cfg.strategy,
model=context.model, model=context.model,
device=device, device=device,
executor=executor, executor=executor,
model_fn=cfg.model_fn, **strategy_kwargs,
**cfg.extra_kwargs,
) )
return context return context
+7 -4
View File
@@ -52,9 +52,11 @@ class Trainer:
if method: if method:
method(context) method(context)
def _trainer_loop(self, resume_dir: Optional[str] = None): def _trainer_loop(self, param_path: Optional[str] = None, resume: bool = False):
context = ( context = (
TrainContextBuilder(self.train_config).with_resume_dir(resume_dir).build() TrainContextBuilder(self.train_config)
.with_param_path(param_path, resume=resume)
.build()
) )
executor = context.executor executor = context.executor
self._call_callbacks("on_train_begin", context) self._call_callbacks("on_train_begin", context)
@@ -95,7 +97,7 @@ class Trainer:
finally: finally:
self._call_callbacks("on_train_end", context) self._call_callbacks("on_train_end", context)
def train(self, resume_dir: Optional[str] = None): def train(self, param_path: Optional[str] = None, resume: bool = False):
cfg = self.train_config cfg = self.train_config
spawn_parallel_fn( spawn_parallel_fn(
self._trainer_loop, self._trainer_loop,
@@ -105,5 +107,6 @@ class Trainer:
master_port=cfg.master_port, master_port=cfg.master_port,
device_type=cfg.device_type, device_type=cfg.device_type,
start_method=cfg.start_method, start_method=cfg.start_method,
resume_dir=resume_dir, param_path=param_path,
resume=resume,
) )
+18 -4
View File
@@ -10,11 +10,19 @@ struct AttentionParams {
int kv_len; int kv_len;
int head_dim; int head_dim;
int use_mask; int use_mask;
int is_causal; int causal_offset; // -1 = non-causal; >=0 = absolute position of first Q token
int causal_offset;
int num_splits; int num_splits;
float scale; float scale;
// Q strides (element offsets for each dim — layout-agnostic)
int q_stride_b, q_stride_h, q_stride_l, q_stride_d;
// KV strides (K and V share the same layout — only base pointers differ)
int kv_stride_b, kv_stride_h, kv_stride_l, kv_stride_d;
// Mask: 2D [batch, kv_len] (mask_q_stride=0) or 3D [batch, q_len, kv_len]
int mask_b_stride; // = kv_len (both 2D and 3D)
int mask_q_stride; // 2D: 0 (all q rows share); 3D: kv_len
const T* __restrict__ q; const T* __restrict__ q;
const T* __restrict__ k; const T* __restrict__ k;
const T* __restrict__ v; const T* __restrict__ v;
@@ -34,7 +42,6 @@ struct PagedAttentionParams {
int kv_len; int kv_len;
int head_dim; int head_dim;
int use_mask; int use_mask;
int is_causal;
int causal_offset; int causal_offset;
float scale; float scale;
@@ -42,12 +49,19 @@ struct PagedAttentionParams {
int page_size; int page_size;
int max_pages; int max_pages;
// Q strides (layout-agnostic)
int q_stride_b, q_stride_h, q_stride_l, q_stride_d;
// Mask strides (2D or 3D)
int mask_b_stride;
int mask_q_stride;
const T* __restrict__ q; const T* __restrict__ q;
const T* __restrict__ k_cache; const T* __restrict__ k_cache;
const T* __restrict__ v_cache; const T* __restrict__ v_cache;
const bool* __restrict__ mask; const bool* __restrict__ mask;
const int64_t* __restrict__ page_table; const int64_t* __restrict__ page_table;
T* __restrict__ o; T* __restrict__ o;
AT* __restrict__ o_part; AT* __restrict__ o_part;
AT* __restrict__ ml_part; AT* __restrict__ ml_part;
+17 -48
View File
@@ -5,24 +5,12 @@
#include "attn_decode_split_kv_mma.cuh" #include "attn_decode_split_kv_mma.cuh"
#endif #endif
static int decode_num_splits(int base_blocks, int tiles_total) {
int sm_count = 0;
cudaDeviceGetAttribute(&sm_count, cudaDevAttrMultiProcessorCount, 0);
int n = (2 * sm_count + base_blocks - 1) / base_blocks;
return std::max(1, std::min(n, std::min(tiles_total, 32)));
}
// Scalar fallback: one warp per query head, split-KV across grid.z. // Scalar fallback: one warp per query head, split-KV across grid.z.
static void launch_scalar_decode(AttentionParams<bf16>& p) { static void launch_scalar_decode(AttentionParams<bf16>& p) {
int group_size = p.q_head / p.kv_head; int group_size = p.q_head / p.kv_head;
int chunks_total = (p.kv_len + DC_CHUNK - 1) / DC_CHUNK; int chunks_total = (p.kv_len + DC_CHUNK - 1) / DC_CHUNK;
p.num_splits = decode_num_splits(p.batch * p.kv_head, chunks_total); p.num_splits = compute_num_splits(p.batch * p.kv_head, chunks_total);
alloc_split_partials(p);
auto fopt = torch::TensorOptions().dtype(torch::kFloat32).device(torch::kCUDA);
auto o_part = torch::empty({p.batch, p.q_head, p.num_splits, p.head_dim}, fopt);
auto ml_part = torch::empty({p.batch, p.q_head, p.num_splits, 2}, fopt);
p.o_part = o_part.data_ptr<float>();
p.ml_part = ml_part.data_ptr<float>();
size_t smem = DC_CHUNK * p.head_dim * sizeof(bf16); size_t smem = DC_CHUNK * p.head_dim * sizeof(bf16);
attn_decode_split_kv_kernel<<<dim3(p.batch * p.kv_head, 1, p.num_splits), dim3(32, group_size), smem>>>(p); attn_decode_split_kv_kernel<<<dim3(p.batch * p.kv_head, 1, p.num_splits), dim3(32, group_size), smem>>>(p);
@@ -38,13 +26,8 @@ static void launch_scalar_decode(AttentionParams<bf16>& p) {
template <int HEAD_DIM, int BC, int STAGES = (HEAD_DIM <= 128) ? 2 : 1> template <int HEAD_DIM, int BC, int STAGES = (HEAD_DIM <= 128) ? 2 : 1>
static void launch_mma_decode(AttentionParams<bf16>& p) { static void launch_mma_decode(AttentionParams<bf16>& p) {
int tiles_total = (p.kv_len + BC - 1) / BC; int tiles_total = (p.kv_len + BC - 1) / BC;
p.num_splits = decode_num_splits(p.batch * p.kv_head, tiles_total); p.num_splits = compute_num_splits(p.batch * p.kv_head, tiles_total);
alloc_split_partials(p);
auto fopt = torch::TensorOptions().dtype(torch::kFloat32).device(torch::kCUDA);
auto o_part = torch::empty({p.batch, p.q_head, p.num_splits, p.head_dim}, fopt);
auto ml_part = torch::empty({p.batch, p.q_head, p.num_splits, 2}, fopt);
p.o_part = o_part.data_ptr<float>();
p.ml_part = ml_part.data_ptr<float>();
attn_decode_split_kv_mma_kernel<HEAD_DIM, BC, STAGES><<<dim3(p.kv_head, p.batch, p.num_splits), 32>>>(p); attn_decode_split_kv_mma_kernel<HEAD_DIM, BC, STAGES><<<dim3(p.kv_head, p.batch, p.num_splits), 32>>>(p);
attn_decode_combine_kernel<<<p.batch * p.q_head, p.head_dim>>>(p); attn_decode_combine_kernel<<<p.batch * p.q_head, p.head_dim>>>(p);
@@ -55,7 +38,7 @@ template <int HEAD_DIM>
static void dispatch_decode(AttentionParams<bf16>& p) { static void dispatch_decode(AttentionParams<bf16>& p) {
#ifndef ASTRAI_NO_MMA #ifndef ASTRAI_NO_MMA
int G = p.q_head / p.kv_head; int G = p.q_head / p.kv_head;
if (!p.use_mask && G >= 1 && G <= 16) { if (G >= 1 && G <= 16) {
launch_mma_decode<HEAD_DIM, 32>(p); launch_mma_decode<HEAD_DIM, 32>(p);
return; return;
} }
@@ -68,35 +51,21 @@ torch::Tensor attn_decode(
torch::Tensor k, torch::Tensor k,
torch::Tensor v, torch::Tensor v,
c10::optional<torch::Tensor> mask, c10::optional<torch::Tensor> mask,
bool is_causal = false, int64_t causal_offset,
int64_t causal_offset = 0, double scale,
c10::optional<double> scale = c10::nullopt int64_t layout
) { ) {
AttentionParams<bf16> p; AttentionParams<bf16> p;
attn_pack_params(q, k, v, mask, is_causal, causal_offset, scale, p); attn_pack_params(q, k, v, mask, causal_offset, scale, layout, p);
TORCH_CHECK(p.q_len == 1, "Q seq_len must be 1"); TORCH_CHECK(p.q_len == 1, "Q seq_len must be 1");
TORCH_CHECK(p.head_dim % 32 == 0, "head_dim must be multiple of 32"); TORCH_CHECK(p.head_dim % 32 == 0, "head_dim must be multiple of 32");
auto O = torch::empty_like(q); // O matches Q's original layout
p.o = (bf16*)O.data_ptr(); auto O = torch::empty_strided(q.sizes(), q.strides(), q.options());
auto O_view = (layout == 1) ? O.transpose(1, 2) : O;
p.o = (bf16*)O_view.data_ptr();
switch (p.head_dim) { DISPATCH_HEAD_DIM(p.head_dim, dispatch_decode, p);
case 32:
dispatch_decode<32>(p);
break;
case 64:
dispatch_decode<64>(p);
break;
case 128:
dispatch_decode<128>(p);
break;
case 256:
dispatch_decode<256>(p);
break;
default:
TORCH_CHECK(false, "decode: unsupported head_dim ", p.head_dim,
" (supported: 32, 64, 128, 256)");
}
return O; return O;
} }
@@ -106,8 +75,8 @@ PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) {
py::arg("k"), py::arg("k"),
py::arg("v"), py::arg("v"),
py::arg("mask") = py::none(), py::arg("mask") = py::none(),
py::arg("is_causal") = false, py::arg("causal_offset") = -1,
py::arg("causal_offset") = 0, py::arg("scale") = 0.0,
py::arg("scale") = py::none(), py::arg("layout") = 0,
"GQA decode (tensor-core head-packing on sm_80+, scalar fallback)"); "GQA decode (tensor-core head-packing on sm_80+, scalar fallback)");
} }
+27 -11
View File
@@ -21,13 +21,16 @@ __global__ void attn_decode_split_kv_kernel(AttentionParams<bf16> p) {
int lane = threadIdx.x; int lane = threadIdx.x;
int hd_per_thread = p.head_dim / 32; int hd_per_thread = p.head_dim / 32;
// Q: [batch, q_head, q_len=1, head_dim] — stride-based
float q_reg[8]; float q_reg[8];
int q_off = ((batch * p.q_head + q_head) * 1) * p.head_dim + lane * hd_per_thread; int q_off = batch * p.q_stride_b + q_head * p.q_stride_h
+ lane * hd_per_thread * p.q_stride_d;
for (int i = 0; i < hd_per_thread; i++) for (int i = 0; i < hd_per_thread; i++)
q_reg[i] = __bfloat162float(p.q[q_off + i]); q_reg[i] = __bfloat162float(p.q[q_off + i * p.q_stride_d]);
int kv_base = ((batch * p.kv_head + kv_head) * p.kv_len) * p.head_dim; // KV: [batch, kv_head, kv_len, head_dim] — stride-based base
int mask_base = batch * p.kv_len; int kv_base = batch * p.kv_stride_b + kv_head * p.kv_stride_h;
int mask_base = batch * p.mask_b_stride;
float m = -FLT_MAX, d = 0.0f, acc_reg[8] = {0.0f}; float m = -FLT_MAX, d = 0.0f, acc_reg[8] = {0.0f};
@@ -43,9 +46,15 @@ __global__ void attn_decode_split_kv_kernel(AttentionParams<bf16> p) {
int chunk_start = ci * DC_CHUNK; int chunk_start = ci * DC_CHUNK;
int this_chunk = min(DC_CHUNK, p.kv_len - chunk_start); int this_chunk = min(DC_CHUNK, p.kv_len - chunk_start);
// Load K into shared memory (gather from strided global)
int total = this_chunk * p.head_dim; int total = this_chunk * p.head_dim;
for (int i = threadIdx.y * 32 + lane; i < total; i += blockDim.x * blockDim.y) for (int i = threadIdx.y * 32 + lane; i < total; i += blockDim.x * blockDim.y) {
k_smem[i] = p.k[kv_base + chunk_start * p.head_dim + i]; int s = i / p.head_dim;
int d_dim = i % p.head_dim;
int kv_idx = chunk_start + s;
int g_off = kv_base + kv_idx * p.kv_stride_l + d_dim * p.kv_stride_d;
k_smem[i] = p.k[g_off];
}
__syncthreads(); __syncthreads();
for (int s = 0; s < this_chunk; s++) { for (int s = 0; s < this_chunk; s++) {
@@ -54,9 +63,10 @@ __global__ void attn_decode_split_kv_kernel(AttentionParams<bf16> p) {
partial += q_reg[i] * __bfloat162float(k_smem[s * p.head_dim + lane * hd_per_thread + i]); partial += q_reg[i] * __bfloat162float(k_smem[s * p.head_dim + lane * hd_per_thread + i]);
partial = warp_reduce_sum(partial) * p.scale; partial = warp_reduce_sum(partial) * p.scale;
if (p.use_mask && p.mask && !p.mask[mask_base + chunk_start + s]) int kv_idx = chunk_start + s;
if (p.use_mask && p.mask && !p.mask[mask_base + kv_idx])
partial = -FLT_MAX; partial = -FLT_MAX;
if (p.is_causal && (chunk_start + s) > p.causal_offset) if (p.causal_offset >= 0 && kv_idx > p.causal_offset)
partial = -FLT_MAX; partial = -FLT_MAX;
float new_m = fmaxf(m, partial); float new_m = fmaxf(m, partial);
@@ -64,9 +74,10 @@ __global__ void attn_decode_split_kv_kernel(AttentionParams<bf16> p) {
float beta = expf(partial - new_m); float beta = expf(partial - new_m);
d = d * alpha + beta; d = d * alpha + beta;
int v_off = kv_base + (chunk_start + s) * p.head_dim + lane * hd_per_thread; // V: stride-based read
int v_off = kv_base + kv_idx * p.kv_stride_l + lane * hd_per_thread * p.kv_stride_d;
for (int i = 0; i < hd_per_thread; i++) for (int i = 0; i < hd_per_thread; i++)
acc_reg[i] = acc_reg[i] * alpha + __bfloat162float(p.v[v_off + i]) * beta; acc_reg[i] = acc_reg[i] * alpha + __bfloat162float(p.v[v_off + i * p.kv_stride_d]) * beta;
m = new_m; m = new_m;
} }
__syncthreads(); __syncthreads();
@@ -94,6 +105,9 @@ __global__ void attn_decode_combine_kernel(AttentionParams<bf16> p) {
int d = threadIdx.x; int d = threadIdx.x;
if (d >= p.head_dim) return; if (d >= p.head_dim) return;
int batch = bh / p.q_head;
int q_head = bh % p.q_head;
size_t split_base = (size_t)bh * p.num_splits; size_t split_base = (size_t)bh * p.num_splits;
const float* mlp = p.ml_part + split_base * 2; const float* mlp = p.ml_part + split_base * 2;
const float* op = p.o_part + split_base * p.head_dim; const float* op = p.o_part + split_base * p.head_dim;
@@ -112,5 +126,7 @@ __global__ void attn_decode_combine_kernel(AttentionParams<bf16> p) {
} }
float inv = (l > 1e-20f) ? (1.0f / l) : 0.0f; float inv = (l > 1e-20f) ? (1.0f / l) : 0.0f;
p.o[(size_t)bh * p.head_dim + d] = __float2bfloat16(acc * inv); // Stride-based output write (q_len=1 for decode, so stride_l not needed)
int o_off = batch * p.q_stride_b + q_head * p.q_stride_h + d * p.q_stride_d;
p.o[o_off] = __float2bfloat16(acc * inv);
} }
+27 -51
View File
@@ -8,39 +8,21 @@ using bf16 = __nv_bfloat16;
// Split-K (FlashDecoding) tensor-core decode via GQA head-packing. // Split-K (FlashDecoding) tensor-core decode via GQA head-packing.
// //
// Decode has q_len == 1, so S = q @ K^T is a GEMV per head — no tensor-core work // Decode has q_len == 1, so S = q @ K^T is a GEMV per head — no tensor-core
// on its own. But GQA gives us G = q_head / kv_head query heads that all share // work on its own. But GQA gives us G = q_head / kv_head query heads that all
// one kv_head. We pack those G heads into the M=16 rows of mma.sync.m16n8k16, // share one kv_head. We pack those G heads into the M=16 rows of
// turning G independent GEMVs into a single GEMM that reuses each loaded K/V tile // mma.sync.m16n8k16, turning G independent GEMVs into a single GEMM that
// across all G heads (K/V load is the decode bottleneck, so the reuse is the win, // reuses each loaded K/V tile across all G heads (K/V load is the decode
// not the flops). The KV sequence is partitioned across gridDim.z blocks so that // bottleneck, so the reuse is the win, not the flops). The KV sequence is
// a decode with only batch*kv_head independent tasks can fill all SMs. Each // partitioned across gridDim.z blocks so that a decode with only
// (batch, kv_head, split) block computes an UN-normalised partial (Oacc, m, l) // batch*kv_head independent tasks can fill all SMs. Each (batch, kv_head,
// over its KV slice; the combine kernel below reduces across splits. Fixes the // split) block computes an UN-normalised partial (Oacc, m, l) over its KV
// "grid too small" bottleneck (0.04 waves/SM → many blocks) for long-context, // slice; the combine kernel below reduces across splits. Fixes the "grid too
// small" bottleneck (0.04 waves/SM → many blocks) for long-context,
// small-batch decode. // small-batch decode.
//
// Partial layout (float, contiguous):
// o_part : [batch, q_head, num_splits, HEAD_DIM]
// ml_part: [batch, q_head, num_splits, 2] (m, l)
//
// Optimizations:
// - cp.async global→shared for K/V (bypasses registers, cuts instruction count)
// - XOR swizzle (swiz_col): LD=HEAD_DIM, zero waste, no bank conflicts
// - Q loaded directly from global into mma A-operand registers (no sQ staging,
// no prologue syncwarp) — frees shared memory for double-buffering
// - Double-buffered KV (STAGES=2): next tile's cp.async overlaps current
// tile's MMA compute — hides global load latency / boosts bandwidth
// utilization for small-batch (low-occupancy) decode
// - Predicated cp.async (cp_async_16_pred) for full AND partial tiles on one
// uniform path — eliminates the scalar fallback branch
//
// Smem footprint (BC=32): STAGES=2 → 2*(sK+sV) = 2*2*32*HEAD_DIM*2 bytes.
// D=128: 16 KB (fits 48 KB static cap). D=256: 32 KB (also fits).
// STAGES=1 fallback (4/8 KB) for smem-constrained configs.
template <int HEAD_DIM, int BC, int STAGES = 2> template <int HEAD_DIM, int BC, int STAGES = 2>
__global__ void attn_decode_split_kv_mma_kernel(AttentionParams<bf16> p) { __global__ void attn_decode_split_kv_mma_kernel(AttentionParams<bf16> p) {
constexpr int BR = 16;
constexpr int KD = HEAD_DIM / 16; constexpr int KD = HEAD_DIM / 16;
constexpr int NC8 = BC / 8; constexpr int NC8 = BC / 8;
constexpr int KT2 = BC / 16; constexpr int KT2 = BC / 16;
@@ -66,25 +48,13 @@ __global__ void attn_decode_split_kv_mma_kernel(AttentionParams<bf16> p) {
__shared__ __align__(16) bf16 sV[STAGES * BC * LD]; __shared__ __align__(16) bf16 sV[STAGES * BC * LD];
// ---- Load Q directly from global into mma A-operand registers ---- // ---- Load Q directly from global into mma A-operand registers ----
// Same layout as prefill: frag[0]/[2] = row gid, frag[1]/[3] = row gid+8 const int q_base = batch * p.q_stride_b + q_head0 * p.q_stride_h;
// cols kt*16 + tid4*2 + {0,1} / +{8,9}. pau[0]=cols c,c+1; pau[4]=c+8,c+9.
const int q_base = (batch * p.q_head + q_head0) * HEAD_DIM;
const int qra = gid; const int qra = gid;
const int qrb = gid + 8; const int qrb = gid + 8;
const bool va = qra < G, vb = qrb < G; const bool va = qra < G, vb = qrb < G;
unsigned Qa[KD][4]; unsigned Qa[KD][4];
#pragma unroll load_q_mma_frags<KD>(p.q + q_base, p.q_stride_h, p.q_stride_d,
for (int kt = 0; kt < KD; kt++) { qra, qrb, va, vb, tid4, Qa);
int c = kt * 16 + tid4 * 2;
const unsigned* pau = reinterpret_cast<const unsigned*>(
&p.q[q_base + qra * HEAD_DIM + c]);
const unsigned* pbu = reinterpret_cast<const unsigned*>(
&p.q[q_base + qrb * HEAD_DIM + c]);
Qa[kt][0] = va ? pau[0] : 0u;
Qa[kt][1] = vb ? pbu[0] : 0u;
Qa[kt][2] = va ? pau[4] : 0u;
Qa[kt][3] = vb ? pbu[4] : 0u;
}
float Oacc[DN8][4]; float Oacc[DN8][4];
#pragma unroll #pragma unroll
@@ -92,8 +62,8 @@ __global__ void attn_decode_split_kv_mma_kernel(AttentionParams<bf16> p) {
Oacc[j][0] = Oacc[j][1] = Oacc[j][2] = Oacc[j][3] = 0.0f; Oacc[j][0] = Oacc[j][1] = Oacc[j][2] = Oacc[j][3] = 0.0f;
float m0 = -FLT_MAX, m1 = -FLT_MAX, l0 = 0.0f, l1 = 0.0f; float m0 = -FLT_MAX, m1 = -FLT_MAX, l0 = 0.0f, l1 = 0.0f;
const int kv_base = (batch * p.kv_head + kv_head) * p.kv_len * HEAD_DIM; // KV: stride-based base — [batch, kv_head, kv_len, head_dim]
const int mask_base = batch * p.kv_len; const int kv_base = batch * p.kv_stride_b + kv_head * p.kv_stride_h;
const int tiles_total = (p.kv_len + BC - 1) / BC; const int tiles_total = (p.kv_len + BC - 1) / BC;
const int tiles_per_split = (tiles_total + p.num_splits - 1) / p.num_splits; const int tiles_per_split = (tiles_total + p.num_splits - 1) / p.num_splits;
const int ti_begin = split * tiles_per_split; const int ti_begin = split * tiles_per_split;
@@ -111,8 +81,10 @@ __global__ void attn_decode_split_kv_mma_kernel(AttentionParams<bf16> p) {
int kc = kv0 + r; int kc = kv0 + r;
bool valid = kc < p.kv_len; bool valid = kc < p.kv_len;
int off = r * LD + swiz_col(d, r, SWIZ_MASK); int off = r * LD + swiz_col(d, r, SWIZ_MASK);
cp_async_16_pred(&dK[off], &p.k[kv_base + kc * HEAD_DIM + d], valid); // KV stride-based: contiguous within head_dim (stride_d == 1 typically)
cp_async_16_pred(&dV[off], &p.v[kv_base + kc * HEAD_DIM + d], valid); int g_off = kv_base + kc * p.kv_stride_l + d * p.kv_stride_d;
cp_async_16_pred(&dK[off], &p.k[g_off], valid);
cp_async_16_pred(&dV[off], &p.v[g_off], valid);
} }
cp_async_commit(); cp_async_commit();
}; };
@@ -148,9 +120,13 @@ __global__ void attn_decode_split_kv_mma_kernel(AttentionParams<bf16> p) {
Sacc[n8][0] *= p.scale, Sacc[n8][1] *= p.scale, Sacc[n8][0] *= p.scale, Sacc[n8][1] *= p.scale,
Sacc[n8][2] *= p.scale, Sacc[n8][3] *= p.scale; Sacc[n8][2] *= p.scale, Sacc[n8][3] *= p.scale;
int maxc = p.is_causal ? min(p.kv_len, p.causal_offset + 1) : p.kv_len; // Decode: q_len=1, so qrow0=qrow1=0, mask_q_stride irrelevant
int maxc = (p.causal_offset >= 0) ? min(p.kv_len, p.causal_offset + 1) : p.kv_len;
mma_softmax_tile<NC8, DN8>(kv0, maxc, maxc, mma_softmax_tile<NC8, DN8>(kv0, maxc, maxc,
mask_base, p.mask, has_mask, 0, 0,
p.mask_b_stride, 0,
batch,
p.mask, has_mask,
Sacc, Oacc, m0, m1, l0, l1, lane); Sacc, Oacc, m0, m1, l0, l1, lane);
mma_pv_accumulate<DN8, KT2>(Sacc, bV, LD, SWIZ_MASK, lane, Oacc); mma_pv_accumulate<DN8, KT2>(Sacc, bV, LD, SWIZ_MASK, lane, Oacc);
+152 -18
View File
@@ -1,43 +1,177 @@
#pragma once #pragma once
#include <torch/extension.h> #include <torch/extension.h>
#include <c10/cuda/CUDAGuard.h>
#include "attn_common.h" #include "attn_common.h"
using bf16 = __nv_bfloat16;
inline int compute_num_splits(int base_blocks, int tiles_total) {
int sm_count = 0;
cudaDeviceGetAttribute(&sm_count, cudaDevAttrMultiProcessorCount, 0);
int n = (2 * sm_count + base_blocks - 1) / base_blocks;
return std::max(1, std::min(n, std::min(tiles_total, 32)));
}
// Dispatch head_dim: shared macro — avoids C++20 lambda template syntax.
// Usage: DISPATCH_HEAD_DIM(hd, fn, arg)
// Expands to: fn<32>(arg); fn<64>(arg); etc.
#define DISPATCH_HEAD_DIM(hd, fn, arg) \
switch (hd) { \
case 32: fn<32>(arg); break; \
case 64: fn<64>(arg); break; \
case 128: fn<128>(arg); break; \
case 256: fn<256>(arg); break; \
default: \
TORCH_CHECK(false, "unsupported head_dim ", hd, \
" (supported: 32, 64, 128, 256)"); \
}
template<typename P>
inline void alloc_split_partials(P& p) {
auto fopt = torch::TensorOptions().dtype(torch::kFloat32).device(torch::kCUDA);
auto o_part = torch::empty({p.batch, p.q_head, p.num_splits, p.head_dim}, fopt);
auto ml_part = torch::empty({p.batch, p.q_head, p.num_splits, 2}, fopt);
p.o_part = (float*)o_part.data_ptr();
p.ml_part = (float*)ml_part.data_ptr();
}
// ---- Shared Q-dims + strides extraction ----
template <typename P>
inline void extract_q_dims_and_strides(torch::Tensor& q, int64_t layout, P& p) {
if (layout == 1) q = q.transpose(1, 2);
p.batch = (int)q.size(0);
p.q_head = (int)q.size(1);
p.q_len = (int)q.size(2);
p.head_dim = (int)q.size(3);
p.q_stride_b = (int)q.stride(0);
p.q_stride_h = (int)q.stride(1);
p.q_stride_l = (int)q.stride(2);
p.q_stride_d = (int)q.stride(3);
}
// ---- Shared mask packing ----
template <typename P>
inline void pack_mask(const c10::optional<torch::Tensor>& mask, P& p) {
if (p.use_mask) {
auto m = mask.value();
TORCH_CHECK(m.is_cuda(), "mask must be on CUDA");
TORCH_CHECK(m.dtype() == torch::kBool, "mask must be bool");
TORCH_CHECK(m.size(0) == p.batch, "mask batch mismatch");
TORCH_CHECK(m.size(m.dim() - 1) == p.kv_len, "mask kv_len mismatch");
if (m.dim() == 2) {
p.mask_b_stride = (int)m.stride(0);
p.mask_q_stride = 0;
} else if (m.dim() == 3) {
TORCH_CHECK(m.size(1) == p.q_len, "mask q_len mismatch");
p.mask_b_stride = (int)m.stride(0);
p.mask_q_stride = (int)m.stride(1);
} else {
TORCH_CHECK(false, "mask must be 2D [batch, kv_len] or 3D [batch, q_len, kv_len]");
}
p.mask = m.data_ptr<bool>();
} else {
p.mask = nullptr;
p.mask_b_stride = 0;
p.mask_q_stride = 0;
}
}
// ---- attn_pack_params (contiguous KV) ----
template<typename T> template<typename T>
inline void attn_pack_params( inline void attn_pack_params(
torch::Tensor q, torch::Tensor q,
torch::Tensor k, torch::Tensor k,
torch::Tensor v, torch::Tensor v,
c10::optional<torch::Tensor> mask, c10::optional<torch::Tensor> mask,
bool is_causal,
int64_t causal_offset, int64_t causal_offset,
c10::optional<double> scale, double scale,
int64_t layout,
AttentionParams<T>& p AttentionParams<T>& p
) { ) {
const at::cuda::OptionalCUDAGuard device_guard(device_of(q));
TORCH_CHECK(q.is_cuda() && k.is_cuda() && v.is_cuda()); TORCH_CHECK(q.is_cuda() && k.is_cuda() && v.is_cuda());
TORCH_CHECK(q.dtype() == torch::kBFloat16); TORCH_CHECK(q.dtype() == torch::kBFloat16);
TORCH_CHECK(k.dtype() == torch::kBFloat16); TORCH_CHECK(k.dtype() == torch::kBFloat16);
TORCH_CHECK(v.dtype() == torch::kBFloat16); TORCH_CHECK(v.dtype() == torch::kBFloat16);
TORCH_CHECK(k.sizes() == v.sizes(), "K and V must have identical shapes");
TORCH_CHECK(q.dim() == 4 && k.dim() == 4, "Q/K/V must be 4D");
extract_q_dims_and_strides(q, layout, p);
if (layout == 1) k = k.transpose(1, 2), v = v.transpose(1, 2);
p.batch = (int)q.size(0);
p.q_head = (int)q.size(1);
p.kv_head = (int)k.size(1); p.kv_head = (int)k.size(1);
p.q_len = (int)q.size(2);
p.kv_len = (int)k.size(2); p.kv_len = (int)k.size(2);
p.head_dim = (int)q.size(3); TORCH_CHECK(k.size(3) == p.head_dim, "K/V head_dim must match Q");
p.use_mask = mask.has_value() ? 1 : 0;
p.is_causal = is_causal ? 1 : 0; p.kv_stride_b = (int)k.stride(0);
p.kv_stride_h = (int)k.stride(1);
p.kv_stride_l = (int)k.stride(2);
p.kv_stride_d = (int)k.stride(3);
p.causal_offset = (int)causal_offset; p.causal_offset = (int)causal_offset;
p.scale = scale.has_value() ? (float)scale.value() : 1.0f / sqrtf((float)p.head_dim); p.use_mask = mask.has_value() ? 1 : 0;
p.scale = (scale > 0.0) ? (float)scale : 1.0f / sqrtf((float)p.head_dim);
p.q = (const T*)q.data_ptr(); p.q = (const T*)q.data_ptr();
p.k = (const T*)k.data_ptr(); p.k = (const T*)k.data_ptr();
p.v = (const T*)v.data_ptr(); p.v = (const T*)v.data_ptr();
if (p.use_mask) { p.o = nullptr;
TORCH_CHECK(mask.value().dtype() == torch::kBool); p.o_part = nullptr;
TORCH_CHECK(mask.value().dim() == 2); p.ml_part = nullptr;
TORCH_CHECK(mask.value().size(0) == p.batch);
TORCH_CHECK(mask.value().size(1) == p.kv_len); pack_mask(mask, p);
p.mask = mask.value().data_ptr<bool>(); }
} else {
p.mask = nullptr; // ---- attn_pack_paged_params ----
} template<typename T>
inline void attn_pack_paged_params(
torch::Tensor q,
torch::Tensor page_table,
torch::Tensor k_cache,
torch::Tensor v_cache,
int64_t page_size,
int64_t kv_len,
c10::optional<torch::Tensor> mask,
int64_t causal_offset,
double scale,
int64_t layout,
PagedAttentionParams<T>& p
) {
const at::cuda::OptionalCUDAGuard device_guard(device_of(q));
TORCH_CHECK(q.is_cuda() && page_table.is_cuda() && k_cache.is_cuda() && v_cache.is_cuda());
TORCH_CHECK(q.dtype() == torch::kBFloat16, "q must be bf16");
TORCH_CHECK(k_cache.dtype() == torch::kBFloat16, "k_cache must be bf16");
TORCH_CHECK(v_cache.dtype() == torch::kBFloat16, "v_cache must be bf16");
TORCH_CHECK(page_table.dtype() == torch::kLong, "page_table must be int64");
TORCH_CHECK(k_cache.sizes() == v_cache.sizes(), "k_cache and v_cache must have identical shapes");
extract_q_dims_and_strides(q, layout, p);
p.kv_head = (int)k_cache.size(2);
p.kv_len = (int)kv_len;
p.page_size = (int)page_size;
p.max_pages = (int)page_table.size(1);
TORCH_CHECK(q.size(2) == 1, "Q seq_len must be 1 (decode)");
TORCH_CHECK(p.head_dim % 32 == 0, "head_dim must be multiple of 32");
TORCH_CHECK(k_cache.size(1) == page_size,
"k_cache dim 1 must equal page_size, got ",
k_cache.size(1), " vs ", page_size);
p.causal_offset = (int)causal_offset;
p.use_mask = (mask.has_value() && mask.value().defined()) ? 1 : 0;
p.scale = (scale > 0.0) ? (float)scale : 1.0f / sqrtf((float)p.head_dim);
p.page_table = page_table.data_ptr<int64_t>();
p.k_cache = (const T*)k_cache.data_ptr();
p.v_cache = (const T*)v_cache.data_ptr();
p.q = (const T*)q.data_ptr();
p.o = nullptr;
p.o_part = nullptr;
p.ml_part = nullptr;
pack_mask(mask, p);
} }
+48 -7
View File
@@ -108,6 +108,36 @@ __device__ __forceinline__ void cp_async_wait_group() {
asm volatile("cp.async.wait_group %0;" :: "n"(N)); asm volatile("cp.async.wait_group %0;" :: "n"(N));
} }
// ---------------------------------------------------------------------------
// Q-load: load query rows directly from global memory into mma A-operand
// register layout. One call replaces ~15 duplicated lines in each MMA kernel.
// stride_row is p.q_stride_h for decode (q_len=1, G heads) or
// p.q_stride_l for prefill (multi-q rows).
// ---------------------------------------------------------------------------
template <int KD>
__device__ inline void load_q_mma_frags(
const bf16* __restrict__ q,
int stride_row,
int stride_d,
int qra, int qrb,
bool va, bool vb,
int tid4,
unsigned Qa[KD][4])
{
#pragma unroll
for (int kt = 0; kt < KD; kt++) {
int c = kt * 16 + tid4 * 2;
const unsigned* pau = reinterpret_cast<const unsigned*>(
&q[qra * stride_row + c * stride_d]);
const unsigned* pbu = reinterpret_cast<const unsigned*>(
&q[qrb * stride_row + c * stride_d]);
Qa[kt][0] = va ? pau[0] : 0u;
Qa[kt][1] = vb ? pbu[0] : 0u;
Qa[kt][2] = va ? pau[4] : 0u;
Qa[kt][3] = vb ? pbu[4] : 0u;
}
}
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Shared MMA compute functions — used by both decode and prefill MMA kernels. // Shared MMA compute functions — used by both decode and prefill MMA kernels.
// Extracted because S=Q@K^T, online softmax, and P@V are structurally identical // Extracted because S=Q@K^T, online softmax, and P@V are structurally identical
@@ -122,7 +152,9 @@ template <int KD, int NC8>
__device__ inline void mma_compute_scores( __device__ inline void mma_compute_scores(
const unsigned Qa[KD][4], const unsigned Qa[KD][4],
const bf16* __restrict__ sK, const bf16* __restrict__ sK,
int LD, int SWIZ_MASK, int lane, int LD,
int SWIZ_MASK,
int lane,
float Sacc[NC8][4]) float Sacc[NC8][4])
{ {
#pragma unroll #pragma unroll
@@ -142,13 +174,20 @@ __device__ inline void mma_compute_scores(
// Online softmax + Oacc rescale for one K/V tile. // Online softmax + Oacc rescale for one K/V tile.
// maxc0/maxc1: per-row KV column bounds (prefill: per-query-row causal limits; // maxc0/maxc1: per-row KV column bounds (prefill: per-query-row causal limits;
// decode: same value for both rows since q_len==1). // decode: same value for both rows since q_len==1).
// qrow0/qrow1: query row indices (for 3D mask indexing; decode passes 0).
// mask_b_stride/mask_q_stride: mask layout (2D: mask_q_stride=0; 3D: =kv_len).
// Reads Sacc (Q@K^T scores), applies causal/mask, computes P = exp(S - nm), // Reads Sacc (Q@K^T scores), applies causal/mask, computes P = exp(S - nm),
// rescales Oacc by exp(m_old - nm), and updates m/l — all in place. // rescales Oacc by exp(m_old - nm), and updates m/l — all in place.
template <int NC8, int DN8> template <int NC8, int DN8>
__device__ inline void mma_softmax_tile( __device__ inline void mma_softmax_tile(
int kv0, int kv0,
int maxc0, int maxc1, int maxc0,
int mask_base, int maxc1,
int qrow0,
int qrow1,
int mask_b_stride,
int mask_q_stride,
int mask_batch,
const bool* __restrict__ mask, const bool* __restrict__ mask,
bool has_mask, bool has_mask,
float Sacc[NC8][4], float Sacc[NC8][4],
@@ -162,14 +201,16 @@ __device__ inline void mma_softmax_tile(
// Mask out-of-bounds / masked columns: set -FLT_MAX so expf → 0 downstream // Mask out-of-bounds / masked columns: set -FLT_MAX so expf → 0 downstream
// without per-element sentinel checks. Compute tile-local row maxima. // without per-element sentinel checks. Compute tile-local row maxima.
float rmax0 = -FLT_MAX, rmax1 = -FLT_MAX; float rmax0 = -FLT_MAX, rmax1 = -FLT_MAX;
int mask_base0 = mask_batch * mask_b_stride + qrow0 * mask_q_stride;
int mask_base1 = mask_batch * mask_b_stride + qrow1 * mask_q_stride;
#pragma unroll #pragma unroll
for (int n8 = 0; n8 < NC8; n8++) { for (int n8 = 0; n8 < NC8; n8++) {
int cc = kv0 + n8 * 8 + 2 * tid4; int cc = kv0 + n8 * 8 + 2 * tid4;
int c1 = cc + 1; int c1 = cc + 1;
bool b0 = (cc >= maxc0) || (has_mask && !mask[mask_base + cc]); bool b0 = (cc >= maxc0) || (has_mask && !mask[mask_base0 + cc]);
bool b1 = (c1 >= maxc0) || (has_mask && !mask[mask_base + c1]); bool b1 = (c1 >= maxc0) || (has_mask && !mask[mask_base0 + c1]);
bool b2 = (cc >= maxc1) || (has_mask && !mask[mask_base + cc]); bool b2 = (cc >= maxc1) || (has_mask && !mask[mask_base1 + cc]);
bool b3 = (c1 >= maxc1) || (has_mask && !mask[mask_base + c1]); bool b3 = (c1 >= maxc1) || (has_mask && !mask[mask_base1 + c1]);
float s0 = b0 ? -FLT_MAX : Sacc[n8][0]; float s0 = b0 ? -FLT_MAX : Sacc[n8][0];
float s1 = b1 ? -FLT_MAX : Sacc[n8][1]; float s1 = b1 ? -FLT_MAX : Sacc[n8][1];
float s2 = b2 ? -FLT_MAX : Sacc[n8][2]; float s2 = b2 ? -FLT_MAX : Sacc[n8][2];
+19 -86
View File
@@ -3,26 +3,13 @@
#include "attn_paged_decode_split_kv_mma.cuh" #include "attn_paged_decode_split_kv_mma.cuh"
#endif #endif
#include <torch/extension.h> #include "attn_entry_utils.cuh"
#include <c10/cuda/CUDAGuard.h>
static int paged_decode_num_splits(int base_blocks, int tiles_total) {
int sm_count = 0;
cudaDeviceGetAttribute(&sm_count, cudaDevAttrMultiProcessorCount, 0);
int n = (2 * sm_count + base_blocks - 1) / base_blocks;
return std::max(1, std::min(n, std::min(tiles_total, 32)));
}
static void launch_paged_scalar_decode(PagedAttentionParams<bf16>& p) { static void launch_paged_scalar_decode(PagedAttentionParams<bf16>& p) {
int group_size = p.q_head / p.kv_head; int group_size = p.q_head / p.kv_head;
int chunks_total = (p.kv_len + PDC_CHUNK - 1) / PDC_CHUNK; int chunks_total = (p.kv_len + PDC_CHUNK - 1) / PDC_CHUNK;
p.num_splits = paged_decode_num_splits(p.batch * p.kv_head, chunks_total); p.num_splits = compute_num_splits(p.batch * p.kv_head, chunks_total);
alloc_split_partials(p);
auto fopt = torch::TensorOptions().dtype(torch::kFloat32).device(torch::kCUDA);
auto o_part = torch::empty({p.batch, p.q_head, p.num_splits, p.head_dim}, fopt);
auto ml_part = torch::empty({p.batch, p.q_head, p.num_splits, 2}, fopt);
p.o_part = o_part.data_ptr<float>();
p.ml_part = ml_part.data_ptr<float>();
size_t smem = PDC_CHUNK * p.head_dim * sizeof(bf16); size_t smem = PDC_CHUNK * p.head_dim * sizeof(bf16);
dim3 grid = dim3(p.batch * p.kv_head, 1, p.num_splits); dim3 grid = dim3(p.batch * p.kv_head, 1, p.num_splits);
@@ -35,13 +22,8 @@ static void launch_paged_scalar_decode(PagedAttentionParams<bf16>& p) {
template <int HEAD_DIM, int BC, int STAGES = (HEAD_DIM <= 128) ? 2 : 1> template <int HEAD_DIM, int BC, int STAGES = (HEAD_DIM <= 128) ? 2 : 1>
static void launch_paged_mma_decode(PagedAttentionParams<bf16>& p) { static void launch_paged_mma_decode(PagedAttentionParams<bf16>& p) {
int tiles_total = (p.kv_len + BC - 1) / BC; int tiles_total = (p.kv_len + BC - 1) / BC;
p.num_splits = paged_decode_num_splits(p.batch * p.kv_head, tiles_total); p.num_splits = compute_num_splits(p.batch * p.kv_head, tiles_total);
alloc_split_partials(p);
auto fopt = torch::TensorOptions().dtype(torch::kFloat32).device(torch::kCUDA);
auto o_part = torch::empty({p.batch, p.q_head, p.num_splits, p.head_dim}, fopt);
auto ml_part = torch::empty({p.batch, p.q_head, p.num_splits, 2}, fopt);
p.o_part = o_part.data_ptr<float>();
p.ml_part = ml_part.data_ptr<float>();
paged_attn_decode_split_kv_mma_kernel<HEAD_DIM, BC, STAGES><<<dim3(p.kv_head, p.batch, p.num_splits), 32>>>(p); paged_attn_decode_split_kv_mma_kernel<HEAD_DIM, BC, STAGES><<<dim3(p.kv_head, p.batch, p.num_splits), 32>>>(p);
paged_attn_decode_combine_kernel<<<p.batch * p.q_head, p.head_dim>>>(p); paged_attn_decode_combine_kernel<<<p.batch * p.q_head, p.head_dim>>>(p);
@@ -52,7 +34,7 @@ template <int HEAD_DIM>
static void dispatch_paged_decode(PagedAttentionParams<bf16>& p) { static void dispatch_paged_decode(PagedAttentionParams<bf16>& p) {
#ifndef ASTRAI_NO_MMA #ifndef ASTRAI_NO_MMA
int G = p.q_head / p.kv_head; int G = p.q_head / p.kv_head;
if (!p.use_mask && G >= 1 && G <= 16 && p.page_size >= 32) { if (G >= 1 && G <= 16 && p.page_size >= 32) {
launch_paged_mma_decode<HEAD_DIM, 32>(p); launch_paged_mma_decode<HEAD_DIM, 32>(p);
return; return;
} }
@@ -68,68 +50,19 @@ torch::Tensor attn_paged_decode(
int64_t page_size, int64_t page_size,
int64_t kv_len, int64_t kv_len,
c10::optional<torch::Tensor> mask, c10::optional<torch::Tensor> mask,
bool is_causal = false, int64_t causal_offset,
int64_t causal_offset = 0, double scale,
c10::optional<double> scale = c10::nullopt int64_t layout
) { ) {
const at::cuda::OptionalCUDAGuard device_guard(device_of(q)); PagedAttentionParams<bf16> p;
attn_pack_paged_params(q, page_table, k_cache, v_cache,
page_size, kv_len, mask, causal_offset, scale, layout, p);
int batch = q.size(0); auto O = torch::empty_strided(q.sizes(), q.strides(), q.options());
int q_head = q.size(1); auto O_view = (layout == 1) ? O.transpose(1, 2) : O;
int head_dim = q.size(3); p.o = (bf16*)O_view.data_ptr();
int kv_head = k_cache.size(2);
int max_pages = page_table.size(1);
TORCH_CHECK(q.is_cuda() && page_table.is_cuda() && k_cache.is_cuda() && v_cache.is_cuda());
TORCH_CHECK(q.dtype() == torch::kBFloat16, "q must be bf16");
TORCH_CHECK(k_cache.dtype() == torch::kBFloat16, "k_cache must be bf16");
TORCH_CHECK(v_cache.dtype() == torch::kBFloat16, "v_cache must be bf16");
TORCH_CHECK(page_table.dtype() == torch::kLong, "page_table must be int64");
TORCH_CHECK(q.size(2) == 1, "Q seq_len must be 1 (decode)");
TORCH_CHECK(head_dim % 32 == 0, "head_dim must be multiple of 32");
TORCH_CHECK(k_cache.size(1) == page_size,
"k_cache dim 1 must equal page_size, got ",
k_cache.size(1), " vs ", page_size);
TORCH_CHECK(k_cache.size(0) >= 0, "k_cache must have at least 0 pages");
float scale_val = scale.has_value()
? static_cast<float>(scale.value())
: 1.0f / std::sqrt(static_cast<float>(head_dim));
auto O = torch::empty_like(q);
PagedAttentionParams<bf16, float> p;
p.batch = batch;
p.q_head = q_head;
p.kv_head = kv_head;
p.q_len = static_cast<int>(q.size(2));
p.kv_len = static_cast<int>(kv_len);
p.head_dim = head_dim;
p.use_mask = (mask.has_value() && mask.value().defined()) ? 1 : 0;
p.is_causal = is_causal ? 1 : 0;
p.causal_offset = static_cast<int>(causal_offset);
p.scale = scale_val;
p.page_size = static_cast<int>(page_size);
p.max_pages = max_pages;
p.page_table = page_table.data_ptr<int64_t>();
p.k_cache = reinterpret_cast<const bf16*>(k_cache.data_ptr());
p.v_cache = reinterpret_cast<const bf16*>(v_cache.data_ptr());
p.q = reinterpret_cast<const bf16*>(q.data_ptr());
p.mask = p.use_mask ? mask.value().data_ptr<bool>() : nullptr;
p.o = reinterpret_cast<bf16*>(O.data_ptr());
p.o_part = nullptr;
p.ml_part = nullptr;
switch (p.head_dim) {
case 32: dispatch_paged_decode<32>(p); break;
case 64: dispatch_paged_decode<64>(p); break;
case 128: dispatch_paged_decode<128>(p); break;
case 256: dispatch_paged_decode<256>(p); break;
default:
TORCH_CHECK(false, "paged_decode: unsupported head_dim ", p.head_dim,
" (supported: 32, 64, 128, 256)");
}
DISPATCH_HEAD_DIM(p.head_dim, dispatch_paged_decode, p);
return O; return O;
} }
@@ -142,8 +75,8 @@ PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) {
py::arg("page_size"), py::arg("page_size"),
py::arg("kv_len"), py::arg("kv_len"),
py::arg("mask") = py::none(), py::arg("mask") = py::none(),
py::arg("is_causal") = false, py::arg("causal_offset") = -1,
py::arg("causal_offset") = 0, py::arg("scale") = 0.0,
py::arg("scale") = py::none(), py::arg("layout") = 0,
"Paged GQA decode — split-KV with direct page-table access."); "Paged GQA decode — split-KV with direct page-table access.");
} }
+13 -6
View File
@@ -22,11 +22,13 @@ __global__ void paged_attn_decode_split_kv_kernel(PagedAttentionParams<bf16> p)
int lane = threadIdx.x; int lane = threadIdx.x;
int hd_per_thread = p.head_dim / 32; int hd_per_thread = p.head_dim / 32;
// Q: stride-based [batch, q_head, q_len=1, head_dim]
float q_reg[8]; float q_reg[8];
int q_off = ((batch * p.q_head + q_head) * 1) * p.head_dim + lane * hd_per_thread; int q_off = batch * p.q_stride_b + q_head * p.q_stride_h
+ lane * hd_per_thread * p.q_stride_d;
#pragma unroll #pragma unroll
for (int i = 0; i < hd_per_thread; i++) for (int i = 0; i < hd_per_thread; i++)
q_reg[i] = __bfloat162float(p.q[q_off + i]); q_reg[i] = __bfloat162float(p.q[q_off + i * p.q_stride_d]);
float m = -FLT_MAX, d = 0.0f, acc_reg[8] = {0.0f}; float m = -FLT_MAX, d = 0.0f, acc_reg[8] = {0.0f};
@@ -37,7 +39,7 @@ __global__ void paged_attn_decode_split_kv_kernel(PagedAttentionParams<bf16> p)
int ch_begin = split * chunks_per_split; int ch_begin = split * chunks_per_split;
int ch_end = min(chunks_total, ch_begin + chunks_per_split); int ch_end = min(chunks_total, ch_begin + chunks_per_split);
const int mask_base = batch * p.kv_len; const int mask_base = batch * p.mask_b_stride;
for (int ci = ch_begin; ci < ch_end; ci++) { for (int ci = ch_begin; ci < ch_end; ci++) {
int chunk_start = ci * PDC_CHUNK; int chunk_start = ci * PDC_CHUNK;
@@ -70,9 +72,10 @@ __global__ void paged_attn_decode_split_kv_kernel(PagedAttentionParams<bf16> p)
partial += q_reg[i] * __bfloat162float(k_smem[s * p.head_dim + lane * hd_per_thread + i]); partial += q_reg[i] * __bfloat162float(k_smem[s * p.head_dim + lane * hd_per_thread + i]);
partial = paged_warp_reduce_sum(partial) * p.scale; partial = paged_warp_reduce_sum(partial) * p.scale;
if (p.use_mask && p.mask && !p.mask[mask_base + chunk_start + s]) int kv_idx = chunk_start + s;
if (p.use_mask && p.mask && !p.mask[mask_base + kv_idx])
partial = -FLT_MAX; partial = -FLT_MAX;
if (p.is_causal && (chunk_start + s) > p.causal_offset) if (p.causal_offset >= 0 && kv_idx > p.causal_offset)
partial = -FLT_MAX; partial = -FLT_MAX;
float new_m = fmaxf(m, partial); float new_m = fmaxf(m, partial);
@@ -118,6 +121,9 @@ __global__ void paged_attn_decode_combine_kernel(PagedAttentionParams<bf16> p) {
int d = threadIdx.x; int d = threadIdx.x;
if (d >= p.head_dim) return; if (d >= p.head_dim) return;
int batch = bh / p.q_head;
int q_head = bh % p.q_head;
size_t split_base = (size_t)bh * p.num_splits; size_t split_base = (size_t)bh * p.num_splits;
const float* mlp = p.ml_part + split_base * 2; const float* mlp = p.ml_part + split_base * 2;
const float* op = p.o_part + split_base * p.head_dim; const float* op = p.o_part + split_base * p.head_dim;
@@ -136,5 +142,6 @@ __global__ void paged_attn_decode_combine_kernel(PagedAttentionParams<bf16> p) {
} }
float inv = (l > 1e-20f) ? (1.0f / l) : 0.0f; float inv = (l > 1e-20f) ? (1.0f / l) : 0.0f;
p.o[(size_t)bh * p.head_dim + d] = __float2bfloat16(acc * inv); int o_off = batch * p.q_stride_b + q_head * p.q_stride_h + d * p.q_stride_d;
p.o[o_off] = __float2bfloat16(acc * inv);
} }
+13 -25
View File
@@ -11,14 +11,9 @@ using bf16 = __nv_bfloat16;
// directly from the page pool through a page table, eliminating the gather // directly from the page pool through a page table, eliminating the gather
// copy. Each tile (BC=32) fits within a single page (page_size >= 32), so // copy. Each tile (BC=32) fits within a single page (page_size >= 32), so
// the page-table lookup happens once per tile for cp.async. // the page-table lookup happens once per tile for cp.async.
//
// Optimizations mirror attn_decode_split_kv_mma_kernel:
// - Q loaded directly from global into mma A-operand registers (no sQ)
// - Double-buffered KV (STAGES=2) for D<=128, single-buffer for D=256
// - Predicated cp.async for unified full/partial tile path
template <int HEAD_DIM, int BC, int STAGES = (HEAD_DIM <= 128) ? 2 : 1> template <int HEAD_DIM, int BC, int STAGES = (HEAD_DIM <= 128) ? 2 : 1>
__global__ void paged_attn_decode_split_kv_mma_kernel(PagedAttentionParams<bf16> p) { __global__ void paged_attn_decode_split_kv_mma_kernel(PagedAttentionParams<bf16> p) {
constexpr int BR = 16;
constexpr int KD = HEAD_DIM / 16; constexpr int KD = HEAD_DIM / 16;
constexpr int NC8 = BC / 8; constexpr int NC8 = BC / 8;
constexpr int KT2 = BC / 16; constexpr int KT2 = BC / 16;
@@ -32,33 +27,23 @@ __global__ void paged_attn_decode_split_kv_mma_kernel(PagedAttentionParams<bf16>
const int gid = lane >> 2; const int gid = lane >> 2;
const int tid4 = lane & 3; const int tid4 = lane & 3;
const int kv_head_idx = blockIdx.x; const int kv_head = blockIdx.x;
const int batch = blockIdx.y; const int batch = blockIdx.y;
const int split = blockIdx.z; const int split = blockIdx.z;
const int G = p.q_head / p.kv_head; const int G = p.q_head / p.kv_head;
const int q_head0 = kv_head_idx * G; const int q_head0 = kv_head * G;
__shared__ __align__(16) bf16 sK[STAGES * BC * LD]; __shared__ __align__(16) bf16 sK[STAGES * BC * LD];
__shared__ __align__(16) bf16 sV[STAGES * BC * LD]; __shared__ __align__(16) bf16 sV[STAGES * BC * LD];
// ---- Load Q directly from global into mma A-operand registers ---- // ---- Load Q directly from global into mma A-operand registers ----
const int q_base = (batch * p.q_head + q_head0) * HEAD_DIM; const int q_base = batch * p.q_stride_b + q_head0 * p.q_stride_h;
const int qra = gid; const int qra = gid;
const int qrb = gid + 8; const int qrb = gid + 8;
const bool va = qra < G, vb = qrb < G; const bool va = qra < G, vb = qrb < G;
unsigned Qa[KD][4]; unsigned Qa[KD][4];
#pragma unroll load_q_mma_frags<KD>(p.q + q_base, p.q_stride_h, p.q_stride_d,
for (int kt = 0; kt < KD; kt++) { qra, qrb, va, vb, tid4, Qa);
int c = kt * 16 + tid4 * 2;
const unsigned* pau = reinterpret_cast<const unsigned*>(
&p.q[q_base + qra * HEAD_DIM + c]);
const unsigned* pbu = reinterpret_cast<const unsigned*>(
&p.q[q_base + qrb * HEAD_DIM + c]);
Qa[kt][0] = va ? pau[0] : 0u;
Qa[kt][1] = vb ? pbu[0] : 0u;
Qa[kt][2] = va ? pau[4] : 0u;
Qa[kt][3] = vb ? pbu[4] : 0u;
}
float Oacc[DN8][4]; float Oacc[DN8][4];
#pragma unroll #pragma unroll
@@ -66,7 +51,6 @@ __global__ void paged_attn_decode_split_kv_mma_kernel(PagedAttentionParams<bf16>
Oacc[j][0] = Oacc[j][1] = Oacc[j][2] = Oacc[j][3] = 0.0f; Oacc[j][0] = Oacc[j][1] = Oacc[j][2] = Oacc[j][3] = 0.0f;
float m0 = -FLT_MAX, m1 = -FLT_MAX, l0 = 0.0f, l1 = 0.0f; float m0 = -FLT_MAX, m1 = -FLT_MAX, l0 = 0.0f, l1 = 0.0f;
const int mask_base = batch * p.kv_len;
const int tiles_total = (p.kv_len + BC - 1) / BC; const int tiles_total = (p.kv_len + BC - 1) / BC;
const int tiles_per_split = (tiles_total + p.num_splits - 1) / p.num_splits; const int tiles_per_split = (tiles_total + p.num_splits - 1) / p.num_splits;
const int ti_begin = split * tiles_per_split; const int ti_begin = split * tiles_per_split;
@@ -76,7 +60,7 @@ __global__ void paged_attn_decode_split_kv_mma_kernel(PagedAttentionParams<bf16>
// Paged strides (constant for the block) // Paged strides (constant for the block)
const int64_t page_stride = (int64_t)p.page_size * p.kv_head * HEAD_DIM; const int64_t page_stride = (int64_t)p.page_size * p.kv_head * HEAD_DIM;
const int64_t pos_stride = (int64_t)p.kv_head * HEAD_DIM; const int64_t pos_stride = (int64_t)p.kv_head * HEAD_DIM;
const int64_t head_off = (int64_t)kv_head_idx * HEAD_DIM; const int64_t head_off = (int64_t)kv_head * HEAD_DIM;
// ---- Load tile lambda: predicated cp.async, paged addressing ---- // ---- Load tile lambda: predicated cp.async, paged addressing ----
auto load_tile = [&](int ti, int buf) { auto load_tile = [&](int ti, int buf) {
@@ -130,9 +114,13 @@ __global__ void paged_attn_decode_split_kv_mma_kernel(PagedAttentionParams<bf16>
Sacc[n8][0] *= p.scale, Sacc[n8][1] *= p.scale, Sacc[n8][0] *= p.scale, Sacc[n8][1] *= p.scale,
Sacc[n8][2] *= p.scale, Sacc[n8][3] *= p.scale; Sacc[n8][2] *= p.scale, Sacc[n8][3] *= p.scale;
int maxc = p.is_causal ? min(p.kv_len, p.causal_offset + 1) : p.kv_len; // Decode: q_len=1, so qrow0=qrow1=0, mask_q_stride irrelevant
int maxc = (p.causal_offset >= 0) ? min(p.kv_len, p.causal_offset + 1) : p.kv_len;
mma_softmax_tile<NC8, DN8>(kv0, maxc, maxc, mma_softmax_tile<NC8, DN8>(kv0, maxc, maxc,
mask_base, p.mask, has_mask, 0, 0,
p.mask_b_stride, 0,
batch,
p.mask, has_mask,
Sacc, Oacc, m0, m1, l0, l1, lane); Sacc, Oacc, m0, m1, l0, l1, lane);
mma_pv_accumulate<DN8, KT2>(Sacc, bV, LD, SWIZ_MASK, lane, Oacc); mma_pv_accumulate<DN8, KT2>(Sacc, bV, LD, SWIZ_MASK, lane, Oacc);
+11 -26
View File
@@ -35,34 +35,19 @@ torch::Tensor attn_prefill(
torch::Tensor k, torch::Tensor k,
torch::Tensor v, torch::Tensor v,
c10::optional<torch::Tensor> mask, c10::optional<torch::Tensor> mask,
bool is_causal = false, int64_t causal_offset,
int64_t causal_offset = 0, double scale,
c10::optional<double> scale = c10::nullopt int64_t layout
) { ) {
AttentionParams<bf16> p; AttentionParams<bf16> p;
attn_pack_params(q, k, v, mask, is_causal, causal_offset, scale, p); attn_pack_params(q, k, v, mask, causal_offset, scale, layout, p);
TORCH_CHECK(p.head_dim % 16 == 0, "head_dim must be multiple of 16"); TORCH_CHECK(p.head_dim % 16 == 0, "head_dim must be multiple of 16");
auto O = torch::empty_like(q); auto O = torch::empty_strided(q.sizes(), q.strides(), q.options());
p.o = (bf16*)O.data_ptr(); auto O_view = (layout == 1) ? O.transpose(1, 2) : O;
p.o = (bf16*)O_view.data_ptr();
switch (p.head_dim) { DISPATCH_HEAD_DIM(p.head_dim, dispatch_prefill, p);
case 32:
dispatch_prefill<32>(p);
break;
case 64:
dispatch_prefill<64>(p);
break;
case 128:
dispatch_prefill<128>(p);
break;
case 256:
dispatch_prefill<256>(p);
break;
default:
TORCH_CHECK(false, "prefill: unsupported head_dim ", p.head_dim,
" (supported: 32,64,128,256)");
}
return O; return O;
} }
@@ -72,8 +57,8 @@ PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) {
py::arg("k"), py::arg("k"),
py::arg("v"), py::arg("v"),
py::arg("mask") = py::none(), py::arg("mask") = py::none(),
py::arg("is_causal") = false, py::arg("causal_offset") = -1,
py::arg("causal_offset") = 0, py::arg("scale") = 0.0,
py::arg("scale") = py::none(), py::arg("layout") = 0,
"GQA prefill (tensor-core mma on sm_80+, scalar fallback)"); "GQA prefill (tensor-core mma on sm_80+, scalar fallback)");
} }
+22 -10
View File
@@ -50,12 +50,14 @@ __global__ void attn_prefill_split_q_kernel_t(AttentionParams<bf16> p) {
__shared__ __align__(16) bf16 sK[P_BC * HEAD_DIM]; __shared__ __align__(16) bf16 sK[P_BC * HEAD_DIM];
__shared__ __align__(16) bf16 sV[P_BC * HEAD_DIM]; __shared__ __align__(16) bf16 sV[P_BC * HEAD_DIM];
// Q: stride-based load [batch, q_head, q_len, head_dim]
float qreg[DPT]; float qreg[DPT];
if (q_row < p.q_len) { if (q_row < p.q_len) {
int q_off = ((batch * p.q_head + q_head) * p.q_len + q_row) * HEAD_DIM + gpos * DPT; int q_off = batch * p.q_stride_b + q_head * p.q_stride_h
+ q_row * p.q_stride_l + gpos * DPT * p.q_stride_d;
#pragma unroll #pragma unroll
for (int i = 0; i < DPT; i++) for (int i = 0; i < DPT; i++)
qreg[i] = __bfloat162float(p.q[q_off + i]) * p.scale; qreg[i] = __bfloat162float(p.q[q_off + i * p.q_stride_d]) * p.scale;
} }
float m = -FLT_MAX, l = 0.0f; float m = -FLT_MAX, l = 0.0f;
@@ -64,7 +66,9 @@ __global__ void attn_prefill_split_q_kernel_t(AttentionParams<bf16> p) {
for (int i = 0; i < DPT; i++) for (int i = 0; i < DPT; i++)
acc[i] = 0.0f; acc[i] = 0.0f;
int kv_base = ((batch * p.kv_head + kv_head) * p.kv_len) * HEAD_DIM; // KV: stride-based base
int kv_base = batch * p.kv_stride_b + kv_head * p.kv_stride_h;
int mask_batch_base = batch * p.mask_b_stride;
int tiles = (p.kv_len + P_BC - 1) / P_BC; int tiles = (p.kv_len + P_BC - 1) / P_BC;
int tt = G * ROWS; int tt = G * ROWS;
int lid = row * G + gpos; int lid = row * G + gpos;
@@ -79,15 +83,19 @@ __global__ void attn_prefill_split_q_kernel_t(AttentionParams<bf16> p) {
int kv0 = ti * P_BC; int kv0 = ti * P_BC;
int tlen = min(P_BC, p.kv_len - kv0); int tlen = min(P_BC, p.kv_len - kv0);
// Load K/V into shared memory from strided global
for (int i = lid; i < tlen * HEAD_DIM; i += tt) { for (int i = lid; i < tlen * HEAD_DIM; i += tt) {
int gidx = kv_base + (kv0 + i / HEAD_DIM) * HEAD_DIM + (i % HEAD_DIM); int s = i / HEAD_DIM;
sK[i] = p.k[gidx]; int d_dim = i % HEAD_DIM;
sV[i] = p.v[gidx]; int kv_idx = kv0 + s;
int g_off = kv_base + kv_idx * p.kv_stride_l + d_dim * p.kv_stride_d;
sK[i] = p.k[g_off];
sV[i] = p.v[g_off];
} }
__syncthreads(); __syncthreads();
int lim = tlen; int lim = tlen;
if (p.is_causal && q_row < p.q_len) { if (p.causal_offset >= 0 && q_row < p.q_len) {
int ep = q_row + p.causal_offset + 1; int ep = q_row + p.causal_offset + 1;
if (kv0 >= ep) if (kv0 >= ep)
lim = 0; lim = 0;
@@ -95,6 +103,7 @@ __global__ void attn_prefill_split_q_kernel_t(AttentionParams<bf16> p) {
lim = ep - kv0; lim = ep - kv0;
} }
int mask_row_base = mask_batch_base + q_row * p.mask_q_stride;
for (int s = 0; s < lim; s++) { for (int s = 0; s < lim; s++) {
const bf16* kr = sK + s * HEAD_DIM + gpos * DPT; const bf16* kr = sK + s * HEAD_DIM + gpos * DPT;
float part = 0.0f; float part = 0.0f;
@@ -108,7 +117,8 @@ __global__ void attn_prefill_split_q_kernel_t(AttentionParams<bf16> p) {
} }
float dot = group_reduce_sum<G>(part, gmask); float dot = group_reduce_sum<G>(part, gmask);
if (p.use_mask && p.mask && !p.mask[batch * p.kv_len + kv0 + s]) int kv_idx = kv0 + s;
if (p.use_mask && p.mask && !p.mask[mask_row_base + kv_idx])
dot = -FLT_MAX; dot = -FLT_MAX;
float nm = fmaxf(m, dot); float nm = fmaxf(m, dot);
@@ -131,10 +141,12 @@ __global__ void attn_prefill_split_q_kernel_t(AttentionParams<bf16> p) {
} }
if (q_row < p.q_len) { if (q_row < p.q_len) {
int o_off = ((batch * p.q_head + q_head) * p.q_len + q_row) * HEAD_DIM + gpos * DPT; // O: stride-based write
int o_off = batch * p.q_stride_b + q_head * p.q_stride_h
+ q_row * p.q_stride_l + gpos * DPT * p.q_stride_d;
float rl = (l > 1e-10f) ? (1.0f / l) : 0.0f; float rl = (l > 1e-10f) ? (1.0f / l) : 0.0f;
#pragma unroll #pragma unroll
for (int i = 0; i < DPT; i++) for (int i = 0; i < DPT; i++)
p.o[o_off + i] = __float2bfloat16(acc[i] * rl); p.o[o_off + i * p.q_stride_d] = __float2bfloat16(acc[i] * rl);
} }
} }
+28 -36
View File
@@ -57,7 +57,7 @@ __global__ void attn_prefill_split_q_mma_kernel(AttentionParams<bf16> p) {
const int kv_head = q_head / (p.q_head / p.kv_head); const int kv_head = q_head / (p.q_head / p.kv_head);
const int qrow0 = (blockIdx.x * WARPS + warp) * BR; const int qrow0 = (blockIdx.x * WARPS + warp) * BR;
// Static shared memory — sized by template parameters at compile time. // ---- Static shared memory: double-buffered K/V ----
// K/V are double-buffered (STAGES=2): the next tile's cp.async load runs // K/V are double-buffered (STAGES=2): the next tile's cp.async load runs
// while the current tile's tensor-core math executes, hiding global-load // while the current tile's tensor-core math executes, hiding global-load
// latency (FA2-style software pipeline). No dynamic smem / carveout opt-in. // latency (FA2-style software pipeline). No dynamic smem / carveout opt-in.
@@ -65,30 +65,16 @@ __global__ void attn_prefill_split_q_mma_kernel(AttentionParams<bf16> p) {
__shared__ __align__(16) bf16 sK[STAGES * BC * LD]; __shared__ __align__(16) bf16 sK[STAGES * BC * LD];
__shared__ __align__(16) bf16 sV[STAGES * BC * LD]; __shared__ __align__(16) bf16 sV[STAGES * BC * LD];
// Load the Q fragments straight from global into the mma A-operand layout // Load Q fragments straight from global into mma A-operand layout.
// (m16n8k16, row-major): no sQ staging area and no serialized per-warp // stride_row = p.q_stride_l for prefill (multi-q rows across q_len).
// prologue barriers. Each lane reads exactly the 8 Q elements ldmatrix // See attn_mma_utils.cuh for the shared template.
// would have produced, pre-scaled by the attention scale. Kept resident in const int q_base = batch * p.q_stride_b + q_head * p.q_stride_h;
// registers across the tile loop.
// frag[0]/[2]: row = qrow0 + gid ; frag[1]/[3]: row = qrow0 + gid + 8
// frag[0]/[1]: cols kt*16 + tid4*2 + {0,1} ; frag[2]/[3]: + 8
const int q_base = ((batch * p.q_head + q_head) * p.q_len) * HEAD_DIM;
const int qra = qrow0 + gid; const int qra = qrow0 + gid;
const int qrb = qrow0 + gid + 8; const int qrb = qrow0 + gid + 8;
const bool va = qra < p.q_len, vb = qrb < p.q_len; const bool va = qra < p.q_len, vb = qrb < p.q_len;
unsigned Qa[KD][4]; unsigned Qa[KD][4];
#pragma unroll load_q_mma_frags<KD>(p.q + q_base, p.q_stride_l, p.q_stride_d,
for (int kt = 0; kt < KD; kt++) { qra, qrb, va, vb, tid4, Qa);
int c = kt * 16 + tid4 * 2;
const unsigned* pau = reinterpret_cast<const unsigned*>(
&p.q[q_base + qra * HEAD_DIM + c]);
const unsigned* pbu = reinterpret_cast<const unsigned*>(
&p.q[q_base + qrb * HEAD_DIM + c]);
Qa[kt][0] = va ? pau[0] : 0u;
Qa[kt][1] = vb ? pbu[0] : 0u;
Qa[kt][2] = va ? pau[4] : 0u;
Qa[kt][3] = vb ? pbu[4] : 0u;
}
float Oacc[DN8][4]; float Oacc[DN8][4];
#pragma unroll #pragma unroll
@@ -96,18 +82,18 @@ __global__ void attn_prefill_split_q_mma_kernel(AttentionParams<bf16> p) {
Oacc[j][0] = Oacc[j][1] = Oacc[j][2] = Oacc[j][3] = 0.0f; Oacc[j][0] = Oacc[j][1] = Oacc[j][2] = Oacc[j][3] = 0.0f;
float m0 = -FLT_MAX, m1 = -FLT_MAX, l0 = 0.0f, l1 = 0.0f; float m0 = -FLT_MAX, m1 = -FLT_MAX, l0 = 0.0f, l1 = 0.0f;
const int kv_base = ((batch * p.kv_head + kv_head) * p.kv_len) * HEAD_DIM; // KV: stride-based base
const int kv_base = batch * p.kv_stride_b + kv_head * p.kv_stride_h;
const int tiles = (p.kv_len + BC - 1) / BC; const int tiles = (p.kv_len + BC - 1) / BC;
const int qr0 = qrow0 + gid; // row for c0/c1 const int qr0 = qrow0 + gid; // row for c0/c1
const int qr1 = qrow0 + gid + 8; // row for c2/c3 const int qr1 = qrow0 + gid + 8; // row for c2/c3
// Causal tile-skip bounds (no-op when is_causal == 0) // Causal tile-skip bounds (no-op when causal_offset < 0)
const int use_skip = p.is_causal; const int use_skip = (p.causal_offset >= 0) ? 1 : 0;
const int max_kv = qrow0 + BR - 1 + p.causal_offset; const int max_kv = qrow0 + BR - 1 + p.causal_offset;
const int block_max_kv = const int block_max_kv =
blockIdx.x * WARPS * BR + WARPS * BR - 1 + p.causal_offset; blockIdx.x * WARPS * BR + WARPS * BR - 1 + p.causal_offset;
const int has_mask = p.use_mask && p.mask; const int has_mask = p.use_mask && p.mask;
const int mb = batch * p.kv_len;
// Last active tile: block-level causal bound (all warps in the block share // Last active tile: block-level causal bound (all warps in the block share
// the K/V load, so the prefetch range is the block max, not per-warp). // the K/V load, so the prefetch range is the block max, not per-warp).
@@ -120,6 +106,7 @@ __global__ void attn_prefill_split_q_mma_kernel(AttentionParams<bf16> p) {
constexpr int VEC = 8; // bf16 per cp.async unit (16 bytes) constexpr int VEC = 8; // bf16 per cp.async unit (16 bytes)
constexpr int TOTAL = BC * HEAD_DIM; constexpr int TOTAL = BC * HEAD_DIM;
// ---- Load tile lambda: predicated cp.async ----
// Issue cp.async loads for tile `ti` into shared buffer `buf`. Predicated // Issue cp.async loads for tile `ti` into shared buffer `buf`. Predicated
// loads zero-fill rows past kv_len, so partial tiles need no scalar path. // loads zero-fill rows past kv_len, so partial tiles need no scalar path.
auto load_tile = [&](int ti, int buf) { auto load_tile = [&](int ti, int buf) {
@@ -132,13 +119,14 @@ __global__ void attn_prefill_split_q_mma_kernel(AttentionParams<bf16> p) {
int kc = kv0 + r; int kc = kv0 + r;
bool valid = kc < p.kv_len; bool valid = kc < p.kv_len;
int off = r * LD + swiz_col(d, r, SWIZ_MASK); int off = r * LD + swiz_col(d, r, SWIZ_MASK);
cp_async_16_pred(&dK[off], &p.k[kv_base + kc * HEAD_DIM + d], valid); int g_off = kv_base + kc * p.kv_stride_l + d * p.kv_stride_d;
cp_async_16_pred(&dV[off], &p.v[kv_base + kc * HEAD_DIM + d], valid); cp_async_16_pred(&dK[off], &p.k[g_off], valid);
cp_async_16_pred(&dV[off], &p.v[g_off], valid);
} }
cp_async_commit(); cp_async_commit();
}; };
// Prologue: kick off the first tile's load. // ---- Prologue: issue first tile load ----
load_tile(0, 0); load_tile(0, 0);
for (int ti = 0; ti <= t_end; ti++) { for (int ti = 0; ti <= t_end; ti++) {
@@ -170,12 +158,15 @@ __global__ void attn_prefill_split_q_mma_kernel(AttentionParams<bf16> p) {
Sacc[n8][0] *= p.scale, Sacc[n8][1] *= p.scale, Sacc[n8][0] *= p.scale, Sacc[n8][1] *= p.scale,
Sacc[n8][2] *= p.scale, Sacc[n8][3] *= p.scale; Sacc[n8][2] *= p.scale, Sacc[n8][3] *= p.scale;
int maxc0 = p.is_causal ? min(p.kv_len, qr0 + p.causal_offset + 1) int maxc0 = (p.causal_offset >= 0) ? min(p.kv_len, qr0 + p.causal_offset + 1)
: p.kv_len; : p.kv_len;
int maxc1 = p.is_causal ? min(p.kv_len, qr1 + p.causal_offset + 1) int maxc1 = (p.causal_offset >= 0) ? min(p.kv_len, qr1 + p.causal_offset + 1)
: p.kv_len; : p.kv_len;
mma_softmax_tile<NC8, DN8>(kv0, maxc0, maxc1, mma_softmax_tile<NC8, DN8>(kv0, maxc0, maxc1,
mb, p.mask, has_mask, qr0, qr1,
p.mask_b_stride, p.mask_q_stride,
batch,
p.mask, has_mask,
Sacc, Oacc, m0, m1, l0, l1, lane); Sacc, Oacc, m0, m1, l0, l1, lane);
mma_pv_accumulate<DN8, KT2>(Sacc, bV, LD, SWIZ_MASK, lane, Oacc); mma_pv_accumulate<DN8, KT2>(Sacc, bV, LD, SWIZ_MASK, lane, Oacc);
@@ -186,19 +177,20 @@ __global__ void attn_prefill_split_q_mma_kernel(AttentionParams<bf16> p) {
// halves store count and removes the uncoalesced scalar-store penalty) // halves store count and removes the uncoalesced scalar-store penalty)
float rl0 = (l0 > 1e-20f) ? (1.0f / l0) : 0.0f; float rl0 = (l0 > 1e-20f) ? (1.0f / l0) : 0.0f;
float rl1 = (l1 > 1e-20f) ? (1.0f / l1) : 0.0f; float rl1 = (l1 > 1e-20f) ? (1.0f / l1) : 0.0f;
const int o_base = ((batch * p.q_head + q_head) * p.q_len) * HEAD_DIM; // O: stride-based write
const int o_base = batch * p.q_stride_b + q_head * p.q_stride_h;
#pragma unroll #pragma unroll
for (int dn8 = 0; dn8 < DN8; dn8++) { for (int dn8 = 0; dn8 < DN8; dn8++) {
int d = dn8 * 8 + 2 * tid4; int d = dn8 * 8 + 2 * tid4;
if (qr0 < p.q_len) { if (qr0 < p.q_len) {
__nv_bfloat162 v = __floats2bfloat162_rn(Oacc[dn8][0] * rl0, __nv_bfloat162 v = __floats2bfloat162_rn(Oacc[dn8][0] * rl0,
Oacc[dn8][1] * rl0); Oacc[dn8][1] * rl0);
*reinterpret_cast<__nv_bfloat162*>(&p.o[o_base + qr0 * HEAD_DIM + d]) = v; *reinterpret_cast<__nv_bfloat162*>(&p.o[o_base + qr0 * p.q_stride_l + d * p.q_stride_d]) = v;
} }
if (qr1 < p.q_len) { if (qr1 < p.q_len) {
__nv_bfloat162 v = __floats2bfloat162_rn(Oacc[dn8][2] * rl1, __nv_bfloat162 v = __floats2bfloat162_rn(Oacc[dn8][2] * rl1,
Oacc[dn8][3] * rl1); Oacc[dn8][3] * rl1);
*reinterpret_cast<__nv_bfloat162*>(&p.o[o_base + qr1 * HEAD_DIM + d]) = v; *reinterpret_cast<__nv_bfloat162*>(&p.o[o_base + qr1 * p.q_stride_l + d * p.q_stride_d]) = v;
} }
} }
} }
+5 -3
View File
@@ -98,8 +98,9 @@ static void bench() {
AttentionParams<bf16> p; AttentionParams<bf16> p;
p.batch = B; p.q_head = Hq; p.kv_head = Hk; p.q_len = 1; p.kv_len = sl; p.batch = B; p.q_head = Hq; p.kv_head = Hk; p.q_len = 1; p.kv_len = sl;
p.head_dim = D; p.use_mask = 0; p.is_causal = 0; p.causal_offset = 0; p.head_dim = D; p.use_mask = 0; p.causal_offset = -1;
p.scale = 1.0f / sqrtf((float)D); p.scale = 1.0f / sqrtf((float)D);
set_default_strides(p);
p.q = dQ; p.k = dK; p.v = dV; p.mask = nullptr; p.o = dO; p.q = dQ; p.k = dK; p.v = dV; p.mask = nullptr; p.o = dO;
DecodeScratch sc; DecodeScratch sc;
@@ -160,8 +161,9 @@ int main() {
AttentionParams<bf16> p; AttentionParams<bf16> p;
p.batch=B; p.q_head=Hq; p.kv_head=Hk; p.q_len=1; p.kv_len=sl; p.head_dim=D; p.batch=B; p.q_head=Hq; p.kv_head=Hk; p.q_len=1; p.kv_len=sl; p.head_dim=D;
p.use_mask=0; p.is_causal=0; p.causal_offset=0; p.use_mask=0; p.causal_offset=-1;
p.scale=1.0f/sqrtf((float)D); p.scale=1.0f/sqrtf((float)D);
set_default_strides(p);
p.q=dQ; p.k=dK; p.v=dV; p.mask=nullptr; p.o=dO; p.q=dQ; p.k=dK; p.v=dV; p.mask=nullptr; p.o=dO;
// Split-K scratch (max 32 splits), sized for the production MMA path. // Split-K scratch (max 32 splits), sized for the production MMA path.
@@ -180,7 +182,7 @@ int main() {
cudaMemcpy(hOut,dO,nQ*2,cudaMemcpyDeviceToHost); cudaMemcpy(hOut,dO,nQ*2,cudaMemcpyDeviceToHost);
float* ref=new float[nQ]; float* ref=new float[nQ];
cpu_attention_ref(hQ, hK, hV, hMask, ref, B, Hq, Hk, 1, sl, D, 0, 0); cpu_attention_ref(hQ, hK, hV, hMask, ref, B, Hq, Hk, 1, sl, D, -1);
float max_err=0; float max_err=0;
for (size_t i=0;i<nQ;i++){ for (size_t i=0;i<nQ;i++){
+5 -3
View File
@@ -138,13 +138,14 @@ static int run_test(int B, int Hq, int Hkv, int kv_len, int page_size, int seed)
} }
float* h_o_ref = (float*)calloc(B * Hq * HEAD_DIM, sizeof(float)); float* h_o_ref = (float*)calloc(B * Hq * HEAD_DIM, sizeof(float));
cpu_attention_ref(h_q_f, h_k_f, h_v_f, nullptr, h_o_ref, B, Hq, Hkv, 1, kv_len, HEAD_DIM, 0, 0); cpu_attention_ref(h_q_f, h_k_f, h_v_f, nullptr, h_o_ref, B, Hq, Hkv, 1, kv_len, HEAD_DIM, -1);
float scale_val = 1.0f / sqrtf((float)HEAD_DIM); float scale_val = 1.0f / sqrtf((float)HEAD_DIM);
PagedAttentionParams<bf16, float> p; PagedAttentionParams<bf16, float> p;
p.batch = B; p.q_head = Hq; p.kv_head = Hkv; p.q_len = 1; p.batch = B; p.q_head = Hq; p.kv_head = Hkv; p.q_len = 1;
p.kv_len = kv_len; p.head_dim = HEAD_DIM; p.kv_len = kv_len; p.head_dim = HEAD_DIM;
p.use_mask = 0; p.is_causal = 0; p.causal_offset = 0; p.use_mask = 0; p.causal_offset = -1;
set_default_paged_strides(p);
p.num_splits = 1; p.scale = scale_val; p.num_splits = 1; p.scale = scale_val;
p.page_size = page_size; p.max_pages = max_pages; p.page_size = page_size; p.max_pages = max_pages;
p.page_table = d_pt; p.page_table = d_pt;
@@ -272,7 +273,8 @@ static void bench_config(int B, int Hq, int Hkv, int kv_len, int page_size) {
PagedAttentionParams<bf16, float> pa; PagedAttentionParams<bf16, float> pa;
pa.batch = B; pa.q_head = Hq; pa.kv_head = Hkv; pa.q_len = 1; pa.batch = B; pa.q_head = Hq; pa.kv_head = Hkv; pa.q_len = 1;
pa.kv_len = kv_len; pa.head_dim = HEAD_DIM; pa.kv_len = kv_len; pa.head_dim = HEAD_DIM;
pa.use_mask = 0; pa.is_causal = 0; pa.causal_offset = 0; pa.use_mask = 0; pa.causal_offset = -1;
set_default_paged_strides(pa);
pa.num_splits = 1; pa.scale = scale_val; pa.num_splits = 1; pa.scale = scale_val;
pa.page_size = page_size; pa.max_pages = max_pages; pa.page_size = page_size; pa.max_pages = max_pages;
pa.page_table = d_pt; pa.page_table = d_pt;
+5 -3
View File
@@ -75,7 +75,8 @@ static void bench() {
AttentionParams<bf16> p; AttentionParams<bf16> p;
p.batch=B; p.q_head=Hq; p.kv_head=Hk; p.q_len=ql; p.kv_len=kl; p.head_dim=D; p.batch=B; p.q_head=Hq; p.kv_head=Hk; p.q_len=ql; p.kv_len=kl; p.head_dim=D;
p.use_mask=0; p.is_causal=causal; p.causal_offset=0; p.use_mask=0; p.causal_offset=causal?0:-1;
set_default_strides(p);
p.scale=1.0f/sqrtf((float)D); p.scale=1.0f/sqrtf((float)D);
p.q=dQ; p.k=dK; p.v=dV; p.mask=nullptr; p.o=dO; p.q=dQ; p.k=dK; p.v=dV; p.mask=nullptr; p.o=dO;
@@ -143,7 +144,8 @@ int main() {
AttentionParams<bf16> p; AttentionParams<bf16> p;
p.batch=B; p.q_head=Hq; p.kv_head=Hk; p.q_len=ql; p.kv_len=kl; p.head_dim=D; p.batch=B; p.q_head=Hq; p.kv_head=Hk; p.q_len=ql; p.kv_len=kl; p.head_dim=D;
p.use_mask=0; p.is_causal=causal; p.causal_offset=0; p.use_mask=0; p.causal_offset=causal?0:-1;
set_default_strides(p);
p.scale=1.0f/sqrtf((float)D); p.scale=1.0f/sqrtf((float)D);
p.q=dQ; p.k=dK; p.v=dV; p.mask=nullptr; p.o=dO; p.q=dQ; p.k=dK; p.v=dV; p.mask=nullptr; p.o=dO;
@@ -158,7 +160,7 @@ int main() {
cudaMemcpy(hOut,dO,nQ*2,cudaMemcpyDeviceToHost); cudaMemcpy(hOut,dO,nQ*2,cudaMemcpyDeviceToHost);
float* ref=new float[nQ]; float* ref=new float[nQ];
cpu_attention_ref(hQ, hK, hV, nullptr, ref, B, Hq, Hk, ql, kl, D, causal, 0); cpu_attention_ref(hQ, hK, hV, nullptr, ref, B, Hq, Hk, ql, kl, D, causal ? 0 : -1);
float max_err=0; float max_err=0;
for (size_t i=0;i<nQ;i++) { for (size_t i=0;i<nQ;i++) {
+29 -2
View File
@@ -101,6 +101,32 @@ void dispatch_by_head_dim(int head_dim, Fn&& fn) {
_HeadSwitch<32, 64, 128, 256>::call(head_dim, fn); _HeadSwitch<32, 64, 128, 256>::call(head_dim, fn);
} }
// Set default strides for contiguous b h l d layout on AttentionParams.
template<typename P>
inline void set_default_strides(P& p) {
p.q_stride_b = p.q_head * p.q_len * p.head_dim;
p.q_stride_h = p.q_len * p.head_dim;
p.q_stride_l = p.head_dim;
p.q_stride_d = 1;
p.kv_stride_b = p.kv_head * p.kv_len * p.head_dim;
p.kv_stride_h = p.kv_len * p.head_dim;
p.kv_stride_l = p.head_dim;
p.kv_stride_d = 1;
p.mask_b_stride = p.kv_len;
p.mask_q_stride = 0;
}
// Set default Q strides for contiguous b h l d layout on PagedAttentionParams.
template<typename P>
inline void set_default_paged_strides(P& p) {
p.q_stride_b = p.q_head * p.q_len * p.head_dim;
p.q_stride_h = p.q_len * p.head_dim;
p.q_stride_l = p.head_dim;
p.q_stride_d = 1;
p.mask_b_stride = p.kv_len;
p.mask_q_stride = 0;
}
// Generic CPU reference for multi-query / grouped-query attention. // Generic CPU reference for multi-query / grouped-query attention.
// Tensor shapes (all float*): // Tensor shapes (all float*):
// Q : [B, Hq, q_len, D] // Q : [B, Hq, q_len, D]
@@ -108,10 +134,11 @@ void dispatch_by_head_dim(int head_dim, Fn&& fn) {
// V : [B, Hk, kv_len, D] // V : [B, Hk, kv_len, D]
// O : [B, Hq, q_len, D] // O : [B, Hq, q_len, D]
// mask: if q_len == 1, shape is [B, kv_len]; otherwise mask is not supported. // mask: if q_len == 1, shape is [B, kv_len]; otherwise mask is not supported.
// causal_offset: -1 = non-causal; >=0 = absolute position of first Q token.
static void cpu_attention_ref( static void cpu_attention_ref(
const float* Q, const float* K, const float* V, const bool* mask, const float* Q, const float* K, const float* V, const bool* mask,
float* O, int B, int Hq, int Hk, int q_len, int kv_len, int D, float* O, int B, int Hq, int Hk, int q_len, int kv_len, int D,
int is_causal, int causal_offset int causal_offset
) { ) {
float scale = 1.0f / sqrtf((float)D); float scale = 1.0f / sqrtf((float)D);
int n_rep = Hq / Hk; int n_rep = Hq / Hk;
@@ -122,7 +149,7 @@ static void cpu_attention_ref(
float mv = -INFINITY, sv = 0.0f; float mv = -INFINITY, sv = 0.0f;
float accum[256] = {0.0f}; float accum[256] = {0.0f};
int lim = kv_len; int lim = kv_len;
if (is_causal) { if (causal_offset >= 0) {
int c = qi + causal_offset + 1; int c = qi + causal_offset + 1;
lim = (c < kv_len) ? c : kv_len; lim = (c < kv_len) ? c : kv_len;
} }
+1 -1
View File
@@ -5,7 +5,7 @@ from huggingface_hub import snapshot_download
PROJECT_ROOT = Path(__file__).resolve().parents[2] PROJECT_ROOT = Path(__file__).resolve().parents[2]
DEFAULT_LOCAL_DIR = Path(PROJECT_ROOT, "params") DEFAULT_LOCAL_DIR = Path(PROJECT_ROOT, "params")
DEFAULT_REPO_ID = "ViperEk/KHAOSZ" DEFAULT_REPO_ID = "ViperEkura/AstrAI-V1-instruct"
if __name__ == "__main__": if __name__ == "__main__":
parser = argparse.ArgumentParser( parser = argparse.ArgumentParser(
+2 -4
View File
@@ -26,11 +26,9 @@ def batch_generate():
prompts = [ prompts = [
tokenizer.apply_chat_template( tokenizer.apply_chat_template(
[ [{"role": "user", "content": q}],
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": q},
],
tokenize=False, tokenize=False,
add_generation_prompt=True,
) )
for q in inputs for q in inputs
] ]
+24 -8
View File
@@ -42,11 +42,24 @@ def parse_args():
default=2048, default=2048,
help="Maximum tokens to generate", help="Maximum tokens to generate",
) )
parser.add_argument(
"--frequency_penalty",
type=float,
default=0.5,
help="Penalty per occurrence for repeated tokens (0.0 disables, "
"range -2.0~2.0, typical 0.3-1.0)",
)
parser.add_argument(
"--rep_window",
type=int,
default=64,
help="Number of recent prompt tokens to include in penalty history",
)
parser.add_argument( parser.add_argument(
"--system_prompt", "--system_prompt",
type=str, type=str,
default="You are a helpful assistant.", default="",
help="Optional system prompt", help="Optional system prompt (default: empty, model not SFT-trained on system role)",
) )
return parser.parse_args() return parser.parse_args()
@@ -60,18 +73,20 @@ def chat():
model.to(device="cuda", dtype=torch.bfloat16) model.to(device="cuda", dtype=torch.bfloat16)
engine = InferenceEngine(model=model, tokenizer=tokenizer) engine = InferenceEngine(model=model, tokenizer=tokenizer)
messages = [{"role": "system", "content": args.system_prompt}]
while True: while True:
query = input(">> ") query = input(">> ")
if query == "!exit": if query == "!exit":
break break
messages.append({"role": "user", "content": query}) msgs = []
if args.system_prompt:
msgs.append({"role": "system", "content": args.system_prompt})
msgs.append({"role": "user", "content": query})
prompt = tokenizer.apply_chat_template(
msgs, tokenize=False, add_generation_prompt=True
)
full_response = "" full_response = ""
prompt = tokenizer.apply_chat_template(messages, tokenize=False)
for token in engine.generate( for token in engine.generate(
prompt=prompt, prompt=prompt,
stream=True, stream=True,
@@ -79,12 +94,13 @@ def chat():
temperature=args.temperature, temperature=args.temperature,
top_p=args.top_p, top_p=args.top_p,
top_k=args.top_k, top_k=args.top_k,
frequency_penalty=args.frequency_penalty,
rep_window=args.rep_window,
): ):
print(token, end="", flush=True) print(token, end="", flush=True)
full_response += token full_response += token
print() print()
messages.append({"role": "assistant", "content": full_response.strip()})
if __name__ == "__main__": if __name__ == "__main__":
+16 -2
View File
@@ -117,7 +117,7 @@ def print_component_summary(results: dict[str, dict], title: str):
r["er_99_norm"] r["er_99_norm"]
for vs in matrix_groups.values() for vs in matrix_groups.values()
for r in vs for r in vs
if "_norm" not in r or not r.get("is_1d") if not r.get("is_1d")
] ]
if all_er: if all_er:
m = sum(all_er) / len(all_er) m = sum(all_er) / len(all_er)
@@ -232,8 +232,16 @@ def main():
action="store_true", action="store_true",
help="Skip SVD analysis, only show weight statistics (mean/std/min/max).", help="Skip SVD analysis, only show weight statistics (mean/std/min/max).",
) )
parser.add_argument(
"--output",
type=str,
default=None,
help="Save results as JSON to this path.",
)
args = parser.parse_args() args = parser.parse_args()
all_results = {}
def analyze_one(ckpt_dir: str, label: str): def analyze_one(ckpt_dir: str, label: str):
ckpt_dir = Path(ckpt_dir) ckpt_dir = Path(ckpt_dir)
weights_path = ckpt_dir / "model.safetensors" weights_path = ckpt_dir / "model.safetensors"
@@ -294,13 +302,19 @@ def main():
) )
print_layer_grid(results) print_layer_grid(results)
print_weight_stats(results) print_weight_stats(results)
all_results[label] = results
return results return results
analyze_one(args.ckpt_dir, "Primary") analyze_one(args.ckpt_dir, "Primary")
if args.compare: if args.compare:
for cdir in args.compare: for cdir in args.compare:
analyze_one(cdir, "Compare") analyze_one(cdir, f"Compare_{cdir}")
if args.output:
with open(args.output, "w", encoding="utf-8") as f:
json.dump(all_results, f, indent=2)
print(f"\nResults saved to {args.output}")
if __name__ == "__main__": if __name__ == "__main__":
+48 -35
View File
@@ -8,7 +8,6 @@ Config is a single dataclass; side effects are isolated at pipeline boundaries.
""" """
import argparse import argparse
import itertools
import json import json
import os import os
import re import re
@@ -21,6 +20,7 @@ from typing import Dict, Iterator, List, Optional, Sequence, Tuple
import numpy as np import numpy as np
import torch import torch
import tqdm import tqdm
from datasets import load_dataset
from astrai.inference import InferenceEngine from astrai.inference import InferenceEngine
from astrai.model import AutoModel from astrai.model import AutoModel
@@ -30,9 +30,7 @@ from astrai.tokenize import AutoTokenizer
# Config # Config
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
HUMANEVAL_URL = ( HUMANEVAL_HF_DATASET = "openai/openai_humaneval"
"https://github.com/openai/human-eval/raw/master/data/HumanEval.jsonl.gz"
)
STOP_SEQUENCES = [ STOP_SEQUENCES = [
"\nclass ", "\nclass ",
@@ -65,21 +63,16 @@ class EvalConfig:
problem_indices: Optional[List[int]] = None problem_indices: Optional[List[int]] = None
def download(url: str, path: str): def download(path: str):
if os.path.exists(path): if os.path.exists(path):
return return
import gzip
import urllib.request
os.makedirs(os.path.dirname(path) or ".", exist_ok=True) os.makedirs(os.path.dirname(path) or ".", exist_ok=True)
print(f"Downloading {url} ...") print(f"Downloading HumanEval from HuggingFace ({HUMANEVAL_HF_DATASET}) ...")
tmp = path + ".tmp" ds = load_dataset(HUMANEVAL_HF_DATASET, split="test")
urllib.request.urlretrieve(url, tmp) with open(path, "w", encoding="utf-8") as f:
with gzip.open(tmp, "rb") as f_in: for item in ds:
with open(path, "wb") as f_out: f.write(json.dumps(item, ensure_ascii=False) + "\n")
f_out.write(f_in.read()) print(f" saved {len(ds)} problems to {path}")
os.remove(tmp)
print(f" saved to {path}")
def load_jsonl(path: str) -> List[dict]: def load_jsonl(path: str) -> List[dict]:
@@ -233,7 +226,7 @@ def execute_one(args: tuple) -> bool:
return False return False
def test_one(item: dict, cfg: EvalConfig) -> Tuple[str, int, int]: def test_one(item: dict, cfg: EvalConfig, pool=None) -> Tuple[str, int, int]:
from concurrent.futures import ProcessPoolExecutor from concurrent.futures import ProcessPoolExecutor
task_id = item["task_id"] task_id = item["task_id"]
@@ -247,11 +240,16 @@ def test_one(item: dict, cfg: EvalConfig) -> Tuple[str, int, int]:
for c in completions for c in completions
] ]
n = len(codes) n = len(codes)
passed = 0
with ProcessPoolExecutor(max_workers=cfg.test_workers) as pool: def _run(p):
for ok in pool.map(execute_one, codes): return sum(1 for ok in p.map(execute_one, codes) if ok)
if ok:
passed += 1 if pool is not None:
passed = _run(pool)
else:
with ProcessPoolExecutor(max_workers=cfg.test_workers) as p:
passed = _run(p)
return task_id, n, passed return task_id, n, passed
@@ -259,8 +257,14 @@ def test_all(
items: Sequence[dict], items: Sequence[dict],
cfg: EvalConfig, cfg: EvalConfig,
) -> Iterator[Tuple[str, int, int]]: ) -> Iterator[Tuple[str, int, int]]:
for item in tqdm.tqdm(items, desc="Testing", unit="problem"): from concurrent.futures import ProcessPoolExecutor
yield test_one(item, cfg)
pool = ProcessPoolExecutor(max_workers=cfg.test_workers)
try:
for item in tqdm.tqdm(items, desc="Testing", unit="problem"):
yield test_one(item, cfg, pool)
finally:
pool.shutdown(wait=True)
def pass_at_k(n: int, c: int, k: int) -> float: def pass_at_k(n: int, c: int, k: int) -> float:
@@ -273,26 +277,32 @@ def score_results(
results: Iterator[Tuple[str, int, int]], results: Iterator[Tuple[str, int, int]],
k_values: Tuple[int, ...], k_values: Tuple[int, ...],
) -> Dict: ) -> Dict:
# filter to k <= n (peek first result to get n) """Score pass@k for each problem.
first = next(results)
results = itertools.chain([first], results)
n = first[1]
k_values = tuple(k for k in k_values if k <= n)
k values are filtered per-problem: if a problem has n < k samples
(e.g. after deduplication), pass@k is not computed for that problem.
The summary averages only over problems where the k was computed.
"""
scores = {k: [] for k in k_values} scores = {k: [] for k in k_values}
output = {} output = {}
for task_id, n, passed in results: for task_id, n, passed in results:
entry = {"task_id": task_id, "n": n, "passed": passed} entry = {"task_id": task_id, "n": n, "passed": passed}
for k in k_values: for k in k_values:
pk = round(pass_at_k(n, passed, k), 4) if k <= n:
entry[f"pass@{k}"] = pk pk = round(pass_at_k(n, passed, k), 4)
scores[k].append(pk) entry[f"pass@{k}"] = pk
scores[k].append(pk)
else:
entry[f"pass@{k}"] = None
output[task_id] = entry output[task_id] = entry
summary = {} summary = {}
for k in k_values: for k in k_values:
vals = scores[k] vals = scores[k]
summary[f"pass@{k}"] = round(float(np.mean(vals)), 4) if vals:
summary[f"pass@{k}"] = round(float(np.mean(vals)), 4)
else:
summary[f"pass@{k}"] = None
output["_summary"] = summary output["_summary"] = summary
return output return output
@@ -302,7 +312,7 @@ def run_pipeline(cfg: EvalConfig) -> Dict:
with open(cfg.test_only, encoding="utf-8") as f: with open(cfg.test_only, encoding="utf-8") as f:
generated = json.load(f) generated = json.load(f)
else: else:
download(HUMANEVAL_URL, cfg.data_path) download(cfg.data_path)
problems = load_jsonl(cfg.data_path) problems = load_jsonl(cfg.data_path)
if cfg.problem_indices: if cfg.problem_indices:
@@ -375,7 +385,10 @@ def report(scored: Dict):
summary = scored.pop("_summary", {}) summary = scored.pop("_summary", {})
print(f"\n{'=' * 60}") print(f"\n{'=' * 60}")
for k, v in summary.items(): for k, v in summary.items():
print(f" {k}: {v:.2%}") if v is not None:
print(f" {k}: {v:.2%}")
else:
print(f" {k}: N/A")
print(f"{'=' * 60}") print(f"{'=' * 60}")
scored["_summary"] = summary scored["_summary"] = summary
+137 -105
View File
@@ -16,7 +16,9 @@ v2 changelog:
""" """
import argparse import argparse
import glob
import json import json
import os
import statistics import statistics
import torch import torch
@@ -24,28 +26,22 @@ import torch.nn.functional as F
import tqdm import tqdm
from astrai.model import AutoModel from astrai.model import AutoModel
from astrai.preprocessing.packing import plan_bfd
from astrai.tokenize import AutoTokenizer from astrai.tokenize import AutoTokenizer
def _pack_bins(pairs, max_len): def _pack_bins(pairs, max_len):
"""BFD bin packing: pack (c+r) into bins of max total length.""" """BFD bin packing: pack (c+r) into bins of max total length.
indexed = sorted(enumerate(pairs), key=lambda x: -(len(x[1][0]) + len(x[1][1])))
bins = [] Reuses :func:`plan_bfd` so the BFD heuristic stays single-sourced.
lengths = [] """
for orig_idx, (c, r) in indexed: # Treat each pair as a single sequence of length len(c)+len(r) for
size = len(c) + len(r) # planning purposes; plan_bfd works on pure lengths.
best_bin = -1 fake_sequences = [[0] * (len(c) + len(r)) for c, r in pairs]
for bi, rem in enumerate(lengths): plan = plan_bfd(fake_sequences, max_len)
if rem >= size: return [
if best_bin < 0 or rem < lengths[best_bin]: [(i, pairs[i][0], pairs[i][1]) for i in bin_indices] for bin_indices in plan
best_bin = bi ]
if best_bin >= 0:
bins[best_bin].append((orig_idx, c, r))
lengths[best_bin] -= size
else:
bins.append([(orig_idx, c, r)])
lengths.append(max_len - size)
return bins
def _resolve_sentinel_ids(tokenizer, sentinel_text): def _resolve_sentinel_ids(tokenizer, sentinel_text):
@@ -65,6 +61,29 @@ def _resolve_sentinel_ids(tokenizer, sentinel_text):
return [0] return [0]
def _collect_input_files(input_path: str) -> list:
"""Resolve *input_path* to a list of JSONL/JSON files."""
if os.path.isdir(input_path):
files = []
for ext in ("*.jsonl", "*.json"):
files.extend(
sorted(glob.glob(os.path.join(input_path, "**", ext), recursive=True))
)
return files
return sorted(glob.glob(input_path))
def _load_items(filepath: str) -> list:
"""Load JSONL or JSON (array / single dict) into a list of dicts."""
with open(filepath, "r", encoding="utf-8") as f:
if filepath.lower().endswith(".json"):
data = json.load(f)
if isinstance(data, dict):
return [data]
return data
return [json.loads(line) for line in f if line.strip()]
@torch.inference_mode() @torch.inference_mode()
def _score_batch( def _score_batch(
pairs, model, device, max_len=2048, sentinel_ids=None, per_token=False pairs, model, device, max_len=2048, sentinel_ids=None, per_token=False
@@ -204,69 +223,9 @@ def _trim(context_ids, resp_ids, max_len):
return context_ids[overflow:], resp_ids return context_ids[overflow:], resp_ids
def score_plain( def process_file(
model, model,
tokenizer, tokenizer,
instruction,
response,
device,
max_len=2048,
sentinel_ids=None,
per_token=False,
):
"""Compute IFD for a single instruction-response pair (plain format)."""
ctx_ids = tokenizer.encode(instruction, add_special_tokens=False)
resp_ids = tokenizer.encode(response, add_special_tokens=False)
ctx_ids, resp_ids = _trim(ctx_ids, resp_ids, max_len)
if not ctx_ids or not resp_ids:
return {
"L_cond": None,
"L_uncond": None,
"ifd": None,
"skip_reason": "empty ctx or resp",
}
return _score_batch(
[(ctx_ids, resp_ids)],
model,
device,
max_len,
sentinel_ids=sentinel_ids,
per_token=per_token,
)[0]
def score_messages(
model, tokenizer, messages, device, max_len=2048, sentinel_ids=None, per_token=False
):
"""Compute IFD for each assistant turn in a messages array."""
turns = []
for i, msg in enumerate(messages):
if msg.get("role") != "assistant":
continue
ctx_text = "\n\n".join(m["content"] for m in messages[:i])
ctx_ids = tokenizer.encode(ctx_text)
resp_ids = tokenizer.encode(msg["content"], add_special_tokens=False)
ctx_ids, resp_ids = _trim(ctx_ids, resp_ids, max_len)
if ctx_ids and resp_ids:
turns.append((ctx_ids, resp_ids))
if not turns:
return None
raw_scores = _score_batch(
turns, model, device, max_len, sentinel_ids=sentinel_ids, per_token=per_token
)
valid = [s for s in raw_scores if s is not None and s.get("ifd") is not None]
if not valid:
return {"ifd": None, "ifd_turns": raw_scores}
avg = sum(s["ifd"] for s in valid) / len(valid)
return {
"ifd": avg,
"ifd_detail": valid[0] if len(valid) == 1 else None,
"ifd_turns": raw_scores,
}
def process_file(
param_path,
input_file, input_file,
output_file, output_file,
instr_key, instr_key,
@@ -275,28 +234,31 @@ def process_file(
data_format="plain", data_format="plain",
batch_size=1, batch_size=1,
device=None, device=None,
sentinel_text="\n", sentinel_ids=None,
per_token=False, per_token=False,
max_samples=None,
): ):
"""Score a single file, write per-sample JSONL, return summary stats."""
if device is None: if device is None:
device = "cuda" if torch.cuda.is_available() else "cpu" device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if "cuda" in device else torch.float32
model = AutoModel.from_pretrained(param_path) if sentinel_ids is None:
tokenizer = AutoTokenizer.from_pretrained(param_path) sentinel_ids = _resolve_sentinel_ids(tokenizer, "\n")
model.to(device=device, dtype=dtype)
model.eval()
sentinel_ids = _resolve_sentinel_ids(tokenizer, sentinel_text) data = _load_items(input_file)
with open(input_file, encoding="utf-8") as f: if max_samples and len(data) > max_samples:
data = [json.loads(line) for line in f if line.strip()] import random
data = random.sample(data, max_samples)
results = [] results = []
all_ifds = [] all_ifds = []
buffer = [] buffer = []
for item in tqdm.tqdm(data, desc="Computing IFD", unit="sample"): label = os.path.splitext(os.path.basename(input_file))[0]
for item in tqdm.tqdm(data, desc=f" {label}", unit="sample", leave=False):
if data_format == "messages": if data_format == "messages":
turns = [] turns = []
for i, msg in enumerate(item.get("messages", [])): for i, msg in enumerate(item.get("messages", [])):
@@ -356,8 +318,22 @@ def process_file(
f.write(json.dumps(item, ensure_ascii=False) + "\n") f.write(json.dumps(item, ensure_ascii=False) + "\n")
valid_ifd = [v for v in all_ifds if v is not None] valid_ifd = [v for v in all_ifds if v is not None]
stats = {
"samples": len(data),
"valid_ifd": len(valid_ifd),
"skipped": len(data) - len(valid_ifd),
}
if valid_ifd: if valid_ifd:
stats["mean_ifd"] = statistics.mean(valid_ifd)
stats["median_ifd"] = statistics.median(valid_ifd)
if len(valid_ifd) > 1:
stats["stdev_ifd"] = statistics.stdev(valid_ifd)
stats["min_ifd"] = min(valid_ifd)
stats["max_ifd"] = max(valid_ifd)
print(f"\n{'=' * 50}") print(f"\n{'=' * 50}")
print(f" [{label}]")
print(f"{'=' * 50}")
print(f" Samples: {len(data)}") print(f" Samples: {len(data)}")
print(f" Valid IFD: {len(valid_ifd)}") print(f" Valid IFD: {len(valid_ifd)}")
print(f" Skipped: {len(data) - len(valid_ifd)}") print(f" Skipped: {len(data) - len(valid_ifd)}")
@@ -368,7 +344,8 @@ def process_file(
print(f" Min IFD: {min(valid_ifd):.4f}") print(f" Min IFD: {min(valid_ifd):.4f}")
print(f" Max IFD: {max(valid_ifd):.4f}") print(f" Max IFD: {max(valid_ifd):.4f}")
print(f"{'=' * 50}") print(f"{'=' * 50}")
print(f"Results saved to {output_file}") print(f" Results saved to {output_file}")
return stats
def _flush_buffer( def _flush_buffer(
@@ -422,8 +399,18 @@ def main():
description="Compute IFD scores for instruction-response data" description="Compute IFD scores for instruction-response data"
) )
parser.add_argument("--param_path", type=str, required=True, help="Model directory") parser.add_argument("--param_path", type=str, required=True, help="Model directory")
parser.add_argument("--input", type=str, required=True, help="Input JSONL file") parser.add_argument(
parser.add_argument("--output", type=str, required=True, help="Output JSONL file") "--input_path",
type=str,
required=True,
help="Input file, glob pattern, or directory.",
)
parser.add_argument(
"--output_dir",
type=str,
required=True,
help="Directory for output files (summary.json + per-file JSONL).",
)
parser.add_argument("--max_len", type=int, default=2048, help="Max token length") parser.add_argument("--max_len", type=int, default=2048, help="Max token length")
parser.add_argument( parser.add_argument(
"--format", "--format",
@@ -442,6 +429,12 @@ def main():
"--batch_size", type=int, default=8, help="Batch size for model forward passes" "--batch_size", type=int, default=8, help="Batch size for model forward passes"
) )
parser.add_argument("--device", type=str, default=None, help="Device (e.g. cuda:0)") parser.add_argument("--device", type=str, default=None, help="Device (e.g. cuda:0)")
parser.add_argument(
"--dtype",
type=str,
default="bfloat16" if torch.cuda.is_available() else "float32",
help="Torch dtype",
)
parser.add_argument( parser.add_argument(
"--sentinel_text", "--sentinel_text",
type=str, type=str,
@@ -453,21 +446,60 @@ def main():
action="store_true", action="store_true",
help="Include per-token IFD breakdown in output", help="Include per-token IFD breakdown in output",
) )
parser.add_argument(
"--max_samples",
type=int,
default=None,
help="Maximum number of samples per file (random subsample). Default: all.",
)
args = parser.parse_args() args = parser.parse_args()
process_file( if args.device is None:
args.param_path, args.device = "cuda" if torch.cuda.is_available() else "cpu"
args.input, dtype = getattr(torch, args.dtype)
args.output,
args.instr_key, print(f"Loading model from {args.param_path} ...")
args.resp_key, model = AutoModel.from_pretrained(args.param_path)
args.max_len, tokenizer = AutoTokenizer.from_pretrained(args.param_path)
data_format=args.format, model.to(device=args.device, dtype=dtype)
batch_size=args.batch_size, model.eval()
device=args.device,
sentinel_text=args.sentinel_text, sentinel_ids = _resolve_sentinel_ids(tokenizer, args.sentinel_text)
per_token=args.per_token,
) input_files = _collect_input_files(args.input_path)
if not input_files:
print(f"No input files found at {args.input_path}")
return
print(f"Found {len(input_files)} file(s) to evaluate")
os.makedirs(args.output_dir, exist_ok=True)
all_stats = {}
for filepath in input_files:
label = os.path.splitext(os.path.basename(filepath))[0]
output_file = os.path.join(args.output_dir, f"{label}_ifd.jsonl")
stats = process_file(
model=model,
tokenizer=tokenizer,
input_file=filepath,
output_file=output_file,
instr_key=args.instr_key,
resp_key=args.resp_key,
max_len=args.max_len,
data_format=args.format,
batch_size=args.batch_size,
device=args.device,
sentinel_ids=sentinel_ids,
per_token=args.per_token,
max_samples=args.max_samples,
)
all_stats[label] = stats
summary_path = os.path.join(args.output_dir, "summary.json")
with open(summary_path, "w", encoding="utf-8") as f:
json.dump(all_stats, f, ensure_ascii=False, indent=2)
print(f"\nSummary saved to {summary_path}")
if __name__ == "__main__": if __name__ == "__main__":
+9 -16
View File
@@ -5,7 +5,7 @@ Supports all IFEval constraint types except language detection.
Usage:: Usage::
python scripts/tools/evaluate_ifeval.py --param_path ./params \ python scripts/eval/evaluate_ifeval.py --param_path ./params \
--data_path ifeval.jsonl --output results.json \ --data_path ifeval.jsonl --output results.json \
--temperature 0.1 --max_tokens 512 --temperature 0.1 --max_tokens 512
""" """
@@ -14,21 +14,17 @@ import argparse
import json import json
import os import os
import re import re
import urllib.request
from typing import Callable, Dict, List, Optional from typing import Callable, Dict, List, Optional
import torch import torch
import tqdm import tqdm
from datasets import load_dataset
from astrai.inference import InferenceEngine from astrai.inference import InferenceEngine
from astrai.model import AutoModel from astrai.model import AutoModel
from astrai.tokenize import AutoTokenizer from astrai.tokenize import AutoTokenizer
IFEVAL_URL = ( IFEVAL_HF_DATASET = "google/IFEval"
"https://raw.githubusercontent.com/google-research/"
"google-research/master/instruction_following_eval/data/input_data.jsonl"
)
CONSTRAINT_VERIFIERS: Dict[str, Callable[[str, dict], bool]] = {} CONSTRAINT_VERIFIERS: Dict[str, Callable[[str, dict], bool]] = {}
@@ -310,15 +306,12 @@ def download_ifeval(data_path: str):
if os.path.exists(data_path): if os.path.exists(data_path):
return return
os.makedirs(os.path.dirname(data_path) or ".", exist_ok=True) os.makedirs(os.path.dirname(data_path) or ".", exist_ok=True)
print(f"Downloading IFEval from {IFEVAL_URL} ...") print(f"Downloading IFEval from HuggingFace ({IFEVAL_HF_DATASET}) ...")
tmp = data_path + ".tmp" ds = load_dataset(IFEVAL_HF_DATASET, split="train")
urllib.request.urlretrieve(IFEVAL_URL, tmp) with open(data_path, "w", encoding="utf-8") as f:
with open(tmp, "rb") as f_in: for item in ds:
content = f_in.read() f.write(json.dumps(item, ensure_ascii=False) + "\n")
with open(data_path, "wb") as f_out: print(f" saved {len(ds)} items to {data_path}")
f_out.write(content)
os.remove(tmp)
print(f" saved to {data_path}")
def load_problems(data_path: str) -> List[dict]: def load_problems(data_path: str) -> List[dict]:
+92 -56
View File
@@ -4,18 +4,18 @@ import argparse
import csv import csv
import json import json
import os import os
import shutil import random
import tarfile from collections import defaultdict
import requests
import torch import torch
import torch.nn.functional as F import torch.nn.functional as F
import tqdm import tqdm
from datasets import load_dataset
from astrai.model import AutoModel from astrai.model import AutoModel
from astrai.tokenize import AutoTokenizer from astrai.tokenize import AutoTokenizer
MMLU_URL = "https://people.eecs.berkeley.edu/~hendrycks/data.tar" MMLU_HF_DATASET = "cais/mmlu"
MMLU_SUBJECTS = [ MMLU_SUBJECTS = [
"abstract_algebra", "abstract_algebra",
"anatomy", "anatomy",
@@ -77,38 +77,40 @@ MMLU_SUBJECTS = [
] ]
def _download_and_extract(url: str, data_dir: str): def _write_subject_csv(data_dir: str, split: str, subject: str, rows: list[dict]):
tar_path = os.path.join(data_dir, "data.tar") split_dir = os.path.join(data_dir, split)
os.makedirs(data_dir, exist_ok=True) os.makedirs(split_dir, exist_ok=True)
print(f"Downloading MMLU data from {url}...") path = os.path.join(split_dir, f"{subject}_{split}.csv")
resp = requests.get(url, stream=True, timeout=300) with open(path, "w", encoding="utf-8", newline="") as f:
resp.raise_for_status() writer = csv.writer(f)
total = int(resp.headers.get("content-length", 0)) for row in rows:
with tqdm.tqdm(total=total, unit="B", unit_scale=True, desc=" Download") as bar: writer.writerow(row)
with open(tar_path, "wb") as f:
for chunk in resp.iter_content(chunk_size=8192):
f.write(chunk)
bar.update(len(chunk))
print("Extracting...")
with tarfile.open(tar_path, "r") as tf:
tf.extractall(data_dir)
os.remove(tar_path)
def download_mmlu(data_dir: str): def download_mmlu(data_dir: str):
_download_and_extract(MMLU_URL, data_dir) print(f"Downloading MMLU from HuggingFace ({MMLU_HF_DATASET}) ...")
src = os.path.join(data_dir, "data") letters = ("A", "B", "C", "D")
if os.path.exists(src): split_map = {"dev": "dev", "val": "validation", "test": "test"}
for item in os.listdir(src): for local_split, hf_split in split_map.items():
src_item = os.path.join(src, item) ds = load_dataset(MMLU_HF_DATASET, "all", split=hf_split)
dst_item = os.path.join(data_dir, item) grouped: dict[str, list[dict]] = defaultdict(list)
if os.path.exists(dst_item): for item in tqdm.tqdm(ds, desc=f" {local_split}", leave=False):
if os.path.isdir(dst_item): subject = item["subject"]
shutil.rmtree(dst_item) choices = item["choices"]
else: ans_letter = letters[item["answer"]]
os.remove(dst_item) grouped[subject].append(
os.rename(src_item, dst_item) [
os.rmdir(src) item["question"],
f"A){choices[0]}",
f"B){choices[1]}",
f"C){choices[2]}",
f"D){choices[3]}",
ans_letter,
]
)
for subject, rows in grouped.items():
_write_subject_csv(data_dir, local_split, subject, rows)
print(f" {local_split}: {len(ds)} items, {len(grouped)} subjects")
print(f"MMLU data saved to {data_dir}") print(f"MMLU data saved to {data_dir}")
@@ -139,17 +141,12 @@ def load_csv(path: str) -> list[dict]:
return data return data
def build_prompt( def build_prompt(question: str, choices: dict, subject: str) -> str:
question: str, choices: dict, subject: str, n_shot: int, dev_data: list[dict] """Build the raw question prompt (without few-shot examples).
) -> str:
prompt = "" Few-shot examples are handled by ``apply_chat`` to avoid duplication.
if n_shot > 0 and dev_data: """
prompt = f"The following are multiple choice questions (with answers) about {subject}.\n\n" prompt = f"The following are multiple choice questions (with answers) about {subject}.\n\n"
for item in dev_data[:n_shot]:
prompt += f"Question: {item['question']}\n"
for k in ("A", "B", "C", "D"):
prompt += f"{k}. {item[k]}\n"
prompt += f"Answer: {item['answer']}\n\n"
prompt += f"Question: {question}\n" prompt += f"Question: {question}\n"
for k in ("A", "B", "C", "D"): for k in ("A", "B", "C", "D"):
prompt += f"{k}. {choices[k]}\n" prompt += f"{k}. {choices[k]}\n"
@@ -158,19 +155,22 @@ def build_prompt(
def apply_chat( def apply_chat(
tokenizer, raw_prompt: str, n_shot: int, dev_data: list[dict] | None tokenizer,
raw_prompt: str,
n_shot: int,
dev_data: list[dict] | None,
subject: str = "",
) -> str: ) -> str:
"""Wrap raw MMLU prompt in the model's chat template format. """Wrap raw MMLU prompt in the model's chat template format.
For few-shot, prepend example Q&A pairs as a second user/assistant exchange. For few-shot, prepend example Q&A pairs as user/assistant exchanges.
Few-shot examples use the same subject preamble as the test question to
keep the format consistent.
""" """
messages = [] messages = []
if n_shot > 0 and dev_data: if n_shot > 0 and dev_data:
for item in dev_data[:n_shot]: for item in dev_data[:n_shot]:
q = f"Question: {item['question']}\n" q = build_prompt(item["question"], item, subject)
for k in ("A", "B", "C", "D"):
q += f"{k}. {item[k]}\n"
q += "Answer:"
messages.append({"role": "user", "content": q}) messages.append({"role": "user", "content": q})
messages.append({"role": "assistant", "content": item["answer"]}) messages.append({"role": "assistant", "content": item["answer"]})
messages.append({"role": "user", "content": raw_prompt}) messages.append({"role": "user", "content": raw_prompt})
@@ -206,6 +206,25 @@ def choice_logprob(
return score return score
def _permute_choices(item: dict, rng: random.Random) -> tuple[dict, str]:
"""Shuffle the option order of a question.
Returns ``(permuted_item, new_answer_letter)``. The question text and
the *content* of each choice are unchanged; only which letter (A/B/C/D)
maps to which content is shuffled. This neutralises the model's
positional bias (e.g. always picking B).
"""
letters = ("A", "B", "C", "D")
contents = [item[k] for k in letters]
perm = list(letters)
rng.shuffle(perm)
permuted = {"question": item["question"]}
for new_letter, orig_letter in zip(letters, perm):
permuted[new_letter] = item[orig_letter]
new_answer = letters[perm.index(item["answer"])]
return permuted, new_answer
def evaluate_subject( def evaluate_subject(
model, model,
tokenizer, tokenizer,
@@ -214,20 +233,24 @@ def evaluate_subject(
dev_data: list[dict] | None, dev_data: list[dict] | None,
device: str, device: str,
n_shot: int, n_shot: int,
seed: int = 0,
) -> tuple[float, int, int]: ) -> tuple[float, int, int]:
rng = random.Random(seed) if seed >= 0 else None
correct = 0 correct = 0
total = 0 total = 0
for item in tqdm.tqdm(test_data, desc=f"{subject:40s}", leave=False): for item in tqdm.tqdm(test_data, desc=f"{subject:40s}", leave=False):
raw_prompt = build_prompt( if rng is not None:
item["question"], item, subject, n_shot, dev_data or [] permuted, answer = _permute_choices(item, rng)
) else:
context = apply_chat(tokenizer, raw_prompt, n_shot, dev_data or []) permuted, answer = item, item["answer"]
raw_prompt = build_prompt(permuted["question"], permuted, subject)
context = apply_chat(tokenizer, raw_prompt, n_shot, dev_data or [], subject)
context_ids = tokenizer.encode(context) context_ids = tokenizer.encode(context)
scores = { scores = {
c: choice_logprob(model, tokenizer, context_ids, c, device) c: choice_logprob(model, tokenizer, context_ids, c, device)
for c in ("A", "B", "C", "D") for c in ("A", "B", "C", "D")
} }
if max(scores, key=scores.get) == item["answer"]: if max(scores, key=scores.get) == answer:
correct += 1 correct += 1
total += 1 total += 1
return correct / total, correct, total return correct / total, correct, total
@@ -262,6 +285,12 @@ def main():
default="bfloat16" if torch.cuda.is_available() else "float32", default="bfloat16" if torch.cuda.is_available() else "float32",
help="Torch dtype", help="Torch dtype",
) )
parser.add_argument(
"--seed",
type=int,
default=0,
help="Seed for option permutation (0 to enable, -1 to disable)",
)
args = parser.parse_args() args = parser.parse_args()
if args.download or not os.path.exists(args.data_dir): if args.download or not os.path.exists(args.data_dir):
@@ -293,7 +322,14 @@ def main():
test_data = load_csv(test_path) test_data = load_csv(test_path)
acc, corr, tot = evaluate_subject( acc, corr, tot = evaluate_subject(
model, tokenizer, subject, test_data, dev_data, device, args.n_shot model,
tokenizer,
subject,
test_data,
dev_data,
device,
args.n_shot,
seed=args.seed,
) )
results[subject] = {"accuracy": round(acc, 4), "correct": corr, "total": tot} results[subject] = {"accuracy": round(acc, 4), "correct": corr, "total": tot}
total_correct += corr total_correct += corr
+422 -69
View File
@@ -1,5 +1,9 @@
import argparse import argparse
import glob
import json import json
import os
import statistics
from typing import Dict, List, Optional, Tuple
import torch import torch
import torch.nn.functional as F import torch.nn.functional as F
@@ -9,95 +13,400 @@ from astrai.model import AutoModel
from astrai.tokenize import AutoTokenizer from astrai.tokenize import AutoTokenizer
def _collect_input_files(input_path: str) -> List[str]:
"""Resolve *input_path* to a list of JSONL/JSON files."""
if os.path.isdir(input_path):
files = []
for ext in ("*.jsonl", "*.json"):
files.extend(
sorted(glob.glob(os.path.join(input_path, "**", ext), recursive=True))
)
return files
return sorted(glob.glob(input_path))
def _load_items(filepath: str) -> List[dict]:
"""Load JSONL or JSON (array / single dict) into a list of dicts."""
with open(filepath, "r", encoding="utf-8") as f:
if filepath.lower().endswith(".json"):
data = json.load(f)
if isinstance(data, dict):
return [data]
return data
return [json.loads(line) for line in f if line.strip()]
def _encode_batch(
tokenizer: AutoTokenizer, texts: List[str], max_length: int
) -> Tuple[List[List[int]], List[List[int]]]:
"""Encode *texts* and return (token_ids, attention_masks).
Each sequence is left-aligned and padded to the batch max length.
"""
encoded = [tokenizer.encode(t)[:max_length] for t in texts]
if not encoded:
return [], []
max_len = max(len(seq) for seq in encoded)
padded_ids = []
masks = []
for seq in encoded:
pad_len = max_len - len(seq)
padded_ids.append(seq + [tokenizer.pad_id] * pad_len)
masks.append([1] * len(seq) + [0] * pad_len)
return padded_ids, masks
def _compute_batch(
model,
input_ids: torch.Tensor,
attention_mask: torch.Tensor,
) -> Tuple[torch.Tensor, torch.Tensor]:
"""Forward pass and return (log_probs, valid_mask) of shape [B, S-1].
log_probs[i, j] = log P(token j+1 | tokens 0..j)
"""
output = model(input_ids, input_mask=attention_mask)
logits = output["logits"][:, :-1, :] # [B, S-1, V]
targets = input_ids[:, 1:] # [B, S-1]
valid = attention_mask[:, 1:].float() # [B, S-1]
log_probs = F.log_softmax(logits.float(), dim=-1) # [B, S-1, V]
token_log_probs = log_probs.gather(2, targets.unsqueeze(-1)).squeeze(-1) # [B, S-1]
return token_log_probs, valid
def _token_type(token_id: int, stop_ids: frozenset, decode_fn) -> str:
"""Classify a token into a coarse type for analysis.
*stop_ids* is a pre-built set of special token IDs.
*decode_fn* is ``tokenizer.decode`` (or a wrapper) for single-token
decoding.
"""
if token_id in stop_ids:
return "special"
decoded = decode_fn([token_id], skip_special_tokens=True)
if any("\u4e00" <= ch <= "\u9fff" for ch in decoded):
return "cjk"
if any(ord(ch) > 127 for ch in decoded):
return "non_ascii"
return "ascii"
def _percentiles(values: List[float]) -> Dict[str, float]:
"""Compute common percentiles from a list of floats.
Uses linear interpolation between closest ranks (same convention
as NumPy's default).
"""
if not values:
return {}
sorted_vals = sorted(values)
n = len(sorted_vals)
def _pct(p: float) -> float:
if n == 1:
return sorted_vals[0]
k = p * (n - 1)
f = int(k)
c = min(f + 1, n - 1)
return sorted_vals[f] + (sorted_vals[c] - sorted_vals[f]) * (k - f)
return {
"p50": _pct(0.50),
"p90": _pct(0.90),
"p95": _pct(0.95),
"p99": _pct(0.99),
}
class LossAccumulator:
"""Accumulate per-token losses with optional streaming mode.
When *stream* is True (token_level=False), losses are not kept
in memory individually only a running sum/count and a histogram
(for approximate percentiles) are maintained. When *stream* is
False, all losses are retained for exact statistics and per-record
output.
"""
_HIST_BINS = 1000
_HIST_MAX = 20.0 # clamp losses above this for histogram
def __init__(self, stream: bool):
self.stream = stream
self.losses: List[float] = [] if not stream else []
self.total: float = 0.0
self.count: int = 0
self.hist = torch.zeros(self._HIST_BINS, dtype=torch.long)
# per-type losses (only populated when not streaming)
self.by_type: Dict[str, List[float]] = {}
self.type_total: Dict[str, float] = {}
self.type_count: Dict[str, int] = {}
def add(self, losses: List[float]):
self.total += sum(losses)
self.count += len(losses)
if self.stream:
clamped = [min(max(l, 0.0), self._HIST_MAX) for l in losses]
idx = torch.tensor(clamped) / self._HIST_MAX * (self._HIST_BINS - 1)
self.hist += torch.bincount(
idx.long().clamp(0, self._HIST_BINS - 1),
minlength=self._HIST_BINS,
)
else:
self.losses.extend(losses)
def add_typed(self, ttype: str, losses: List[float]):
if not self.stream:
self.by_type.setdefault(ttype, []).extend(losses)
self.type_total[ttype] = self.type_total.get(ttype, 0.0) + sum(losses)
self.type_count[ttype] = self.type_count.get(ttype, 0) + len(losses)
def stats(self) -> Dict:
result: Dict = {}
if self.count == 0:
return result
mean_loss = self.total / self.count
result["overall"] = {
"num_tokens": self.count,
"mean_loss": mean_loss,
"ppl": float(torch.exp(torch.tensor(mean_loss))),
}
if self.stream:
result["overall"].update(self._hist_percentiles())
else:
result["overall"]["median_loss"] = statistics.median(self.losses)
result["overall"].update(_percentiles(self.losses))
if self.type_count:
result["by_token_type"] = {}
for ttype in sorted(self.type_count.keys()):
cnt = self.type_count[ttype]
tmean = self.type_total[ttype] / cnt
entry: Dict = {
"num_tokens": cnt,
"mean_loss": tmean,
"ppl": float(torch.exp(torch.tensor(tmean))),
}
if not self.stream and ttype in self.by_type:
entry["median_loss"] = statistics.median(self.by_type[ttype])
entry.update(_percentiles(self.by_type[ttype]))
result["by_token_type"][ttype] = entry
return result
def _hist_percentiles(self) -> Dict[str, float]:
"""Approximate percentiles from the histogram."""
total = self.hist.sum().item()
if total == 0:
return {}
cum = torch.cumsum(self.hist.float(), dim=0)
result = {}
for label, p in [("p50", 0.5), ("p90", 0.9), ("p95", 0.95), ("p99", 0.99)]:
target = p * total
idx = int(torch.searchsorted(cum, target).item())
idx = min(idx, self._HIST_BINS - 1)
result[label] = (idx + 0.5) / self._HIST_BINS * self._HIST_MAX
return result
def process_file( def process_file(
param_path: str, input_file: str, output_file: str, batch_size: int, text_key: str model,
tokenizer: AutoTokenizer,
items: List[dict],
text_key: str,
batch_size: int,
max_length: int,
token_level: bool,
max_samples: Optional[int],
output_file: Optional[str],
label: str,
device: str = "cuda",
) -> Dict:
"""Evaluate a single dataset (list of items), return summary stats.
If *token_level* is True and *output_file* is set, per-record token_ids
and log_probs are written as JSONL alongside the summary.
"""
if max_samples and len(items) > max_samples:
import random
items = random.sample(items, max_samples)
texts = [item[text_key] for item in items if text_key in item]
print(f" [{label}] {len(texts)} samples, text_key='{text_key}'")
acc = LossAccumulator(stream=not token_level)
per_sample: List[dict] = []
if token_level:
stop_ids = frozenset(tokenizer.stop_ids)
decode_fn = tokenizer.decode
num_batches = (len(texts) + batch_size - 1) // batch_size
for i in tqdm.tqdm(
range(0, len(texts), batch_size),
total=num_batches,
desc=f" {label}",
leave=False,
):
batch_texts = texts[i : i + batch_size]
padded_ids, masks = _encode_batch(tokenizer, batch_texts, max_length)
input_ids = torch.tensor(padded_ids, device=device, dtype=torch.long)
attention_mask = torch.tensor(masks, device=device, dtype=torch.bool)
token_log_probs, valid = _compute_batch(model, input_ids, attention_mask)
for b in range(len(batch_texts)):
seq_len = int(valid[b].sum().item())
lps = token_log_probs[b, :seq_len].tolist()
losses = [-lp for lp in lps]
acc.add(losses)
if token_level:
# log_probs correspond to positions 1..seq_len (predicted
# from position 0..seq_len-1), so token_ids must skip BOS
# at position 0 to stay aligned with log_probs.
ids = padded_ids[b][1 : seq_len + 1]
per_sample.append(
{
"text": batch_texts[b][:200],
"token_ids": ids,
"log_probs": [round(lp, 4) for lp in lps],
"ppl": float(torch.exp(torch.tensor(statistics.mean(losses))))
if losses
else None,
}
)
typed_losses: Dict[str, List[float]] = {}
for tid, loss in zip(ids, losses):
ttype = _token_type(tid, stop_ids, decode_fn)
typed_losses.setdefault(ttype, []).append(loss)
for ttype, tl in typed_losses.items():
acc.add_typed(ttype, tl)
stats = acc.stats()
if token_level and output_file:
with open(output_file, "w", encoding="utf-8") as f:
for item in per_sample:
f.write(json.dumps(item, ensure_ascii=False) + "\n")
return stats
def print_stats(label: str, stats: Dict):
"""Pretty-print summary statistics."""
print(f"\n{'=' * 60}")
print(f" {label}")
print(f"{'=' * 60}")
ov = stats.get("overall", {})
if ov:
print(f" tokens: {ov['num_tokens']:,}")
print(f" mean loss: {ov['mean_loss']:.4f}")
if "median_loss" in ov:
print(f" median loss: {ov['median_loss']:.4f}")
print(f" ppl: {ov['ppl']:.2f}")
if "p50" in ov:
print(
f" p50/p90/p95/p99: "
f"{ov['p50']:.2f} / {ov['p90']:.2f} / {ov['p95']:.2f} / {ov['p99']:.2f}"
)
by_type = stats.get("by_token_type", {})
if by_type:
print(f"\n by token type:")
print(f" {'type':<12} {'count':>8} {'mean_loss':>10} {'ppl':>8}")
print(f" {'-' * 12} {'-' * 8} {'-' * 10} {'-' * 8}")
for ttype, s in by_type.items():
print(
f" {ttype:<12} {s['num_tokens']:>8,} "
f"{s['mean_loss']:>10.4f} {s['ppl']:>8.2f}"
)
def main(
param_path: str,
input_path: str,
output_dir: str,
text_key: str,
batch_size: int,
max_length: int,
token_level: bool,
max_samples: Optional[int],
device: str = "cuda",
dtype: str = "bfloat16",
): ):
# Load model and tokenizer print(f"Loading model from {param_path} ...")
model = AutoModel.from_pretrained(param_path) model = AutoModel.from_pretrained(param_path)
tokenizer = AutoTokenizer.from_pretrained(param_path) tokenizer = AutoTokenizer.from_pretrained(param_path)
model.to(device="cuda", dtype=torch.bfloat16) torch_dtype = getattr(torch, dtype)
model.to(device=device, dtype=torch_dtype)
model.eval()
with open(input_file, "r", encoding="utf-8") as f: input_files = _collect_input_files(input_path)
input_data = [json.loads(line) for line in f] if not input_files:
print(f"No input files found at {input_path}")
return
texts = [item[text_key] for item in input_data] print(f"Found {len(input_files)} file(s) to evaluate")
os.makedirs(output_dir, exist_ok=True)
# Encode all texts all_stats = {}
print(f"Encoding {len(texts)} texts...") for filepath in input_files:
encoded_texts = [tokenizer.encode(text) for text in texts] label = os.path.splitext(os.path.basename(filepath))[0]
items = _load_items(filepath)
if not items:
print(f" [{label}] empty, skipping")
continue
output_data = [] token_output = (
total_batches = (len(encoded_texts) + batch_size - 1) // batch_size os.path.join(output_dir, f"{label}_tokens.jsonl") if token_level else None
for i in tqdm.tqdm(
range(0, len(encoded_texts), batch_size),
total=total_batches,
desc="Computing perplexity",
):
batch_encoded = encoded_texts[i : i + batch_size]
batch_texts = texts[i : i + batch_size]
# Find max length in batch and pad
max_len = max(len(seq) for seq in batch_encoded)
padded_ids = []
masks = []
for seq in batch_encoded:
pad_len = max_len - len(seq)
padded_seq = seq + [tokenizer.pad_id] * pad_len
mask = [True] * len(seq) + [False] * pad_len
padded_ids.append(padded_seq)
masks.append(mask)
# Convert to tensors
input_ids = torch.tensor(padded_ids, device="cuda", dtype=torch.long)
input_mask = torch.tensor(masks, device="cuda", dtype=torch.bool)
# Compute perplexity
output = model(input_ids, input_mask=input_mask)
logits = output["logits"]
# Shift for causal language modeling
shifted_logits = logits[:, :-1, :] # [batch_size, seq_len-1, vocab_size]
shifted_input_ids = input_ids[:, 1:] # [batch_size, seq_len-1]
shifted_mask = input_mask[:, 1:] # [batch_size, seq_len-1]
# Compute cross entropy loss
loss = F.cross_entropy(
shifted_logits.flatten(0, 1),
shifted_input_ids.flatten(0, 1),
reduction="none",
) )
loss = loss.view(shifted_input_ids.shape) # [batch_size, seq_len-1] stats = process_file(
loss = loss * shifted_mask model=model,
sentence_loss = loss.sum(dim=1) / shifted_mask.sum(dim=1).clamp(min=1) tokenizer=tokenizer,
perplexity = torch.exp(sentence_loss) # [batch_size] items=items,
text_key=text_key,
batch_size=batch_size,
max_length=max_length,
token_level=token_level,
max_samples=max_samples,
output_file=token_output,
label=label,
device=device,
)
all_stats[label] = stats
print_stats(label, stats)
for text, ppl in zip(batch_texts, perplexity): if token_output:
output_data.append({text_key: text, "ppl": float(ppl.item())}) print(f" token-level output: {token_output}")
# Write results summary_path = os.path.join(output_dir, "summary.json")
with open(output_file, "w", encoding="utf-8") as f: with open(summary_path, "w", encoding="utf-8") as f:
for item in output_data: json.dump(all_stats, f, ensure_ascii=False, indent=2)
f.write(json.dumps(item, ensure_ascii=False) + "\n") print(f"\nSummary saved to {summary_path}")
print(f"Perplexity computation complete. Results saved to {output_file}")
if __name__ == "__main__": if __name__ == "__main__":
parser = argparse.ArgumentParser(description="Perplexity evaluation on JSONL text.") parser = argparse.ArgumentParser(
description="Perplexity and token-level loss evaluation on JSONL/JSON data."
)
parser.add_argument( parser.add_argument(
"--param_path", type=str, required=True, help="Path to the model directory." "--param_path", type=str, required=True, help="Path to the model directory."
) )
parser.add_argument( parser.add_argument(
"--input_file", type=str, required=True, help="Path to the input file." "--input_path",
type=str,
required=True,
help="Path to input file, glob pattern, or directory.",
) )
parser.add_argument( parser.add_argument(
"--output_file", type=str, required=True, help="Path to the output file." "--output_dir",
) type=str,
parser.add_argument( required=True,
"--batch_size", type=int, default=4, help="Batch size for evaluation." help="Directory for output files (summary.json + per-file token JSONL).",
) )
parser.add_argument( parser.add_argument(
"--text_key", "--text_key",
@@ -105,7 +414,51 @@ if __name__ == "__main__":
default="text", default="text",
help="Key for the text field in the input data.", help="Key for the text field in the input data.",
) )
parser.add_argument(
"--batch_size", type=int, default=4, help="Batch size for evaluation."
)
parser.add_argument(
"--max_length",
type=int,
default=2048,
help="Maximum sequence length (tokens). Longer sequences are truncated.",
)
parser.add_argument(
"--token_level",
action="store_true",
help="Store per-token log_probs and token type analysis. "
"Default: off (only aggregate stats).",
)
parser.add_argument(
"--max_samples",
type=int,
default=None,
help="Maximum number of samples per file (random subsample). Default: all.",
)
parser.add_argument(
"--device",
type=str,
default="cuda" if torch.cuda.is_available() else "cpu",
help="Device for model inference.",
)
parser.add_argument(
"--dtype",
type=str,
default="bfloat16" if torch.cuda.is_available() else "float32",
help="Torch dtype for model weights.",
)
args = parser.parse_args() args = parser.parse_args()
with torch.inference_mode(): with torch.inference_mode():
process_file(**vars(args)) main(
param_path=args.param_path,
input_path=args.input_path,
output_dir=args.output_dir,
text_key=args.text_key,
batch_size=args.batch_size,
max_length=args.max_length,
token_level=args.token_level,
max_samples=args.max_samples,
device=args.device,
dtype=args.dtype,
)
+96 -23
View File
@@ -1,8 +1,10 @@
import argparse import argparse
import json import json
import time
from typing import Optional from typing import Optional
import torch import torch
from tqdm import tqdm
from astrai.inference import InferenceEngine from astrai.inference import InferenceEngine
from astrai.model import AutoModel from astrai.model import AutoModel
@@ -20,55 +22,102 @@ def processor(
response_key: str, response_key: str,
max_tokens: Optional[int], max_tokens: Optional[int],
batch_size: int, batch_size: int,
num_samples: int = 1,
cache_len: int = 2048,
frequency_penalty: float = 0.0,
rep_window: int = 64,
): ):
# Load model and tokenizer print(f"Loading model from {param_path} ...")
t0 = time.time()
model = AutoModel.from_pretrained(param_path) model = AutoModel.from_pretrained(param_path)
tokenizer = AutoTokenizer.from_pretrained(param_path) tokenizer = AutoTokenizer.from_pretrained(param_path)
model.to(device="cuda", dtype=torch.bfloat16) model.to(device="cuda", dtype=torch.bfloat16)
print(f" model loaded in {time.time() - t0:.1f}s")
# Create inference engine
engine = InferenceEngine( engine = InferenceEngine(
model=model, tokenizer=tokenizer, max_batch_size=batch_size model=model,
tokenizer=tokenizer,
max_batch_size=batch_size * num_samples,
max_seq_len=cache_len,
max_prompt_len=cache_len,
) )
print(f"Reading {input_json_file} ...")
with open(input_json_file, "r", encoding="utf-8") as f: with open(input_json_file, "r", encoding="utf-8") as f:
input_data = [json.loads(line) for line in f] input_data = [json.loads(line) for line in f]
# Check input format: chat messages or raw text
if input_data and "messages" in input_data[0]: if input_data and "messages" in input_data[0]:
# Chat format: [{"messages": [...]}]
prompts = [ prompts = [
tokenizer.apply_chat_template(item["messages"], tokenize=False) tokenizer.apply_chat_template(item["messages"], tokenize=False)
for item in input_data for item in input_data
] ]
else: else:
# Raw text format: [{"question": "..."}]
prompts = [item[question_key] for item in input_data] prompts = [item[question_key] for item in input_data]
print(f" {len(prompts)} prompts loaded\n")
# Use provided max_tokens or default to model config max_len
if max_tokens is None: if max_tokens is None:
max_tokens = model.config.max_len max_tokens = model.config.max_len
# Generate responses (batch) chunk_size = max(1, batch_size)
responses = engine.generate(
prompt=prompts,
stream=False,
max_tokens=max_tokens,
temperature=temperature,
top_p=top_p,
top_k=top_k,
)
# Write results
with open(output_json_file, "w", encoding="utf-8") as f: with open(output_json_file, "w", encoding="utf-8") as f:
for prompt, response in zip(prompts, responses): pbar = tqdm(
if input_data and "messages" in input_data[0]: total=len(prompts) * num_samples,
output_item = {"response": response} unit="gen",
desc=f" Generating ({num_samples}x/prompt)",
)
for chunk_start in range(0, len(prompts), chunk_size):
chunk = prompts[chunk_start : chunk_start + chunk_size]
if num_samples > 1:
chunk_expanded = [p for p in chunk for _ in range(num_samples)]
resp_chunk = engine.generate(
prompt=chunk_expanded,
stream=False,
max_tokens=max_tokens,
temperature=temperature,
top_p=top_p,
top_k=top_k,
frequency_penalty=frequency_penalty,
rep_window=rep_window,
)
resp_chunk = [
resp_chunk[i * num_samples : (i + 1) * num_samples]
for i in range(len(chunk))
]
else: else:
output_item = {question_key: prompt, response_key: response} resp_chunk = engine.generate(
f.write(json.dumps(output_item, ensure_ascii=False) + "\n") prompt=chunk,
stream=False,
max_tokens=max_tokens,
temperature=temperature,
top_p=top_p,
top_k=top_k,
frequency_penalty=frequency_penalty,
rep_window=rep_window,
)
for i, prompt in enumerate(chunk):
if input_data and "messages" in input_data[0]:
orig = input_data[chunk_start + i]
output_item = {**orig, response_key: resp_chunk[i]}
else:
output_item = {
question_key: prompt,
response_key: resp_chunk[i],
}
f.write(json.dumps(output_item, ensure_ascii=False) + "\n")
pbar.update(len(chunk) * num_samples)
pbar.close()
elapsed = time.time() - t0
print(
f"\nDone! {len(prompts)} prompts x {num_samples} samples -> {output_json_file}"
)
print(f"Total time: {elapsed:.1f}s ({elapsed / len(prompts):.2f}s/prompt)")
# Cleanup
engine.shutdown() engine.shutdown()
@@ -126,12 +175,36 @@ if __name__ == "__main__":
default=1, default=1,
help="Batch size for generating responses (default: 1).", help="Batch size for generating responses (default: 1).",
) )
parser.add_argument(
"--num_samples",
type=int,
default=1,
help="Number of responses per prompt (expands batch internally, default: 1).",
)
parser.add_argument( parser.add_argument(
"--max_tokens", "--max_tokens",
type=int, type=int,
default=None, default=None,
help="Maximum tokens to generate (default: model config max_len).", help="Maximum tokens to generate (default: model config max_len).",
) )
parser.add_argument(
"--cache_len",
type=int,
default=2048,
help="KV cache & prompt truncation length (default: 2048, lower = less memory).",
)
parser.add_argument(
"--frequency_penalty",
type=float,
default=0.0,
help="Frequency penalty to reduce repetition (default: 0.0, try 0.5-1.0).",
)
parser.add_argument(
"--rep_window",
type=int,
default=64,
help="Window size for frequency penalty (default: 64).",
)
args = parser.parse_args() args = parser.parse_args()
+21 -11
View File
@@ -8,7 +8,7 @@ import torch.optim as optim
from torch import Tensor, nn from torch import Tensor, nn
from astrai.config import AutoRegressiveLMConfig, TrainConfig from astrai.config import AutoRegressiveLMConfig, TrainConfig
from astrai.dataset import DatasetFactory from astrai.dataset import DatasetFactory, dpo_collate_fn, grpo_collate_fn
from astrai.model import AutoRegressiveLM from astrai.model import AutoRegressiveLM
from astrai.model.components.decoder_block import DecoderBlock from astrai.model.components.decoder_block import DecoderBlock
from astrai.trainer import SchedulerFactory, Trainer from astrai.trainer import SchedulerFactory, Trainer
@@ -90,6 +90,7 @@ class MuonMix(optim.Optimizer):
def load_state_dict(self, state_dict: Dict[str, Any]): def load_state_dict(self, state_dict: Dict[str, Any]):
self.muon.load_state_dict(state_dict["muon"]) self.muon.load_state_dict(state_dict["muon"])
self.adamw.load_state_dict(state_dict["adamw"]) self.adamw.load_state_dict(state_dict["adamw"])
self.param_groups = [*self.muon.param_groups, *self.adamw.param_groups]
def parse_args() -> argparse.Namespace: def parse_args() -> argparse.Namespace:
@@ -115,6 +116,13 @@ def parse_args() -> argparse.Namespace:
required=True, required=True,
help="Path to the model parameters or resume checkpoint.", help="Path to the model parameters or resume checkpoint.",
) )
parser.add_argument(
"--resume",
action="store_true",
default=False,
help="Resume training from checkpoint at --param_path "
"(restore epoch, consumed_samples, optimizer & scheduler state).",
)
parser.add_argument( parser.add_argument(
"--n_epoch", type=int, default=1, help="Number of epochs to train." "--n_epoch", type=int, default=1, help="Number of epochs to train."
@@ -140,8 +148,8 @@ def parse_args() -> argparse.Namespace:
parser.add_argument( parser.add_argument(
"--max_grad_norm", "--max_grad_norm",
type=float, type=float,
default=1.0, default=None,
help="Max gradient norm for clipping.", help="Max gradient norm for clipping. None disables clipping.",
) )
parser.add_argument( parser.add_argument(
"--weight_decay", "--weight_decay",
@@ -252,12 +260,6 @@ def parse_args() -> argparse.Namespace:
default="checkpoint/logs", default="checkpoint/logs",
help="Directory for metric logs.", help="Directory for metric logs.",
) )
parser.add_argument(
"--grpo_sync_interval",
type=int,
default=200,
help="GRPO ref model sync interval (steps).",
)
parser.add_argument( parser.add_argument(
"--start_epoch", type=int, default=0, help="Start epoch for training." "--start_epoch", type=int, default=0, help="Start epoch for training."
) )
@@ -390,6 +392,7 @@ def train(
train_type: str, train_type: str,
param_path: str, param_path: str,
data_root_path: str, data_root_path: str,
resume: bool,
n_epoch: int, n_epoch: int,
batch_per_device: int, batch_per_device: int,
start_epoch: int, start_epoch: int,
@@ -444,7 +447,6 @@ def train(
"clip_eps": kwargs.pop("grpo_clip_eps"), "clip_eps": kwargs.pop("grpo_clip_eps"),
"kl_coef": kwargs.pop("grpo_kl_coef"), "kl_coef": kwargs.pop("grpo_kl_coef"),
"group_size": kwargs.pop("group_size"), "group_size": kwargs.pop("group_size"),
"sync_interval": kwargs.pop("grpo_sync_interval"),
} }
executor_kwargs = { executor_kwargs = {
@@ -458,6 +460,7 @@ def train(
load_path=data_root_path, load_path=data_root_path,
window_size=window_size, window_size=window_size,
stride=stride, stride=stride,
tokenizer_path=param_path,
) )
optimizer_fn = partial( optimizer_fn = partial(
@@ -502,6 +505,12 @@ def train(
grad_ckpt_modules = [DecoderBlock] if gradient_checkpointing else [] grad_ckpt_modules = [DecoderBlock] if gradient_checkpointing else []
collate_fn = None
if train_type == "dpo":
collate_fn = dpo_collate_fn
elif train_type == "grpo":
collate_fn = grpo_collate_fn
train_config = TrainConfig( train_config = TrainConfig(
model_fn=model_fn, model_fn=model_fn,
strategy=train_type, strategy=train_type,
@@ -534,10 +543,11 @@ def train(
executor_kwargs=executor_kwargs, executor_kwargs=executor_kwargs,
extra_kwargs=strategy_kwargs, extra_kwargs=strategy_kwargs,
neftune_alpha=neftune_alpha, neftune_alpha=neftune_alpha,
collate_fn=collate_fn,
) )
trainer = Trainer(train_config) trainer = Trainer(train_config)
trainer.train(resume_dir=param_path) trainer.train(param_path=param_path, resume=resume)
if __name__ == "__main__": if __name__ == "__main__":
+642 -54
View File
@@ -1,14 +1,16 @@
import json import json
import os import os
import tempfile
import numpy as np import numpy as np
import pytest import pytest
import torch import torch
from astrai.config.preprocess_config import PipelineConfig from astrai.config.preprocess_config import PipelineConfig
from astrai.dataset.dataset import DatasetFactory, SEQDataset from astrai.dataset.dataset import DatasetFactory, dpo_tokenize
from astrai.dataset.storage import ( from astrai.dataset.storage import (
H5Store, H5Store,
JsonlStore,
StoreFactory, StoreFactory,
detect_format, detect_format,
) )
@@ -56,6 +58,13 @@ def _write_jsonl_dataset(test_dir, tokenizer_path, records, config_overrides=Non
return data_dir return data_dir
def _fake_fetch_record(self, idx, keys):
"""FakeStore.fetch_record matching real Store semantics."""
if isinstance(keys, str):
return self._data[keys][idx]
return {k: self._data[k][idx] for k in keys}
def _make_seq_dataset( def _make_seq_dataset(
test_dir, name="data", seq_length=200, train_type="seq", data=None, **load_kwargs test_dir, name="data", seq_length=200, train_type="seq", data=None, **load_kwargs
): ):
@@ -109,7 +118,7 @@ def test_dpo_strategy_with_random_data(base_test_env):
) )
assert dpo_dataset is not None assert dpo_dataset is not None
assert dpo_dataset.storage is not None assert dpo_dataset.store is not None
assert len(dpo_dataset) > 0 assert len(dpo_dataset) > 0
# Test that we can get DPO items without errors # Test that we can get DPO items without errors
@@ -138,7 +147,7 @@ def test_sft_dataset_with_random_data(base_test_env):
) )
assert sft_dataset is not None assert sft_dataset is not None
assert sft_dataset.storage is not None assert sft_dataset.store is not None
assert len(sft_dataset) > 0 assert len(sft_dataset) > 0
# Test that we can get SFT items without errors # Test that we can get SFT items without errors
@@ -169,39 +178,37 @@ def test_dataset_with_custom_stride(base_test_env):
assert len(dataset) > len(default_stride_dataset) assert len(dataset) > len(default_stride_dataset)
def test_dataset_count_property(base_test_env): def test_dataset_token_count_property(base_test_env):
"""dataset.token_count exposes the raw stream token length."""
test_dir = base_test_env["test_dir"] test_dir = base_test_env["test_dir"]
dataset = _make_seq_dataset(test_dir, "count_test_data") dataset = _make_seq_dataset(test_dir, "count_test_data")
assert dataset.count == 200 assert dataset.token_count == 200
assert dataset.count > len(dataset) assert dataset.token_count > len(dataset)
assert len(dataset) == (200 - 1 - 64) // 64 + 1 assert len(dataset) == (200 - 1 - 64) // 64 + 1
def test_empty_dataset_count():
"""Test count returns 0 when no data is loaded"""
dataset = SEQDataset(window_size=64, stride=32)
assert dataset.count == 0
assert dataset.keys == []
def test_dataset_too_short_for_window(base_test_env): def test_dataset_too_short_for_window(base_test_env):
test_dir = base_test_env["test_dir"] test_dir = base_test_env["test_dir"]
dataset = _make_seq_dataset(test_dir, "short", seq_length=30) dataset = _make_seq_dataset(test_dir, "short", seq_length=30)
assert len(dataset) == 0 assert len(dataset) == 0
assert dataset.count == 30 assert dataset.token_count == 30
def test_unloaded_dataset_getitem_raises(): def test_unloaded_sample_window_raises():
"""__getitem__ without load() should fail clearly""" """Store.sample_window before load raises RuntimeError."""
dataset = SEQDataset(window_size=64, stride=32) from astrai.dataset.storage import H5Store
with pytest.raises(RuntimeError, match="not loaded"):
dataset.get_index(0) store = H5Store(window_size=64, stride=64)
with pytest.raises(IndexError, match="Data too short"):
store.sample_window(0)
def test_unloaded_dataset_len(): def test_unloaded_dataset_len():
"""__len__ without load() returns 0""" """__len__ on a store with no data returns 0."""
dataset = SEQDataset(window_size=64, stride=32) from astrai.dataset.storage import H5Store
assert len(dataset) == 0
store = H5Store(window_size=64, stride=64)
assert len(store) == 0
def test_store_unloaded_len(): def test_store_unloaded_len():
@@ -214,7 +221,7 @@ def test_store_unloaded_len():
def test_store_fetch_begin_equals_end(base_test_env): def test_store_fetch_begin_equals_end(base_test_env):
test_dir = base_test_env["test_dir"] test_dir = base_test_env["test_dir"]
dataset = _make_seq_dataset(test_dir, "empty_fetch", seq_length=100, window_size=32) dataset = _make_seq_dataset(test_dir, "empty_fetch", seq_length=100, window_size=32)
result = dataset.storage.fetch(10, 10, "sequence") result = dataset.store.fetch(10, 10, "sequence")
assert result.numel() == 0 assert result.numel() == 0
@@ -264,7 +271,7 @@ def test_store_multi_segment_concat(base_test_env):
store = StoreFactory.create("h5") store = StoreFactory.create("h5")
store.load(data_dir) store.load(data_dir)
assert len(store) == 9 assert store.token_count == 9
result = store.fetch(2, 7, "sequence") result = store.fetch(2, 7, "sequence")
assert result.tolist() == [3, 4, 5, 6, 7] assert result.tolist() == [3, 4, 5, 6, 7]
@@ -293,7 +300,9 @@ def test_mmap_store_load_and_fetch(base_test_env):
store = StoreFactory.create("bin") store = StoreFactory.create("bin")
store.load(test_dir) store.load(test_dir)
assert len(store) == 200 assert store.token_count == 200
assert store.num_records == 0
assert len(store) == 0 # no window configured, no records → 0 samples
assert "sequence" in store.keys assert "sequence" in store.keys
result = store.fetch(10, 20, "sequence") result = store.fetch(10, 20, "sequence")
@@ -306,64 +315,110 @@ def test_mmap_dataset_load(base_test_env):
save_bin(test_dir, data) save_bin(test_dir, data)
dataset = DatasetFactory.load("seq", test_dir, window_size=64) dataset = DatasetFactory.load("seq", test_dir, window_size=64)
assert len(dataset) > 0 assert len(dataset) > 0
assert dataset.count == 200 assert dataset.token_count == 200
assert dataset[0]["input_ids"].shape[0] == 64 assert dataset[0]["input_ids"].shape[0] == 64
def test_normalize_empty_key(): def test_normalize_empty_key():
"""_normalize with empty tensor list does not crash""" """_normalize with empty tensor list does not crash."""
store = H5Store() store = H5Store()
store._normalize({"sequence": []}) store._normalize({"sequence": []})
assert len(store) == 0 assert len(store) == 0
assert store.num_records == 0 # empty key forces num_records=0
assert store.keys == ["sequence"] assert store.keys == ["sequence"]
def test_normalize_mixed_empty_key(): def test_normalize_mixed_empty_key():
"""_normalize with empty + non-empty keys returns min=0""" """_normalize with empty + non-empty keys returns min=0 records."""
store = H5Store() store = H5Store()
store._normalize({"sequence": [torch.tensor([1, 2, 3])], "loss_mask": []}) store._normalize({"sequence": [torch.tensor([1, 2, 3])], "loss_mask": []})
assert len(store) == 0 assert len(store) == 0
assert store.num_records == 0
assert store.token_count == 0 # min() over keys
assert set(store.keys) == {"sequence", "loss_mask"} assert set(store.keys) == {"sequence", "loss_mask"}
def test_grpo_dataset_dtype(base_test_env): def test_grpo_dataset_dtype(base_test_env):
test_dir = base_test_env["test_dir"] """GRPO dataset returns correct dtypes for per-record structured data."""
dummy_data = { from astrai.dataset.dataset import GRPODataset
"prompts": [torch.randint(0, 100, (100,), dtype=torch.int32)],
"responses": [torch.randint(0, 100, (100,), dtype=torch.int32)], G = 4
"masks": [torch.ones(100, dtype=torch.int32)], store = type(
"rewards": [torch.ones(100, dtype=torch.float32)], "FakeStore",
} (),
dataset = _make_seq_dataset( {
test_dir, "grpo_dtype", train_type="grpo", data=dummy_data, window_size=32 "keys": ["prompts", "responses", "masks", "rewards"],
) "num_records": 1,
"token_count": 0,
"_data": {
"prompts": [torch.randint(0, 100, (10,), dtype=torch.int32)],
"responses": [
[torch.randint(0, 100, (5,), dtype=torch.int32) for _ in range(G)]
],
"masks": [[torch.ones(5, dtype=torch.int32) for _ in range(G)]],
"rewards": [torch.rand(G, dtype=torch.float32)],
},
"fetch_record": _fake_fetch_record,
"__len__": lambda self: self.num_records,
},
)()
dataset = GRPODataset(store=store)
item = dataset[0] item = dataset[0]
assert item["prompts"].dtype == torch.long assert item["prompts"].dtype == torch.long
assert item["responses"].dtype == torch.long assert all(r.dtype == torch.long for r in item["responses"])
assert item["masks"].dtype == torch.bool assert all(m.dtype == torch.bool for m in item["masks"])
assert item["rewards"].dtype == torch.float32 assert item["rewards"].dtype == torch.float32
def test_grpo_dataset_load(base_test_env): def test_grpo_dataset_load(base_test_env):
test_dir = base_test_env["test_dir"] """GRPO dataset loads record-structured data with per-response boundaries."""
dummy_data = { from astrai.dataset.dataset import GRPODataset
"prompts": [_rand_seq(200)],
"responses": [_rand_seq(200)], G = 3
"masks": [torch.ones(200, dtype=torch.int64)], prompt_len = 8
"rewards": [torch.rand(200, dtype=torch.float32)], resp_lens = [5, 7, 4]
} store = type(
dataset = _make_seq_dataset( "FakeStore",
test_dir, "grpo_test", train_type="grpo", data=dummy_data (),
) {
assert len(dataset) > 0 "keys": ["prompts", "responses", "masks", "rewards"],
"num_records": 1,
"token_count": 0,
"_data": {
"prompts": [torch.randint(0, 100, (prompt_len,))],
"responses": [[torch.randint(0, 100, (rl,)) for rl in resp_lens]],
"masks": [[torch.ones(rl, dtype=torch.int64) for rl in resp_lens]],
"rewards": [torch.tensor([0.9, 0.3, 0.7], dtype=torch.float32)],
},
"fetch_record": _fake_fetch_record,
"__len__": lambda self: self.num_records,
},
)()
dataset = GRPODataset(store=store)
assert len(dataset) == 1
item = dataset[0] item = dataset[0]
assert "prompts" in item assert "prompts" in item
assert "responses" in item assert "responses" in item
assert "masks" in item assert "masks" in item
assert "rewards" in item assert "rewards" in item
assert item["prompts"].shape[0] == 64
assert item["responses"].shape[0] == 64 # Prompts is 1-D
assert item["prompts"].shape == (prompt_len,)
# Responses is a list of G tensors with correct lengths
assert len(item["responses"]) == G
for i, r in enumerate(item["responses"]):
assert r.shape == (resp_lens[i],)
# Masks align with responses
assert len(item["masks"]) == G
for i, m in enumerate(item["masks"]):
assert m.shape == (resp_lens[i],)
# Rewards has G elements
assert item["rewards"].shape == (G,)
def test_detect_format_bin_dir(base_test_env): def test_detect_format_bin_dir(base_test_env):
@@ -408,7 +463,34 @@ def test_dataset_load_explicit_storage_type(base_test_env):
test_dir = base_test_env["test_dir"] test_dir = base_test_env["test_dir"]
dataset = _make_seq_dataset(test_dir, "explicit", storage_type="h5") dataset = _make_seq_dataset(test_dir, "explicit", storage_type="h5")
assert len(dataset) > 0 assert len(dataset) > 0
assert dataset.count == 200 assert dataset.token_count == 200
def _write_json_dataset(test_dir, tokenizer_path, records, config_overrides=None):
"""Write JSONL dataset — one JSON object per line."""
data_dir = os.path.join(test_dir, "json_data")
os.makedirs(data_dir, exist_ok=True)
with open(os.path.join(data_dir, "data.jsonl"), "w", encoding="utf-8") as f:
for rec in records:
f.write(json.dumps(rec, ensure_ascii=False) + "\n")
config = {
"tokenizer_path": tokenizer_path,
"version": 1,
"input": {"sections": [{"field": "text", "action": "train"}]},
"preprocessing": {"max_seq_len": 128, "min_chars": 0},
"output": {"position_ids_mode": "continuous"},
}
if config_overrides:
config.update(config_overrides)
with open(
os.path.join(data_dir, "dataset_config.json"), "w", encoding="utf-8"
) as f:
json.dump(config, f, ensure_ascii=False, indent=2)
return data_dir
def test_detect_format_jsonl_dir(base_test_env): def test_detect_format_jsonl_dir(base_test_env):
@@ -422,6 +504,90 @@ def test_detect_format_jsonl_dir(base_test_env):
assert detect_format(data_dir) == "jsonl" assert detect_format(data_dir) == "jsonl"
def test_detect_format_json_dir(base_test_env):
"""detect_format returns 'jsonl' for directory with .json files."""
test_dir = base_test_env["test_dir"]
tokenizer_path = _save_test_tokenizer(test_dir, base_test_env["tokenizer"])
data_dir = _write_json_dataset(
test_dir,
tokenizer_path,
[{"text": "hello world"}, {"text": "foo bar baz qux"}],
)
assert detect_format(data_dir) == "jsonl"
def test_json_store_seq(base_test_env):
"""JsonlStore loads .json array correctly."""
test_dir = base_test_env["test_dir"]
tokenizer_path = _save_test_tokenizer(test_dir, base_test_env["tokenizer"])
data_dir = _write_json_dataset(
test_dir,
tokenizer_path,
[{"text": "hello world"}, {"text": "foo bar baz qux"}],
)
store = StoreFactory.create("jsonl")
store.load(data_dir)
assert len(store) > 0
assert "sequence" in store.keys
dataset = DatasetFactory.load("seq", data_dir, window_size=8)
assert len(dataset) > 0
item = dataset[0]
assert "input_ids" in item
assert "target_ids" in item
def test_json_store_no_tokenizer_path(base_test_env):
"""JsonlStore uses dataset dir as tokenizer_path when omitted."""
test_dir = base_test_env["test_dir"]
tokenizer = base_test_env["tokenizer"]
tokenizer.set_chat_template(
"{% for message in messages %}{{ message['role'] }}:{{ message['content'] }}\n{% endfor %}"
)
data_dir = os.path.join(test_dir, "self_contained")
os.makedirs(data_dir, exist_ok=True)
# Save tokenizer files directly in the dataset directory
tokenizer.save_pretrained(data_dir)
# Write .jsonl data
records = [
{
"messages": [
{"role": "user", "content": "hi"},
{"role": "assistant", "content": "hello"},
]
}
]
with open(os.path.join(data_dir, "data.jsonl"), "w", encoding="utf-8") as f:
for rec in records:
f.write(json.dumps(rec, ensure_ascii=False) + "\n")
# dataset_config.json WITHOUT tokenizer_path
config = {
"version": 1,
"input": {
"sections": [{"field": "messages", "action": "$role", "template": True}]
},
"mask": {"user": "mask", "assistant": "train"},
"mask_default": "mask",
"preprocessing": {"max_seq_len": 128, "min_chars": 0},
"output": {"position_ids_mode": "continuous"},
}
with open(
os.path.join(data_dir, "dataset_config.json"), "w", encoding="utf-8"
) as f:
json.dump(config, f, ensure_ascii=False, indent=2)
store = StoreFactory.create("jsonl")
store.load(data_dir)
assert len(store) > 0
assert "sequence" in store.keys
assert "loss_mask" in store.keys
def test_jsonl_store_seq(base_test_env): def test_jsonl_store_seq(base_test_env):
test_dir = base_test_env["test_dir"] test_dir = base_test_env["test_dir"]
tokenizer_path = _save_test_tokenizer(test_dir, base_test_env["tokenizer"]) tokenizer_path = _save_test_tokenizer(test_dir, base_test_env["tokenizer"])
@@ -512,3 +678,425 @@ def test_jsonl_store_pipeline_config_roundtrip(base_test_env):
config = PipelineConfig.from_dict(raw) config = PipelineConfig.from_dict(raw)
assert config.output.position_ids_mode == "doc_reset" assert config.output.position_ids_mode == "doc_reset"
assert config.preprocessing.max_seq_len == 64 assert config.preprocessing.max_seq_len == 64
# ---------------------------------------------------------------------------
# GRPO end-to-end: builder → JsonlStore → GRPODataset → collate_fn
# ---------------------------------------------------------------------------
def _write_grpo_jsonl(test_dir, tokenizer_path, records):
"""Write a GRPO JSONL dataset directory with config."""
data_dir = os.path.join(test_dir, "grpo_jsonl")
os.makedirs(data_dir, exist_ok=True)
with open(os.path.join(data_dir, "data.jsonl"), "w", encoding="utf-8") as f:
for rec in records:
f.write(json.dumps(rec, ensure_ascii=False) + "\n")
config = {
"tokenizer_path": tokenizer_path,
"version": 1,
"input": {
"sources": {
"prompts": {
"sections": [
{
"field": "prompt",
"action": "mask",
"add_special_tokens": True,
}
]
},
"responses": {
"sections": [{"field": "responses", "action": "train"}],
"list_field": True,
"mask_key": "masks",
},
"rewards": {
"sections": [{"field": "rewards", "action": "value"}],
},
}
},
"mask": {"user": "mask", "assistant": "train"},
"mask_default": "mask",
"preprocessing": {"max_seq_len": 128},
"output": {"position_ids_mode": "none"},
}
with open(
os.path.join(data_dir, "dataset_config.json"), "w", encoding="utf-8"
) as f:
json.dump(config, f, ensure_ascii=False, indent=2)
return data_dir
def test_grpo_builder_preserves_response_boundaries(base_test_env):
"""MultiOutputMaskBuilder with list_field returns List[List[int]] for responses."""
from astrai.preprocessing.builder import SectionedMaskBuilder
from tests.data.conftest import make_grpo_no_template_config
tokenizer = base_test_env["tokenizer"]
tokenizer_path = _save_test_tokenizer(base_test_env["test_dir"], tokenizer)
builder = SectionedMaskBuilder()
config = make_grpo_no_template_config()
config.preprocessing.max_seq_len = 128
item = {
"prompt": "What is 2+2?",
"responses": ["4", "four", "2+2=4"],
"rewards": [0.9, 0.1, 0.5],
}
result = builder.build(item, config, tokenizer)
assert result is not None
# prompts should be flat list of ints
assert isinstance(result["prompts"], list)
assert isinstance(result["prompts"][0], int)
# responses should be list of lists (one per response)
assert isinstance(result["responses"], list)
assert isinstance(result["responses"][0], list)
assert isinstance(result["responses"][0][0], int)
assert len(result["responses"]) == 3
# masks should match responses structure
assert isinstance(result["masks"], list)
assert len(result["masks"]) == 3
for i in range(3):
assert len(result["masks"][i]) == len(result["responses"][i])
# rewards should be flat list of floats
assert isinstance(result["rewards"], list)
assert all(isinstance(r, float) for r in result["rewards"])
assert len(result["rewards"]) == 3
def test_grpo_end_to_end_jsonl(base_test_env):
"""Full GRPO pipeline: JSONL → JsonlStore → GRPODataset → collate_fn."""
from astrai.dataset.dataset import grpo_collate_fn
test_dir = base_test_env["test_dir"]
tokenizer = base_test_env["tokenizer"]
tokenizer_path = _save_test_tokenizer(test_dir, tokenizer)
records = [
{
"prompt": "What is 2+2?",
"responses": ["4", "four", "The answer is 4"],
"rewards": [0.9, 0.1, 0.5],
},
{
"prompt": "Write a haiku",
"responses": ["Leaves fall", "Cherry blossoms bloom in spring"],
"rewards": [0.3, 0.8],
},
]
data_dir = _write_grpo_jsonl(test_dir, tokenizer_path, records)
dataset = DatasetFactory.load("grpo", data_dir, window_size=0)
assert len(dataset) == 2
# Item 0: 3 responses
item0 = dataset[0]
assert item0["prompts"].ndim == 1
assert len(item0["responses"]) == 3
assert len(item0["masks"]) == 3
assert item0["rewards"].shape == (3,)
for r, m in zip(item0["responses"], item0["masks"]):
assert r.shape == m.shape
# Item 1: 2 responses (different group size)
item1 = dataset[1]
assert len(item1["responses"]) == 2
assert item1["rewards"].shape == (2,)
# Collate: batch records with same G (item0 has G=3)
batch = grpo_collate_fn([item0, item0])
assert batch["prompts"].shape[0] == 2
assert batch["responses"].ndim == 3
assert batch["responses"].shape[0] == 2
assert batch["responses"].shape[1] == 3 # G=3
assert batch["masks"].shape == batch["responses"].shape
assert batch["rewards"].shape == (2, 3)
def test_grpo_collate_variable_lengths():
"""collate_fn pads variable-length responses to [B, G, R_max]."""
from astrai.dataset.dataset import grpo_collate_fn
batch = [
{
"prompts": torch.tensor([1, 2, 3]),
"responses": [torch.tensor([4, 5]), torch.tensor([6, 7, 8, 9])],
"masks": [torch.tensor([1, 1]), torch.tensor([1, 1, 1, 1])],
"rewards": torch.tensor([0.9, 0.1]),
},
{
"prompts": torch.tensor([10, 11]),
"responses": [torch.tensor([12]), torch.tensor([13, 14, 15])],
"masks": [torch.tensor([1]), torch.tensor([1, 1, 1])],
"rewards": torch.tensor([0.5, 0.5]),
},
]
result = grpo_collate_fn(batch)
assert result["prompts"].shape == (2, 3) # B=2, P_max=3
assert result["responses"].shape == (2, 2, 4) # B=2, G=2, R_max=4
assert result["masks"].shape == (2, 2, 4)
assert result["rewards"].shape == (2, 2)
# Check padding: item 1 prompt is length 2, padded to 3
assert result["prompts"][1, 2] == 0
# Check response content: item 0, response 0 is [4,5] padded to 4
assert result["responses"][0, 0, 0] == 4
assert result["responses"][0, 0, 1] == 5
assert result["responses"][0, 0, 2] == 0 # padded
assert not result["masks"][0, 0, 2] # padded
# Check response content: item 0, response 1 is [6,7,8,9] no padding
assert result["responses"][0, 1, 3] == 9
assert result["masks"][0, 1, 3]
def test_grpo_multiple_records(base_test_env):
"""GRPODataset loads multiple records with correct structure."""
from astrai.dataset.dataset import GRPODataset
G = 4
n_records = 5
dummy_responses = [
[torch.randint(0, 100, (np.random.randint(3, 8),)) for _ in range(G)]
for _ in range(n_records)
]
store = type(
"FakeStore",
(),
{
"keys": ["prompts", "responses", "masks", "rewards"],
"num_records": n_records,
"token_count": 0,
"_data": {
"prompts": [torch.randint(0, 100, (10,)) for _ in range(n_records)],
"responses": dummy_responses,
"masks": [
[torch.ones(r.shape[0], dtype=torch.int64) for r in resps]
for resps in dummy_responses
],
"rewards": [
torch.rand(G, dtype=torch.float32) for _ in range(n_records)
],
},
"fetch_record": _fake_fetch_record,
"__len__": lambda self: self.num_records,
},
)()
dataset = GRPODataset(store=store)
assert len(dataset) == n_records
for i in range(n_records):
item = dataset[i]
assert len(item["responses"]) == G
assert len(item["masks"]) == G
assert item["rewards"].shape == (G,)
for g in range(G):
assert item["responses"][g].shape == item["masks"][g].shape
def _write_dpo_jsonl(test_dir, records):
"""Write a raw DPO JSONL file (no dataset_config.json)."""
path = os.path.join(test_dir, "dpo.jsonl")
with open(path, "w", encoding="utf-8") as f:
for rec in records:
f.write(json.dumps(rec, ensure_ascii=False) + "\n")
return path
def test_dpo_tokenize_pure_function():
"""dpo_tokenize returns flat lists with correct mask alignment."""
class FakeTokenizer:
def apply_chat_template(
self, messages, tokenize=True, add_generation_prompt=True
):
ids = []
for m in messages:
ids.append(len(m["content"]))
ids.append(-1)
if add_generation_prompt:
ids.append(99)
return ids
record = {"prompt": "ab", "chosen": "xyz", "rejected": "w"}
result = dpo_tokenize(record, FakeTokenizer(), max_len=64)
assert set(result.keys()) == {"chosen", "rejected", "chosen_mask", "rejected_mask"}
assert len(result["chosen"]) == len(result["chosen_mask"])
assert len(result["rejected"]) == len(result["rejected_mask"])
assert result["chosen_mask"][0] == 0
assert any(m == 1 for m in result["chosen_mask"])
assert result["rejected_mask"][0] == 0
def test_dpo_tokenize_malformed_record():
"""dpo_tokenize returns None for missing fields."""
class FakeTokenizer:
def apply_chat_template(
self, messages, tokenize=True, add_generation_prompt=True
):
return [1]
assert dpo_tokenize({}, FakeTokenizer()) is None
assert dpo_tokenize({"prompt": "a"}, FakeTokenizer()) is None
assert dpo_tokenize({"prompt": "a", "chosen": "b"}, FakeTokenizer()) is None
def test_dpo_jsonl_lazy_load(base_test_env):
"""DPODataset loads raw JSONL with tokenizer_path → lazy processor."""
test_dir = base_test_env["test_dir"]
tokenizer_path = _save_test_tokenizer(test_dir, base_test_env["tokenizer"])
records = [
{"input": "Hello", "chosen": "world", "rejected": "earth"},
{"input": "Foo", "chosen": "bar", "rejected": "baz"},
]
path = _write_dpo_jsonl(test_dir, records)
ds = DatasetFactory.load(
train_type="dpo",
load_path=path,
window_size=0,
tokenizer_path=tokenizer_path,
)
assert len(ds) == 2
assert ds.store.num_records == 2
assert ds.store._processor is not None
item = ds[0]
assert set(item.keys()) == {"chosen", "rejected", "chosen_mask", "rejected_mask"}
assert item["chosen"].dtype == torch.long
assert item["chosen_mask"].dtype == torch.bool
assert item["chosen"].shape == item["chosen_mask"].shape
assert item["chosen"].shape == item["rejected"].shape
def test_dpo_jsonl_lazy_no_tokenizer():
"""DPODataset on jsonl without tokenizer_path falls back to eager
(which requires dataset_config.json, so it should raise)."""
with tempfile.TemporaryDirectory() as d:
path = os.path.join(d, "dpo.jsonl")
with open(path, "w") as f:
f.write(json.dumps({"input": "a", "chosen": "b", "rejected": "c"}) + "\n")
with pytest.raises(FileNotFoundError, match="dataset_config.json"):
DatasetFactory.load(
train_type="dpo",
load_path=path,
window_size=0,
)
def test_jsonl_store_lazy_len_returns_record_count(base_test_env):
"""JsonlStore in lazy mode: len() returns record count, not tokens."""
test_dir = base_test_env["test_dir"]
records = [{"input": str(i), "chosen": "c", "rejected": "r"} for i in range(5)]
path = _write_dpo_jsonl(test_dir, records)
store = JsonlStore()
store.load(path, processor=lambda r: {"chosen": torch.tensor([1, 2])})
assert len(store) == 5
assert store.num_records == 5
def test_jsonl_store_eager_len_returns_token_count(base_test_env):
"""JsonlStore in eager mode: num_records reflects per-record count."""
test_dir = base_test_env["test_dir"]
tokenizer_path = _save_test_tokenizer(test_dir, base_test_env["tokenizer"])
data_dir = _write_jsonl_dataset(
test_dir,
tokenizer_path,
[{"text": "hello world"}, {"text": "foo bar"}],
config_overrides={
"preprocessing": {"max_seq_len": 128, "min_chars": 0},
"output": {"position_ids_mode": "none"},
},
)
store = JsonlStore()
store.load(data_dir)
assert store.num_records == 2
assert len(store.keys) > 0
def test_h5_store_dual_mode(base_test_env):
"""H5Store supports both fetch (stream) and fetch_record (record).
No window configured ``len(store)`` reflects the record count
(2). ``token_count`` retains the legacy stream length (128), and
token-stream access via :meth:`fetch` is still available for
callers that want explicit begin/end control.
"""
test_dir = base_test_env["test_dir"]
seq_length = 64
dummy_data = {
"chosen": [_rand_seq(seq_length), _rand_seq(seq_length)],
"rejected": [_rand_seq(seq_length), _rand_seq(seq_length)],
}
save_h5(test_dir, "dpo_data", dummy_data)
store = H5Store()
store.load(test_dir)
assert store.token_count == seq_length * 2
assert store.num_records == 2
assert len(store) == 2 # no window configured → record count
rec0 = store.fetch_record(0, "chosen")
assert rec0.shape == (seq_length,)
stream = store.fetch(0, 10, "chosen")
assert stream.shape == (10,)
# Window-configured view of the same data uses stream sample count:
# token_count=128, window_size=64 → num_samples = (128-1-64)//64 + 1 = 1
stream_view = H5Store(window_size=seq_length, stride=seq_length)
stream_view.load(test_dir)
assert len(stream_view) == 1
def test_mmap_store_stream_only_no_offsets(base_test_env):
"""MmapStore without offsets: num_records == 0, stream works.
No window configured ``len(store)`` is 0 (no iterate units).
``token_count`` remains 128 for raw token slicing, and ``fetch``
provides direct token-range access.
"""
test_dir = base_test_env["test_dir"]
seq_length = 128
dummy_data = {"sequence": [_rand_seq(seq_length)]}
save_bin(test_dir, dummy_data)
store = StoreFactory.create("bin")
store.load(test_dir)
assert store.token_count == seq_length
assert store.num_records == 0
assert len(store) == 0
chunk = store.fetch(0, 32, "sequence")
assert chunk.shape == (32,)
+16 -3
View File
@@ -349,7 +349,17 @@ def test_grpo_basic(chat_tokenizer, builder):
assert "responses" in result assert "responses" in result
assert "masks" in result assert "masks" in result
assert "rewards" in result assert "rewards" in result
assert len(result["responses"]) == len(result["masks"])
# responses is List[List[int]] — one per response
assert len(result["responses"]) == 4
assert all(isinstance(r, list) for r in result["responses"])
assert all(isinstance(r[0], int) for r in result["responses"])
# masks is List[List[int]] — one per response, matching length
assert len(result["masks"]) == 4
for i in range(4):
assert len(result["masks"][i]) == len(result["responses"][i])
assert result["rewards"] == [1.0, 0.5, 0.8, 0.2] assert result["rewards"] == [1.0, 0.5, 0.8, 0.2]
@@ -362,8 +372,11 @@ def test_grpo_response_tokens_all_trained(chat_tokenizer, builder):
} }
result = builder.build(item, config, chat_tokenizer) result = builder.build(item, config, chat_tokenizer)
masks = result["masks"] masks = result["masks"]
assert all(m == 1 for m in masks) # masks is List[List[int]] — each response's mask should be all 1s
assert len(masks) == len(result["responses"]) assert len(masks) == 2
for m in masks:
assert all(v == 1 for v in m)
assert len(m) == len(result["responses"][masks.index(m)])
def test_grpo_single_reward(chat_tokenizer, builder): def test_grpo_single_reward(chat_tokenizer, builder):
+67
View File
@@ -7,6 +7,7 @@ from astrai.config.preprocess_config import (
PipelineConfig, PipelineConfig,
ProcessingConfig, ProcessingConfig,
) )
from astrai.preprocessing.packing import PackingStrategyFactory
from astrai.preprocessing.pipeline import Pipeline, filter_by_length from astrai.preprocessing.pipeline import Pipeline, filter_by_length
from tests.data.conftest import ( from tests.data.conftest import (
_CHAT_SECTIONS, _CHAT_SECTIONS,
@@ -262,3 +263,69 @@ def test_grpo_pipeline(temp_dir, tokenizer_dir):
assert "masks" in meta assert "masks" in meta
assert "rewards" in meta assert "rewards" in meta
assert "sequence" not in meta assert "sequence" not in meta
# ---------------------------------------------------------------------------
# BFD split packing
# ---------------------------------------------------------------------------
_TRU = "keep_start"
def _total_tokens(keys, key="sequence"):
return sum(len(s) for s in keys[key])
def test_bfd_split_preserves_all_tokens():
"""No tokens are lost — split chunks are kept, not truncated away."""
packer = PackingStrategyFactory.create("bfd_split")
max_len = 10
keys = {
"sequence": [list(range(25)), list(range(3))],
"loss_mask": [[1] * 25, [1] * 3],
}
result = packer.apply(keys, max_len, _TRU)
assert _total_tokens(result) == 28
for seq in result["sequence"]:
assert len(seq) <= max_len
def test_bfd_split_chunk_alignment():
"""loss_mask chunks must align with sequence chunks."""
packer = PackingStrategyFactory.create("bfd_split")
max_len = 10
keys = {
"sequence": [list(range(25))],
"loss_mask": [[0] * 5 + [1] * 20],
}
result = packer.apply(keys, max_len, _TRU)
for seq, mask in zip(result["sequence"], result["loss_mask"]):
assert len(seq) == len(mask)
def test_bfd_split_short_unchanged():
"""Sequences under max_packed_len should not be split."""
packer = PackingStrategyFactory.create("bfd_split")
max_len = 10
keys = {"sequence": [list(range(5))], "loss_mask": [[1] * 5]}
result = packer.apply(keys, max_len, _TRU)
assert _total_tokens(result) == 5
assert len(result["sequence"]) >= 1
def test_bfd_split_vs_bfd():
"""bfd loses tokens from over-length sequences; bfd_split does not."""
max_len = 10
keys = {
"sequence": [list(range(25)), list(range(8))],
"loss_mask": [[1] * 25, [1] * 8],
}
bfd = PackingStrategyFactory.create("bfd").apply(keys, max_len, _TRU)
split = PackingStrategyFactory.create("bfd_split").apply(keys, max_len, _TRU)
assert _total_tokens(bfd) < 33
assert _total_tokens(split) == 33
+6 -6
View File
@@ -1,4 +1,4 @@
from astrai.dataset import ResumableDistributedSampler from astrai.dataset import RDSampler
def test_random_sampler_consistency(random_dataset): def test_random_sampler_consistency(random_dataset):
@@ -6,8 +6,8 @@ def test_random_sampler_consistency(random_dataset):
dataset = random_dataset dataset = random_dataset
# Create two samplers with same seed # Create two samplers with same seed
sampler1 = ResumableDistributedSampler(dataset, seed=42) sampler1 = RDSampler(dataset, seed=42)
sampler2 = ResumableDistributedSampler(dataset, seed=42) sampler2 = RDSampler(dataset, seed=42)
indices1 = list(iter(sampler1)) indices1 = list(iter(sampler1))
indices2 = list(iter(sampler2)) indices2 = list(iter(sampler2))
@@ -20,8 +20,8 @@ def test_random_sampler_different_seeds(random_dataset):
dataset = random_dataset dataset = random_dataset
# Create two samplers with different seeds # Create two samplers with different seeds
sampler1 = ResumableDistributedSampler(dataset, seed=42) sampler1 = RDSampler(dataset, seed=42)
sampler2 = ResumableDistributedSampler(dataset, seed=123) sampler2 = RDSampler(dataset, seed=123)
indices1 = list(iter(sampler1)) indices1 = list(iter(sampler1))
indices2 = list(iter(sampler2)) indices2 = list(iter(sampler2))
@@ -35,7 +35,7 @@ def test_sampler_across_epochs(random_dataset):
dataset = random_dataset dataset = random_dataset
n = len(dataset) n = len(dataset)
sampler = ResumableDistributedSampler(dataset, seed=42) sampler = RDSampler(dataset, seed=42)
# Get indices for first epoch # Get indices for first epoch
epoch1_indices = list(iter(sampler)) epoch1_indices = list(iter(sampler))
+106
View File
@@ -3,6 +3,7 @@
import torch import torch
from astrai.inference.sample import ( from astrai.inference.sample import (
FrequencyPenaltyStrategy,
SamplingPipeline, SamplingPipeline,
TemperatureStrategy, TemperatureStrategy,
TopKStrategy, TopKStrategy,
@@ -125,3 +126,108 @@ def test_module_sample_batch():
assert tokens.shape == (2,) assert tokens.shape == (2,)
for t in tokens: for t in tokens:
assert 0 <= t < logits.size(-1) assert 0 <= t < logits.size(-1)
def test_frequency_penalty_noop_when_zero():
logits = torch.tensor([[1.0, 2.0, 3.0]])
input_ids = torch.tensor([[0, 2]])
s = FrequencyPenaltyStrategy(penalty=0.0)
result = s.apply(logits.clone(), input_ids=input_ids)
assert torch.equal(result, logits)
def test_frequency_penalty_noop_when_no_input_ids():
logits = torch.tensor([[1.0, 2.0, 3.0]])
s = FrequencyPenaltyStrategy(penalty=0.5)
result = s.apply(logits.clone())
assert torch.equal(result, logits)
def test_frequency_penalty_single_occurrence():
logits = torch.tensor([[4.0, 1.0, 2.0]])
input_ids = torch.tensor([[0, 2]])
input_mask = torch.tensor([[True, True]])
s = FrequencyPenaltyStrategy(penalty=0.5)
result = s.apply(logits.clone(), input_ids=input_ids, input_mask=input_mask)
assert result[0, 0] == 3.5
assert result[0, 1] == 1.0
assert result[0, 2] == 1.5
def test_frequency_penalty_multiple_occurrences():
logits = torch.tensor([[4.0, 1.0, 2.0]])
input_ids = torch.tensor([[0, 2, 0]])
input_mask = torch.tensor([[True, True, True]])
s = FrequencyPenaltyStrategy(penalty=0.5)
result = s.apply(logits.clone(), input_ids=input_ids, input_mask=input_mask)
assert result[0, 0] == 3.0
assert result[0, 1] == 1.0
assert result[0, 2] == 1.5
def test_frequency_penalty_respects_padding_mask():
logits = torch.tensor([[4.0, 1.0, 2.0]])
input_ids = torch.tensor([[0, 2, 0]])
input_mask = torch.tensor([[True, True, False]])
s = FrequencyPenaltyStrategy(penalty=0.5)
result = s.apply(logits.clone(), input_ids=input_ids, input_mask=input_mask)
assert result[0, 0] == 3.5
assert result[0, 1] == 1.0
assert result[0, 2] == 1.5
def test_frequency_penalty_batch_tensor():
logits = torch.tensor(
[
[4.0, 1.0, 2.0],
[3.0, 5.0, 1.0],
]
)
input_ids = torch.tensor([[0, 2, 0], [1, 1, 0]])
input_mask = torch.tensor([[True, True, True], [True, True, False]])
s = FrequencyPenaltyStrategy(penalty=torch.tensor([0.5, 1.0]))
result = s.apply(logits.clone(), input_ids=input_ids, input_mask=input_mask)
assert result[0, 0] == 3.0
assert result[0, 2] == 1.5
assert result[1, 1] == 3.0
def test_frequency_penalty_negative_penalty_boosts_repeats():
logits = torch.tensor([[4.0, 1.0, 2.0]])
input_ids = torch.tensor([[0, 0]])
input_mask = torch.tensor([[True, True]])
s = FrequencyPenaltyStrategy(penalty=-0.5)
result = s.apply(logits.clone(), input_ids=input_ids, input_mask=input_mask)
assert result[0, 0] == 5.0
def test_frequency_penalty_in_pipeline():
logits = torch.tensor([[5.0, 1.0, 2.0, 3.0]])
input_ids = torch.tensor([[0, 2, 0]])
input_mask = torch.tensor([[True, True, True]])
pipeline = SamplingPipeline(
[
TemperatureStrategy(1.0),
FrequencyPenaltyStrategy(0.5),
]
)
result = pipeline.apply(logits.clone(), input_ids=input_ids, input_mask=input_mask)
assert result[0, 0] == 4.0
assert result[0, 2] == 1.5
def test_sample_with_frequency_penalty():
logits = torch.tensor([[5.0, 1.0, 2.0, 3.0]])
input_ids = torch.tensor([[0, 2, 0]])
input_mask = torch.tensor([[True, True, True]])
tokens = sample(
logits,
temperature=1.0,
top_k=0,
top_p=1.0,
frequency_penalty=0.5,
input_ids=input_ids,
input_mask=input_mask,
)
assert tokens.shape == (1,)
assert 0 <= tokens[0] < logits.size(-1)
+1 -1
View File
@@ -46,7 +46,7 @@ def test_early_stopping_simulation(base_test_env, early_stopping_dataset):
# Resume from latest checkpoint # Resume from latest checkpoint
load_dir = os.path.join(base_test_env["test_dir"], "epoch_0_step_1") load_dir = os.path.join(base_test_env["test_dir"], "epoch_0_step_1")
trainer = Trainer(train_config) trainer = Trainer(train_config)
trainer.train(resume_dir=load_dir) trainer.train(param_path=load_dir, resume=True)
# Verify checkpoint was saved at expected step # Verify checkpoint was saved at expected step
load_dir = os.path.join(base_test_env["test_dir"], "epoch_1_step_5") load_dir = os.path.join(base_test_env["test_dir"], "epoch_1_step_5")
+224
View File
@@ -0,0 +1,224 @@
import pytest
import torch
from astrai.config.model_config import AutoRegressiveLMConfig
from astrai.model.transformer import AutoRegressiveLM
from astrai.trainer.strategy import GRPOStrategy
class _FakeExecutor:
"""Minimal executor stub providing ``unwrap_model`` for ref model creation."""
def unwrap_model(self, model):
return model.state_dict()
def _make_config(vocab_size=200, max_len=64):
return AutoRegressiveLMConfig(
vocab_size=vocab_size,
dim=16,
n_heads=2,
n_kv_heads=1,
dim_ffn=32,
max_len=max_len,
n_layers=2,
norm_eps=1e-5,
)
def _make_model(device):
config = _make_config()
model = AutoRegressiveLM(config).to(device=device)
return model, config
def _make_batch(
batch_size=2, group_size=4, prompt_len=8, response_len=12, device="cpu"
):
"""Construct a GRPO batch with deterministic shapes.
Returns dict with prompts [B, P], responses [B, G, R], masks [B, G, R],
rewards [B, G].
"""
prompts = torch.randint(0, 200, (batch_size, prompt_len), device=device)
responses = torch.randint(
0, 200, (batch_size, group_size, response_len), device=device
)
# All response tokens valid.
masks = torch.ones(batch_size, group_size, response_len, device=device)
# Distinct rewards per group member so std > 0.
rewards = torch.randn(batch_size, group_size, device=device)
return {
"prompts": prompts,
"responses": responses,
"masks": masks,
"rewards": rewards,
}
def _make_frozen_copy(model, device):
"""Create a frozen copy of ``model`` with independent weights loaded."""
config = _make_config()
copy = AutoRegressiveLM(config).to(device=device)
copy.load_state_dict(model.state_dict())
copy.requires_grad_(False)
copy.eval()
return copy
@pytest.fixture
def grpo_strategy():
"""Build a GRPOStrategy with a small real model and fake executor."""
device = "cuda" if torch.cuda.is_available() else "cpu"
model, config = _make_model(device)
old_model = _make_frozen_copy(model, device)
ref_model = _make_frozen_copy(model, device)
strategy = GRPOStrategy(
model=model,
device=device,
old_model=old_model,
ref_model=ref_model,
clip_eps=0.2,
kl_coef=0.01,
group_size=4,
model_fn=lambda c=config: AutoRegressiveLM(c).to(device=device),
executor=_FakeExecutor(),
)
return strategy, device
def test_grpo_loss_is_finite(grpo_strategy):
"""compute_loss returns a finite scalar."""
strategy, device = grpo_strategy
batch = _make_batch(device=device)
loss = strategy.compute_loss(batch)
assert loss.dim() == 0
assert torch.isfinite(loss).item()
def test_grpo_loss_backward(grpo_strategy):
"""Loss is differentiable w.r.t. policy model parameters."""
strategy, device = grpo_strategy
batch = _make_batch(device=device)
loss = strategy.compute_loss(batch)
loss.backward()
# At least some parameter should receive a gradient.
has_grad = any(
p.grad is not None and p.grad.abs().sum().item() > 0
for p in strategy.model.parameters()
)
assert has_grad
def test_grpo_ref_model_not_updated(grpo_strategy):
"""Backward should not populate gradients on ref_model."""
strategy, device = grpo_strategy
batch = _make_batch(device=device)
loss = strategy.compute_loss(batch)
loss.backward()
for p in strategy.ref_model.parameters():
assert p.grad is None
def test_grpo_old_model_not_updated(grpo_strategy):
"""Backward should not populate gradients on old_model."""
strategy, device = grpo_strategy
batch = _make_batch(device=device)
loss = strategy.compute_loss(batch)
loss.backward()
for p in strategy.old_model.parameters():
assert p.grad is None
def test_grpo_prompt_tokens_masked(grpo_strategy):
"""When only prompt-equivalent tokens are unmasked (response mask all 0),
the policy loss should be zero (no valid tokens contribute)."""
strategy, device = grpo_strategy
batch = _make_batch(device=device)
# Zero out all response masks → no response token contributes.
batch["masks"] = torch.zeros_like(batch["masks"])
loss = strategy.compute_loss(batch)
# With no valid tokens, policy_loss term is 0 and KL term is 0.
assert loss.item() == pytest.approx(0.0, abs=1e-6)
def test_grpo_identical_rewards_zero_advantage(grpo_strategy):
"""When all group rewards are identical, advantage is 0 → policy_loss is 0.
Only the KL term remains (which is 0 when policy == ref at init)."""
strategy, device = grpo_strategy
batch = _make_batch(device=device)
batch["rewards"] = torch.ones(batch["rewards"].shape, device=device)
loss = strategy.compute_loss(batch)
# At init policy == old == ref, so ratio == 1, KL == 0; advantage == 0.
assert loss.item() == pytest.approx(0.0, abs=1e-5)
def test_grpo_sync_old_model(grpo_strategy):
"""sync_old_model copies current policy weights into old_model."""
strategy, device = grpo_strategy
# Perturb policy model so it differs from old.
with torch.no_grad():
for p in strategy.model.parameters():
p.add_(0.05)
# old_model should still hold original weights (differ from policy).
policy_sd = strategy.model.state_dict()
old_sd = strategy.old_model.state_dict()
differs_before = any(
not torch.allclose(policy_sd[k], old_sd[k]) for k in policy_sd if k in old_sd
)
assert differs_before
strategy.sync_old_model()
old_sd_after = strategy.old_model.state_dict()
matches = all(
torch.allclose(policy_sd[k], old_sd_after[k])
for k in policy_sd
if k in old_sd_after
)
assert matches
def test_grpo_partial_mask(grpo_strategy):
"""Only the first half of response tokens are valid."""
strategy, device = grpo_strategy
batch = _make_batch(device=device)
B, G, R = batch["masks"].shape
half = R // 2
batch["masks"][:, :, half:] = 0.0
loss = strategy.compute_loss(batch)
assert torch.isfinite(loss).item()
def test_grpo_clipping_effect(grpo_strategy):
"""After diverging policy from ref, ratio should be clipped to [1-eps, 1+eps]
on the surrogate. Verify loss is finite and non-zero for distinct rewards."""
strategy, device = grpo_strategy
# Diverge policy from ref.
with torch.no_grad():
for p in strategy.model.parameters():
p.add_(0.3)
batch = _make_batch(device=device)
loss = strategy.compute_loss(batch)
assert torch.isfinite(loss).item()
# With distinct rewards and diverged policy, loss should be non-trivial.
assert loss.abs().item() > 1e-4
def test_grpo_no_reduction_param():
"""GRPOStrategy.__init__ must not accept ``reduction`` (removed)."""
import inspect
sig = inspect.signature(GRPOStrategy.__init__)
assert "reduction" not in sig.parameters
def test_grpo_shapes_3d_batch(grpo_strategy):
"""Verify compute_loss handles non-square prompt/response lengths."""
strategy, device = grpo_strategy
batch = _make_batch(
batch_size=3, group_size=4, prompt_len=10, response_len=8, device=device
)
loss = strategy.compute_loss(batch)
assert torch.isfinite(loss).item()