- _default_backend lazy init protected with threading.Lock - attention() raises when explicit backend cannot handle call - FlashAttnBackend rejects prefill with non-None attn_mask - training test uses TORCH_NATIVE backend directly
- _default_backend lazy init protected with threading.Lock - attention() raises when explicit backend cannot handle call - FlashAttnBackend rejects prefill with non-None attn_mask - training test uses TORCH_NATIVE backend directly