refactor: harden inference cache state and attention dispatch
- split KVCache into phase-specific PrefillKVCache/DecodeKVCache types selected by start_pos - unify steady-state detection in TaskCacheManager - guard decode steady-state reuse with the cached task signature so recycled req slots cannot replay a prior generation's tokens and positions - collapse attention backend fwd_decode/fwd_prefill into a single subclass-owned forward with a shared _check_fwd guard - fix thread-safety gap in weight update and validate prefill inputs before KV allocation - centralize magic constants in InferenceConfig and align docs with behavior
This commit is contained in:
@@ -150,6 +150,7 @@ class InferenceScheduler:
|
||||
)
|
||||
|
||||
def _ensure_weight_update_ready(self) -> None:
|
||||
"""Check weight update preconditions. Must be called under _weight_lock."""
|
||||
if self._loop_thread is not None and self._loop_thread.is_alive():
|
||||
raise RuntimeError("Stop the scheduler before updating model weights")
|
||||
if self._task_mgr.get_active_tasks() or self._task_mgr.get_waiting_tasks():
|
||||
|
||||
Reference in New Issue
Block a user