[Bug]: AssertionError at kv_cache_utils.py:1042 — dense draft model + hybrid-attention main (DeltaNet+SWA) fails in unify_kv_cache_spec_page
![[Bug]: AssertionError at kv_cache_utils.py:1042 — dense draft model + hybrid-attention main (DeltaNet+SWA) fails in unify_kv_cache_spec_page](https://www.chat-gpts.plus/wp-content/uploads/2026/07/43626-9ece450b-768x403.jpg)
用户在使用 vLLM 的推测解码功能(speculative decoding)时,主模型为混合注意力架构(Qwen3-Coder-Next-80B-A3B,包含 DeltaNet 和 SWA),draft 模型为密集注意力架构(LocoOperator-4B)。引擎初始化(engine init)
![[Bug]: Triton block quantized (e.g. MXFP4) MoE kernels producing NaNs due to OOB reads on scale values](https://www.chat-gpts.plus/wp-content/uploads/2026/07/47303-d9edb114-768x403.jpg)







