[Bug]: Triton block quantized (e.g. MXFP4) MoE kernels producing NaNs due to OOB reads on scale values
![[Bug]: Triton block quantized (e.g. MXFP4) MoE kernels producing NaNs due to OOB reads on scale values](https://www.chat-gpts.plus/wp-content/uploads/2026/07/47303-d9edb114-768x403.jpg)
用户在 vLLM(v0.24.0 及 main 分支)上运行 MXFP4(例如 GPT-OSS 20B)MoE 模型,并在 Hopper 架构 GPU(H100、B200)上使用 Triton 默认内核时触发。问题在启用约束解码(constrained decoding)或其它内存分配操作(如位掩码


![[Bug]: After Raptor run, created chunks have wrong docnm_kwd](https://www.chat-gpts.plus/wp-content/uploads/2026/07/13393-64fd7d89-768x403.jpg)


![Eval bug: [SYCL] Fence expiration time out - Sudden hang at exact same point in prompt processing](https://www.chat-gpts.plus/wp-content/uploads/2026/07/25350-c256f83d-768x403.jpg)


![[Bug]: meta-llama/Llama-3.2-1B-Instruct Fails With ROCM_ATTN Due To Seg Fault](https://www.chat-gpts.plus/wp-content/uploads/2026/07/36180-1b1e4943-768x403.jpg)