[Bug]: FlashInfer fused allreduce + residual RMSNorm + quant produces corrupted output with FP32 norm weights
![[Bug]: FlashInfer fused allreduce + residual RMSNorm + quant produces corrupted output with FP32 norm weights](https://www.chat-gpts.plus/wp-content/uploads/2026/07/48324-bc8edd74-768x403.jpg)
用户在使用 vLLM 服务( vllm serve )加载 FP4 量化模型(如 nvidia/Qwen3.6-27B-NVFP4 )时,启用了 --tensor-parallel-size 2 (多 GPU),并传入编译配置 {"pass_config":{"fuse_allreduce_rms"






![[程序员] LLM 知识库,如何统一本地与云端修改](https://www.chat-gpts.plus/wp-content/uploads/2026/07/ai_cover_4-389-768x403.jpg)

