快速结论:在 Ascend NPU(torch_npu)上使用 Diffusers 的 RMSNorm 且设置 elementwise_affine=False(即 weight=None)时,NPU 分支会把 None 传给 torch_npu.npu_rms_norm,而该算子要求 gamma 必须是真实 Tensor,因此报错。优先排查 Diffusers 版本是否已包含修复补丁。
适用环境:Issue 已确认环境:Diffusers 0.40.0.dev0;Linux x86_64(glibc 2.39,内核 5.15.0-119);Python 3.11.15;PyTorch 2.9.0+cpu;torch_npu 2.9.0.dev20260207;Accelerator 为 Ascend NPU;使用 accelerate 多 NPU 并行。其他依赖:huggingface_hub 1.27.0、Transformers 5.3.0、Accelerate 1.10.1、PEFT 0.18.0、bitsandbytes 0.49.2、optimum-quanto 0.2.7、Safetensors 0.8.0;xFormers 未安装。
最快修复方案:升级 Diffusers 到包含 PR #14288 的版本(或从对应分支安装)。Issue 中维护者确认该问题“should be fixed after #14288”,并在合并后关闭了该 Issue。
注意事项:该修复的具体实现细节在 Issue 讨论中未展开,若升级后仍然复现,建议附上 torch_npu 版本和最小复现代码重新开 Issue。若暂时无法升级,可优先尝试的规避方式是构造 RMSNorm 时使用 elementwise_affine=True(让 weight 成为真实 Tensor),但这会改变模型结构与权重加载行为,需自行评估,且该规避方式在 Issue 中未经验证。
问题场景
用户在 Ascend NPU 上运行 Diffusers,调用 diffusers.models.normalization.RMSNorm。当 RMSNorm 以 elementwise_affine=False 构造时,self.weight 为 None。Issue 中提到这类似 LTX-2 风格的 block norm(无可学习仿射参数),即 gamma 恒为 1、不做仿射变换。此时 RMSNorm.forward 的 NPU 分支会把 self.weight(为 None)直接传入融合算子 torch_npu.npu_rms_norm,导致前向直接报错。
报错原文
RMSNorm crashes on NPU when elementwise_affine=False (weight=None): npu_rms_norm requires a real gamma tensor
npu::npu_rms_norm(Tensor input, Tensor gamma, float epsilon=1e-6) -> (Tensor, Tensor)
On Ascend NPU the call to `torch_npu.npu_rms_norm(hidden_states, None, eps)` violates
the op schema and the forward aborts.
原因分析
最可能的原因:torch_npu.npu_rms_norm 的算子 schema 将 gamma 声明为必需参数(required Tensor,而非 Optional)。而 RMSNorm 在 elementwise_affine=False 时 self.weight 合法地为 None(表示 gamma 恒为 1)。NPU 分支未对 weight=None 做特殊处理,直接把 None 传给该算子,违反算子 schema,前向中断。非 NPU 的 else 分支本身能正确处理 weight=None,因此该问题仅在 NPU 路径上暴露。
环境排查
- 确认 Diffusers 版本是否已包含 PR #14288 的修复(Issue 报告版本为 0.40.0.dev0)。
- 确认 Python 版本(报告为 3.11.15)。
- 确认 PyTorch 版本(报告为 2.9.0+cpu)与 torch_npu 版本(报告为 2.9.0.dev20260207)。
- 确认 Accelerator 为 Ascend NPU,且
is_torch_npu_available()==True。 - 确认是否使用 accelerate 多 NPU 并行。
- 确认触发层:RMSNorm 是否以
elementwise_affine=False构造(self.weight is None)。 - 注意:此问题需要真实 Ascend NPU 设备才能复现,非 NPU 构建会在更早阶段报后端可用性错误,而非 gamma=None 错误。
解决步骤
- 先确认当前 Diffusers 版本:
python -c "import diffusers; print(diffusers.__version__)"。 - 确认是否已包含修复 PR #14288 对应的改动。
- 若未包含,升级 Diffusers 到包含 #14288 的版本(或安装对应分支),然后重新运行最小复现脚本。
- 若已包含修复但仍复现,收集 Ascend NPU 环境信息(torch_npu 版本、CANN 相关版本)、最小复现代码和完整报错堆栈,在 Issue 中补充或重新开 Issue。
- 在无法升级时,可优先尝试的规避方式是在构造
RMSNorm时使用elementwise_affine=True,使weight成为真实 Tensor;需注意这会改变模型结构,需要相应权重文件支持。
验证方法
在 Ascend NPU 环境中重新运行 Issue 提供的复现脚本(RMSNorm(dim=4096, eps=1e-6, elementwise_affine=False).npu().to(torch.bfloat16) 后对 torch.randn(1, 256, 4096, device="npu", dtype=torch.bfloat16) 做前向)。若前向能正常返回输出、不再抛出 gamma=None 相关的算子 schema 错误,则说明修复生效。建议同时确认输出 dtype 与形状符合预期。
参考来源
AI 工具推荐
想把多个 AI 模型放在一个入口?
GamsGo AI 集成 ChatGPT、DeepSeek、Gemini、Claude、Midjourney、Veo 等常用模型,适合写作、绘图、视频和日常 AI 工作流。
推广链接:通过此链接购买,我可能获得佣金,不影响你的价格。
这个方案解决了吗?
可以继续搜索完整报错,或查看同一工具的其他排查指南。
![[BUG] o1, o1-pro, and o3 reasoning models fallback to 8k default context window](https://www.chat-gpts.plus/wp-content/uploads/2026/10/7303-58ff742f-768x403.jpg)
![[Bug]: pip install litellm still exceeds Windows MAX_PATH under Microsoft Store Python](https://www.chat-gpts.plus/wp-content/uploads/2026/10/43851-e38f0a83-768x403.jpg)
