快速结论:这个报错通常发生在 AMD 显卡(如 9060 XT 16G)上用 ROCm 版 PyTorch 跑 ComfyUI 的文本编码/注意力相关流程时,本质是 AOTriton kernel 缺失导致 SDPA 静默失败,随后在无关的 CUDA/HIP 调用点抛出 No module named 'OpenGL_accelerate' 之类的误导性报错。优先排查 PyTorch 是否带 AOTriton 图像,并尝试切换到 --use-split-cross-attention。
适用环境:Windows;AMD 9060 XT 16G;Python 3.12.13(Anaconda);ROCm 版 PyTorch(报错表现为 torch.AcceleratorError / hipErrorInvalidValue);ComfyUI 配合 comfy-aimdo、comfy-kitchen 等组件。Issue 中未明确给出 ROCm 与 PyTorch 的具体版本号。
最快修复方案:升级到已包含 PR #15648 修复的 ComfyUI 版本。Issue 报告者确认合并 PR #15648 后问题解决。
注意事项:PR #15648 是根本修复来源,若仍使用旧版代码,仅靠改命令行参数可能只是绕过而非真正修复。Issue 中关于 pip install comfy-aimdo --force-reinstall 与检查 aotriton.images 目录的建议属于排查推测,并非已验证修复。
问题场景
用户在 Windows 上以 AMD 9060 XT 16G 运行 ComfyUI,启动参数包含 --windows-standalone-build --enable-dynamic-vram --disable-smart-memory --disable-pinned-memory --reserve-vram 3 --use-ck-attention --disable-auto-launch --cuda-device 1,在运行 minimaxh3 官方 t2v(文生视频)工作流时触发崩溃。用户已尝试禁用自定义节点,问题仍然存在。
报错原文
torch.AcceleratorError: CUDA error: invalid argument on 9060xt 16g
...
[INFO] comfy-aimdo failed to load: Could not find module 'E:\miniconda3\envs\comfyui\Lib\site-packages\comfy_aimdo\aimdo_rocm.dll' (or one of its dependencies).
[INFO] NOTE: comfy-aimdo currently only supports Nvidia and AMD GPUs
...
Traceback ... llama.py:616 / get_cast_buffer
hipErrorInvalidValue
原因分析
根据讨论,comfy-aimdo 的 DLL 加载失败几乎可以确定是干扰项(red herring):ComfyUI 会显式回退到旧版 ModelPatcher,而崩溃实际发生在核心的 get_cast_buffer。真正报出的 hipErrorInvalidValue 来自另一条路径。
最可能的原因是 PyTorch/ROCm 构建中缺少 AOTriton kernel 图像,但 torch 仍将 flash attention 报告为可用。缺失的 AOTriton kernel 可能静默失败,直到下一次不相关的启动才报错,这与本 Issue 的崩溃形态一致。讨论认为这可能与 #15647 同源,最终由 PR #15648 修复。
环境排查
- 确认 ComfyUI 版本是否已合并 PR #15648。
- 确认 PyTorch 是否为带 AOTriton 图像的 ROCm 构建。
- 确认是否设置了可能影响注意力后端的环境变量。
- 确认
comfy-aimdo、comfy-kitchen、hipBLASLt 等组件的可用状态(DLL 能否加载)。 - 确认启动参数中哪些是必需的;
--enable-dynamic-vram在相关提交后已不再需要。
解决步骤
- 升级 ComfyUI 到包含 PR #15648 的版本,这是 Issue 报告中确认的修复方式。
- 若暂无法升级,可优先尝试在启动命令中加入
--use-split-cross-attention,绕过 AOTriton 路径、强制文本编码器使用attention_basic。这是 Issue 中提出的决定性测试,可判断是否为同一根因。 - 在同一 conda 环境中执行以下命令,检查 PyTorch 是否带 AOTriton 图像(Issue 建议的排查命令):
python -c "import torch, os; d=os.path.join(torch.__path__[0],'lib','aotriton.images'); print('exists:', os.path.isdir(d)); print('entries:', sorted(os.listdir(d)) if os.path.isdir(d) else None)"
若输出exists: False,说明 PyTorch 构建缺少 AOTriton kernel 图像,与 #15647 情形吻合。 - 可优先尝试
pip install comfy-aimdo --force-reinstall处理 DLL 依赖问题,但讨论明确指出这大概率不是崩溃根因。 - 逐个移除启动参数中的
--use-ck-attention、--disable-smart-memory、--disable-pinned-memory、--reserve-vram 3进行对比测试,定位是否有参数触发问题。
验证方法
重新运行原 minimaxh3 官方 t2v 工作流,不再出现 hipErrorInvalidValue / torch.AcceleratorError 即表示已修复。Issue 报告者确认合并 PR #15648 后问题解决,可作为最终验证依据。
参考来源
Comfy-Org/ComfyUI #15653 | PR #15648 | 相关 Issue #15647
AI 工具推荐
想把多个 AI 模型放在一个入口?
GamsGo AI 集成 ChatGPT、DeepSeek、Gemini、Claude、Midjourney、Veo 等常用模型,适合写作、绘图、视频和日常 AI 工作流。
推广链接:通过此链接购买,我可能获得佣金,不影响你的价格。
这个方案解决了吗?
可以继续搜索完整报错,或查看同一工具的其他排查指南。

![ValueError: You have modified the pretrained model configuration to control generation We detected the following values set - {'suppress_tokens': [1, 2, 7, 8, 9, ...]}. This strategy to contro](https://www.chat-gpts.plus/wp-content/uploads/2026/10/49456-8601d486-768x403.jpg)
