快速结论:这个报错通常出现在你用 DistributedConfig(fsdp_size=N) 在加载时就把模型 FSDP2 分片,然后把已分片的模型交给 Trainer 训练的场景。此时 accelerate 的 prepare_model 仍按 DDP 分支处理,撞上 DTensor 检查而报错。优先排查 accelerate 侧 guard 与 FSDP 信号的分工,而不是模型加载本身。
适用环境:Issue 已确认环境为 transformers main(d56c55bf56)、torch 2.13.0+cu130、accelerate 1.15.0.dev0、Python 3.13.13、Linux、2x H100;模型 Qwen/Qwen3-0.6B,通过 torchrun --nproc_per_node 2 运行。
最快修复方案:暂无确认的一步修复方案。Issue 讨论指向需要 accelerate 侧调整 prepare_model 的 guard(放宽到不只检查 tp_enabled),并厘清“已分片”信号由 ParallelismConfig.dp_shard_enabled 还是 FSDP plugin 承载;在方向明确前,不建议强行绕过。
注意事项:讨论中的代码行号是作者对着当时的 main 阅读源码得出的,作者也明确说明“这是源码阅读,不是复现,本机没有 2×H100 验证”,因此结论属于分析推断而非已验证修复。此外,若最终走“只靠 ParallelismConfig、不带 plugin”的路线,Trainer 中 is_fsdp_enabled = accelerator.state.fsdp_plugin is not None 会变成 False,可能静默绕过 FSDP 相关 checkpoint 路径,是一种更隐蔽的失败。
问题场景
用户在使用 Transformers 的 Trainer 训练时,先用 AutoModelForCausalLM.from_pretrained(..., distributed_config=DistributedConfig(fsdp_size=2)) 在加载阶段就把模型做 FSDP2 分片,然后把该模型直接传给 Trainer 并调用 trainer.train()。也就是想让 Trainer 接住一个“加载时已分片”的模型,而不是让 FSDP 走 accelerate plugin 的常规流程。
报错原文
Traceback (most recent call last):
File "repro.py", line 37, in <module>
trainer.train()
File ".../transformers/trainer.py", line 1452, in train
return inner_training_loop(...)
File ".../transformers/trainer.py", line 1491, in _inner_training_loop
self._prepare_for_training(...)
File ".../transformers/trainer.py", line 1627, in _prepare_for_training
... self.accelerator.prepare(...)
File ".../accelerate/accelerator.py", line 1557, in prepare
File ".../accelerate/accelerator.py", line 1879, in prepare_model
raise ValueError(
ValueError: Your model contains `DTensor` parameters, which is incompatible with DDP. Maybe you loaded your model with `device_map='auto'`? Specify `device_map='cuda'` or 'xpu' or 'cpu' instead.
原因分析
可能原因在于 accelerate 的 prepare_model 分支判断。讨论中指出,accelerate/accelerator.py 中 prepare_model(约 L1877)的判断是:
if self.multi_device and not (self.parallelism_config and self.parallelism_config.tp_enabled):
if model_has_dtensor(model):
raise ValueError("Your model contains `DTensor` parameters, which is incompatible with DDP...")
这个 guard 只在 tp_enabled 时短路,ParallelismConfig.dp_shard_enabled 在该文件中仅被引用一次(约 L827 的 mesh accessor),从未在这里参与判断。因此用 fully_shard 在加载时分片的模型仍会落入 DDP 分支并触发 DTensor 检查,无论 ParallelismConfig 里带了什么。
更关键的是,讨论认为“接住已分片模型”的逻辑其实已经实现——fsdp2_prepare_model(utils/fsdp_utils.py,约 L733)开头判断 is_type_fsdp 后直接 return model,而 apply_fully_sharded_data_parallelism 用的是可组合的 fully_shard,所以模型确实是 FSDPModule,本应走“不再分片、直接采用”。但该路径不可达:上面的 guard 先把它导向 DDP,且 is_fsdp2 依赖 distributed_type == FSDP and fsdp_plugin.fsdp_version == 2,在没有 FSDP plugin 的情况下,仅携带 dp_shard_size 的 ParallelismConfig 不会路由到那里。
环境排查
- 确认 transformers 版本/commit:Issue 报告为 main(
d56c55bf56)。 - 确认 torch:2.13.0+cu130。
- 确认 accelerate:1.15.0.dev0。
- 确认 Python:3.13.13,Linux。
- 确认硬件:2x H100(Issue 的复现环境;讨论者本人无 2×H100,未做验证)。
- 确认模型加载方式:是否使用了
device_map='auto',或使用了DistributedConfig(fsdp_size=N)在加载阶段分片。 - 确认是否设置了 FSDP plugin:
accelerator.state.fsdp_plugin是否为 None,这会直接影响is_fsdp_enabled与 checkpoint 路径的判断。
解决步骤
- 先区分场景:如果你只是普通 DDP 训练却看到这个报错,按报错提示检查是否误用了
device_map='auto';若确实如此,改用device_map='cuda'、’xpu’ 或 ‘cpu’(报错原文给出的方向)。 - 如果你确实是“加载时 FSDP2 分片 + Trainer”的场景,Issue 给出的结论是这属于当前设计尚未明确覆盖的组合,而不是靠改一行配置就能修好的问题,暂无确认的一步修复方案。
- 可优先尝试的方向(来自讨论,未经 2×H100 验证):在 accelerate 侧放宽
prepare_model的 guard,使其不再只对tp_enabled短路,从而让已分片的FSDPModule能走到fsdp2_prepare_model的“直接采用”逻辑。 - 同时需要决定“已分片”信号由谁承载:是
ParallelismConfig.dp_shard_enabled,还是必须存在 FSDP plugin。这一选择是本 Issue 的核心分歧点,不是纯代码改动。 - 若选择 plugin-free 路线,需同步检查
Trainer中is_fsdp_enabled = accelerator.state.fsdp_plugin is not None(约 L840)的影响,避免 FSDP 相关 checkpoint 路径被静默绕过。 - 在 transformers 侧与 accelerate 侧的方向明确前,建议不要自行包装 DTensor 模型绕过检查,以免掩盖真实的并行配置错误。
验证方法
由于 Issue 中讨论者明确表示没有 2×H100 环境、结论仅为源码阅读,目前没有已验证的通过标准。可参考的验证方式是:在 torchrun --nproc_per_node 2 下重跑 Issue 中的复现脚本,确认 trainer.train() 不再抛出该 DTensor ValueError;若走 FSDP plugin-free 路线,还需额外确认 FSDP 相关的 checkpoint 保存/加载路径仍被正确启用,而不是因为 is_fsdp_enabled 为 False 被跳过。
参考来源
huggingface/transformers #48210
相关背景 PR:#48204(FSDP2 x expert parallelism)
AI 工具推荐
想把多个 AI 模型放在一个入口?
GamsGo AI 集成 ChatGPT、DeepSeek、Gemini、Claude、Midjourney、Veo 等常用模型,适合写作、绘图、视频和日常 AI 工作流。
推广链接:通过此链接购买,我可能获得佣金,不影响你的价格。
这个方案解决了吗?
可以继续搜索完整报错,或查看同一工具的其他排查指南。


