ValueError: Your model contains `DTensor` parameters, which is incompatible with DDP. Maybe you loaded your model with `device_map=’auto’`? Specify `device_map=’cuda’` or ‘xpu’ or ‘cpu’ instea

这个报错通常出现在你用 DistributedConfig(fsdp_size=N) 在加载时就把模型 FSDP2 分片,然后把已分片的模型交给 Trainer 训练的场景。此时 accelerate 的 prepare_model 仍按 DDP 分支处理,撞上 DTensor 检查而报错。优先排查

快速结论:这个报错通常出现在你用 DistributedConfig(fsdp_size=N) 在加载时就把模型 FSDP2 分片,然后把已分片的模型交给 Trainer 训练的场景。此时 accelerate 的 prepare_model 仍按 DDP 分支处理,撞上 DTensor 检查而报错。优先排查 accelerate 侧 guard 与 FSDP 信号的分工,而不是模型加载本身。

适用环境:Issue 已确认环境为 transformers main(d56c55bf56)、torch 2.13.0+cu130、accelerate 1.15.0.dev0、Python 3.13.13、Linux、2x H100;模型 Qwen/Qwen3-0.6B,通过 torchrun --nproc_per_node 2 运行。

最快修复方案:暂无确认的一步修复方案。Issue 讨论指向需要 accelerate 侧调整 prepare_model 的 guard(放宽到不只检查 tp_enabled),并厘清“已分片”信号由 ParallelismConfig.dp_shard_enabled 还是 FSDP plugin 承载;在方向明确前,不建议强行绕过。

注意事项:讨论中的代码行号是作者对着当时的 main 阅读源码得出的,作者也明确说明“这是源码阅读,不是复现,本机没有 2×H100 验证”,因此结论属于分析推断而非已验证修复。此外,若最终走“只靠 ParallelismConfig、不带 plugin”的路线,Trainer 中 is_fsdp_enabled = accelerator.state.fsdp_plugin is not None 会变成 False,可能静默绕过 FSDP 相关 checkpoint 路径,是一种更隐蔽的失败。

问题场景

用户在使用 Transformers 的 Trainer 训练时,先用 AutoModelForCausalLM.from_pretrained(..., distributed_config=DistributedConfig(fsdp_size=2)) 在加载阶段就把模型做 FSDP2 分片,然后把该模型直接传给 Trainer 并调用 trainer.train()。也就是想让 Trainer 接住一个“加载时已分片”的模型,而不是让 FSDP 走 accelerate plugin 的常规流程。

报错原文

Traceback (most recent call last):
  File "repro.py", line 37, in <module>
    trainer.train()
  File ".../transformers/trainer.py", line 1452, in train
    return inner_training_loop(...)
  File ".../transformers/trainer.py", line 1491, in _inner_training_loop
    self._prepare_for_training(...)
  File ".../transformers/trainer.py", line 1627, in _prepare_for_training
    ... self.accelerator.prepare(...)
  File ".../accelerate/accelerator.py", line 1557, in prepare
  File ".../accelerate/accelerator.py", line 1879, in prepare_model
    raise ValueError(
ValueError: Your model contains `DTensor` parameters, which is incompatible with DDP. Maybe you loaded your model with `device_map='auto'`? Specify `device_map='cuda'` or 'xpu' or 'cpu' instead.

原因分析

可能原因在于 accelerate 的 prepare_model 分支判断。讨论中指出,accelerate/accelerator.py 中 prepare_model(约 L1877)的判断是:

if self.multi_device and not (self.parallelism_config and self.parallelism_config.tp_enabled):
    if model_has_dtensor(model):
        raise ValueError("Your model contains `DTensor` parameters, which is incompatible with DDP...")

这个 guard 只在 tp_enabled 时短路,ParallelismConfig.dp_shard_enabled 在该文件中仅被引用一次(约 L827 的 mesh accessor),从未在这里参与判断。因此用 fully_shard 在加载时分片的模型仍会落入 DDP 分支并触发 DTensor 检查,无论 ParallelismConfig 里带了什么。

更关键的是,讨论认为“接住已分片模型”的逻辑其实已经实现——fsdp2_prepare_model(utils/fsdp_utils.py,约 L733)开头判断 is_type_fsdp 后直接 return model,而 apply_fully_sharded_data_parallelism 用的是可组合的 fully_shard,所以模型确实是 FSDPModule,本应走“不再分片、直接采用”。但该路径不可达:上面的 guard 先把它导向 DDP,且 is_fsdp2 依赖 distributed_type == FSDP and fsdp_plugin.fsdp_version == 2,在没有 FSDP plugin 的情况下,仅携带 dp_shard_size 的 ParallelismConfig 不会路由到那里。

环境排查

  • 确认 transformers 版本/commit:Issue 报告为 main(d56c55bf56)。
  • 确认 torch:2.13.0+cu130。
  • 确认 accelerate:1.15.0.dev0。
  • 确认 Python:3.13.13,Linux。
  • 确认硬件:2x H100(Issue 的复现环境;讨论者本人无 2×H100,未做验证)。
  • 确认模型加载方式:是否使用了 device_map='auto',或使用了 DistributedConfig(fsdp_size=N) 在加载阶段分片。
  • 确认是否设置了 FSDP plugin:accelerator.state.fsdp_plugin 是否为 None,这会直接影响 is_fsdp_enabled 与 checkpoint 路径的判断。

解决步骤

  1. 先区分场景:如果你只是普通 DDP 训练却看到这个报错,按报错提示检查是否误用了 device_map='auto';若确实如此,改用 device_map='cuda'、’xpu’ 或 ‘cpu’(报错原文给出的方向)。
  2. 如果你确实是“加载时 FSDP2 分片 + Trainer”的场景,Issue 给出的结论是这属于当前设计尚未明确覆盖的组合,而不是靠改一行配置就能修好的问题,暂无确认的一步修复方案。
  3. 可优先尝试的方向(来自讨论,未经 2×H100 验证):在 accelerate 侧放宽 prepare_model 的 guard,使其不再只对 tp_enabled 短路,从而让已分片的 FSDPModule 能走到 fsdp2_prepare_model 的“直接采用”逻辑。
  4. 同时需要决定“已分片”信号由谁承载:是 ParallelismConfig.dp_shard_enabled,还是必须存在 FSDP plugin。这一选择是本 Issue 的核心分歧点,不是纯代码改动。
  5. 若选择 plugin-free 路线,需同步检查 Trainer 中 is_fsdp_enabled = accelerator.state.fsdp_plugin is not None(约 L840)的影响,避免 FSDP 相关 checkpoint 路径被静默绕过。
  6. 在 transformers 侧与 accelerate 侧的方向明确前,建议不要自行包装 DTensor 模型绕过检查,以免掩盖真实的并行配置错误。

验证方法

由于 Issue 中讨论者明确表示没有 2×H100 环境、结论仅为源码阅读,目前没有已验证的通过标准。可参考的验证方式是:在 torchrun --nproc_per_node 2 下重跑 Issue 中的复现脚本,确认 trainer.train() 不再抛出该 DTensor ValueError;若走 FSDP plugin-free 路线,还需额外确认 FSDP 相关的 checkpoint 保存/加载路径仍被正确启用,而不是因为 is_fsdp_enabled 为 False 被跳过。

参考来源

huggingface/transformers #48210

相关背景 PR:#48204(FSDP2 x expert parallelism)

GamsGo AI

AI 工具推荐

想把多个 AI 模型放在一个入口?

GamsGo AI 集成 ChatGPT、DeepSeek、Gemini、Claude、Midjourney、Veo 等常用模型,适合写作、绘图、视频和日常 AI 工作流。

了解 GamsGo AI

推广链接:通过此链接购买,我可能获得佣金,不影响你的价格。

这个方案解决了吗?

celebrityanime
celebrityanime
文章: 26254

发表回复

您的邮箱地址不会被公开。 必填项已用 * 标注