Consider forking and maintaining pyctcdecode or switch to torchaudio.models.decoder

这个报错通常发生在同时安装 transformers[torch-speech] 和 pyannote-audio==4.0.0 时,原因是 pyctcdecode 限制 numpy=2.0 。优先排查 numpy 版本,或改用维护者提供的

这个报错通常发生在同时安装 transformers[torch-speech] 和 pyannote-audio==4.0.0 时,原因是 pyctcdecode 限制 numpy=2.0 。优先排查 numpy 版本,或改用维护者提供的

该问题发生在你在 Transformers 中想要为 T5 接入 Linformer、Performer、Nystromformer 等高效自注意力机制时,核心提示为 “Implementing efficient self attention in T5”。优先排查当前 Transformers

这个 Issue 标题 Moving in a folder & `push_to_hub` for a `trust_remote_code=True` model 对应两个诉求:自定义模型通过 trust_remote_code=True 加载时,无法使用 .. 向上引用仓库里的其他 Pytho

这个报错发生在使用 mlx_lm.generate 加载模型仓库时,因为 tokenizer_config.json 中 tokenizer_class 被设置为 "TokenizersBackend" ,而当前 transformers 版本(5.14.1)无法识别该 class。优先排查 tra

当用户需要向 Trainer 传递一个已创建的自定义 Accelerator 实例时(核心英文报错:Allow users to pass through `Accelerate` instance), 无需修改 Trainer 源码 ,直接利用 TrainingArguments 中的 accel

当使用 Hugging Face Transformers 的 model.generate() 对 Llama 等语言模型进行批量解码时,即使 do_sample=False 使用 greedy 搜索,不同 batch size 下相同输入的第一个样本输出结果可能不一致。优先确认 padding_
![[testing] making network tests more reliable](https://www.chat-gpts.plus/wp-content/uploads/2026/07/12061-08a35669-768x403.jpg)
该报错通常发生在 Transformers 测试套件或下载 Hub 资源时遇到网络瞬时故障(如 502/500/403 错误)。优先在下载 API 中启用自动重试机制,设置 try_times=3 和 try_sleep_secs=1 。
![[RFC]: Hidden States Extraction](https://www.chat-gpts.plus/wp-content/uploads/2026/07/33118-296a083d-768x403.jpg)
用户在使用 vLLM 进行模型推理时,需要获取中间层的隐藏状态(hidden states)用于以下目的:

用户在 transformers 5.0.0.dev0(安装自 main 分支)中使用 AutoModelForCausalLM 配合 BitsAndBytesConfig 加载 4-bit 量化模型时,遇到 OOM。相同配置在 transformers 4.48 上可正常加载(约 9GB 显存),

用户使用 Transformers 5.13.0.dev0 从预训练加载 Qwen/Qwen3-30B-A3B 模型,设置了 `device_map="auto"` 并指定 `offload_folder`(磁盘卸载路径)。加载后调用 `model.save_pretrained("save_pat