Flaky MellumModelTest.test_load_balancing_loss — the only copy without @is_flaky

这个报错出现在 Transformers 合并队列(merge queue)运行 MellumModelTest.test_load_balancing_loss 时,断言对 aux_loss 的标量比较失败(Expected 2.0 but got 2.037735939025879),并导致 P

这个报错出现在 Transformers 合并队列(merge queue)运行 MellumModelTest.test_load_balancing_loss 时,断言对 aux_loss 的标量比较失败(Expected 2.0 but got 2.037735939025879),并导致 P
![[Bug]: Kimi-K3-NVFP4 on 8xB300 produces degenerate, incoherent output in the reasoning channel on v0.27.0](https://www.chat-gpts.plus/wp-content/uploads/2026/09/51798-e8ba3af0-768x403.jpg)
该问题通常出现在 8x B300/B200 节点上用 vLLM v0.27.0(含 v0.27.1)serving Kimi-K3-NVFP4、GLM-5.2-FP8 等大模型时,进程不崩溃、latency 正常,但 reasoning channel 输出退化成重复无意义 token。优先排查点在

这个问题通常出现在用 torch.compile (inductor 后端)跑 Zyphra/Zamba-7B-v1 时,表现为编译/运行极其缓慢,容易被误判为“卡死”。优先排查 ZambaMambaMixer.forward 中写死的 use_associative_scan=False ,它会把

这个报错通常出现在安装或升级 Transformers 5.12.1 时,依赖解析器因 tokenizers 版本范围冲突而失败。核心问题是 tokenizers 0.23.0 实际上只是 rc0、并未正式发布,但约束里仍写了 >=0.22.0,<=0.23.0 ,优先排查 tokenizers 版

该报错通常出现在更新或重新拉取 Fooocus 后,`models/prompt_expansion/fooocus_expansion/` 目录下只剩 `pytorch_model.bin`,缺少 tokenizer 需要的 `config.json` 等文件,导致启动时 `FooocusExpa

当你用 Transformers v4.48.0 及以上版本加载原始 Gemma 1.0 系列 checkpoint(如 google/gemma-2b 、 google/gemma-7b 、 google/codegemma-2b )时,模型会静默地使用精确 erf GELU( GELUActiv

这个报错通常出现在 Kohya_ss 里用 BLIP 给训练集图片打标(captioning)时,只要 num_beams 大于 1 就会在 BLIP beam search 阶段崩掉,优先排查 transformers 版本以及 BLIP 与 beam search 的兼容性。部分用户在同一阶段还

这个报错通常出现在 TextGen WebUI 加载 GGUF 模型(llama.cpp loader)时,根因可能是当前安装环境缺少部分依赖,其中评论中直接点名的就是 ModuleNotFoundError: No module named 'auto_gptq' 。优先排查依赖是否完整安装,再考

这个报错来自一个 Feature request:把 Transformers 中所有模型 Config 类统一迁移为 dataclass,通常出现在你升级到包含该改动的版本后,自定义 Config 子类或老式 __init__ 写法与新基类不兼容的场景;优先排查自定义 Config 是否仍按旧的

这不是普通的“代码跑错”报错,而是在用 torch.export 导出的文本生成模型(例如适配 ExecuTorch 工作流)做推理时,发现导出产物里只有“单步预测下一个 token”的 inner transformer,缺少自回归生成逻辑,因此无法像 eager / torch.compile