WhisperFeatureExtractor: one non-finite input sample makes the entire feature matrix NaN, and the pipeline transcribes it silently

当输入音频波形里出现哪怕一个非有限值(NaN/Inf),部分音频特征提取器会因对整个波形统计量做归一化,把全矩阵变成 NaN,而 ASR pipeline 仍会静默输出形如 '0' 的转写结果,不报错也不警告。优先排查数据管线里是否混入了损坏样本,并确认所用 checkpoint 的特征提取配置是否
![RuntimeError: The shape of the mask [1, 9476] at index 1 does not match the shape of the indexed tensor [1, 9477, 3584] at index 1](https://www.chat-gpts.plus/wp-content/uploads/2026/09/48835-104181fb-768x403.jpg)







