ASR pipeline silently destroys stereo audio in channels-last layout (what soundfile/librosa return): mean(axis=0) averages across time

该报错发生在 Transformers 的自动语音识别(ASR)pipeline 处理双声道(stereo)音频时。若音频由 soundfile 或 librosa 加载(形状为 samples × channels),pipeline 内部用 mean(axis=0) 降混音会错误地按时间轴求均值

![[Bug] All models hang on GB300 (SM103) with FlashInfer 0.6.7](https://www.chat-gpts.plus/wp-content/uploads/2026/08/38729-54a22b97-768x403.jpg)

![ValueError: The following `model_kwargs` are not used by the model: ['condition_on_prev_tokens', 'logprob_threshold', 'compression_ratio_threshold'] (note: typos in the generate arguments will](https://www.chat-gpts.plus/wp-content/uploads/2026/08/30740-9d2a0ee2-768x403.jpg)




