标签: Python

RuntimeError: Could not find nvcc and default cuda_home=’/usr/local/cuda’ doesn’t exist`. Note that pointing `CUDA_HOME` at the pip-bundled `nvidia/cu13` nvcc is not a fix — the JIT then fails o

RuntimeError: Could not find nvcc and default cuda_home='/usr/local/cuda' doesn't exist`. Note that pointing `CUDA_HOME` at the pip-bundled `nvidia/cu13` nvcc is not a fix — the JIT then fails o

该报错通常发生在 Blackwell 架构(SM120)GPU 上启动 vLLM 并启用 FlashInfer sampler 时,根因是系统 CUDA 环境与 Python 环境内 CUDA 工具链版本不一致(混用 12.8 与 13.x 组件),导致 FlashInfer JIT 编译时头文件与

ValueError: The following `model_kwargs` are not used by the model: [‘condition_on_prev_tokens’, ‘logprob_threshold’, ‘compression_ratio_threshold’] (note: typos in the generate arguments will

ValueError: The following `model_kwargs` are not used by the model: ['condition_on_prev_tokens', 'logprob_threshold', 'compression_ratio_threshold'] (note: typos in the generate arguments will

这个报错通常发生在使用 Transformers 的 Whisper 模型进行批处理推理时,输入音频小于 30 秒但未按 Whisper 的固定输入长度要求进行填充或截断。优先排查特征提取器(Feature Extractor)是否在短音频场景下正确设置了 padding 和 truncation