标签: LLM

RuntimeError: Could not find nvcc and default cuda_home=’/usr/local/cuda’ doesn’t exist`. Note that pointing `CUDA_HOME` at the pip-bundled `nvidia/cu13` nvcc is not a fix — the JIT then fails o

RuntimeError: Could not find nvcc and default cuda_home='/usr/local/cuda' doesn't exist`. Note that pointing `CUDA_HOME` at the pip-bundled `nvidia/cu13` nvcc is not a fix — the JIT then fails o

该报错通常发生在 Blackwell 架构(SM120)GPU 上启动 vLLM 并启用 FlashInfer sampler 时,根因是系统 CUDA 环境与 Python 环境内 CUDA 工具链版本不一致(混用 12.8 与 13.x 组件),导致 FlashInfer JIT 编译时头文件与

为 LLM 生成的 GPU Kernel 打造合约级验证器

为 LLM 生成的 GPU Kernel 打造合约级验证器

两位研究者构建了一套比“跑几个随机输入”严格得多的 GPU 内核合约级验证器,用它审计了 2638 个已被 LLM 生成系统验收为“正确”的 kernel,发现 62.1% 至少存在一项违规,原先行业标准测试严重高估了生成 kernel 的正确性。

Eval bug: WIP

Eval bug: WIP

该问题发生在 llama-server 的 per-sequence slot 状态恢复失败后,同一进程内的后续 completion 会持续损坏或返回空结果,重启进程才能恢复正常。优先排查 KV cache 的 slot 恢复/回滚路径,而不是模型或硬件配置。