Eval bug: NextN/MTP tensors now load by default for existing GGUFs, no load-time opt-out (regression from #25980)

该报错发生在升级 llama.cpp 后,旧有的 GLM-5.2(GLM_DSA)等包含 NextN/MTP 张量(如 blk.78 )的 GGUF 模型开始默认加载这些张量,即使没有指定 --spec-type draft-mtp ,导致显存/内存占用上升,在接近上限时触发 cudaMalloc





![[Bug]: RayExecutorV2 should sanitize inherited Ray runtime_env before creating model workers](https://www.chat-gpts.plus/wp-content/uploads/2026/08/50413-9594759d-768x403.jpg)


