[Bug][DCP] GLM-5.3 dense prefill consumes uninitialized K rows on non-owner ranks
![[Bug][DCP] GLM-5.3 dense prefill consumes uninitialized K rows on non-owner ranks](https://www.chat-gpts.plus/wp-content/uploads/2026/09/54907-d7bce83d-768x403.jpg)
该报错发生在 vLLM 使用 DCP(解码上下文并行)运行 GLM-5.3 等模型的 dense prefill 阶段,非 owner rank 因 fused_norm_rope 内核的 slot_mapping 负值早退条件错误,导致 K 矩阵(kv_c_out、k_pe_out)未初始化就被
![[Bug][DCP] GLM-5.3 dense prefill consumes uninitialized K rows on non-owner ranks](https://www.chat-gpts.plus/wp-content/uploads/2026/09/54907-d7bce83d-768x403.jpg)
该报错发生在 vLLM 使用 DCP(解码上下文并行)运行 GLM-5.3 等模型的 dense prefill 阶段,非 owner rank 因 fused_norm_rope 内核的 slot_mapping 负值早退条件错误,导致 K 矩阵(kv_c_out、k_pe_out)未初始化就被

该报错发生在 RAGFlow 0.26.0 的 Agent(智能体)画布中,当添加“问题分类节点”或“路由分类与条件节点”时,工作流状态在序列化进入 SSE 响应流时泄漏了内部使用的 functools.partial 对象,导致 json.dumps() 无法序列化。优先排查 RAGFlow 版本
![[Bug]: minimum_should_match fraction is truncated, not rounded, in every full-text doc-store connector](https://www.chat-gpts.plus/wp-content/uploads/2026/09/19027-c05031a9-768x403.jpg)
RAGFlow 全文检索连接器把 minimum_should_match 小数比例转成百分比字符串时,用 int() 截断导致阈值比预期少 1%(如 0.29 被转成 "28%")。优先检查 rag/utils/ 与 memory/utils/ 下六个连接器 search() 方法中的 str(i
![[Vulkan] FA f16-scratch fast path never enabled for hybrid-model (non-unified) KV caches](https://www.chat-gpts.plus/wp-content/uploads/2026/09/28135-2a4ef7fa-768x403.jpg)
该报错通常出现在混合架构模型(如 Qwen3.8,包含 DeltaNet/SSM 层)在 Vulkan 后端运行量化 KV 缓存时,Flash Attention 的 f16-scratch 快速路径未被触发,导致 prefill 性能显著下降。优先检查 `ggml_vk_flash_attn()`

该报错发生在 llama.cpp(含 llama-server)通过 Vulkan 后端加载 GGUF 模型时,在内存适配(-fit)或模型张量加载阶段立即触发分段错误(segfault)。首要排查方向是回退到已知可用的 cb295bf 版本,并检查 Vulkan 驱动(RADV)与 llama.c
![[Bug]: [SpecDecode] Hybrid Mamba (align) corrupts under speculative decoding when a KV connector is attached — even with zero retrieved toke](https://www.chat-gpts.plus/wp-content/uploads/2026/09/53505-3ea51de1-768x403.jpg)
该报错发生在 vLLM 使用 Mamba 混合架构模型(如 Qwen3-8B、Qwen3-27B 的 Hybrid Mamba 版本)启用投机解码(SpecDecode)并挂载 KV Connector(如 LMCacheMPConnector)时,即使检索 token 数为 0 也会出现生成内容损
![[Bug]: Failed to abort requests when killing client process.](https://www.chat-gpts.plus/wp-content/uploads/2026/09/10806-837d7c30-768x403.jpg)
该问题发生在客户端进程被强制终止后,vLLM 流式请求未及时中断并继续占用显存/算力。优先排查 FastAPI/Starlette 的 is_disconnected 检测是否失效,并考虑对该函数进行 monkey-patch 修复。

该报错通常发生在 RAGFlow 智能体画布中的 Variable Assigner(变量赋值器)节点配置了合法的数值 0 参数,或使用了 clear、remove_first、remove_last 等无需参数的操作符时。优先排查你填写的 parameter 值是否为 0,或者操作符是否确实不需要
![[Question]: Error: module 'xgboost' has no attribute 'Booster' during PDF parsing](https://www.chat-gpts.plus/wp-content/uploads/2026/09/13568-9cdfb087-768x403.jpg)
该报错通常发生在 RAGFlow 解析大 PDF 时,内部调用的 xgboost 版本与项目要求的 1.6.0 不兼容或安装损坏。优先检查并固定 xgboost 版本为 1.6.0。

这个报错通常发生在 macOS 26.3+ 系统下,使用 llama.cpp 将模型层数(-ngl > 0)卸载到 Apple GPU(Metal)时触发,表现为 llama-cli 或 llama-server 在提示词处理或长上下文生成中崩溃。优先排查系统显卡驱动与 llama.cpp 的兼容性