Eval bug: Tensor parallelism crashes when combined with `-ncmoe` with Qwen 3.5 397B.

用户在 llama.cpp(b9672 版本)中,使用 `llama-completion` 或 `llama-cli` 工具,通过 `--sm tensor` 启用张量并行,指定 `--device ROCm1,ROCm3`(AMD GPU),配合 `-ncmoe 58` 参数加载 Qwen 3.

用户在 llama.cpp(b9672 版本)中,使用 `llama-completion` 或 `llama-cli` 工具,通过 `--sm tensor` 启用张量并行,指定 `--device ROCm1,ROCm3`(AMD GPU),配合 `-ncmoe 58` 参数加载 Qwen 3.
![[Bug]: GLM tool-call streaming final chunks repeat metadata and combine arguments with finish_reason](https://www.chat-gpts.plus/wp-content/uploads/2026/07/44098-762f8462-768x403.jpg)
用户在 vLLM main 分支(commit 6bdabbad5 / 023808c23 )上,通过 OpenAIServingChat 服务层,使用 GLM 工具解析器(如 glm45 / glm47 )进行工具调用的流式输出时触发。问题不依赖具体 GLM 模型权重加载,可通过 unit tes

用户在使用 vLLM 0.23.1rc1.dev788+gfa4321de3 部署 deepseek-ai/DeepSeek-V4-Flash-DSpark 模型时,启用了 DSpark 推测解码( --spec-method dspark --spec-tokens 5 ),在 KV-cache
![[Question]: build docker images error on Mac M4](https://www.chat-gpts.plus/wp-content/uploads/2026/07/10073-6f796378-768x403.jpg)
用户按照 RAGFlow 官方文档在 Mac M4 上执行 docker build --build-arg LIGHTEN=1 -f Dockerfile -t infiniflow/ragflow:nightly-slim . 命令时,Dockerfile 中的 apt install 阶段无法
![[Bug]: Streaming output segmentation (Qwen3-ASR)](https://www.chat-gpts.plus/wp-content/uploads/2026/07/47421-9196bf15-768x403.jpg)
用户在 vLLM 中部署 Qwen3-ASR 模型,并开启流式推理(streaming output)。期望获得连续的 ASR 转录文本,但实际输出却被切割为多个 5 秒长度的分段片段,需要额外的后处理才能合并完整结果。
![[Bug]: presentation parsing bug](https://www.chat-gpts.plus/wp-content/uploads/2026/07/13060-7fd73526-768x403.jpg)
用户在使用 RAGFlow 的 ragflow:nightly 镜像(commit ID: 26d,image version: v0.23.1-312-g38289084a)解析 PPTX 演示文稿时,发现输出的解析结果中缺少所有幻灯片中的图片。
![[Bug]: Tool schema marks **kwargs as a required (untyped) parameter, forcing the LLM to fill it](https://www.chat-gpts.plus/wp-content/uploads/2026/07/22134-dfdd0514-768x403.jpg)
用户在 LlamaIndex 中使用 FunctionTool.from_defaults() 包装一个带有 **kwargs 的函数,然后通过 llm.chat_with_tools() 调用 OpenAI 驱动并启用 strict=True 模式时触发此错误。该错误同样适用于其他带有 *args

用户在 Windows 系统下,使用 llama-server 加载 Gemma-4-26B-A4B 模型及其对应的 mmproj 文件,通过 Vulkan 后端在 AMD Radeon 780M Graphics(AMD 8840U)上运行。文本推理正常,但一旦输入图片,视觉编码器(vision

用户在 Linux 上使用 llama.cpp (版本 3808, debug 模式)时,通过 llama-gguf-split 合并了 Hugging Face 上的 Qwen2-57B-A14B-Instruct-GGUF 分卷文件,然后使用 llama-cli 加载生成的 qwen2-57b-
![[Bug]: Vllm + Gemma 4 + claude code: tool calling problems](https://www.chat-gpts.plus/wp-content/uploads/2026/07/39043-5bb1c48d-768x403.jpg)
用户使用 vLLM 部署 Gemma 4 系列模型,通过 OpenAI 兼容 API 提供给 Claude Code 等工具进行多轮 agentic 工作流。在持续 20–30 次工具调用后,出现以下两类失败模式: