[Bug]: GLM-5.3 reasoning leaks into content when clients pass enable_thinking/thinking=false — parser gates on kwargs the GLM-5.3 template n

当客户端对 GLM-5.3-Flash 传入 chat_template_kwargs: {"enable_thinking": false} (或 thinking: false )时,vLLM 的 GLM reasoning parser 会关闭推理提取,但 GLM-5.3 的 chat tem

快速结论:当客户端对 GLM-5.3-Flash 传入 chat_template_kwargs: {"enable_thinking": false}(或 thinking: false)时,vLLM 的 GLM reasoning parser 会关闭推理提取,但 GLM-5.3 的 chat template 根本不读这个变量,模型仍会输出思考内容,导致原始 scratchpad 和游离的 </think> 泄漏进 message.content。优先排查客户端是否在向 GLM-5.3 发送 enable_thinking/thinking 关闭参数。

适用环境:Issue 已验证环境为 GLM-5.3-Flash 模型,启动参数 --reasoning-parser glm45,运行在跟踪 base commit 55969c16 的 fork 上;报告者指出当前 main 的 gating 逻辑未变,但因硬件限制无法在 main 上做运行时验证。Issue 未提供操作系统、Python、CUDA、显卡型号等具体版本信息。

最快修复方案:暂无确认的一步修复方案。Issue 中维护者给出的临时建议是:当前不要对 GLM-5.3 使用 enable_thinking 选项(即从请求中去掉 chat_template_kwargs 里的关闭参数),也可以自行修改 chat_template 让模板在 enable_thinking=false 时往生成提示里写入 <think></think>。这两个办法属于 workaround,不是 parser 侧正式修复。

注意事项:不能简单地在 parser 里忽略 thinking/enable_thinking,因为 GLM-4.5/4.6 的模板确实会响应 enable_thinking: false,对它们关闭提取是正确的,必须保留。Issue 中讨论的修复方向是输出驱动:当 kwargs 已关闭提取但输出仍带 think 结束标签时,按标签切分而不是把 scratchpad 当正文透传;无标签输出保持透传,从而对 GLM-4.5/4.6 行为逐字节不变。评论明确指出流式场景无法干净地自动检测(</think> 到达时 scratchpad 已经流出),该场景在 PR 中被显式排除,不要当作已覆盖。

问题场景

用户在 vLLM 上部署 GLM-5.3-Flash 并启用 GLM reasoning parser(--reasoning-parser glm45),通过 OpenAI 兼容的 /v1/chat/completions 接口调用。很多客户端习惯性地对任何 GLM 系列模型发送 chat_template_kwargs: {"enable_thinking": false}(或 thinking: false)来关闭思考——这是 GLM-4.5/4.6 文档化的做法。此时模型侧和 parser 侧对“是否思考”的判断出现分歧,最终返回的 content 里混入了推理过程和孤立的 </think>。

报错原文

{"role": "assistant", "content": "Simple question.2 + 2 = **4**\n\nIs there anything else you'd like help with?"}

原因分析

GLM-5.3 的 chat template 里没有 thinking 开关,它只认 reasoning_effort('low'/'high',其它值回落到 'max')和 clear_thinking,enable_thinking 在模板中完全不存在:

{%- set effective_reasoning_effort = reasoning_effort if reasoning_effort is defined and reasoning_effort in ['low', 'high'] else 'max' -%}

但 parser 侧的 vllm/parser/glm47_moe.py 仍然沿用 GLM-4.5/4.6 时代的 kwargs 来门控推理提取:

chat_kwargs = kwargs.get("chat_template_kwargs", {}) or {}
thinking = chat_kwargs.get("thinking", None)
enable_thinking = chat_kwargs.get("enable_thinking", None)
self.thinking_enabled = (
    True
    if thinking is None and enable_thinking is None
    else bool(thinking) or bool(enable_thinking)
)

于是当客户端传入 enable_thinking: false 时,模板不受影响、模型照常思考,而 parser 把 thinking_enabled 置为 False 并关闭提取,两边行为不一致,scratchpad 直接落到 content。同样请求不带该 kwarg 时,content 干净、思考内容进入 reasoning 字段,可作为对照。

环境排查

  • 确认模型为 GLM-5.3-Flash;GLM-4.5/4.6 模板确实响应 enable_thinking: false,不应混为一谈。
  • 确认服务端 reasoning parser 参数,Issue 验证场景为 --reasoning-parser glm45。
  • 确认客户端请求体里是否带有 chat_template_kwargs,尤其是 enable_thinking 或 thinking 字段。
  • 确认 vLLM 版本/commit;Issue 在跟踪 base 55969c16 的 fork 上实测,并指出当前 main 的 gating 逻辑未变,但未在 main 上做运行时验证。
  • Issue 未提供 Python、CUDA、PyTorch、显卡型号等版本信息,排查时不必据此假设。

解决步骤

  1. 先用 curl 复现,观察 content 中是否出现 </think> 和原始推理文本:chat_template_kwargs 设为 {"enable_thinking": false},模型 GLM-5.3-Flash,提问 “What is 2+2?”。
  2. 发出对照请求,去掉 chat_template_kwargs,确认此时 content 干净、思考内容位于 reasoning。这能证明问题由该参数触发。
  3. 可优先尝试:从客户端请求中移除 enable_thinking/thinking 参数,不对 GLM-5.3 使用该关闭选项。这是 Issue 中维护者给出的临时建议。
  4. 可优先尝试:按 Issue 评论中维护者的示例修改 chat_template,在 add_generation_prompt 分支中,当 enable_thinking 已定义且为 false 时写入 <think></think>:{%- if add_generation_prompt -%} <|assistant|>{{- '<think></think>' if (enable_thinking is defined and not enable_thinking) else '<think>' -}} {%- endif -%}。
  5. 如果希望从 parser 侧根治,需等待或跟进输出驱动的修复:kwargs 关闭提取但输出仍带 think 结束标签时按标签切分,extract_content_ids 在 id 空间做同样处理;无标签输出保持透传,以确保 GLM-4.5/4.6 行为不变。注意该修复在流式场景下未覆盖。

验证方法

用上述 curl 对照请求验证:不带 enable_thinking/thinking 时,content 应为干净答案,思考内容出现在 reasoning;带该参数时若仍出现 scratchpad 或 </think>,说明 parser 侧修复尚未生效。若采用修改 chat_template 的 workaround,可检查生成提示中是否正确插入了 <think></think>,并观察 content 是否不再泄漏。另有回归场景需要关注:kwarg 关闭且带游离结束标签(泄漏场景)、kwarg 关闭且输出干净(GLM-4.5/4.6 行为保持)、默认行为不变。

参考来源

vllm-project/vllm #54744

GamsGo AI

AI 工具推荐

想把多个 AI 模型放在一个入口?

GamsGo AI 集成 ChatGPT、DeepSeek、Gemini、Claude、Midjourney、Veo 等常用模型,适合写作、绘图、视频和日常 AI 工作流。

了解 GamsGo AI

推广链接:通过此链接购买,我可能获得佣金,不影响你的价格。

这个方案解决了吗?

celebrityanime
celebrityanime
文章: 26267

发表回复

您的邮箱地址不会被公开。 必填项已用 * 标注