快速结论:当客户端对 GLM-5.3-Flash 传入 chat_template_kwargs: {"enable_thinking": false}(或 thinking: false)时,vLLM 的 GLM reasoning parser 会关闭推理提取,但 GLM-5.3 的 chat template 根本不读这个变量,模型仍会输出思考内容,导致原始 scratchpad 和游离的 </think> 泄漏进 message.content。优先排查客户端是否在向 GLM-5.3 发送 enable_thinking/thinking 关闭参数。
适用环境:Issue 已验证环境为 GLM-5.3-Flash 模型,启动参数 --reasoning-parser glm45,运行在跟踪 base commit 55969c16 的 fork 上;报告者指出当前 main 的 gating 逻辑未变,但因硬件限制无法在 main 上做运行时验证。Issue 未提供操作系统、Python、CUDA、显卡型号等具体版本信息。
最快修复方案:暂无确认的一步修复方案。Issue 中维护者给出的临时建议是:当前不要对 GLM-5.3 使用 enable_thinking 选项(即从请求中去掉 chat_template_kwargs 里的关闭参数),也可以自行修改 chat_template 让模板在 enable_thinking=false 时往生成提示里写入 <think></think>。这两个办法属于 workaround,不是 parser 侧正式修复。
注意事项:不能简单地在 parser 里忽略 thinking/enable_thinking,因为 GLM-4.5/4.6 的模板确实会响应 enable_thinking: false,对它们关闭提取是正确的,必须保留。Issue 中讨论的修复方向是输出驱动:当 kwargs 已关闭提取但输出仍带 think 结束标签时,按标签切分而不是把 scratchpad 当正文透传;无标签输出保持透传,从而对 GLM-4.5/4.6 行为逐字节不变。评论明确指出流式场景无法干净地自动检测(</think> 到达时 scratchpad 已经流出),该场景在 PR 中被显式排除,不要当作已覆盖。
问题场景
用户在 vLLM 上部署 GLM-5.3-Flash 并启用 GLM reasoning parser(--reasoning-parser glm45),通过 OpenAI 兼容的 /v1/chat/completions 接口调用。很多客户端习惯性地对任何 GLM 系列模型发送 chat_template_kwargs: {"enable_thinking": false}(或 thinking: false)来关闭思考——这是 GLM-4.5/4.6 文档化的做法。此时模型侧和 parser 侧对“是否思考”的判断出现分歧,最终返回的 content 里混入了推理过程和孤立的 </think>。
报错原文
{"role": "assistant", "content": "Simple question.2 + 2 = **4**\n\nIs there anything else you'd like help with?"}
原因分析
GLM-5.3 的 chat template 里没有 thinking 开关,它只认 reasoning_effort('low'/'high',其它值回落到 'max')和 clear_thinking,enable_thinking 在模板中完全不存在:
{%- set effective_reasoning_effort = reasoning_effort if reasoning_effort is defined and reasoning_effort in ['low', 'high'] else 'max' -%}
但 parser 侧的 vllm/parser/glm47_moe.py 仍然沿用 GLM-4.5/4.6 时代的 kwargs 来门控推理提取:
chat_kwargs = kwargs.get("chat_template_kwargs", {}) or {}
thinking = chat_kwargs.get("thinking", None)
enable_thinking = chat_kwargs.get("enable_thinking", None)
self.thinking_enabled = (
True
if thinking is None and enable_thinking is None
else bool(thinking) or bool(enable_thinking)
)
于是当客户端传入 enable_thinking: false 时,模板不受影响、模型照常思考,而 parser 把 thinking_enabled 置为 False 并关闭提取,两边行为不一致,scratchpad 直接落到 content。同样请求不带该 kwarg 时,content 干净、思考内容进入 reasoning 字段,可作为对照。
环境排查
- 确认模型为 GLM-5.3-Flash;GLM-4.5/4.6 模板确实响应
enable_thinking: false,不应混为一谈。 - 确认服务端 reasoning parser 参数,Issue 验证场景为
--reasoning-parser glm45。 - 确认客户端请求体里是否带有
chat_template_kwargs,尤其是enable_thinking或thinking字段。 - 确认 vLLM 版本/commit;Issue 在跟踪 base
55969c16的 fork 上实测,并指出当前main的 gating 逻辑未变,但未在main上做运行时验证。 - Issue 未提供 Python、CUDA、PyTorch、显卡型号等版本信息,排查时不必据此假设。
解决步骤
- 先用 curl 复现,观察
content中是否出现</think>和原始推理文本:chat_template_kwargs设为{"enable_thinking": false},模型GLM-5.3-Flash,提问 “What is 2+2?”。 - 发出对照请求,去掉
chat_template_kwargs,确认此时content干净、思考内容位于reasoning。这能证明问题由该参数触发。 - 可优先尝试:从客户端请求中移除
enable_thinking/thinking参数,不对 GLM-5.3 使用该关闭选项。这是 Issue 中维护者给出的临时建议。 - 可优先尝试:按 Issue 评论中维护者的示例修改
chat_template,在add_generation_prompt分支中,当enable_thinking已定义且为 false 时写入<think></think>:{%- if add_generation_prompt -%} <|assistant|>{{- '<think></think>' if (enable_thinking is defined and not enable_thinking) else '<think>' -}} {%- endif -%}。 - 如果希望从 parser 侧根治,需等待或跟进输出驱动的修复:kwargs 关闭提取但输出仍带 think 结束标签时按标签切分,
extract_content_ids在 id 空间做同样处理;无标签输出保持透传,以确保 GLM-4.5/4.6 行为不变。注意该修复在流式场景下未覆盖。
验证方法
用上述 curl 对照请求验证:不带 enable_thinking/thinking 时,content 应为干净答案,思考内容出现在 reasoning;带该参数时若仍出现 scratchpad 或 </think>,说明 parser 侧修复尚未生效。若采用修改 chat_template 的 workaround,可检查生成提示中是否正确插入了 <think></think>,并观察 content 是否不再泄漏。另有回归场景需要关注:kwarg 关闭且带游离结束标签(泄漏场景)、kwarg 关闭且输出干净(GLM-4.5/4.6 行为保持)、默认行为不变。
参考来源
AI 工具推荐
想把多个 AI 模型放在一个入口?
GamsGo AI 集成 ChatGPT、DeepSeek、Gemini、Claude、Midjourney、Veo 等常用模型,适合写作、绘图、视频和日常 AI 工作流。
推广链接:通过此链接购买,我可能获得佣金,不影响你的价格。
这个方案解决了吗?
可以继续搜索完整报错,或查看同一工具的其他排查指南。
![[Bug]: Rust frontend chat-template rendering fails with 500 "unknown function: raise_exception" for templates that legally use it (e.g. Qwen](https://www.chat-gpts.plus/wp-content/uploads/2026/09/59009-5c4e83a4-768x403.jpg)
![[Bug]: Host memory is not reducing after the model is loaded into Intel XPU](https://www.chat-gpts.plus/wp-content/uploads/2026/09/50269-2f2deaf5-768x403.jpg)
