[Bug]: Qwen3.5 structured output doesn’t work
![[Bug]: Qwen3.5 structured output doesn't work](https://www.chat-gpts.plus/wp-content/uploads/2026/09/35700-bd9ee480-768x403.jpg)
在 vLLM 上通过 OpenAI 兼容接口为 Qwen3.5 提供 structured output 时,thinking 模式下返回的 JSON 可能被 markdown 代码块包裹(如 ```json ... ```),或在启用 MTP 投机解码时违反 response_format 约束。
![[Bug]: Qwen3.5 structured output doesn't work](https://www.chat-gpts.plus/wp-content/uploads/2026/09/35700-bd9ee480-768x403.jpg)
在 vLLM 上通过 OpenAI 兼容接口为 Qwen3.5 提供 structured output 时,thinking 模式下返回的 JSON 可能被 markdown 代码块包裹(如 ```json ... ```),或在启用 MTP 投机解码时违反 response_format 约束。
![[Model Support] DeepSeek-V4.1 Tracking Issue](https://www.chat-gpts.plus/wp-content/uploads/2026/09/56400-8f0f3572-768x403.jpg)
这是 DeepSeek-V4.1(DSv4.1)模型支持的 Tracking Issue,不是单一报错。若你在 vLLM 加载 DeepSeek-V4.1-Flash 时遇到模型无法解析(如 [Model Support] DeepSeek-V4.1 Tracking Issue )、ROCm 段错

这个报错通常出现在 Open WebUI 的 New Chat 页面已经带有临时聊天参数、但页面是通过应用内导航(例如浏览器 Back 按钮)返回而没有发生整页加载时。此时地址栏显示临时聊天参数,但临时模式没有开启,输入内容会被正常保存;优先排查临时聊天参数的读取时机是否只在布局挂载时执行。

这个报错通常发生在 Ollama 以 Docker 容器方式运行、同时接入了 Intel iGPU 的机器上——容器只看到 CPU,日志里明确写出 "Not utilizing Intel QuickSync iGPU",原因是集显被 Ollama 默认丢弃。优先排查是否设置了 OLLAMA_IGP

这个报错通常出现在 Linux 上安装 SwarmUI、系统 Python 被标记为 externally managed(如 Ubuntu 24+)导致 pip 无法全局安装依赖时。优先确认是否可以直接用官方 Docker 方案,或系统里是否存在可用的 Python 二进制供 SwarmUI 创建

当你在多线程环境(例如 Web 服务)中反复调用 OpenAI.responses.parse 并传入 Pydantic 模型作为 text_format 时,出现过 Unrestricted caching keyed by generated types causes memory leak i

这个报错通常出现在调用 OpenAI Python SDK v2.21.0 的 responses.parse() 并传入自定义 Pydantic 模型(例如 text_format=GuardrailDecision )时,序列化解析后的响应会触发 PydanticSerializationUne

当你在 openai-python 的 Structured Outputs(如 client.beta.chat.completions.parse )里使用 Pydantic 模型,且模型包含带默认值的可选字段或 datetime / 日期字段时,Pydantic 生成的 JSON Schema
![[Bug]: Refine/CompactAndRefine streaming collapsed to a single chunk since 0.14.22 (#21374)](https://www.chat-gpts.plus/wp-content/uploads/2026/09/22831-856a571a-768x403.jpg)
当你在 LlamaIndex 0.14.22 及之后版本中使用 Refine 或 CompactAndRefine 并开启 streaming=True 时,流式输出会退化成一次性返回单个 chunk(完整答案在生成结束后才吐出),核心报错表现为:[Bug]: Refine/CompactAndRe

这个问题通常出现在用 vLLM 加载 Aria 的 MoE 专家权重时——checkpoint 里的专家名是 experts.fc1.weight / experts.fc2.weight ,但 vLLM 的融合专家加载器却去查询 w13_weight.weight / w2_weight.weig