[Bug]: Refine/CompactAndRefine streaming collapsed to a single chunk since 0.14.22 (#21374)
![[Bug]: Refine/CompactAndRefine streaming collapsed to a single chunk since 0.14.22 (#21374)](https://www.chat-gpts.plus/wp-content/uploads/2026/09/22831-856a571a-1-768x403.jpg)
当你在 LlamaIndex 0.14.22 及以上版本中,用 Refine 或 CompactAndRefine 并设置 streaming=True 时,流式输出会退化成一次性返回单个 chunk。优先确认是否调用了受影响版本,并在能改动 response mode 的前提下先切到 TREE_S
![[Bug]: Refine/CompactAndRefine streaming collapsed to a single chunk since 0.14.22 (#21374)](https://www.chat-gpts.plus/wp-content/uploads/2026/09/22831-856a571a-1-768x403.jpg)
当你在 LlamaIndex 0.14.22 及以上版本中,用 Refine 或 CompactAndRefine 并设置 streaming=True 时,流式输出会退化成一次性返回单个 chunk。优先确认是否调用了受影响版本,并在能改动 response mode 的前提下先切到 TREE_S

这个报错通常出现在用 LlamaIndex 的 FunctionAgent 构建工作流、并在 workflow.run(...) 阶段执行时,属于 llama-index-workflows 与 llama-index-core 版本不匹配导致的兼容问题。优先排查这两个包的版本组合,而不是先去改 A
![[Bug]: Streaming responses broken since 0.14.0 for ContextChatEngine and similar classes](https://www.chat-gpts.plus/wp-content/uploads/2026/09/22749-8310a4fe-1-768x403.jpg)
这个报错通常出现在使用 LlamaIndex ContextChatEngine 、 CondenseQuestionChatEngine 或 CondensePlusContextChatEngine 进行流式对话时:调用 astream_chat() 或 stream_chat() 后,异步/同

当你已经配好 OpenAI 兼容的 FLUX 图片生成端点、Playground → Images 能出图,但普通聊天里看不到 Image/Image Generation 入口、模型只把 dalle.text2im 当纯文本吐出来时,先把排查重点放在聊天输入框的 Integrations 菜单里

该问题通常出现在 macOS 上运行 Ollama 桌面应用 0.34.x(含 0.34.1、0.34.2,以及 0.34.3-rc1)时,应用启动后会卡死甚至让整个系统失去响应。优先排查方向不是 MLX 或模型加载,而是应用启动时对 ChatGPT/Codex 的检测逻辑在主线程上执行 osasc

这个报错通常出现在用 llama-server 跑 MiMo-V2.6-Distill-Qwen-9B(GGUF)并调用 OpenAI 兼容 tools API 时,模型自带的 chat template 被误判为 Qwen3-Coder 模板,导致工具调用永远走不到正常结束。优先排查 llama.

该报错通常发生在 llama.cpp 使用 Vulkan 后端、开启多设备张量并行( -sm tensor )并加载模型时,进程在模型加载阶段触发 SIGSEGV;日志中会先出现 ggml-backend-meta: multi buffers are unsupported leading to

这个报错通常出现在 Linux 上使用 llama.cpp 的 SYCL 后端搭配 Intel Arc A770 启动 llama-server / llama-cli 时,启动阶段在获取设备显存大小失败后触发 ggml_abort 并 core dump。优先排查 SYCL 构建是否缺少 Inte
![[Bug][ROCm/gfx942]: DeepSeek-V4-Flash silent retrieval corruption for prompts ≥ ~4-5k tokens (AITER sparse indexer)](https://www.chat-gpts.plus/wp-content/uploads/2026/09/52109-0c477c70-768x403.jpg)
这类静默检索损坏通常出现在 ROCm/gfx942 上运行 DeepSeek-V4-Flash/Pro 并启用 AITER 稀疏索引器(DSA indexer)的长上下文场景中;当压缩后的候选 token 超过 index_topk 、开始真正发生 top-k 选择后,服务器不会报错,但 needl

当你在 Ollama 中通过 top_logprobs 请求超过 20 个候选 token 时,请求会在进入后端前被 Ollama 自己的请求校验拦截,报错信息为 top_logprobs is capped at 20, but nothing downstream requires that 。