[Bug]: as_query_engine(streaming=True) buffers entire response into a single-item generator when using local LLM
![[Bug]: as_query_engine(streaming=True) buffers entire response into a single-item generator when using local LLM](https://www.chat-gpts.plus/wp-content/uploads/2026/06/22183-748ac896-768x403.jpg)
用户在 macOS 上使用 LlamaIndex 0.14 版本,通过 Ollama 封装本地 8B 模型( llama3.1:8b-instruct-q4_K_M ),调用 index.as_query_engine(streaming=True) 时触发。依赖包括 llama-index (>=
![[Bug]: ExcelParser drops cells with value 0 or False](https://www.chat-gpts.plus/wp-content/uploads/2026/06/16421-f8e0440e-768x403.jpg)


![[Bug]: test_flashinfer_cutlass_mxfp4_fused_moe accuracy mismatch on H20 (sm90) — 89% mismatch vs 20% threshold](https://www.chat-gpts.plus/wp-content/uploads/2026/06/46585-55e7134d-768x403.jpg)
![[Question]: Knowledge graph extraction problem](https://www.chat-gpts.plus/wp-content/uploads/2026/06/7119-a01a8c11-768x403.jpg)
![Misc. bug: [ROCm] Significantly lower token generation performance vs Vulkan on RX 7900 XTX (gfx1100)](https://www.chat-gpts.plus/wp-content/uploads/2026/06/20934-97b8a216-768x403.jpg)
![[Feature Request]: Restart Failed Subtasks Instead of Restarting Entire Process](https://www.chat-gpts.plus/wp-content/uploads/2026/06/8682-f9e725d4-768x403.jpg)

