Eval bug: on HIP, a sequence whose prompt shares a llama_decode() batch with another sequence’s decode row gets corrupted logits, and every

在 HIP(ROCm)后端下,当一个新序列的 prompt 行与另一个正在生成序列的 decode 行被放进同一个 llama_decode() batch 时,加入方序列(seq 1)的首个采样 token 就会出错,并退化成 3–7 个 token 的循环,但所有调用仍返回成功。优先确认是否在使

![Eval bug: [Vulkan] GGML_ASSERT(wg0 device->properties.limits.maxComputeWorkGroupCount on Intel Arc A770 when running Qwen 3.8 flash next](https://www.chat-gpts.plus/wp-content/uploads/2026/09/28247-20007be1-768x403.jpg)

![[v0.20.0] cp38-abi3 wheels contains cp312 bindings](https://www.chat-gpts.plus/wp-content/uploads/2026/09/41487-1b3c2932-768x403.jpg)
![[Bug]: MiMo-V2.5 chat template is auto-detected as string, reordering multimodal content](https://www.chat-gpts.plus/wp-content/uploads/2026/09/53820-a0da3df8-768x403.jpg)
![[Bug]: tool_choice="required" not enforced with Qwen3.8-Flash-Next when enable_thinking=false (streaming); xgrammar "Failed to advance FSM"](https://www.chat-gpts.plus/wp-content/uploads/2026/09/55552-a09a0a5f-768x403.jpg)

![[RFC] RL CI Matrix for vLLM: Behavioral + Physical + Protocol Coverage](https://www.chat-gpts.plus/wp-content/uploads/2026/09/45585-41fefa5f-768x403.jpg)
![[RFC]: Partial Cache Hits for Hybrid Models](https://www.chat-gpts.plus/wp-content/uploads/2026/09/45702-8c7caa82-768x403.jpg)