Eval bug: Llama-server on Mac starting at b5478 only producing 2-3 streaming tokens on Qwen3 and Deepseek R1 0528

该“Eval bug: Llama-server on Mac starting at b5478 only producing 2-3 streaming tokens on Qwen3 and Deepseek R1 0528”现象并非 llama.cpp 本身的推理错误,而是从 b5478 版



![[Bug][KV Offload][P2P] EngineCore crash reconnecting to peer: stale dead ZmqConnection remains registered](https://www.chat-gpts.plus/wp-content/uploads/2026/07/49809-64a65152-768x403.jpg)
![[Bug]: MiniMax-M3 Multimodal Model Crashes During Inference](https://www.chat-gpts.plus/wp-content/uploads/2026/07/49940-e0372e77-768x403.jpg)


![[Bug]: Ragflow would NOT create the es index if the document no parsing and throw error: no chunk found](https://www.chat-gpts.plus/wp-content/uploads/2026/07/8221-f99094b8-768x403.jpg)
![[Question]: In v0.19.0, In addition to the General chunking method, other methods(like Presentation) still cannot be parsed using the VLM mo](https://www.chat-gpts.plus/wp-content/uploads/2026/07/8186-d319b8b0-768x403.jpg)