Misc. bug: UMA detection incorrectly limits available memory on AMD APUs with large TTM allocations

该问题发生在 AMD APU(如 Strix Halo、Ryzen AI Max+ 系列)上,llama.cpp 的 UMA 检测逻辑误将“集成显卡”识别为共享内存系统,导致可用显存被错误地限制为系统内存的 MemAvailable,而忽略了驱动实际报告的 TTM/专用 VRAM 分配。优先检查 l
![[Bug] All models hang on GB300 (SM103) with FlashInfer 0.6.7](https://www.chat-gpts.plus/wp-content/uploads/2026/08/38729-54a22b97-768x403.jpg)


![[Bug]:amd mi308x gpu, vllm 0.27.0~0.27.1, rocm 7.2.3, Kimi-K2.7-Coder start fails:AssertionError: mla_gluon requires gfx950 (CDNA4), got gfx](https://www.chat-gpts.plus/wp-content/uploads/2026/08/51964-339b289b-768x403.jpg)


![[Question]: Getting all 0 vector similarity in retrieval testing by adding a reranker model](https://www.chat-gpts.plus/wp-content/uploads/2026/08/9388-f8d94761-768x403.jpg)

