[Bug]: ~7× MTP (K=3) decode-throughput regression on Qwen3.6-35B-A3B (GB10 / sm_121) in recent nightlies
![[Bug]: ~7× MTP (K=3) decode-throughput regression on Qwen3.6-35B-A3B (GB10 / sm_121) in recent nightlies](https://www.chat-gpts.plus/wp-content/uploads/2026/07/47297-ed34241c-768x403.jpg)
用户在 DGX Spark (GB10, sm_121, aarch64) 上,使用 vLLM 0.23.x nightly 版本(如 commit a16dbd5b8 和 93d8f834d),对 `nvidia/Qwen3.6-35B-A3B-2.06GB-per-token` 模型进行单请求






![[Summary] Regarding memory issue in tests](https://www.chat-gpts.plus/wp-content/uploads/2026/07/18525-586377bc-768x403.jpg)

