[Bug]: [structured outputs] speculative decoding + `VLLM_ENFORCE_STRICT_TOOL_CALLING=1` failed to advance FSM
![[Bug]: [structured outputs] speculative decoding + `VLLM_ENFORCE_STRICT_TOOL_CALLING=1` failed to advance FSM](https://www.chat-gpts.plus/wp-content/uploads/2026/09/44006-23963f83-768x403.jpg)
该报错发生在 vLLM 开启结构化输出(structured outputs / JSON schema)并同时启用 speculative decoding(MTP)与 VLLM_ENFORCE_STRICT_TOOL_CALLING=1 时,导致 FSM(有限状态机)无法推进。优先确认 vLLM
![[Bug]: Illegal CUDA memory access in flashinfer_trtllm MoE autotune on aarch64/Grace](https://www.chat-gpts.plus/wp-content/uploads/2026/09/46861-76e7a5eb-768x403.jpg)

![[Question]: why my RAPTOR always use the wrong embedding model?](https://www.chat-gpts.plus/wp-content/uploads/2026/09/7679-06f85f40-768x403.jpg)



![[Bug][DCP] NVIDIA DeepSeek-V3.2 / GLM-5.2 fused attention bypasses DCP handling](https://www.chat-gpts.plus/wp-content/uploads/2026/09/50095-245de1c7-768x403.jpg)

![[Feature Request]: Add MonkeyOCRv2 as a document parsing backend (working adapter included)](https://www.chat-gpts.plus/wp-content/uploads/2026/09/18671-00b219e6-768x403.jpg)