快速结论:该报错通常出现在用 vLLM serve 的 OpenAI-compatible Embeddings API 跑 Qwen3-VL-Embedding-8B 等多模态 embedding 模型时,请求里的图像特征未命中多模态缓存(mm cache),触发 `Expected a cached item for mm_hash` 断言,随后 output handler 又触发 `assert req_state.detokenizer is not None` 导致服务端 crash。优先排查多模态缓存状态是否被异常请求污染,以及 /tokenize、400 校验失败请求与正式请求混用的情况。
适用环境:Issue 中确认的环境包括 vLLM(含 v0.30.0)、OpenAI-compatible Embeddings endpoint、模型 Qwen3-VL-Embedding-8B、`–runner pooling`、`–dtype bfloat16`、`–tensor-parallel-size 2`、`–max-model-len 65536`(后续复现中为 32768)、`–trust-remote-code`,以及云厂商托管推理服务和 Kubernetes 部署。Python、CUDA、显卡型号在 Issue 中未明确给出。
最快修复方案:在 v0.30.0 复现场景中,Issue 明确验证可用的规避方式是给服务加上 `–disable-mm-preprocessor-cache`,禁用多模态预处理器缓存后不再出现 P0/P1 cache drift,坏请求单独失败而服务保持在线。旧版本可优先尝试升级到包含 #38545 的 vLLM 版本。
注意事项:`–disable-mm-preprocessor-cache` 只是绕过缓存不一致,会牺牲多模态特征缓存带来的性能收益;升级版本在 v0.30.0 中只是把 mm_hash 断言改为可重试的 warning,embedding(pooling)请求仍会因为 `detokenizer` 断言崩溃,因此不能单独视为彻底修复。此外,Issue 中“/tokenize 污染缓存”的说法来自用户个案,不是官方确认结论。
问题场景
用户在 vLLM 上以 OpenAI-compatible Embeddings API 部署 Qwen3-VL-Embedding-8B,使用 pooling runner 处理带图像的多模态 embedding 请求。服务在真实流量下间歇性崩溃:先是多模态缓存取不到对应 `mm_hash` 的条目,报 `Expected a cached item for mm_hash`,随后 output handler 处理到不完整的请求状态,再报 `assert req_state.detokenizer is not None`,最终 EngineCore 或 APIServer 出错。
后续复现进一步确认了一个确定性触发路径:先发送包含图像但格式不合法的 embedding 请求,返回 400 Bad Request;再发送包含同一张图片的正常请求,就会触发 P0/P1 多模态缓存漂移,进入 retryable abort 路径,并在 pooling 请求上把服务打崩。
报错原文
AssertionError: Expected a cached item for mm_hash='1bdd9fe1df7949177eff89badb62b8d7904c75ea8d60d6826388faa1f105531d'
(APIServer pid=30) ERROR ... AsyncLLM output_handler failed.
Traceback (most recent call last):
File ".../vllm/v1/engine/async_llm.py", line 664, in output_handler
processed_outputs = output_processor.process_outputs(...)
File ".../vllm/v1/engine/output_processor.py", line 629, in process_outputs
assert req_state.detokenizer is not None
AssertionError
在 v0.30.0 中,第一个断言被替换为可重试的 warning:
WARNING [core.py:1868] Multi-modal cache miss for request embd-... (mm_hashes=['a9a6596b...']): P0/P1 cache drift; returning a retryable response so the items are resent with data.
但 pooling 请求仍会在 output handler 处崩溃:
AssertionError: assert req_state.detokenizer is not None
File ".../vllm/v1/engine/output_processor.py", line 696, in process_outputs
原因分析
核心问题是多模态特征缓存不一致。EngineCore 在 `preprocess_add_request()` 中通过 `mm_receiver_cache.get_and_update_features()` 按 `mm_hash` 取特征,但缓存里没有这个条目,于是触发断言。可能原因包括:多模态缓存被异常请求提前写入或污染、P0/P1 之间的缓存视图不一致、LRU 淘汰与正在进行的请求竞争,或者多进程/多 worker 路由导致缓存非 sticky。
第二个 `detokenizer is not None` 断言是连锁反应:请求在多模态缓存阶段已经失败,但请求状态没有完全清理,output handler 仍然尝试处理它。Issue 评论指出这属于 #25568 修复未覆盖的新路径:客户端 abort 路径被修过,但“retryable mm-cache-miss”的 abort 路径没有覆盖到 pooling/embedding 请求。
另一个用户个案认为,`/tokenize` 请求不做推理、不写 mm cache,随后 `/chat/completions` 尝试使用 LRU cache 时发现缓存为空而失败;首次 `/tokenize` → `/chat/completions` 可以工作,第二次就失败。这一条属于用户观察,不是官方最终结论。
环境排查
- 确认 vLLM 版本:旧版本可优先升级到包含 #38545 的版本;v0.30.0 可通过 `–disable-mm-preprocessor-cache` 规避缓存漂移。
- 确认启动参数是否使用 `–runner pooling`、`–dtype bfloat16`、`–tensor-parallel-size 2`、`–max-model-len` 以及 `–trust-remote-code`。
- 确认是否部署了多个 worker 或使用请求路由,导致同一 `mm_hash` 的缓存不 sticky。
- 确认崩溃前是否有 `/tokenize` 请求、返回 400 的 malformed embedding 请求,或包含图像的超大请求。
- 确认是否在 Kubernetes 或云厂商托管推理服务上运行,内部 worker 设置是否可调。
- Python、CUDA、PyTorch、显卡型号在 Issue 中未给出,需要结合自身环境另行确认。
解决步骤
- 先检查服务日志,确认崩溃前是否出现多模态缓存 miss 的 warning、400 Bad Request,或来自 `/tokenize` 的请求。若存在,先隔离这类请求来源。
- 如果使用旧版本 vLLM,先升级到包含 #38545 的版本,观察 `Expected a cached item for mm_hash` 是否消失。
- 如果升级后仍出现 pooling 请求的 `assert req_state.detokenizer is not None`,按 Issue 验证过的规避方式,在启动命令中加入 `–disable-mm-preprocessor-cache`,例如:
vllm serve /mnt/data --runner pooling --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 65536 --trust-remote-code --disable-mm-preprocessor-cache - 如果缓存状态已经被污染,按用户个案中的做法,移除或避免 `/tokenize` 请求并重启服务,让缓存回到干净状态。
- 如果暂时不能改启动参数,至少避免“先发含图 400 请求、再发同图正常请求”的调用模式,降低触发 P0/P1 cache drift 的概率。
验证方法
在加入 `–disable-mm-preprocessor-cache` 后,重复 Issue 中的确定性触发步骤:先发送一个包含图像但校验失败的 embedding 请求(预期返回 400),再发送包含同一张图片的正常 embedding 请求。如果服务不再出现 `Expected a cached item for mm_hash`,也不再因 `assert req_state.detokenizer is not None` 崩溃,并且 `/health` 持续返回 200、正常 embedding 请求可继续处理,即可认为规避生效。若升级到 v0.30.0 后只看到可重试的 mm-cache-miss warning 但服务不崩,也说明第一个断言已被处理,但仍需关注 pooling 请求是否触发第二个断言。
参考来源
AI 工具推荐
想把多个 AI 模型放在一个入口?
GamsGo AI 集成 ChatGPT、DeepSeek、Gemini、Claude、Midjourney、Veo 等常用模型,适合写作、绘图、视频和日常 AI 工作流。
推广链接:通过此链接购买,我可能获得佣金,不影响你的价格。
这个方案解决了吗?
可以继续搜索完整报错,或查看同一工具的其他排查指南。
![[Vulkan] GGML_ASSERT(neq0 == HSK) failed in ggml-vulkan.cpp during speculative draft decoding (MTP) with tensor split](https://www.chat-gpts.plus/wp-content/uploads/2026/10/29418-35687e0f-768x403.jpg)

