Misc. bug: /slots save/restore silently loses all prompt reuse on hybrid/recurrent models — checkpoints are never persisted

当你在 llama-server 上对 hybrid/recurrent 模型(如 Qwen3.5/Qwen3.6、DeltaNet、Mamba 等)使用 --slot-save-path 做磁盘 /slots save/restore 时,restore 会报告成功( n_restored 等于完

快速结论:当你在 llama-server 上对 hybrid/recurrent 模型(如 Qwen3.5/Qwen3.6、DeltaNet、Mamba 等)使用 --slot-save-path 做磁盘 /slots save/restore 时,restore 会报告成功(n_restored 等于完整 token 数),但下一次相同前缀请求仍然 cache_n = 0 并全量重算。优先排查 restore 后 checkpoint 是否被持久化,以及前缀匹配后是否因缺少 checkpoint 触发 forcing full prompt re-processing。

适用环境:Apple macOS 26.3,Apple M5 Pro,Metal;llama.cpp version 10068 (571d0d540),AppleClang 21.0.0.21000101 for Darwin arm64;受影响模块 llama-server;模型示例 Qwen3.6-35B-A3B-MXFP4_MOE.gguf。Issue 中另有人报告在 Qwen3.5 122B A10B with MTP、Qwen3.6-27B、Gemma4-26B-A4B 上观察到同类现象。

最快修复方案:Issue 中暂无合并进主线的官方一步修复方案。社区 fork 提交(headbouyJB 的 commit c369f24ea004e8cecc354e5afbf3dffea17c970d)被报告可让 restored session 从 cache_n = 0 变为接近全量复用;但该提交是 AI 辅助生成、作者本人未按项目规则提 PR,是否采用由维护者决定,请自行评估后再使用。

注意事项:主仓库尚未确认合并该修复;使用 fork/补丁存在与当前 master 后续改动冲突的风险。另有报告指出即使打上补丁,某些情况下仍可能 cache miss,后来确认是构建了错误分支;但也有人反馈修复后 checkpoint 能正常恢复(日志出现 restored N context checkpoint(s) from sidecar)。此外还有用户报告部分模型在内存中也不创建 checkpoint,这可能与 chat template(如 preserve thinking 相关参数)、--cache-ram 容量或客户端 harness 改写历史有关,属于可能原因而非本 Issue 已确认结论。

问题场景

用户使用 llama-server 加载 hybrid/recurrent 模型(例如 Qwen3.6-35B-A3B-MXFP4_MOE.gguf),启动参数包含 --slot-save-path ./kv,并通过 HTTP 接口调用 /slots/{id}?action=save 与 /slots/{id}?action=restore 执行 KV/上下文存档与恢复。典型流程是:先冷启动一个长提示填满某个 slot,保存该 slot,擦除所有 slot,再 restore,然后发送相同前缀的请求并指定同一个 id_slot,期望复用前缀。

报错原文

Misc. bug: /slots save/restore silently loses all prompt reuse on hybrid/recurrent models — checkpoints are never persisted

POST /slots/{id}?action=restore reports complete success — n_restored equals
the full token count and the slot subsequently reports those tokens in
/slots — but the restored state is never reused. The next request with an
identical prefix reprocesses every token: cache_n = 0.

forcing full prompt re-processing due to lack of cache data (likely due to
SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)

cached n_tokens = 0, memory_seq_rm [0, end)

原因分析

前缀匹配本身是正确的:server-context.cpp 中 n_past = slot.prompt.tokens.get_common_prefix(input_tokens); 能返回完整 token 数,说明 restore 确实重新填充了 slot.prompt.tokens。

问题出在随后几行的处理逻辑:当 pos_min >= pos_min_thold 时,代码会在 slot.prompt.checkpoints 中反向查找可用 checkpoint;如果找不到(it == slot.prompt.checkpoints.rend()),就会设置 do_reset = true,打印 forcing full prompt re-processing due to lack of cache data,并把 pos_next = 0; n_past = 0;,导致前面恢复出来的前缀被丢弃、全量重算。

对 hybrid/recurrent 模型而言,无法对 recurrent state 做部分回退,context checkpoint 是回滚到某个前缀的唯一手段(该要求由 SWA-only 扩展到 hybrid/recurrent,见 #16382)。因此如果 on-disk save/restore 路径没有把 checkpoint 一并持久化,restore 之后 checkpoint 列表为空,恢复出来的 token 前缀就完全无法被复用。这也是本 Issue 与已知的内存 checkpoint 问题不同的地方:该环境下内存复用正常(命中率约 97.9%),只有磁盘路径失效。

环境排查

  • 确认 llama.cpp 版本与 commit:Issue 报告 version 10068 (571d0d540),并称在 master 178a6c44 上仍存在;tools/server/server-context.cpp 在两个 commit 之间无变更。
  • 确认操作系统与硬件:macOS 26.3,Apple M5 Pro,Metal。
  • 确认为 hybrid/recurrent 模型(Qwen3.5/Qwen3.6、DeltaNet、Mamba 等),这是触发条件之一。
  • 确认启动参数包含 --slot-save-path,并记录保存文件名。
  • 确认请求是否显式指定 cache_prompt: true 与 id_slot,以便观察 cache_n。
  • 如果使用社区 fork 修复,确认构建来源分支正确(有用户因构建了错误分支而误判补丁无效)。
  • 如果同时存在内存 checkpoint 不创建的情况,可检查 chat template 相关设置、--cache-ram 容量、以及客户端 harness 是否改写历史(如某些 harness 需要关闭 attribution header)。

解决步骤

  1. 先按 Issue 提供的脚本复现,确认 /slots/{id}?action=restore 返回 n_restored 为完整 token 数,但下一次相同前缀请求 cache_n = 0,即问题属于本 Issue 描述的磁盘路径失效。
  2. 观察 server 日志中是否出现 forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory...) 以及 cached n_tokens = 0, memory_seq_rm [0, end),用于确认是 checkpoint 缺失导致重置。
  3. 检查保存目录中是否生成 sidecar checkpoint 文件。修复后日志应出现类似 restored N context checkpoint(s) from sidecar <path>.ckpt 的信息。
  4. 可优先尝试社区 fork 提交:headbouyJB/llama.cpp commit c369f24ea004e8cecc354e5afbf3dffea17c970d,其 diff 为 ggml-org/llama.cpp:master...headbouyJB:fix-25913。该提交被报告可在 Qwen3.5 122B A10B with MTP 上让 restored session 的 cache_n 从 0 变为接近全量复用。
  5. 如果自行构建补丁后仍 cache miss,先确认构建的是正确分支,重新构建后再测。
  6. 该修复尚未进入主仓库,不要在未评估风险的生产环境中直接替换;如维护者后续合并,应改用主线版本。

验证方法

用相同前缀发起第二次请求并指定同一个 id_slot:修复前 .timings 中 {"cache_n":0,"prompt_n":5529};修复后应看到 cache_n 明显上升、prompt_n 显著下降。同时 server 日志中不再出现 forcing full prompt re-processing due to lack of cache data,并出现 restored N context checkpoint(s) from sidecar ...。Issue 中报告的对照数据是:14,906-token 前缀 restore 约 0.12s,而全量重算约 75.9s;修复后不应再发生全量重算。

参考来源

ggml-org/llama.cpp #25913

社区 fork 提交:headbouyJB/llama.cpp c369f24;对比:ggml-org/llama.cpp compare master…headbouyJB:fix-25913

复现材料:WinPooh32/llama.cpp-save-restore-cache-miss-issue

GamsGo AI

AI 工具推荐

想把多个 AI 模型放在一个入口?

GamsGo AI 集成 ChatGPT、DeepSeek、Gemini、Claude、Midjourney、Veo 等常用模型,适合写作、绘图、视频和日常 AI 工作流。

了解 GamsGo AI

推广链接:通过此链接购买,我可能获得佣金,不影响你的价格。

这个方案解决了吗?

celebrityanime
celebrityanime
文章: 28308

发表回复

您的邮箱地址不会被公开。 必填项已用 * 标注