MLX runner: KV cache not released between requests causes memory accumulation and severe tgTPS collapse (v0.30.8)

用户在 Apple M4 Max(64 GB 统一内存)macOS 系统上,使用 Ollama v0.30.8 运行 MLX 模型 qwen3.6:35b-mlx (上下文窗口 262144)。通过 bench.py 脚本按递增 prompt 长度(1024, 4096, 8192, 16384,








