RuntimeError: Already borrowed

该报错出现在 vLLM 服务化部署 Qwen3.5-122B-A10B-FP8 等大型 MoE 模型时,当大量并发请求(尤其是包含 image_url 内容的多模态请求)通过 /v1/chat/completions 接口发送,导致 EngineCore 进程崩溃。同一并发量下纯文本请求可以稳定运行

该报错出现在 vLLM 服务化部署 Qwen3.5-122B-A10B-FP8 等大型 MoE 模型时,当大量并发请求(尤其是包含 image_url 内容的多模态请求)通过 /v1/chat/completions 接口发送,导致 EngineCore 进程崩溃。同一并发量下纯文本请求可以稳定运行

用户在安装了 PyTorch 2.4.0 和最新版 `transformers` 后,尝试导入 `from transformers.distributed.sharding_utils import DtensorShardOperation` 时触发。该问题还会级联影响到所有间接导入 `shar
![[activations] pytorch-1.11+ Tanh Gelu Approximation](https://www.chat-gpts.plus/wp-content/uploads/2026/07/15397-f170e494-768x403.jpg)
用户在 HuggingFace Transformers 库中使用 gelu_new (即 ACT2FN["gelu_new"] )激活函数时,发现 PyTorch 1.11+ 已内置了 Tanh GELU 近似实现。用户希望在 Transformers 中检测到 PyTorch ≥ 1.11 时,
![[CI Failure]: LM Eval PCP (4xB200)](https://www.chat-gpts.plus/wp-content/uploads/2026/07/49334-644d5265-768x403.jpg)
该问题在 vLLM 项目的 CI 流水线中被触发,具体场景为运行 evals/gsm8k 中的 test_gsm8k_correctness 测试用例,测试环境配置为 4 块 B200 GPU。用户在执行 PCP (Precision Calibration Profile) 结合 MLA (Mul
![[Feature] Will there be any integration of using Flex-attention (and Paged attention)?](https://www.chat-gpts.plus/wp-content/uploads/2026/07/34527-339024db-768x403.jpg)
用户在 GitHub Issue 中询问 Hugging Face Transformers 库是否会集成 FlexAttention(基于 torch.compile 的高性能注意力机制,支持因果掩码、相对位置编码、Alibi、滑动窗口、PrefixLM、Tanh Soft-Capping 及 P

用户在使用 Hugging Face Accelerate 的命令行工具估算第三方模型的内存占用时触发。具体命令为: accelerate estimate-memory stefan-it/span-marker-gelectra-large-germeval14 --dtypes float32

用户在使用 vLLM 启动 API 服务(通过 Docker 命令 vllm/vllm-openai:qwen3_5 镜像)部署通过 LLaMA-Factory 或 SWIFT 微调后的 Qwen3.5 模型时触发此错误。无论微调方式是 LoRA(后合并)还是全参微调,均会复现。服务启动参数中包含
![[Summary] Regarding memory issue in tests](https://www.chat-gpts.plus/wp-content/uploads/2026/07/18525-586377bc-768x403.jpg)
在对 Hugging Face Transformers 库执行测试时,用户在 CircleCI 持续集成环境中发现了内存泄漏问题。具体涉及以下框架和测试:

用户使用 transformers 5.12.0 加载 google/gemma-4-E2B-it 模型,并传入 assistant_model 启用 MTP(Multi-Token Prediction)加速生成,同时使用了 DynamicCache 和 use_cache=True 。生成过程中
![[LLaMA3] 'add_bos_token=True, add_eos_token=True' seems not taking effect](https://www.chat-gpts.plus/wp-content/uploads/2026/07/30947-cde5267c-768x403.jpg)
用户使用 Transformers 加载 LLaMA 3(如 llama3-8b)的 AutoTokenizer ,在 from_pretrained 中或通过 tokenizer.add_bos_token = True / tokenizer.add_eos_token = True 尝试控制