Clarify Gemma4 vision bidirectional mask behavior for full vs sliding attention

用户在 HuggingFace Transformers 库中使用 Gemma4 视觉语言模型(VLM)时,关注到 use_bidirectional_attention="vision" 配置下,图像 token 的双向注意力掩码在全局(full)与滑动(sliding)注意力层中的应用方式与官方

用户在 HuggingFace Transformers 库中使用 Gemma4 视觉语言模型(VLM)时,关注到 use_bidirectional_attention="vision" 配置下,图像 token 的双向注意力掩码在全局(full)与滑动(sliding)注意力层中的应用方式与官方
![[Bug]: Dynamic NTK RoPE scaling is wrong](https://www.chat-gpts.plus/wp-content/uploads/2026/06/41236-6f0d9c97-768x403.jpg)
用户在 vLLM 推理框架中加载 nomic-embed-text-v1 等嵌入模型时触发问题。该模型官方支持 8192 上下文长度,但 vLLM 仅支持 2048。问题源于 vLLM 中 DynamicNTKScalingRotaryEmbedding 实现的计算公式错误。可能受影响的范围包括所有

用户在使用 transformers 库(版本 5.8.0.dev0 或当前 main 分支,提交 a0fb01c )通过 LlamaConfig 初始化模型时触发。典型场景是自定义脚本中为 LlamaForCausalLM 传递自定义 head_dim 参数,但 hidden_size 与 num

在 Huggingface Diffusers 库中使用 Flux2 模型(transformer_flux2.py)进行训练或微调,启用了梯度检查点(gradient checkpointing)功能。具体发生在调制(modulation)计算的输出以元组(tuple of Tensors)形式传

用户在 Hugging Face Transformers 中使用 SinqConfig 量化模型(如 google/gemma-4-E4B-it ),执行 model.save_pretrained(save_dst) 后,再通过 AutoModelForCausalLM.from_pretrai
![[Bug]: test_flashinfer_cutlass_mxfp4_fused_moe accuracy mismatch on H20 (sm90) — 89% mismatch vs 20% threshold](https://www.chat-gpts.plus/wp-content/uploads/2026/06/46585-55e7134d-768x403.jpg)
运行 vLLM 内置的单元测试 test_flashinfer_cutlass_mxfp4_fused_moe 时,在单卡 NVIDIA H20 (sm90) 上触发。该测试被条件 gate HOPPER_MXFP4_BF16_AVAILABLE 收集(sm90 + flashinfer 存在)。使

用户在阅读 Hugging Face Transformers 官方文档中 DiaForConditionalGeneration 模型(DIA 文本到对话模型)的文本生成示例时,发现示例代码无法直接运行。

用户在使用 HuggingFace Transformers 库加载 GPT-2 模型(或其他 CausalLM 模型)时,设置 generation_config.cache_implementation = "static" 并开启了 torch.compile(model.forward, b

用户在 HuggingFace Transformers 框架下使用 Qwen3.5 模型( Qwen3_5TextModel )进行 padding-free 变长序列推理。典型场景包括:将多个样本打包成一个 batch(packing),每个样本长度不同,通过 position_ids 在每个样

长江证券研报指出,中控技术推出的 TPT 工业大模型在技术路线和商业落地速度上,已做到国内第一、全球领先水平,并将 AI 能力从化工石化拓展至全流程工业。2026年第一季度,公司工业AI收入已达1.84亿元,超过2025年前三季度总和,证明其产品已进入快速放量阶段。