快速结论:当使用 T5Tokenizer 这类基于 SentencePiece 的 tokenizer 加载没有 precompiled_charsmap 的 protobuf 序列化 spiece.model 时,会触发 “Cannot parse precompiled_charsmap” 报错;优先排查 spiece.model 是否缺少该字段,以及转换链路是否把空 bytes 传给了 normalizers.Precompiled。
适用环境:Issue 中已确认的环境为:transformers 5.18.0.dev0、Linux(glibc2.43)、Python 3.14.4、huggingface_hub 1.31.0、safetensors 0.8.0、accelerate 1.15.0、PyTorch 2.12.0+rocm7.14.1、AMD Radeon RX 7800 XT,未使用分布式或 GPU。
最快修复方案:暂无确认的一步修复方案。Issue 讨论中给出了一个可优先尝试的方向:在 convert_slow_tokenizer.py 解析 protobuf 时,对空的 precompiled_charsmap 做真值判断,把空 bytes 转成 None,从而跳过 normalizers.Precompiled 的实例化。该方向在本地验证中能解决加载崩溃,但尚未合并为官方修复。
注意事项:上述方法只解决转换路径的加载崩溃,Issue 讨论明确指出它并不保证与原生 SentencePiece 完全等价,例如带换行的文本 token ID 仍可能不同。维护者也明确不接受直接由代码代理提交的 PR。
问题场景
用户在自有任务中加载 T5 系列 tokenizer 时触发问题。典型场景是使用 T5Tokenizer.from_pretrained(local_dir),目录中只有 special_tokens_map.json、spiece.model 和 tokenizer_config.json,没有 tokenizer.json。Issue 使用 google/umt5-xxl 的 spiece.model 复现,移除 tokenizer.json 后仅保留上述文件即可触发。
报错原文
Traceback (most recent call last):
File "", line 1, in
tok = T5Tokenizer.from_pretrained("/src/google_umt5-xxl")
File "/src/transformers/src/transformers/tokenization_utils_base.py", line 1736, in from_pretrained
return cls._from_pretrained(
~~~~~~~~~~~~~~~~~~~~^
resolved_vocab_files,
^^^^^^^^^^^^^^^^^^^^^
......
**kwargs,
^^^^^^^^^
)
^
File "/src/transformers/src/transformers/tokenization_utils_base.py", line 1932, in _from_pretrained
tokenizer = cls(*init_inputs, **init_kwargs)
File "/src/transformers/src/transformers/models/t5/tokenization_t5.py", line 121, in __init__
self._tokenizer.normalizer = normalizers.Precompiled(_spm_precompiled_charsmap)
~~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^
Exception: Error while attempting to build Precompiled normalizer: Cannot parse precompiled_charsmap
原因分析
最可能的原因是:当 protobuf 序列化的 spiece.model 缺少 precompiled_charsmap 时,解析代码会返回该 bytes 字段的默认值,即空 bytes b"",而不是 None。随后这个 b"" 被直接传入 T5Tokenizer 并用于实例化 normalizers.Precompiled,Rust 后端无法解析 0 长度的字节序列,于是抛出 “T5 sentencepiece tokenizer without precompiled_charsmap cannot be loaded” 对应的 Cannot parse precompiled_charsmap 异常。
Issue 中的定位涉及 tokenization_t5.py 第 120-121 行和 convert_slow_tokenizer.py 第 195 行附近。讨论中维护者还提出一个边界问题:修复目标究竟是得到由 SentencePiece 模型转换而来的 TokenizersBackend,还是 SentencePieceBackend。报告者确认目标是前者,即基于 TokenizersBackend 的 tokenizer。
环境排查
- 确认
transformers版本,Issue 中为 5.18.0.dev0。 - 确认 Python 版本,Issue 中为 3.14.4。
- 确认 PyTorch 版本与加速环境,Issue 中为 2.12.0+rocm7.14.1,AMD Radeon RX 7800 XT,未使用 GPU。
- 确认问题目录中是否存在
tokenizer.json;Issue 复现时移除了该文件。 - 确认
spiece.model是否缺少precompiled_charsmap,或该字段是否为空 bytes。 - 确认 tokenizer 期望的后端类型是
TokenizersBackend还是SentencePieceBackend;Issue 中报告者确认目标是TokenizersBackend。 - 确认复现所用模型与 revision;Issue 讨论中使用了
google/umt5-xxlrevision66cb9e7e85526fe440a945569e42c72fb6cbc0ad。
解决步骤
- 先用最小目录复现:仅保留
special_tokens_map.json、spiece.model和tokenizer_config.json,移除tokenizer.json,再执行T5Tokenizer.from_pretrained(local_dir),确认报错可稳定复现。 - 检查
convert_slow_tokenizer.py中解析 protobuf 后得到的precompiled_charsmap值;如果该字段缺失,预期值为空 bytesb""。 - 可优先尝试的修复方向:在
convert_slow_tokenizer.py第 195 行附近加入真值判断,把空 bytes 映射为None,例如讨论中提到的charsmap if charsmap else None,使T5Tokenizer不再把空 bytes 传给normalizers.Precompiled。 - 另一种 Issue 正文提出的思路是在
T5Tokenizer侧判断长度:当len(self.proto.normalizer_spec.precompiled_charsmap) > 0时才传入_spm_precompiled_charsmap,否则传None。该写法来自 Issue 正文的 expected behavior,属于建议而非已验证的官方方案。 - 如果目标是获得
SentencePieceBackend,不要直接套用上述转换路径修复;Issue 讨论明确说明SentencePieceBackend加载路径未被测试。
验证方法
重新执行 T5Tokenizer.from_pretrained(local_dir),确认不再出现 Exception: Error while attempting to build Precompiled normalizer: Cannot parse precompiled_charsmap,且 tokenizer 能成功加载。Issue 讨论中还建议做 save/reload 回归检查,并用若干文本对比原生 SentencePiece token ID。注意:讨论中报告即使加载成功,"line one\nline two" 的 token ID 仍可能与原生 SentencePiece 不一致,因此加载成功不等于完全等价。另需注意,Issue 中的本地修复尚未合并,维护者已表示不接受代码代理直接提交的 PR,最终修复应以官方仓库后续变更为准。
参考来源
huggingface/transformers #48942
AI 工具推荐
想把多个 AI 模型放在一个入口?
GamsGo AI 集成 ChatGPT、DeepSeek、Gemini、Claude、Midjourney、Veo 等常用模型,适合写作、绘图、视频和日常 AI 工作流。
推广链接:通过此链接购买,我可能获得佣金,不影响你的价格。
这个方案解决了吗?
可以继续搜索完整报错,或查看同一工具的其他排查指南。
![[Feature]: Integer token IDs for logprobs in `/inference/v1/generate` responses (`GenerateLogProbs`)](https://www.chat-gpts.plus/wp-content/uploads/2026/10/57574-ffe54962-768x403.jpg)

