快速结论:在 LiteLLM 1.100.0 上以 stream=True 调用时,重复的失败通知会绕过 has_run_logging 的去重保护,导致每次失败都触发一次 failure callback;若启用了 s3_v2,同一对象 key 会被重复入队并调度多次并发上传。优先排查 Logging.has_run_logging 对 async_failure/sync_failure 的 stream 提前返回逻辑。
适用环境:Issue 已确认:LiteLLM litellm[proxy]==1.100.0,Python 3.13,pytest,干净虚拟环境;行为在 1.102.1 上复现,且 main 分支 e484a7c 代码未变。测试不依赖 provider 调用、AWS 凭据或仓库 fixture。
最快修复方案:暂无确认的一步修复方案。Issue 建议的最小改动是:把 Logging.has_run_logging 中的 stream 豁免限制为仅 async_success 和 sync_success,其余事件(含失败)继续走既有的 per-event、per-Logging-instance 去重标记;该限制方案在隔离环境中已验证 4 项测试全部通过。
注意事项:不要全局去重 S3 key 或 request ID,因为不同尝试和请求仍需各自独立的记录。Issue 中的复现只证明了重复上传调度,并未模拟或断言 S3 限流;修复限制在失败回调路径,成功流式 chunk 必须继续到达 stream callback。
问题场景
用户在 LiteLLM 1.100.0 上启用 stream=True,并在同一个 Logging 实例上触发多次失败通知时,每个 configured failure callback 都会被重复调用。当同时使用 s3_v2 logger 时,三次通知会为同一个对象 key 排入三条记录,batcher 会调度三次并发上传;而 stream=False 的相同场景只产生一次上传。一个典型触发路径是 Router.async_get_available_deployment:每次失败的 deployment 选择尝试都会对请求共享的 logging 对象调度 logging_obj.async_failure_handler(...),因此即使没有发生任何 provider 请求,重试选择失败也会重复触发 failure callback。
报错原文
[Bug]: Streaming failure callbacks bypass duplicate logging guard and enqueue repeated S3 uploads
# 问题代码位置
if self.stream is not None and self.stream is True:
# Ignore check on stream, as there can be multiple chunks
return
self.model_call_details[f"has_logged_{event_type}"] = True
# 测试结果
Current result: 2 failed, 2 passed
Both streaming cases fail with `assert 3 == 1`
After applying only the suggested success-event restriction: 4 passed
# 行为确认输出
stream=False: 2 failure notifications -> 1 failure-callback dispatch # guard works
stream=True : 2 failure notifications -> 2 failure-callback dispatches # guard bypassed
原因分析
最可能的原因是 Logging.has_run_logging 对所有事件类型都执行了流式提前返回。should_run_logging 读取的 has_logged_{event_type} 标记由 has_run_logging 设置,但当 stream is True 时,函数在设置该标记之前就 return 了,导致标记永远不会被写入,后续每次重复的失败通知都会通过去重检查。chunk 豁免的本意(“可能有多个 chunk”)只适用于成功事件;失败通知每次尝试只应产生一次,因此该豁免被错误地应用到了 async_failure/sync_failure。Issue 中还指出一个二阶放大效应:Router.async_get_available_deployment 会在共享 logging 对象上按失败选择次数调度 failure handler,因此一个包含 K 个 deployment 的 fallback 链即使没有 provider 请求,也会 fan out 出 K 个重复失败回调;配合 s3_v2 就是同一 key 的 K 条队列记录。
环境排查
- 确认 LiteLLM 版本:Issue 在
litellm[proxy]==1.100.0上复现,并确认 1.102.1 行为一致、main分支e484a7c代码未变。 - 确认 Python 版本:Issue 使用 Python 3.13。
- 确认是否启用
stream=True:非流式路径的去重保护正常,流式路径才会绕过。 - 确认是否使用
s3_v2logger:该场景下重复 callback 会转化为同一 key 的重复入队和并发上传调度。 - 确认是否存在 Router 失败选择重试:
Router.async_get_available_deployment的失败尝试会在共享 logging 对象上重复调度 failure handler。 - 确认测试环境是否干净:Issue 使用独立虚拟环境、pytest,无需 AWS 凭据或 provider 调用。
解决步骤
- 在
litellm/litellm_core_utils/litellm_logging.py中定位Logging.has_run_logging(Issue 指出位于 L2065-2075)以及should_run_logging(L2052-2063)。 - 检查
has_run_logging中针对self.stream is not None and self.stream is True的提前返回分支,确认它当前对所有event_type生效。 - 按 Issue 建议的最小改动,将该 stream 豁免限制为仅
async_success和sync_success;对async_failure/sync_failure继续执行self.model_call_details[f"has_logged_{event_type}"] = True,保留既有的 per-event、per-Logging-instance 去重保护。 - 不要全局去重 S3 key 或 request ID,避免丢失不同尝试和请求各自的记录。
- 在隔离环境中运行 Issue 提供的回归测试
test_stream_failure_once.py,覆盖stream=False/True与concurrent=False/True组合。 - 如无法立即修改库代码,可优先尝试确认是否必须开启
stream=True;Issue 已确认非流式路径的去重保护正常。
验证方法
在干净虚拟环境中按 Issue 提供的步骤安装依赖并运行回归测试:LITELLM_LOCAL_MODEL_COST_MAP=True /tmp/litellm-duplicate-failure/bin/python -m pytest -q test_stream_failure_once.py。修复前预期结果为 2 failed, 2 passed,两个流式用例以 assert 3 == 1 失败,两个非流式对照组通过;按建议仅限制成功事件的 stream 豁免后,预期结果为 4 passed。同时确认成功流式 chunk 仍能到达 stream callback,且 sync/async failure callback 路径保持独立。
参考来源
AI 工具推荐
想把多个 AI 模型放在一个入口?
GamsGo AI 集成 ChatGPT、DeepSeek、Gemini、Claude、Midjourney、Veo 等常用模型,适合写作、绘图、视频和日常 AI 工作流。
推广链接:通过此链接购买,我可能获得佣金,不影响你的价格。
这个方案解决了吗?
可以继续搜索完整报错,或查看同一工具的其他排查指南。


