快速结论:当 Chat Completions 请求携带 tools 且走 LiteLLM 的 Responses API 桥接(如 gpt-5.4+ / gpt-6-luna),模型在同一轮里既输出说明文字又调用工具时,非流式响应会被拆成两个 choice,导致客户端只读 choices[0] 而丢失工具调用。优先排查你使用的 LiteLLM 版本是否已包含 #44346 的修复。
适用环境:Issue 中确认:LiteLLM 1.102.1(proxy)、1.104.0、1.104.2、v1.105.0-rc.3 仍存在;main 已修复;相关 openai 2.54.0;涉及 Google ADK 的 LiteLlm 客户端。Issue 未提供操作系统、Python、CUDA、显卡信息。
最快修复方案:升级到包含 #44346 修复的 LiteLLM 版本(即 main 上的合并形态 if accumulated_tool_calls and choices: 折叠进 choices[-1]);若暂时无法升级,改用 stream: true 的流式路径规避。
注意事项:Issue 明确说明该修复未包含在 v1.104.2 和 v1.105.0-rc.3 两个最新 tag 中,属于发布时序问题;本问题仅影响非流式、走 Responses-to-Chat 桥接、且同一轮同时返回 message item 与 function_call item 的请求,直连 Chat Completions 与流式路径不受影响。
问题场景
用户在 LiteLLM(proxy 或 SDK)中,向被路由到 Responses API 桥接的模型(如 gpt-5.4+、gpt-6-luna)发起 非流式 POST /chat/completions 请求,附带 tools 和 reasoning_effort,且 system prompt 要求模型先说明再调用工具。当模型在同一轮里输出一段文字后又调用工具时,_convert_response_output_to_choices(litellm/completion_extras/litellm_responses_transformation/transformation.py)会为 n=1 的请求返回两个 choice,导致 Chat Completions 客户端(例如 Google ADK 的 LiteLlm,google/adk/models/lite_llm.py)只读取 choices[0],把这一轮当成最终答案,工具调用被静默丢弃。
报错原文
[Bug]: Responses bridge returns a narrated tool call as TWO chat choices (text in choices[0] with finish_reason stop, function_call in choices[1]) so chat clients lose the tool call
Multiple choices found in response but only the first one will be used
{"choices": [
{"index": 0, "finish_reason": "stop",
"message": {"role": "assistant", "content": "Voy a consultar el Código del Trabajo chileno para localizar la definición legal de «contrato de trabajo».", "tool_calls": null}},
{"index": 1, "finish_reason": "tool_calls",
"message": {"role": "assistant", "content": null, "tool_calls": [{"type": "function", "function": {"name": "buscar_legislacion", "arguments": "{\"consulta\": \"...\"}"}}]}}
]}
原因分析
最可能的原因是转换函数 _convert_response_output_to_choices 在处理「文本 message item」与「function_call item」共存时,把累积到的工具调用追加成了独立的新 choice,而不是折叠合并进 choices[-1]。Issue 正文指出,工具调用累加器旁的注释(“a choice per call would hide every call after choices[0] from chat clients”)说明同类问题此前已针对「多个工具调用」修复,但没有覆盖「文字 + 工具调用」这一组合。结果 choices[0] 只有文本、finish_reason: "stop"、tool_calls: null,客户端只读第一个 choice 时工具调用即丢失。流式路径 translate_responses_chunk_to_openai_stream 把文本 delta 与工具调用 delta 放在同一个 choice 上,因此不受影响,这也让非流式 bug 更不易被察觉。
环境排查
- 确认 LiteLLM 版本:Issue 中 1.102.1(proxy)、1.104.0、1.104.2、v1.105.0-rc.3 仍可复现;
main已修复。 - 确认
openai依赖版本(Issue 复现环境为 2.54.0)。 - 确认请求是否为非流式(
stream未开启或为false)。 - 确认请求是否经过 Responses API 桥接路由(如
gpt-5.4+ /gpt-6-luna等模型)。 - 确认请求是否同时携带
tools和reasoning_effort。 - 确认客户端是否只读取
choices[0](如 Google ADK 的LiteLlm)。 - Issue 未提供操作系统、Python、CUDA、显卡信息,无需额外核对。
解决步骤
- 先确认当前命中版本:检查
litellm/completion_extras/litellm_responses_transformation/transformation.py。若看到如下「后置追加」形态,说明仍是未修复版本:if accumulated_tool_calls: msg = Message( content=None, tool_calls=accumulated_tool_calls, ... ) choices.append(Choices(message=msg, finish_reason="tool_calls", index=index)) - 升级到包含 #44346 修复的版本。合并后的形态为
if accumulated_tool_calls and choices:折叠进choices[-1],即工具调用并入最后一个已有 choice。 - 注意不要只根据版本号判断:Issue 评论确认 v1.104.2 和 v1.105.0-rc.3 两个最新 tag 仍为修复前代码,属于发布时序问题,不是重新打开该 Issue。
- 若无法立即升级,可优先尝试改用流式路径(
stream: true)作为规避手段,因为流式转换把文本与工具调用放在同一个 choice 上。 - 若你的桥接场景只涉及多
message变体,可参考 #37299;与 #41123 的 message-item merge 相关,但该 PR 保留尾部的 tool_calls choice,单独使用不能修复本 case。
验证方法
用 Issue 提供的离线段级复现(无需网络或凭证):构造包含 ResponseReasoningItem、ResponseOutputMessage(文本为 “Let me check the weather.”)和 ResponseFunctionToolCall(get_weather)的 output_items,调用 LiteLLMResponsesTransformationHandler._convert_response_output_to_choices(output_items)。修复前输出为:
2
0 stop 'Let me check the weather.' None
1 tool_calls None [get_weather]
修复后应只返回 一个 choice,其 content="Let me check the weather."、tool_calls 内含 get_weather、finish_reason="tool_calls"。也可在 proxy 上重发原始非流式请求,确认不再出现两个 choice、工具调用不再丢失。
参考来源
AI 工具推荐
想把多个 AI 模型放在一个入口?
GamsGo AI 集成 ChatGPT、DeepSeek、Gemini、Claude、Midjourney、Veo 等常用模型,适合写作、绘图、视频和日常 AI 工作流。
推广链接:通过此链接购买,我可能获得佣金,不影响你的价格。
这个方案解决了吗?
可以继续搜索完整报错,或查看同一工具的其他排查指南。
![[Security]: litellm PyPI package (v1.82.7 + v1.82.8) compromised — full timeline and status](https://www.chat-gpts.plus/wp-content/uploads/2026/10/24518-7eaf0753-768x403.jpg)
![[Bug]: `litellm_settings.max_budget` ignores `budget_duration`; global proxy budget is a hardcoded trailing-30-day cap](https://www.chat-gpts.plus/wp-content/uploads/2026/10/31292-3ad334c3-768x403.jpg)
