OpenAI says using its Responses API harness with GPT-5.6 Sol tripled its ARC-AGI-3 score and used fewer tokens, after Sol with the official harness scored 7.8% (OpenAI)

OpenAI 通过其 Responses API 的调用框架(harness)与 GPT-5.6 Sol 模型搭配,在 ARC-AGI-3 基准测试上将得分从 7.8% 提升至约 23.4%(即翻了三倍),同时使用的 token 更少。这表明 API 层面的调用策略对模型实际推理效果影响巨大,而不仅是模型本身…








