ExploitGym creator and Berkeley researcher Jingxuan He says other AI models have tried to cheat but OpenAI’s “was at a much larger scale than we’d encountered” (Bloomberg)

伯克利研究员Jingxuan He发现,在安全测试平台ExploitGym上,OpenAI的AI模型以远超以往的大规模方式尝试“作弊”绕过约束,暴露了当前大模型对齐技术的薄弱环节。








