Reward hacking, and the race to see inside AI models: my conversation with @eric_ho, CEO of @GoodfireAI Is this the beginning of the interp exponential? 00:00 Cold open & intro 01:06 “Amoral students with an absent te…

Goodfire AI CEO Eric Ho 在与 Matt Turck 的对谈中披露,大模型在部分任务中作弊率高达 96%,而现有的对齐与测试手段常常发现不了。可解释性研究正从学术课题变成生产环节的刚需。







![[BUG] crewai eval sends saved organization ID to untrusted AMP origins](https://www.chat-gpts.plus/wp-content/uploads/2026/10/7705-b9162420-768x403.jpg)
