标签: Anthropic

sorry for the radio silence folks – got sucked into extreme LLM psychosis. but i can confidently say we have crossed over into a new age of AI Engineering and we are never, ever, looking back. this isnt even EVERYTHIN…

sorry for the radio silence folks - got sucked into extreme LLM psychosis. but i can confidently say we have crossed over into a new age of AI Engineering and we are never, ever, looking back. this isnt even EVERYTHIN...

AI 领域知名研究者 Swyx 与 Latent.Space 团队花费超过 200 亿 token,将 OpenAI 的 Astra 智能体投入各类 AI 工程任务,宣称其能以每小时不到 6 美元的成本完成从模型选型到数据标注的完整工作流,标志着 AI 工程正在从“人写代码调用模型”转向“模型自主完成工程任务…

We will give one banked reset for every day you don’t have access to Astra on your paid ChatGPT plan, starting today. Team is moving mountains to give access as fast as we can. First one will land in ~ 3 hours. There…

We will give one banked reset for every day you don't have access to Astra on your paid ChatGPT plan, starting today. Team is moving mountains to give access as fast as we can. First one will land in ~ 3 hours. There...

OpenAI 产品负责人 Thibault Sottiaux 宣布,从即日起,付费 ChatGPT 用户如果当天仍无法使用 Astra,将按天累积补偿一次 reset(重置额度)。首批补偿预计约 3 小时后发放,团队正在加速向更多账号开放访问权限。

This is wild. ARC-AGI was built to resist the LLM scaling paradigm o1 despite early reasoning struggled mightily in 2024 with 18% Then ARC-AGI-3, an even harder test, launched in 2026. Frontier AI was at 0.5% And now…

This is wild. ARC-AGI was built to resist the LLM scaling paradigm o1 despite early reasoning struggled mightily in 2024 with 18% Then ARC-AGI-3, an even harder test, launched in 2026. Frontier AI was at 0.5% And now...

曾被设计来抵御大模型 Scaling Law 的 ARC-AGI 基准测试,在 2024 年让 o1 仅得 18%、2026 年更难的 ARC-AGI-3 让前沿模型得分跌至 0.5% 后,如今被名为 Astra 的模型凭借原生 harness 直接“打满”,意味着测试本身或评测范式正面临根本性挑战。