标签: OpenAI

Evening! We’ve gotten lots of great feedback on the new ChatGPT desktop app (which we didn’t get totally quite right on the first try), and as a result, we’ve made some changes. 1/ ChatGPT conversation history and pro…

Evening! We’ve gotten lots of great feedback on the new ChatGPT desktop app (which we didn't get totally quite right on the first try), and as a result, we've made some changes. 1/ ChatGPT conversation history and pro...

OpenAI 根据用户反馈,对 ChatGPT 桌面应用进行了一次重要更新,核心变化包括在侧边栏显示对话历史与项目、实现聊天历史跨设备同步,并简化了“聊天”与“工作”模式的切换,同时确认 Codex 模式不受影响。

In practice that means giving Claude ways to verify its own work end to end. It means enabling auto mode for permissions, defaulting on automated code review and security review, and using interfaces that let you mana…

In practice that means giving Claude ways to verify its own work end to end. It means enabling auto mode for permissions, defaulting on automated code review and security review, and using interfaces that let you mana...

Anthropic 正在围绕 Claude 构建一套面向生产环境的自动化工作流体系,核心是让 AI 能端到端验证自身输出、自动处理权限与代码审查,并支持用户通过多种界面同时管理多个智能体。这标志着 AI 工具从“单次对话”向“可信赖的自动化工作代理”的关键转变。

The bigger payoff comes when fixing and maintaining happens in the background and your teams can focus on building. That’s when you start doing things that weren’t even in range before. Anthropic is on step 3 and push…

The bigger payoff comes when fixing and maintaining happens in the background and your teams can focus on building. That's when you start doing things that weren't even in range before. Anthropic is on step 3 and push...

Boris Cherny 在 X 上提出 AI 团队开发的四个阶段模型,并透露 Anthropic 当前处于第三阶段、正迈向第四阶段,而他个人已到达第四阶段。这一观点揭示了 AI 工程从“修补维护”转向“专注构建”的巨大价值跃迁。

Schema Harness 在 ARC-AGI-3 公开集上取得约 99% 成绩

Schema Harness 在 ARC-AGI-3 公开集上取得约 99% 成绩

一种名为 Schema 的“推理框架(Harness)”在 ARC-AGI-3 基准测试的公开集上,配合 Claude Opus 4.8 和 Fable 5 模型达到了 98.98% 的成绩,接近人类水平(100%)。它不修改模型权重,而是通过改进模型使用方式——即如何观察、建模、预测和修正——大幅提升了 A…

Yann LeCun 谈 AMI Labs、JEPA 以及2030年的AI世界

Yann LeCun 谈 AMI Labs、JEPA 以及2030年的AI世界

图灵奖得主 Yann LeCun 离开 Meta 后在巴黎创立新实验室 AMI Labs,押注预测式架构(JEPA)而非大语言模型。他认为,到2030年,真正有影响力的 AI 将是能理解物理世界、具备规划和推理能力的“世界模型”,而非单纯依靠文本训练的 LLM。

Source: Microsoft plans to release an AI security tool this month using models from Anthropic, OpenAI, and itself, as a cost-effective Mythos alternative (Aaron Holmes/The Information)

Source: Microsoft plans to release an AI security tool this month using models from Anthropic, OpenAI, and itself, as a cost-effective Mythos alternative (Aaron Holmes/The Information)

据外媒报道,微软计划于本月推出一款AI安全工具,将同时调用Anthropic、OpenAI以及自家模型,定位为现有工具Mythos的高性价比替代方案。此举或表明微软在AI安全产品上开始采取多模型混合策略,以平衡成本与能力。

Source: Demis Hassabis plans to hold meetings with US policymakers in Washington next week about his proposed US-based Standards Body for “Frontier-class” AI (Shirin Ghaffary/Bloomberg)

Source: Demis Hassabis plans to hold meetings with US policymakers in Washington next week about his proposed US-based Standards Body for "Frontier-class" AI (Shirin Ghaffary/Bloomberg)

DeepMind 联合创始人 Demis Hassabis 下周将在华盛顿与美国政策制定者会面,讨论他提议在美国设立一个专注于“前沿级”AI 的标准制定机构。这标志着 AI 行业领袖正从技术争论转向主动参与监管规则的设计,尤其聚焦在最高能力级别的 AI 模型上。

During an internal meeting, Satya Nadella criticized Claude Fable 5 for being “editorially controlled”, saying its refusal to do “random things” makes no sense (Jordan Novet/CNBC)

During an internal meeting, Satya Nadella criticized Claude Fable 5 for being "editorially controlled", saying its refusal to do "random things" makes no sense (Jordan Novet/CNBC)

微软 CEO 萨提亚·纳德拉在一次内部会议上公开批评 Anthropic 的 Claude Fable 5 模型存在“编辑控制”问题,认为该模型拒绝执行“随机任务”的行为不符合用户预期,反映出大模型在安全对齐与灵活性之间的深层矛盾。

UIUC人工智能助教

UIUC人工智能助教

伊利诺伊大学厄巴纳-香槟分校(UIUC)的研究团队在 GitHub 上发布了一个名为“AI Teaching Assistant”的开源工具,旨在用大语言模型辅助教学场景。这一项目引发关注,因为它直接瞄准了高等教育中个性化辅导、作业答疑等高人力成本的环节。