标签: AI

turns out you can get indirect prompt injection to ~0 on unseen attacks if you stack enough layers (model training + input probes + a classifier checking intent). didn’t expect that a year ago. auto mode is default in…

turns out you can get indirect prompt injection to ~0 on unseen attacks if you stack enough layers (model training + input probes + a classifier checking intent). didn't expect that a year ago. auto mode is default in...

Anthropic 开发者关系负责人 Boris Cherny 透露,通过组合模型训练、输入探测和意图分类器三层防线,Claude 已将间接提示注入攻击在未见攻击样本上的成功率压到接近 0,且该方案将随 Claude Code 自动模式在下周成为默认配置。

dear openai just make a new phone everyone wants openaiphone we can read 2-4x faster than we talk and speak openai alexa reachy hybrid is fine but pls just be a stepping stone to phone we want phone signed, everybody…

dear openai just make a new phone everyone wants openaiphone we can read 2-4x faster than we talk and speak openai alexa reachy hybrid is fine but pls just be a stepping stone to phone we want phone signed, everybody...

OpenAI 被曝正在开发一款形似甜甜圈、尺寸如冰球的人类化智能音箱,价格预计 300-400 美元;与此同时,知名开发者 Swyx 公开呼吁 OpenAI 直接做手机,认为音箱只是过渡品。硬件传闻与社区期待之间的落差,正在成为 AI 行业的新话题。

YouTube 的 AI 检测让我们栽了跟头

YouTube 的 AI 检测让我们栽了跟头

知名科普频道 Kurzgesagt 被 YouTube 的 AI 检测系统误判为“AI 生成内容”,导致最新视频流量被异常限制,成为该频道 2013 年以来表现最差的投稿。事件暴露了平台用 AI 治理 AI 内容时的误伤风险。

v2.1.226

v2.1.226

Anthropic 于 8 月 8 日发布 Claude Code v2.1.226,官方说明仍是“Bug 修复和可靠性改进”。值得关注的是,这款开发者工具正以高频小版本迭代,重心明显从功能堆叠转向稳定性打磨。

AI实验室该像危险动物的主人一样被对待吗?

AI实验室该像危险动物的主人一样被对待吗?

《经济学人》近期提出一个值得认真对待的治理思路:与其争论大模型是否具备人格或能力边界,不如先像管理危险动物主人那样,让AI实验室对其模型造成的损害承担严格责任。这个话题把AI安全讨论从“模型是否危险”转向“谁为损害负责”。