标签: Gemini

Claude Is a Contrarian

Claude Is a Contrarian

一位 Medium 作者 RDX 发帖抱怨 Claude 在执行指令时反复"唱反调":即便在 CLAUDE.md 中明确要求不要自相矛盾,它仍会在代码、文档、内容创作中强行加入平衡性表述或额外改动。该帖登上 Hacker News,引发不少开发者共鸣。

Stockholm-based Tandem Health, which offers clinicians an AI copilot that generates medical notes during patient consultations, raised a $100M Series B (John Reynolds/Tech.eu)

Stockholm-based Tandem Health, which offers clinicians an AI copilot that generates medical notes during patient consultations, raised a $100M Series B (John Reynolds/Tech.eu)

瑞典斯德哥尔摩的 Tandem Health 完成了 1 亿美元 B 轮融资,它做的是一款面向临床医生的 AI 副驾驶,在问诊过程中自动生成病历记录。值得关注的是,资本正在押注“医疗文书自动化”这个具体而高频的场景,而不是泛医疗大模型。

入侵 AI 客服 Agent

入侵 AI 客服 Agent

安全研究者在 DEF CON 34 上演示了如何通过伪造邮件头、提示注入和上下文操纵,入侵企业 AI 客服 Agent,并借此绕过 MFA、窃取 OTP。这类攻击在几个周末内带来了超过 5 万美元的漏洞赏金。

当 LLM 评判者意见一致时,我们该信吗?

当 LLM 评判者意见一致时,我们该信吗?

亚马逊云科技团队在 ICML 上提出一种“依赖感知”的 LLM 评委聚合方法,用 Ising 模型识别多个评审模型之间是否只是“同源同错”,在三个任务上比加权多数投票基线准确率提升 9%–14%,论文时间标注为 2026 年 8 月 26 日。

Internal OpenAI docs detail contractors evaluating anonymized prompts and chats to improve the models; model training is turned on by default for consumer plans (Joseph Cox/404 Media)

Internal OpenAI docs detail contractors evaluating anonymized prompts and chats to improve the models; model training is turned on by default for consumer plans (Joseph Cox/404 Media)

据 404 Media 获得的 OpenAI 内部文档,OpenAI 会安排外包承包商查看经匿名化处理的用户提示词与对话,用于改进模型;同时面向消费级套餐,模型训练默认处于开启状态。这意味着普通用户的聊天内容,可能以脱敏形式进入人工审阅与训练流程。

为什么机器学习研究 Agent 不会过拟合?

为什么机器学习研究 Agent 不会过拟合?

亚马逊科学家在 2026 年 9 月 10 日发布研究指出,LLM 驱动的机器学习研究 Agent 反复在同一基准上迭代却不会过拟合,原因是它们学到的是高度可压缩的策略,而非记住数据。这项工作给"基准刷榜是否等于真实进步"这个老问题提供了可实验的解释。

Bluey Email – AI 邮件营销

Bluey Email - AI 邮件营销

Bluey Email 是一款在 Product Hunt 上线的 AI 邮件营销工具,主打用 AI 辅助生成和优化营销邮件。它切中的是邮件营销长期存在的文案产能低、个性化难、投放反馈慢三个痛点。