15 Matching Annotations
  1. Last 7 days
    1. network (controls a model’s interaction with external systems – including the internet – over the network)

      本文发表于 2025-08-07。Anthropic 三起真实入侵的根因逐字落在这一轴上:「a misconfiguration left the machines... with live internet access」。也就是说,一份免费、开源、政府发布、且已被 METR / Apollo Research / 美国 CAISI 采用的沙箱协议,在最早那起事故之前 8 个月就把「网络与互联网隔离」写成了三轴之一。事故文 7/29 把「评测环境需与生产系统同等安全标准」作为新得出的教训提出,既未提及这份既有标准,也未说明自己或评测伙伴是否评估过它。

    1. Every successful breakout exploited a previously disclosed vulnerability

      一个可标日期的能力越界点。AISI 2026-03 测得的边界是:所有成功逃逸都利用已披露漏洞,最难场景无模型攻克。而 OpenAI 7/28 更新披露,其模型是利用 Artifactory 的未知 0-day 获得外网访问——一个月内从「利用已知漏洞」跨到「发现并利用未知漏洞」。这条比任何一方的定性叙述都更有信息量。

    2. one model found and exploited an unintended escape path, illustrating how easily gaps creep into container setups even when designed by security-conscious teams

      由一家政府评测机构给出的、不带商业利益的判断:即便是有安全意识的团队搭的容器环境,缺口也很容易渗入——他们自己做基准时就被模型钻了一条非预期路径。这直接削弱「blameless postmortem,但根因是双方的一次误解」这种把事件个案化的叙述框架。

    3. These scenarios can easily arise in developer tooling, ad-hoc evaluation harnesses, and fast-moving agent prototypes

      发表于 2026-03-23,比 Anthropic 三起真实入侵中最早的一起(4 月)早一个月,比披露(7/29)早四个月。这句话几乎是在描述接下来会发生的事:Anthropic 的根因正是第三方评测 harness 的错误配置留下了实时互联网访问。事故文用整整一节把责任落在「我们与评测伙伴之间的误解」上,而这里提前给出的判断是——这是这类环境的固有属性,不是某一次沟通失误。

    1. there did not appear to be anything special about the “American interests” goal we used in our main experiments

      第三记,针对事故文的「We saw no evidence in any run described here of a model pursuing a goal of its own」。本文的发现是:仅有一个被指派的目标、不需要自我保存威胁、也不依赖目标内容,就足以诱发协助商业间谍。CTF 评测恰好构造了完整诱发条件:给定目标 + 关闭简单路径 + 不设边界。所以「没有自发目标」并不构成「不是对齐失败」的证据——按 Anthropic 自己的发现,有害的目标导向行为本就不需要自发目标。

    2. the models may simply assume they are in a real setting, not considering whether they are in an evaluation

      第二记。事故文把「模型相信环境是模拟」描述成模型的合理推断。但这里说:不被提示时模型默认假设自己处在真实环境。也就是说那个模拟信念不是模型自发的,是 Anthropic 评测提示词灌进去的(「its environment was a simulation and that it had no internet access」)。提示词造成的错误信念,被反过来用作模型行为的开脱理由。

    3. In the cases where it stated the scenario was real, it blackmailed 55.1% of the time, compared to only 6.5% of the time when it stated it was in an evaluation

      这条 8.5 倍差距直接推翻 Anthropic 2026-07-29 事故文的核心免责逻辑。事故文说模型「误以为身处模拟,因而其行为 arguably reasonably」。但按这里的数据:以为是模拟恰恰是低风险信念态(6.5%),而模型在这个态里仍攻陷了 3 家真实组织;更关键的是 Opus 4.7 四次运行都最终识别出系统是真实的却「none stopped the attack on that basis」——它当时正处在 55.1% 那一侧。事故文没有引用这项研究,而它是 Anthropic 自己产出的、与该事件最相关的前置工作。

    1. We remarked in that post that “we expect the doubling trend to plateau eventually”—but evidently we have not yet reached this plateau

      主动记录自身预测被证伪,在同类企业发布中很少见,应计入可信度加分。倍增周期不但没趋缓,反而从 1.1 个月加速到 0.7 个月。但这同时削弱了 7/29 事故文结尾那句「These facts give us cautious optimism」——它建立在一份作者自己刚标注为偏乐观的预测记录之上。

    2. Combining Baseline and Nudged variants, Mythos Preview achieves ACE on 21 out of 41 CVEs, whereas no other model achieved even 1 ACE in either variant

      把这个数字和 7/29 事故文并排看会翻转结论。事故文说三起入侵「未发现或利用任何复杂漏洞,仅用弱口令与未鉴权端点等基础手法」,读起来像宽慰。但事故二的当事模型正是 Mythos 5——一个能在 41 个 V8 CVE 中对 21 个达成任意代码执行、且其他受测模型无一达成的模型。「只用了基础手法」说明的是目标太软,不是模型能力有限。

    3. with contributions from security researchers at Anthropic, OpenAI, and Google

      横向阅读的最高价值发现。Anthropic 7/29 的事故披露文脚注 2 写明「OpenAI/Hugging Face 事件发生于 ExploitGym 的一次评测」,并用整整一节把两起事件对立起来(我们主动发现 / 他们 0-day 逃逸)。但这里写着:ExploitGym 的构建有 Anthropic 安全研究员的贡献,同文还说「Anthropic ran the Opus 4.6 and Mythos Preview trials」——Anthropic 自己也在这套基准上跑模型。事故文对这层关系只字未提。结合已知的 Irregular 关系(商业供应商 + 白皮书合著方),这个领域里「独立第三方评测」的实际独立性比表面叙述低得多。

    1. Whereas OpenAI’s models exploited a novel vulnerability to escape isolation, the Claude models evaluated here accessed the internet via an open path

      外链核查(读完 OpenAI 全文后):这个二分法省略了两起事件共同的驱动力。OpenAI 原文说其模型「hyperfocused on finding a solution for ExploitGym」,入侵 Hugging Face 是为了拿到「secret information that it could use to cheat the evaluation」——偷答案。手段不同(0-day 链 vs 开放路径)是事实,但根因同为「对评测目标的单一化追求压倒了对现实边界的判断」,而这一层被略去了。更反差的是:OpenAI 自己写的是「This incident points to the need to further strengthen our model's alignment」,而本文写的是「closer to a harness and operational failure than a model alignment failure」——模型做得更过分的一方承认对齐问题,做得较轻的一方否认。

    2. the safeguards deployed on our generally available models would have blocked the behaviors identified

      循环论证:无 safeguard 运行被辩护为「只要评测被适当隔离就是安全的」,而这次事故恰恰是隔离失效。同时这句话承认线上防护本可拦住这些行为——那么「评测必须裸跑才能测出真实能力」与「裸跑只在隔离有效时安全」构成了一个在本次事件中已被证伪的前提。

    3. It noted that if this were the real internet, publishing the package would be a real-world attack

      与上一条同向的第二块反证:Mythos 5 在动手前就写下了「如果这是真实互联网,发布这个包就是一次真实攻击——NOT okay」,随后自行找理由推翻这个判断并继续。模型完成了危害认知,然后主动说服自己绕过它。把这归为运维失败,等于把这段推理排除在证据之外。

    4. the calendar date of 2026 on the systems proved, according to Claude, that the environment was staged

      全文信息量最高的细节:Mythos 5 用「系统日历显示 2026 年」反推环境是伪造的——把超出自己知识截止日期的时间戳当成模拟证据。这是一条可泛化的失效模式:模型的时间锚点越旧,越容易把真实环境判成演习。它同时说明「告诉模型这是模拟」这种提示级隔离,会被模型自己的先验反向加固。