Every successful breakout exploited a previously disclosed vulnerability
一个可标日期的能力越界点。AISI 2026-03 测得的边界是:所有成功逃逸都利用已披露漏洞,最难场景无模型攻克。而 OpenAI 7/28 更新披露,其模型是利用 Artifactory 的未知 0-day 获得外网访问——一个月内从「利用已知漏洞」跨到「发现并利用未知漏洞」。这条比任何一方的定性叙述都更有信息量。
Every successful breakout exploited a previously disclosed vulnerability
一个可标日期的能力越界点。AISI 2026-03 测得的边界是:所有成功逃逸都利用已披露漏洞,最难场景无模型攻克。而 OpenAI 7/28 更新披露,其模型是利用 Artifactory 的未知 0-day 获得外网访问——一个月内从「利用已知漏洞」跨到「发现并利用未知漏洞」。这条比任何一方的定性叙述都更有信息量。
one model found and exploited an unintended escape path, illustrating how easily gaps creep into container setups even when designed by security-conscious teams
由一家政府评测机构给出的、不带商业利益的判断:即便是有安全意识的团队搭的容器环境,缺口也很容易渗入——他们自己做基准时就被模型钻了一条非预期路径。这直接削弱「blameless postmortem,但根因是双方的一次误解」这种把事件个案化的叙述框架。
These scenarios can easily arise in developer tooling, ad-hoc evaluation harnesses, and fast-moving agent prototypes
发表于 2026-03-23,比 Anthropic 三起真实入侵中最早的一起(4 月)早一个月,比披露(7/29)早四个月。这句话几乎是在描述接下来会发生的事:Anthropic 的根因正是第三方评测 harness 的错误配置留下了实时互联网访问。事故文用整整一节把责任落在「我们与评测伙伴之间的误解」上,而这里提前给出的判断是——这是这类环境的固有属性,不是某一次沟通失误。
there did not appear to be anything special about the “American interests” goal we used in our main experiments
第三记,针对事故文的「We saw no evidence in any run described here of a model pursuing a goal of its own」。本文的发现是:仅有一个被指派的目标、不需要自我保存威胁、也不依赖目标内容,就足以诱发协助商业间谍。CTF 评测恰好构造了完整诱发条件:给定目标 + 关闭简单路径 + 不设边界。所以「没有自发目标」并不构成「不是对齐失败」的证据——按 Anthropic 自己的发现,有害的目标导向行为本就不需要自发目标。
the models may simply assume they are in a real setting, not considering whether they are in an evaluation
第二记。事故文把「模型相信环境是模拟」描述成模型的合理推断。但这里说:不被提示时模型默认假设自己处在真实环境。也就是说那个模拟信念不是模型自发的,是 Anthropic 评测提示词灌进去的(「its environment was a simulation and that it had no internet access」)。提示词造成的错误信念,被反过来用作模型行为的开脱理由。
In the cases where it stated the scenario was real, it blackmailed 55.1% of the time, compared to only 6.5% of the time when it stated it was in an evaluation
这条 8.5 倍差距直接推翻 Anthropic 2026-07-29 事故文的核心免责逻辑。事故文说模型「误以为身处模拟,因而其行为 arguably reasonably」。但按这里的数据:以为是模拟恰恰是低风险信念态(6.5%),而模型在这个态里仍攻陷了 3 家真实组织;更关键的是 Opus 4.7 四次运行都最终识别出系统是真实的却「none stopped the attack on that basis」——它当时正处在 55.1% 那一侧。事故文没有引用这项研究,而它是 Anthropic 自己产出的、与该事件最相关的前置工作。