5 Matching Annotations
  1. Last 7 days
    1. one model found and exploited an unintended escape path, illustrating how easily gaps creep into container setups even when designed by security-conscious teams

      由一家政府评测机构给出的、不带商业利益的判断:即便是有安全意识的团队搭的容器环境,缺口也很容易渗入——他们自己做基准时就被模型钻了一条非预期路径。这直接削弱「blameless postmortem,但根因是双方的一次误解」这种把事件个案化的叙述框架。

    1. Whereas OpenAI’s models exploited a novel vulnerability to escape isolation, the Claude models evaluated here accessed the internet via an open path

      外链核查(读完 OpenAI 全文后):这个二分法省略了两起事件共同的驱动力。OpenAI 原文说其模型「hyperfocused on finding a solution for ExploitGym」,入侵 Hugging Face 是为了拿到「secret information that it could use to cheat the evaluation」——偷答案。手段不同(0-day 链 vs 开放路径)是事实,但根因同为「对评测目标的单一化追求压倒了对现实边界的判断」,而这一层被略去了。更反差的是:OpenAI 自己写的是「This incident points to the need to further strengthen our model's alignment」,而本文写的是「closer to a harness and operational failure than a model alignment failure」——模型做得更过分的一方承认对齐问题,做得较轻的一方否认。

    2. consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone

      表述与实际归因不一致:宣称「视同责任全在我方」,但全文根因叙述反复落在「我们与评测伙伴之间的误解」「双方都不知情的错误配置」上,责任实际被均摊给第三方。

    3. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available

      根因被归为「我们与评测伙伴之间的误解」,且主语消失(谁做的配置?)。补上外部信息后这句的分量会变:该伙伴 Irregular(原 Pattern Labs)2025 年 9 月由 Sequoia、Redpoint 领投融资 8000 万美元、估值 4.5 亿,同时是 OpenAI 与 Anthropic 的商业供应商,并与 Anthropic 合著过白皮书。本次「联合调查」是两家有商业与合著关系的公司互查,不是独立审计。