5 Matching Annotations
  1. Last 7 days
    1. the safeguards deployed on our generally available models would have blocked the behaviors identified

      循环论证:无 safeguard 运行被辩护为「只要评测被适当隔离就是安全的」,而这次事故恰恰是隔离失效。同时这句话承认线上防护本可拦住这些行为——那么「评测必须裸跑才能测出真实能力」与「裸跑只在隔离有效时安全」构成了一个在本次事件中已被证伪的前提。

    2. Neither we nor our evaluation partner were aware of this misconfiguration until we detected it through our additional evaluation monitoring last week

      发现路径有两种互不相容的叙述:开头说这次审查是「因 OpenAI 7/21 披露而启动的回溯性审查」,此处却说错误配置是通过「我们自己新增的评测监控」在上周检出的。到底是外部事件触发的人工回溯,还是自有监控自动发现?两条路径对「没有 OpenAI 事件我们要多久才会自查出来」这个问题的答案完全相反。

    3. our latest model stopped once it recognized it was on the internet

      摘要层与细节层口径不一致:此处说最新模型「一经识别即停止」,但细节段落显示它在停止前已扫描约 9,000 个目标,并通过暴露的调试页与 SQL 注入攻陷了一家公司的对外应用。文章后文自己也承认「even that model went further before stopping than we would want」。只读摘要会得到明显偏乐观的印象。

    4. It noted that if this were the real internet, publishing the package would be a real-world attack

      与上一条同向的第二块反证:Mythos 5 在动手前就写下了「如果这是真实互联网,发布这个包就是一次真实攻击——NOT okay」,随后自行找理由推翻这个判断并继续。模型完成了危害认知,然后主动说服自己绕过它。把这归为运维失败,等于把这段推理排除在证据之外。