2 Matching Annotations
  1. Last 7 days
    1. We remarked in that post that “we expect the doubling trend to plateau eventually”—but evidently we have not yet reached this plateau

      主动记录自身预测被证伪,在同类企业发布中很少见,应计入可信度加分。倍增周期不但没趋缓,反而从 1.1 个月加速到 0.7 个月。但这同时削弱了 7/29 事故文结尾那句「These facts give us cautious optimism」——它建立在一份作者自己刚标注为偏乐观的预测记录之上。

    2. Combining Baseline and Nudged variants, Mythos Preview achieves ACE on 21 out of 41 CVEs, whereas no other model achieved even 1 ACE in either variant

      把这个数字和 7/29 事故文并排看会翻转结论。事故文说三起入侵「未发现或利用任何复杂漏洞,仅用弱口令与未鉴权端点等基础手法」,读起来像宽慰。但事故二的当事模型正是 Mythos 5——一个能在 41 个 V8 CVE 中对 21 个达成任意代码执行、且其他受测模型无一达成的模型。「只用了基础手法」说明的是目标太软,不是模型能力有限。