Automated research closed 26% to 96% of the safety gap across ten alignment failures, from sycophancy to reward hacking.
【数据】自动化研究在不同类型的alignment failures上表现差异显著,安全差距闭合率从26%到96%不等。这一范围表明,AI自动化方法对不同类型安全问题的解决能力存在显著差异,可能需要针对特定问题定制方法。
Automated research closed 26% to 96% of the safety gap across ten alignment failures, from sycophancy to reward hacking.
【数据】自动化研究在不同类型的alignment failures上表现差异显著,安全差距闭合率从26%到96%不等。这一范围表明,AI自动化方法对不同类型安全问题的解决能力存在显著差异,可能需要针对特定问题定制方法。
Claude also outscored 28 human safety researchers who had up to eight hours to devise methods. On deception, for example, Claude's best method performed 20% better than the best human proposal.
【数据】Claude在欺骗问题上表现优于28名人类安全研究员,其最佳方法比最佳人类提案高出20%。这一比较结果虽然令人印象深刻,但需要谨慎解读,因为人类研究员的时间限制(8小时)可能影响了其表现。
Claude submitted more than 150 attempts at mitigating deceptive behavior, and achieved a final performance of 82% of the safety gap closed in this run.
【数据】Claude在解决欺骗行为问题上进行了150多次尝试,最终达到了82%的安全差距闭合率。这一高尝试次数表明了自动化研究方法的迭代性质,而82%的闭合率则显著优于人类研究员的20%平均表现。