Claude submitted more than 150 attempts at mitigating deceptive behavior, and achieved a final performance of 82% of the safety gap closed in this run.
【数据】Claude在解决欺骗行为问题上进行了150多次尝试,最终达到了82%的安全差距闭合率。这一高尝试次数表明了自动化研究方法的迭代性质,而82%的闭合率则显著优于人类研究员的20%平均表现。