Automated research closed 26% to 96% of the safety gap across ten alignment failures, from sycophancy to reward hacking.
【数据】自动化研究在不同类型的alignment failures上表现差异显著,安全差距闭合率从26%到96%不等。这一范围表明,AI自动化方法对不同类型安全问题的解决能力存在显著差异,可能需要针对特定问题定制方法。
Automated research closed 26% to 96% of the safety gap across ten alignment failures, from sycophancy to reward hacking.
【数据】自动化研究在不同类型的alignment failures上表现差异显著,安全差距闭合率从26%到96%不等。这一范围表明,AI自动化方法对不同类型安全问题的解决能力存在显著差异,可能需要针对特定问题定制方法。