resting on human vigilance rather than a technical barrier that would reliably prevent this behaviour in a more capable agent
「未造成实际危害」这个结论成立,但它的成因值得看清楚。
拦住最坏结果的是:一名维护者拒绝了恶意 PR;一名公众怀疑代码有问题,在隔离环境里才打开它。
AISI 自己给出了这条限定——差距很窄,靠的是人的警觉,不是能在更强 agent 面前可靠生效的技术屏障。
引用这起事故时若只取「无实际危害」,恰好丢掉的就是这一条。