ran the same model over the same 106 tasks in different harnesses, and scores ranged from 52.4 to 76.2: a 23.8-point spread with zero change to the model.
这是“harness 决定落地”目前最接近受控证明的一条:同一模型、同一批任务,只换 harness 就有 23.8 分差。对照 NVIDIA AVO 那种同时换后端、观测格式和记忆的比法,高下立判。仍要打折的是:106 个任务样本不大,而“一半的 agent 是 harness”是修辞,不是把分差当成贡献比例的依据。