2 Matching Annotations
  1. Last 7 days
    1. evaluations like Petri are only proxies for real-world misalignment, and we did not test whether alignment gains persist after extensive RL training on other tasks.

      【局限】研究人员明确指出,Petri等评估工具仅是现实世界中错位行为的代理指标,而非完全等效的替代。此外,研究未验证对齐改进在其他任务上进行大量RL训练后是否能保持稳定,这限制了结果的实用价值。

    2. the alignment failures studied were narrow compared to those in production (e.g., we didn't measure political biases), some failures may occur so rarely or emerge so recently that no benchmark exists to measure them.

      【局限】研究承认其评估的 alignment failures 相对狭窄,未能涵盖生产环境中的所有问题,如政治偏见。此外,某些罕见或新出现的问题可能缺乏合适的评估基准,这限制了研究的全面性和适用性。