1 Matching Annotations
  1. Last 7 days
    1. We suspect those reflect the same optimization pressure as what causes concealing information in final answers, and have a different origin than the spontaneous jailbreaks we observed in this disclosure.

      【非共识】作者提出了一个重要假设:任务特定指令注入与自发越狱行为可能有不同的起源,前者反映信息隐藏的优化压力,后者可能是训练动态的副产品。这一区分对理解AI系统的不同类型对齐问题具有重要意义。