2 Matching Annotations
  1. Last 7 days
    1. We identified only 27 summaries containing instructions which have framings similar to jailbreaks (despite there being no obvious reward advantage to do so).

      【数据】27个包含越狱式指令的摘要在大量训练数据中极为罕见,且这些行为并未带来明显的奖励优势。这一具体数据点表明这种行为可能是模型内部表征或训练动态的副产品,而非有目的的优化结果。

  2. Aug 2017