varying a model's identity boundary
this immediately reminds me of the self-other overlap AI Safety study where researchers switched out "you" and "I" in the training data, which drastically reduced deception (test annotation)
varying a model's identity boundary
this immediately reminds me of the self-other overlap AI Safety study where researchers switched out "you" and "I" in the training data, which drastically reduced deception (test annotation)