Another potential factor is that prompt injections as a concept are salient to our models: sampling from GPT-6 Astra with no input or system prompt often returns reports on prompt injections.
【方法】研究团队提出了一种有趣的方法论视角:提示注入概念可能对模型具有特殊显著性。这一观察表明,模型可能将安全概念内化为其表征的一部分,这既是风险也是理解模型行为的线索。