This suggests that the models can identify which dataset an audio sample comes from and select the spelling convention that benchmark expects, even though both forms sound identical.
最关键的一层不是模型记住了答案,而是它先认出「我正在被哪个基准测试」,再切换成那套书写规范。这意味着污染检测不能只查文本重合,还得防声学指纹;同域但训练截止之后新采的音频一喂进去,这种行为大多消失,正好反证了它的存在。