On LibriSpeech, some of the strongest benchmark-performing models reproduced masked numbers in roughly 30–40% of examples, even though the number itself had been removed.
把数字从音频里静音掉,模型照样把原文的数字填回来,这是本文最干净的证伪设计——正确答案在声学上根本不存在,答对只可能来自记忆。这种「构造一个不可能答对的题」的思路,比事后统计污染率有力得多,值得搬到其他模态的评测里。