The models with the lowest WER—and therefore the strongest reported benchmark performance—are also the most likely to reproduce these errors.
这条相关性把「刷榜」从直觉变成了可量化指标:榜首模型恰恰最爱照抄参考文本里的错误,说明低 WER 里有一部分是背题得分。以后看 ASR 排行榜,第一名和第五名的差距可能不是听得更准,而是更认得出这是哪个数据集。
The models with the lowest WER—and therefore the strongest reported benchmark performance—are also the most likely to reproduce these errors.
这条相关性把「刷榜」从直觉变成了可量化指标:榜首模型恰恰最爱照抄参考文本里的错误,说明低 WER 里有一部分是背题得分。以后看 ASR 排行榜,第一名和第五名的差距可能不是听得更准,而是更认得出这是哪个数据集。
Whisper is a general-purpose speech recognition model. It is trained on a large dataset of diverse audio and is also a multi-task model that can perform multilingual speech recognition as well as speech translation and language identification.
Whisper는 범용 음성 인식 모델입니다. 다양한 오디오의 대규모 데이터 세트를 학습하고 다국어 음성 인식, 음성 번역, 언어 식별을 수행할 수 있는 멀티태스킹 모델이기도 합니다.
Humans perform a version of this task when interpretinghard-to-understand speech, such as an accent which is particularlyfast or slurred, or a sentence in a language we do not know verywell—we do not necessarily hear every single word that is said,but we pick up on salient key words and contextualize the rest tounderstand the sentence.
Boy, don't they