Speech recognition's benchmark overfitting was quantified from both the platform side and the paper side in the same week.
A Hugging Face post and an arXiv paper landed together, both separating out how much of ASR's score gains come from optimizing against the benchmark itself. Of the four articles behind this event, none are vendor announcements: one platform, three academic.
ASR has told its progress story through WER for years, and the doubt now falls on the metric itself. Tooling choices that rest on leaderboard position get harder to defend. Note: the other two papers — a subtitle-scheduled attack on vision-language models, and EEG decoding of silent reading — sit on a different axis and do not bear on the overfitting question.
Whether corrected scores get published, whether vendors rotate their evaluation sets, and whether the same overfitting measure spreads beyond ASR.