I Re-Tested My Malayalam ASR Model on Clean Data — Part 2
In July I wrote I Checked My Own ASR Dataset for Leakage — Here’s What I Found. The short version: 41.6% of the test rows in my Malayalam speech corpus had an exact-duplicate transcript in the training split. Any number I computed on that test set was partly measuring memorisation. I filtered the test set
I Re-Tested My Malayalam ASR Model on Clean Data — Part 2 Read More »








