In Part 1 I found that my Malayalam speech model’s test set leaked into its training data. In Part 2 I re-tested the model on clean data only, and promised a better v2.
This post reports two things. First, v2 is trained, and on the same clean test set it cuts word errors by 12 points and character errors almost in half. Second, while checking that result, I found a bug in my own evaluation. It had been making both models look worse than they are, including the numbers I published in Part 2.
The new model is sajilck/whisper-small-malayalam-v2. Same size as v1, same speed, trained entirely on free Kaggle GPUs.
You can try v2 yourself in a Malayalam voice assistant demo that runs entirely on a free CPU.
The result
I scored v1 and v2 on exactly the same 231 clean clips, with the same notebook and settings. v2 brings word error rate down from 43.5% to 31.2% and character error rate from 7.3% to 4.0%. The 95% confidence intervals don’t overlap, so this is not noise.
The biggest gains are on the hardest audio: crowdsourced phone recordings (CommonVoice) and radio news (Shrutilipi). The share of sentences transcribed perfectly nearly doubled, from 15% to 28%. At 4% CER, v2 gets about 96 of every 100 characters right.
What changed in training
v2 starts again from the original whisper-small, not from v1. Three changes made the difference.
A higher learning rate. v1 trained at 1e-5. The Vividh-ASR work suggested that was far too cautious. Before committing a full run, I tried 300 steps each at 1e-4 and 2e-4. Both were stable and scored within noise of each other, so I chose the safer 1e-4 for the full 10,000 steps.
Cleaner training data. Part 1 showed that IMaSC and SMC read the same fixed sentences again and again. I capped every sentence at three copies, which removed 16,771 rows. I also removed 703 rows whose transcripts were too long for Whisper’s decoder. That left 69,437 training rows.
Evaluating from the start. Every 1,000 steps, the training script scored a fixed sample of the clean test set. I could watch word error rate fall from 54% to 31% and flatten out near the end, so I knew the final checkpoint was the one to keep.
The whole run took about nine hours in one free Kaggle session. Getting there took a few rounds of fixing my own training script first, which is a story for another day.
The second bug: my evaluation was cutting off long clips
The first time I scored v2 in my evaluation notebook, it got 11% character error rate. The training log, on similar clean clips, said 4%. Two honest measurements of the same model should not disagree by that much.
I first suspected the library version, since training and evaluation used different versions of transformers. Re-running with the training version changed nothing.
The answer was sitting in the list of worst transcripts. One clip ended mid-word: …മുഖംമൂടി ധരിച്ച നൂറോളം പ. The model hadn’t failed. My notebook had stopped it.
The notebook limited each transcript to 225 tokens, a setting I’d copied without thinking about Malayalam. Whisper’s tokenizer breaks Malayalam into about 25 tokens per second of speech, so 225 tokens covers only about 9 seconds. Anything longer was cut off, and every missing word counted as an error. Long clips suffered most, and Shrutilipi’s radio news sentences are the longest.
Key insight: a default setting tuned for English can quietly break evaluation in another script.
I raised the limit to Whisper’s maximum, added a check that counts any transcript hitting it, and re-scored both models. That check now reads zero for both.
This changes numbers I published in Part 2. The corrected v1 figures are:
| v1, clean test set | Published in Part 2 | Corrected |
|---|---|---|
| Word error rate | 47.8% | 43.5% |
| Character error rate | 13.6% | 7.3% |
| Shrutilipi (WER / CER) | 56.8% / 22.3% | 48.2% / 9.4% |
| IMaSC on clean rows | 42.5% | 39.5% |
Two conclusions in Part 2 change. Broadcast news was not v1’s weakest source; crowdsourced CommonVoice audio was. And the leakage gap on IMaSC was about 16 points, not 19. The main finding stands: leaked test data made the model look better than it was. I’ve added a correction note to Part 2 and to the v1 model card.
What’s still open
v2 is better, but not finished. Three things I know are still weak:
- Speakers may overlap between train and test. My filter removes repeated sentences, not repeated voices. On truly new speakers and microphones, accuracy will probably be lower. Both models share this problem, so the comparison is fair, but the absolute numbers may be optimistic.
- Crowdsourced phone audio is the hardest. CommonVoice is still the weakest source at 38% WER.
- It is not real-time on a CPU. On 2 CPU cores it takes about twice as long as the audio itself. The next step is a whisper.cpp build of v2 for my CPU-only Malayalam voice assistant (now live in the demo).
The model, its card with every number in this post, and the evaluation notebook are all public. If you work on speech for Indian languages, I’d value one piece of advice: what’s the best way you’ve found to build a test set that is clean of both repeated sentences and repeated speakers?