I Re-Tested My Malayalam ASR Model on Clean Data — Part 2

In July I wrote I Checked My Own ASR Dataset for Leakage — Here’s What I Found. The short version: 41.6% of the test rows in my Malayalam speech corpus had an exact-duplicate transcript in the training split. Any number I computed on that test set was partly measuring memorisation.

I filtered the test set down to 2,684 rows that the model had never seen, and I promised to report results on that clean set. This is that post.

The model is sajilck/whisper-small-malayalam, a whisper-small fine-tuned on six Malayalam speech sources. I re-ran the same CPU evaluation notebook as before, this time only on clean rows.

The result: 47.8% WER on audio the model never saw

On clean rows, the model scores 47.8% WER (95% CI 44.0–51.6%) and 13.6% CER. On the old, leaky test split it scored 42.9% WER. So leakage was flattering the model by about 5 points overall.

The overall number hides where the change happened:

wer-clean-vs-leaky

IMaSC moves the most because almost all of its test rows were leaked: only 44 of 1,656 survived filtering. The sources that were mostly clean to begin with barely move.

The 15-point gap is not all memorisation

I also scored 60 of the leaked rows: test rows whose transcript appears in training. The model got 32.5% WER on those, against 47.8% on clean rows.

It is tempting to call that 15-point gap “memorisation”. I don’t think that is fully honest. The leaked rows are mostly IMaSC, which is clean read speech and easy for the model. The clean rows lean towards harder sources. So part of the gap is simply a different mix of audio.

The fair comparison is one source against itself. IMaSC on the leaky split scored 23.9% WER. IMaSC on clean rows scored 42.5%. Same corpus, same recording conditions, one main difference: whether the model had seen the sentence before. That 18.6-point jump is the clearest evidence of memorisation I have.

Broadcast news is the real weakness

Shrutilipi is All India Radio news audio. It is the largest source in my training data, and it is also where the model does worst: 56.8% WER and 22.3% CER, the highest character error rate of any source.

That surprised me. More data usually means better results. My best guesses at why:

  • Radio audio is noisier than studio recordings, with music, compression and varied microphones.
  • News sentences are long and full of names, places and English loanwords.
  • I trained v1 with a learning rate of 1e-5, which the Vividh-ASR work suggests is far too low. A model that is under-trained will struggle most on its hardest data.

This matters because Shrutilipi is 88% of the clean test set. If I had reported one number over the whole clean set, it would have been about 55% WER and mostly a radio-news score. So the 47.8% headline samples each source evenly instead, and I report Shrutilipi separately.

Why 48% WER but only 14% CER

The model gets about 86% of characters right but only about half of words. That gap is normal for Malayalam. Words are long and agglutinative, so a single wrong letter makes the whole word wrong.

I used to think most of those word errors were just spacing: the model writing മുഖ്യ മന്ത്രി where the reference has മുഖ്യമന്ത്രി. So this time I measured it. Only 15.2% of utterances were exactly right. Ignoring spaces entirely raised that to just 20.3%. Spacing is part of the story, but not most of it.

Most errors are real recognition mistakes, often one sound:

Reference Model output
കുടുംബം ഉടുംബം
ആറാമത്തെ ആരാമത്തെ
പിന്നതാരുടേതാണ് പിന്നത് ആറുടേതാണ്

The last one is both: a split word and a wrong consonant. For anyone reading Malayalam ASR numbers, I’d suggest looking at CER alongside WER.

What’s still open, and what v2 changes

In Part 1 I said v2 training was underway. It isn’t finished. Before running it properly, I reviewed my own training script and found problems that would have broken a multi-session Kaggle run. I’ve fixed those first.

v2 changes, based on what this evaluation showed:

  • Learning rate: up from 1e-5, tested with short runs at 1e-4 and 2e-4 before the full run.
  • Cleaner training data: clips over 30 seconds and over-long transcripts removed, and repeated sentences capped. IMaSC and SMC repeat the same sentences many times.
  • Evaluation from the start: v2 is scored on the same clean set, so v1 and v2 can be compared directly.

Two honest limitations remain. The clean set has only 44 IMaSC, 32 OpenSLR 63 and 5 SMC rows, so I can’t measure read speech well. And I haven’t yet checked the training split for duplicates within itself.

Everything is public: the model and its updated card, with full per-source results and limitations. I’ll post v2 results against these numbers when it’s trained.

If you build speech datasets by combining sources, check your split before you trust your numbers. Then check again after you fix it.

Leave a Comment

Your email address will not be published. Required fields are marked *