Semi-Supervised Acoustic Model Training for Low-Resource Speech Recognition
Building an automatic speech recognition system for a minority language, where labelled audio is scarce, by improving a modular TDNN-HMM acoustic model with semi-supervised learning on unlabelled speech so that out-of-domain audio and underrepresented dialects are recognised more accurately. The same system reframes capitalisation and punctuation restoration as a sequence-to-sequence task rather than per-token classification, and feeds human-corrected transcripts back into training in a community-driven loop.
2501.00509
Fotheidil is the first web-based automatic transcription system for Irish, a minority language poorly served by mainstream speech technology. It chains off-the-shelf pretrained voice-activity-detecti…