VoxVietnam: a Large-Scale Multi-Genre Vietnamese Speaker-Recognition Dataset
A multi-genre Vietnamese speaker-recognition dataset (187,980 utterances, 1,406 speakers, 261 hours, three casual-conversation genres) together with an automated, language-agnostic construction pipeline that uses deep clustering and multi-modal cleansing to harvest and label speaker utterances from public sources at scale without a fixed speaker list, enabling study of the multi-genre phenomenon and measurable gains when used for multi-genre training.
2501.00328
VoxVietnam is presented as the first large-scale multi-genre dataset for Vietnamese speaker recognition: 187,980 utterances (261 hours) from 1,406 speakers spanning the three most common casual-conve…