Sample-Size Requirements for Depression Detection from Spoken Language
Empirical reference results on how training- and test-set size govern the stability of depression-risk classifiers built from spoken language. Through a fully crossed design over five training sizes and six test sizes on a ~65K-response PHQ-8-labeled corpus, it establishes that test sets below ~1K samples and training sets below ~2K samples yield unreliable AUC, that NLP and acoustic models share the same size-scaling behavior, and that these patterns persist under train/test population mismatch.
2501.00617
An empirical study quantifying how training- and test-set size affect the stability of depression-risk classifiers built from spoken language. Using a proprietary corpus of ~65K PHQ-8-labeled spoken …