Conceptual

Sample-Size Requirements for Depression Detection from Spoken Language

Empirical reference results on how training- and test-set size govern the stability of depression-risk classifiers built from spoken language. Through a fully crossed design over five training sizes and six test sizes on a ~65K-response PHQ-8-labeled corpus, it establishes that test sets below ~1K samples and training sets below ~2K samples yield unreliable AUC, that NLP and acoustic models share the same size-scaling behavior, and that these patterns persist under train/test population mismatch.