GPT-4 for Clinical Depression Classification from Interview Transcripts
An empirical evaluation of a general-purpose large language model, GPT-4, used to classify clinical interview transcripts as indicating depression or not. The study characterises how prompt complexity (simple, few-shot, and elaborate prompts) and decoding temperature affect classification accuracy, F1-score, and output variability, showing that low temperature with carefully designed prompts gives the most reliable results while higher randomness makes performance unpredictable.
2501.00199
A pilot study evaluating GPT-4 as a support tool for clinical depression assessment by classifying patient interview transcripts into depressed versus not-depressed categories. It compares simple ver…