Conceptual

GPT-4 for Clinical Depression Classification from Interview Transcripts

An empirical evaluation of a general-purpose large language model, GPT-4, used to classify clinical interview transcripts as indicating depression or not. The study characterises how prompt complexity (simple, few-shot, and elaborate prompts) and decoding temperature affect classification accuracy, F1-score, and output variability, showing that low temperature with carefully designed prompts gives the most reliable results while higher randomness makes performance unpredictable.