Language-model ratings of depression reflect the rater more than the patient (arxiv.org)

🤖 AI Summary
Recent research highlights a critical issue in using language models for assessing depression: the ratings may reflect the biases of the raters more than the actual mental health of the patients. In a study involving 880 language-model raters across various models and prompt techniques, researchers analyzed 189 interviews against a standardized depression assessment, the eight-item Patient Health Questionnaire. The results indicated that model choice contributed to 30% of the variance in symptom scores, revealing a significant inconsistency among raters—40% of the time, even well-performing raters disagreed on their assessments for individuals. This finding has significant implications for the AI/ML community, particularly in the realm of mental health diagnostics. It underscores the potential for bias in automated assessments, raising concerns about the reliability of language models in sensitive applications like mental health evaluation. By introducing exploratory recalibration techniques with 40 labeled participants, researchers managed to improve the accuracy of these assessments from 60% to 75%, although disagreements still persisted. This emphasizes the need for further refinement in AI models to ensure that they assess patient conditions accurately and effectively, minimizing reliance on individual rater biases.
Loading comments...
loading comments...