Spatiotemporal Deep Learning for Robust Depression Severity Classification in Clinical Settings
Abstract
Depression is among the most prevalent and disabling mental health disorders worldwide, yet its clinical diagnosis still depends largely on subjective self-report instruments and expert judgment, which are difficult to scale and prone to inter-rater variability. Automated, objective assessment of depression severity from observable behavior could substantially extend access to screening, but most existing video-based approaches rely on static images or shallow temporal modelling and are validated on small or demographically homogeneous datasets, limiting their clinical generalizability. We propose a deep learning model that classifies depression severity from facial videos. It is evaluated on a real clinical dataset of 177 participants, including healthy controls and patients with mild, moderate, and severe depression. The framework couples a ConvNeXt-Tiny spatial backbone, enhanced with Squeeze-and-Excitation and Convolutional Block Attention modules, with a 256-unit Long Short-Term Memory (LSTM) network applied frame-wise through a time-distributed design to model the temporal dynamics of facial behavior. Severe class imbalance is mitigated through weighted categorical cross-entropy and targeted data augmentation, and patient-level stratification is enforced so that no participant contributes frames to both the training and evaluation splits. Under participant-level five-fold cross-validation, the proposed model achieves strong classification performance across all four severity levels, attaining 82.7% accuracy, a macro-averaged F1-score of 78.9%, and a macro-averaged AUC of 0.917 on held-out participants, and outperforming the strongest single-backbone baseline (Swin Transformer, 71.2% accuracy) by 11.5 percentage points. We further examine the validation protocol and the role of strict patient-level separation in interpreting these results, situating them against the leakage risk inherent in densely sampled video. The proposed framework offers a lightweight, behaviorally grounded, and reproducible approach to automated depression severity assessment, with direct relevance to scalable clinical screening.
Keywords
Depression severity classification; spatiotemporal deep learning; facial video analysis; ConvNeXt; LSTM; attention mechanisms; affective computing; clinical mental health screening.