Cross-Modal Benchmarking of Acoustic Prosody and Ventral Striatal BOLD for Depression-Related Anhedonia Classification: A Pre-Registered Study with the ClinicalWhisper Pipeline
Can computational analysis of a brief voice recording classify depression-related anhedonia as effectively as task-based fMRI? We address this question through a pre-registered cross-modal benchmarking study (osf.io/bsvrj) that evaluates two independent classification pipelines against depression-related anhedonia operationalized via self-report. Anhedonia-- the diminished capacity to experience pleasure or motivation to pursue rewards--is a transdiagnostic marker of reward-system dysfunction that predicts treatment resistance in major depressive disorder, yet current assessment requires either expensive functional neuroimaging or subjective self-report scales, neither of which scales to routine screening. We benchmarked two independent classification pipelines: Stream A extracted 88 acoustic prosody features from DAIC-WOZ clinical interviews (n = 142) using the ClinicalWhisper pipeline (Whisper Large-v3, pyannote diarization, OpenSMILE eGeMAPS v02); Stream B extracted nucleus accumbens BOLD activation during a reward task from the UCLA ds000030 dataset (n = 272). Three classifiers (logistic regression, random forest, gradient-boosted trees) were evaluated under stratified 5-fold cross-validation with fixed, pre-registered hyperparameters. Stream A achieved a best AUC-ROC of 0.63 (random forest; permutation p = .049), confirming H1 that acoustic prosody classifies PHQ-8-defined anhedonia above chance at = .05 (uncorrected), though this result does not survive Bonferroni correction for three classifiers (/3 = .017). Bootstrap analysis of {Delta}AUC confirmed non-inferiority relative to Stream B (H2: 95% CI lower bound > -0.10). However, Stream B itself did not achieve above-chance classification (p = .057), so this non-inferiority finding reflects comparable, modest performance across both modalities rather than equivalence to a validated neural biomarker. Pitch variability features (F0) ranked among the top 5 predictors by SHAP value in the gradient-boosted trees model (H3: partially supported). An exploratory combined model (eGeMAPS + recurrence quantification analysis) reached AUC = 0.65 (GBT; p = .032), though this lift was not statistically significant by DeLong test (z = -0.44, p = .66). These results provide initial evidence that vocal prosody, acquired through a standardized, open-source pipeline, carries depression-related information comparable to fMRI-derived ventral striatal activation for binary classification. However, neither streams operationalization isolates anhedonia from general depressive symptomatology, and the cross-modal design compares different constructs across different cohorts. We refer to the classification target throughout as the "anhedonia composite" to acknowledge that PHQ-8 Items 1+2 conflate anhedonia with dysphoria. We discuss these constraints and their implications for the causal framework motivating this work.