Skip to main content
Frontiers in Psychiatry logoLink to Frontiers in Psychiatry
. 2026 Jun 19;17:1825268. doi: 10.3389/fpsyt.2026.1825268

A gender-emotion interaction multi-task network for depression recognition via transformer-based multimodal fusion

Yujuan Xing 1,*, Ruifang He 2, Xiaoli Cao 1, Ping Tan 1, Li Chen 1
PMCID: PMC13328178  PMID: 42404723

Abstract

Depression is characterized by high prevalence, high recurrence, high disability and high mortality, which seriously affects people’s work and life. Among various behavioral biomarkers, speech-based features have gained increasing attention in depression detection due to their non-invasive nature, affordability, and rich capacity for conveying affective states. However, conventional depression recognition approaches rely solely on unimodal acoustic representations and largely overlook the influence of emotion and gender. To address this limitation, this study proposed a gender-emotion interaction multi-task network(G-EIMTNet) for depression recognition via transformer-based cross modal fusion. In the feature fusion stage, the deep representations of Mel-spectrograms were extracted using convolutional neural networks(CNN), and then the Maximum Correlation Minimum Redundancy (MRMR) algorithm was employed to select acoustic higher-order statistical features that were highly correlated with emotions and depressive states. These two types of features were then fused through the transformer attention mechanism. In the depression recognition stage, a depression recognition network for the interaction between gender and emotion was constructed based on a multi-task framework. Experiments on the AVEC2014 dataset showed that this approach outperformed the baseline model by 15.88% and 14.73% in accuracy and F1 score, respectively. Ablation experiments verify the effectiveness of multi-modal fusion and gender-emotion interaction.

Keywords: depression recognition, emotion valence, gender-emotion interaction, multimodal fusion, transformer attention

1. Introduction

Depression is a prevalent affective disorder characterized by persistent sadness, fatigue, and feelings of hopelessness, often leading to diminished interest in work and daily activities. It also adversely affects sleep patterns and appetite, further contributing to impaired concentration, and in severe cases, it may result in suicidal behavior. According to the World Health Organization’s 2021 report, the COVID-19 pandemic has significantly impacted global mental health, with healthcare workers, frontline personnel, students, individuals living alone, and those with pre-existing mental health conditions being disproportionately affected (1). Studies have shown that there is a significant positive correlation between depression and the risk of death (2). Currently, the diagnosis of depression predominantly relies on patients’ subjective reports, psychiatric evaluations, and rating scales. These approaches are limited by low efficiency and strong subjectivity (3). The adoption of simple, efficient, and non-invasive automated depression recognition methods not only facilitates early intervention but also reduces post-onset resource allocation and losses. Furthermore, such approaches help protect patients’ privacy, alleviate their concerns, and increase treatment-seeking rates. Speech signals offer unique advantages including easy acquisition, non-contact, minimal invasiveness, device portability, and low usage restrictions. Additionally, speech contains rich emotional cues that can directly and conveniently convey affective states. Years of research have demonstrated the feasibility of identifying depression through the speech analysis.

In the current research on speech-based depression recognition, feature extraction is the key. Early studies mainly relied on hand-crafted acoustic features such as fundamental frequency (F0), energy, jitter and shimmer, which are prosodic features. These directly reflect the pathological manifestations of depressed patients, such as slower speech rate and flat intonation, and have clear physical meanings (4, 5). The study further confirmed that specific acoustic statistical indicators are of great significance for clinical interpretability in long speech segments (6). However, hand-crafted features often have high dimensionality and contain substantial redundant information; direct input to the model can easily lead to the curse of dimensionality. Therefore, feature selection algorithms such as maximum relevance and minimum redundancy (MRMR) are widely used in biomedical signal processing to remove irrelevant noise and retain the most discriminative biomarkers (7). With the development of deep learning, researchers have leveraged convolutional neural networks for processing Mel spectrograms, thereby effective capturing complex time-frequency features in speech (8). However, deep spectral features and hand-crafted features are heterogeneous information, and their fusion is a challenge. Simple concatenation ignores the temporal dependencies and nonlinear correlations among features. In contrast to simple concatenation, cross-modal attention and bidirectional interaction are introduced, which is more robust to small samples, long sequences and modal noise. Multi-modal Transformer fusion network (9) proved that Cross-Attention can effectively align heterogeneous modal information in speech emotion recognition task (10). applied this mechanism to depression detection, and achieved complementary enhancement between features through a text-guided cross-modal Transformer. These studies show that the attention mechanism can adaptively weight different feature streams and significantly enhance the recognition effect.

In addition, depression is associated with atypical processing of emotional stimuli, including perception, response, and memory encoding. It can also be conceptualized as a mood disorder characterized by negative attentional bias (11). Moreover, there are gender differences in depression, and women are at a higher risk of depression (12). Most of the current researches ignore the influence of gender and emotion on the depression representation, which affects the robustness of the model in complex scenarios (13). pointed out that gender bias can seriously weaken the fairness and accuracy of the depression recognition model (14). constructed a joint emotion-depression detection network and proved that emotion recognition task can help the model extract more robust features. Ali’s research results in multi-task learning (MTL) prove that joint modeling of emotional states can significantly improve the generalization ability of the model through shared representation (15).

Despite promising advances in deep representation learning, feature fusion, and multi-task learning for depression recognition, several limitations remain: (i) the adaptive fusion of heterogeneous features is still suboptimal, and (ii) the integration of gender and emotion-related information within a multi-task framework has been largely underexplored. Motivated by these gaps, this study proposes G-EIMTNet, a gender–emotion interaction-aware multi-task network designed specifically to overcome these two limitations. Specifically, a dual-branch feature extraction module is designed: one branch employs a convolutional neural network (CNN) to learn deep semantic representations from Mel-spectrograms; the other applies minimum redundancy maximum relevance (mRMR) feature selection to identify features strongly associated with depression. These two complementary feature streams are then dynamically fused via a transformer-based cross-modal attention module. Finally, a multi-task learning framework explicitly modeling the interplay between gender and emotion is incorporated to enhance depression recognition performance.

The main contributions of this paper are summarized as follows:

  • A transformer-based cross-modal attention module is designed to effectively solve the problem that it is difficult to efficiently fuse the deep spectral features extracted by CNN and the hand-craft features selected by mRMR.

  • G-EIMTNet is constructed, which significantly enhances the robustness and generalization ability of the model in complex scenes by mining the potential association between gender, emotion and depression.

The remainder of this paper is organized as follows. Related works for multimodal fusion of heterogeneous acoustic features and the influence of emotion and gender on depression recognition are given in Section 2. Section 3 introduces our proposed method. Experimental results and discussion are described in Sections 4. Conclusions are drawn in section 5.

2. Related work

2.1. Multimodal fusion of heterogeneous acoustic features

Multimodal fusion considers not only the interactivity across modalities but also the correlations among heterogeneous features within each modality. Current modal fusion strategies primarily fall into three categories: feature-level fusion (early fusion) (16), model-level fusion (mid-stage fusion) (17), and decision-level fusion (late fusion) (18). directly concatenated the deep spectral features extracted by CNN with the MFCC along the feature dimension. The experimental results on DAIC-WOZ and MODMA were significantly better than those obtained by using only MFCC or only deep spectral features (19). concatenated LPC and MFCC into a feature sequence and fed it into 1D-CNN and long short-term memory (LSTM) respectively to capture local spectral structure and long-term temporal dependencies. The study employed a simulated annealing metaheuristic search to learn optimal weights for each feature subset, thereby achieving weighted feature fusion. Experimental results demonstrate that, within the spectral domain, complementary information exists among the original spectrum, Mel-spectrogram, and MFCCs (20). The attention-based acoustic feature fusion network, named ABAFNet (21), jointly processed four complementary acoustic representations: the upper envelope, linear spectrogram features, Mel-spectrogram, and high-level statistical features (HSFs), feeding into CNN and LSTM. Late weighted fusion was performed via a weight adjustment module (WAM), enabling ABAFNet to outperform both single-feature baselines and simple concatenation on two clinical depression speech databases (22). used the attention mechanism to fuse the MFCC and the spectrogram at an early stage. The JTA framework, proposed by Li (23), fuses the log-Mel spectrogram with features extracted from a pre-trained model (WavLM) via transactive attention. The multimodal spatiotemporal attention framework proposed by Niu (24) demonstrated that jointly audio and video features, and explicitly capturing complementary information across modalities, yields superior depression severity prediction performance compared to unimodal baselines. Saji (25) proposed a hybrid framework that combines facial action units (AU) with interview texts, using the transformer architecture for multi-instance learning (MIL), and effectively leveraging the complementarity between visual features and text features through late fusion.

Collectively, these studies indicated that attention mechanisms have become a central focus in heterogeneous feature fusion research (26, 27). In depression speech analysis, particularly under data-limited conditions, attention-based fusion of deep spectral features and hand-crafted acoustic features demonstrates greater robustness than spectrum-only models (28). However, the existing fusion strategies are difficult to achieve fine-grained feature alignment between heterogeneous features. Especially when dealing with the significant differences in physical semantics and data dimensions between high-order statistical features and deep spectral features, traditional methods are unable to achieve deep semantic fusion. Therefore, this paper proposes a cross-fusion method based on transformer, which utilizes the self-attention mechanism to capture cross-modal heterogeneous correlations, thereby solving the problem of deep fusion.

2.2. Research on the influence of emotion and gender on depression recognition

Prior work has found that the speech of individuals with depression in different situations consistently exhibits the characteristics of emotional passivation such as loudness reduction, low fundamental frequency, slow speech speed and prolonged pause duration (29). The meta-analysis conducted by Dell’Acqua (30) provides the physiological underpinnings for emotional blunting. By validating the emotion context insensitivity (ECI) hypothesis, their study revealed a generalized reduction in the amplitude of the late positive potential (LPP) when depressed individuals were exposed to various emotional stimuli. These findings suggest that the behavioral coldness observed in patients is deeply rooted in a significant impairment of the brain’s capacity to process emotional information. Recently, there has been much research interest in enhancing the performance of depression recognition by analyzing emotion influence. Xing (31) proposed to combine mRMR feature selection and multi-task Bagging ensemble for depression recognition. It was found that some HSFs more reflected valence, while others were more stably related to depression. After the explicit introduction of emotion recognition task into depression recognition model, the stability and accuracy of depression recognition were improved. The MoE model proposed by Guo (32) extracted speaker related features and emotion related features respectively, and then migrates them to the depression corpus for fusion. On the self-built Chinese depression corpus and AVEC2014 corpus, MoE performs better than only using acoustic features. This research further indicated that conflating emotion-related features with speaker-specific features may obscure depression-relevant acoustic cues. Other studies also demonstrated that incorporating emotion-related task into depression detection can improve the accuracy and generalization performance of the model (33, 34).

The research of depression recognition revealed that acoustic feature differences, such as MFCC3, were statistically significant only in male participants, suggesting that depression’s acoustic manifestations are not fully consistent across genders (29). Yang systematically analysis gender and racial biases in machine learning models for mental health assessment based on speech behavior, finding substantial disparities in both acoustic feature distributions and self-reported symptom scores (e.g., PHQ-9) across gender groups in anxiety and depression detection tasks (35). A domain-adversarial training framework, proposed by Kim (36), treated gender as a nuisance domain and mitigated the depression detection model’s reliance on gender-specific cues via adversarial learning. Experiments on the E-DAIC demonstrated that this approach not only improved the overall F1 score by approximately 13% but also substantially narrowed the performance gap between male and female. Some studies, from the perspective of privacy protection, have reduced the influence of demographic information such as gender through speaker decoupling technology, in order to improve the performance of depression detection (37, 38). A recent review noted that in terms of indicators such as F0 and formants, female depressives often show a decrease in peak frequency and amplitude, while male depressives tend to exhibit an increase in amplitude variation and formant changes. When age and gender are explicitly controlled in modeling, the prediction accuracy of speech features for depression is often higher and more stable (39). Gender and emotions have a significant impact on the depression recognition. However, most current studies have conducted separate research on these two factors, without delving deeply into their correlation. Therefore, this paper proposes a gender-emotion interaction multi-task depression recognition framework, which enhances the model’s recognition performance by modeling the correlation between gender and emotions.

3. Method

3.1. Description of dataset

This paper adopts the AVEC2014 dataset, which consists of 300 task-oriented depression videos recorded in human-computer interaction scenarios, involving 84 participants (40). The AVEC2014 dataset is oriented towards two tasks: participants read an excerpt of the fable “The North Wind and the Sun” in German; and the participants answer one of a series of questions in German, such as: “What is your favorite dish?”, “What is your best gift and why?”, or discuss a sad childhood memory, etc. Each record in this dataset has a BDI-II depression score (ranging from 0 to 63) and arousal/valence/dominance scores. Arousal, valence and dominance are three dimensions that describe emotions. Valence evaluates an individual’s positive or negative emotional response to a specific object, person or situation. Arousal reflects the intensity of an individual’s emotional response, while dominance indicates the influence an individual has over their environment. Arousal, valence and dominance are annotated for each speech recording by 3–5 raters using FeelTrace (41), which are time-continuous and each frame or second of the recording was assigned an emotional score.

The main objective of this study is depression recognition, which is a binary classification. Therefore, we mark the records with a depression score higher than 14 as depressed and those with a score lower than 14 as normal. Early studies conducted by Valstar and Schuller revealed that valence exhibits the strongest correlation with emotional states, followed sequentially by arousal and dominance (40). Among these dimensions, valence is the most readily comprehensible. We calculated the average valence value [-1, 1] of time continuity for each audio recording with multiple raters’ annotations. If the score was greater than 0, the emotional valence label of recording was positive; otherwise, it was negative. We manually annotated gender labels based on the recordings. We enhance the speech data by adding noise to improve the robustness of the model.

3.2. Feature extraction

The input features of the method proposed in this paper fall into two categories, namely, deep features and hand-crafted features. The feature extraction process is shown in Figure 1. In the stage of speech preprocessing, we employed a pre-emphasis function of H(z)=1−αz−1, which α=0.95. The overlaying framing approach was employed, with a Hamming window applied. The frame length and frame shift were set to 25ms and 10ms, respectively.

Figure 1.

Flowchart depicting an audio analysis process: audio and a spectrogram are preprocessed, generating a mel spectrogram for CNN-based deep feature extraction and handcrafted features via opensmile and mRMR analysis.

The procession of feature extraction.

We obtained mel-spectrograms from audio data and scaled them to 256×128. We utilized a CNN to extract deep features, with the activation function being ReLU. For the extraction of hand-crafted features, this study employed OpenSmile and Librosa to derive low level descriptors(LLDs) from preprocessed audio data. These features included loudness, Mel-frequency cepstral coefficients (MFCCs), log Mel-frequency band power (logMelFreqBand), line spectral pair frequencies (lspFreq), fundamental frequency contour (F0finEnv), voicing probability of fundamental frequency candidates (voicingFinalUnclipped), fundamental frequency (F0), logarithmic energy(logEnergy), short-time energy(STE), zero-crossing rate(ZCR), formant center frequencies and bandwidths(F1/F2/F3/B1/B2/B3) along with their first- and second-order differences, as well as voice quality-related features such as jitter, shimmer, harmonic amplitudes and differences (A1, A2, A3, H1, H2, H4, H1-H2, H2-H4, H1-A1, H1-A2, H1-A3, HNR15), and cepstral peak prominence (CPP). High-level statistical function features (HSFs) were derived by applying functional statistics to these low-level descriptors (LLDs). The distribution of the LLDs and the corresponding HSFs adopted in this paper is shown in Table 1, where △ and △△ respectively represent the first-order and second-order differences of the relevant LLDs. Table 2 lists the statistical functions used.

Table 1.

Selected LLDs and their first-order and second-order differences.

LLDs first-order and second-order differences
Spectrum related MFCC(15)+△ lspFreq(8)+△ logMelFreqBand(8)+△
Energy related Log energy(1) +△+△△ short-time energy(1)+△+△△
short-time zero crossing rate(1)+△+△△
F0 related F0finEnv(1)+△ voicingFinalUnclipped(1)+△
F0final(1)+△ voiced sound (1) unvoiced sound (1)
Formant related Formant center frequency (F1/F2/F3) (3)+△+△△
Formant center frequency bandwidth(B1/B2/B3) (3)+△+△△
Voice quality jitterLocal(1)+△ jitterDDP(1)+△ shimmerLocal(1)+△
Harmonics to Noise Ratio (HNR) (1) loudness(1) +△ CPP (1)
A1/A2/A3(3) H1/H2/H4(3) sound pressure level(SPL)(1)
H1-H2/H2-H4/H1-A1/H1-A2/H1-A3(5)

Table 2.

Statistical functions.

Category Function
Statistic max, min, mean, stddev, skewness, kurtosis, Max-Min,
iqr2-3(quartile3-quartile2), quartile1,quartile2,quartile3,percentile1.0,
percentile99.0, iqr1-3(quartile3-quartile1), pctlrange0-1, upleveltime75,
upleveltime90, iqr1-2(quartile2-quartile1)
Regression linregc1, linregc2, linregerrA, linregerrQ

Due to the high dimensionality of the extracted HSFs, we adopted mRMR (42) for feature selection. The core idea of the mRMR algorithm is to maximize the correlation between features and class labels while minimizing the correlation between features. That is say that, it selects features that are highly correlated with the class labels and have the least redundancy among themselves. We input the depression labels and valence labels into mRMR respectively to select features highly relevant to depression and emotion valence.

3.3. Proposed method

The proposed G-EIMTNet in this paper is mainly divided into two stages: multi-modal feature fusion and depression recognition. The framework is shown in Figure 2. Suppose that, the deep features extracted from the mel-spectrogram using CNN are denoted as FMel−CNN, the hand-crafted features are denoted as Fstat, and the features selected by the mRMR algorithm are denoted as Fstat−mRMR.

Figure 2.

Flowchart illustrating a two-stage neural network pipeline: Stage 1 fuses Mel spectrograms and acoustic features using a transformer attention mechanism; Stage 2 performs depression, emotion, and gender recognition using convolutional, pooling, and fully connected layers with auxiliary tasks feeding intermediate representations.

The multimodal depression recognition framework based on G-EIMTNet.

3.3.1. Transformer-based multimodal fusion

The objective of this stage is to fuse deep spectral features with handcrafted acoustic features. In our transformer-based cross-modal fusion frame, Query (Q) is derived from FMel−CNN and serves to search the focus. Key (K) is obtained from Fstat−mRMR, functioning as the indexing key. Value (V) is formed by the concatenation of FMel−CNN and Fstat−mRMR, containing the complete original information. The calculations of Q, K and V are as shown in Equations 1–3:

 Query (Q)=Wq·FMel−CNN (1)
Key (K)=Wk·Fstat−mRMR (2)
Value (V)=Wv·Concat(FMel−CNN,Fstat−mRMR) (3)

where Wq,Wk,Wv denote the weight matrices for linear mappings.

Scaled Dot-Product Attention is used for computing attention weights, the formula is shown in Equation 4:

Attention(Q,K,V)=Softmax(QKTdk)V (4)

where dk is the dimension of linear projection. In the feature fusion mechanism proposed in this paper, deep features are leveraged to query the importance of statistical features, followed by a weighting operation applied to the fused features. In fact, the Transformer uses multi-head attention. It does not perform the above calculation just once, but h times. The calculation is as shown in Equation 5.

Attentionmulti−head(Q,K,V)=Concat(head1,head2,⋯,headh) (5)

where headi=Attention(QWiQ,KWiK,VWiV),  (i=1,2,⋯,h). After the weighted addition, the features are processed sequentially through the Add&Norm and Feed Forward layers, yielding the final feature Ffused, which acts as the input to the subsequent stage. The formula is as shown in Equation 6:

Ffused=Norm(Add(Attentionmulti−head(Q,K,V)·V)) (6)

3.3.2. Gender-emotion interaction multi-task network

At this stage, emotion recognition and gender recognition are incorporated as auxiliary tasks to integrate emotion and gender information into the primary task (depression recognition), as illustrated in Figure 2. G-EIMTNet comprises three task branches, with information flow interaction facilitated via dedicated pathways.

In the auxiliary task of emotion recognition, the input feature Ffused is fed into a convolutional network, which outputs the emotional valence (positive/negative) while extracting the feature map Hemo subsequent to the ReLU layer. This feature map, which implies deep emotional information, is then transmitted to the gender recognition and depression recognition branches, as marked by red lines in Figure 2.

In the auxiliary task of gender recognition, the feature Ffused is processed via a fully connected layer, followed by an element-wise addition with the emotion feature Hemo. The corresponding calculation formula is given as in Equation 7:

Hin_gen=FC(Ffused)⊕αHemo (7)

The purpose of doing this is to demonstrate the influence of emotional states on the acoustic manifestations of gender, such as anger or sadness which may mask the fundamental frequency characteristics of gender. In equation (7), α is a regulating parameter. While performing gender recognition, the feature map Hgen after the Pool layer is extracted. This feature contains gender-specific acoustic patterns and, like Hemo, will be fed to the main task of depression recognition, as indicated by the blue line in Figure 2.

In the main depression recognition task, the fused feature Ffused is first processed via a fully connected layer and then element-wise added to the emotional feature Hemo to yield Hin_dep. Subsequently, after passing through a convolutional layer and a pooling layer, the resulting feature is concatenated with the gender feature Hgen to generate Hmid_dep, which is ultimately fed into a fully connected layer for depression recognition (depression/normal). The calculation is as shown in Equations 8 and 9.

Hin_dep=FC(Ffused)⊕βHemo (8)
Hmid_dep=Concat(Pool(Conv(Hin_dep)),Hgen) (9)

Prior to generating the final decision, G-EIMTNet explicitly leverages the gender feature Hgen to calibrate features, thereby mitigating the issue of model gender bias. The training of the multi-task model is optimized using a joint loss function. The total loss Ltotal consists of three parts, as shown in Equation 10:

Ltotal=Ldep+λ1Lemo+λ2Lgen (10)

where λ1 and λ2 are hyperparameters used to balance the contribution weights of the auxiliary tasks to the main task. Ldep, Lemo and Lgen denote the cross-entropy losses corresponding to the main task and the two auxiliary tasks, respectively. The calculation formula of cross-entropy loss is as shown in Equation 11:

L=−∑i=1Cyilog(y^i) (11)

3.3.3. The training process of G-EIMTNet

By introducing the two stages of feature fusion and multi-task classification, the overall training process of G-EIMTNet was described as Table 3.

Table 3.

The training process of G-EIMTNet.

Input Xmel: Mel spectrograms
Xacoustic: Acoustic features
Output Ddepression: Depression/Normal
Demotion: Positive/Negative
Dgender: Male/Female
Process # Stage 1: Multimodal Fusion based on Transformer Attention
# 1. Feature extraction FMel−CNN=CNN_Extractor( Xmel)
Fstat−mRMR= mRMR_Selector( Xacoustic)
# 2. Mapped to the Transformer space
Query = FC_Layer( FMel−CNN)
Key = FC_Layer( Fstat−mRMR)
Value = Concatenate( FMel−CNN, Fstat−mRMR)
# 3. Feature fusion
X = Dropout( Query, Key, Value)
X = MultiHeadAttention(X)
X= Add_and_Norm(X)
X = FeedForward(X)
Ffused= Add(X, Residual_Connection)
# Stage 2: Depression Recognition based on G-EIMTNet
# 1. Auxiliary Task 1: emotion recognition
Hemo_pre = Conv_Block( Ffused)
Hemo_pre = BatchNormalization( Hemo_pre)
Hemo = ReLU( Hemo_pre)
Demotion = Softmax(FC_Layer( Hemo))
# 2. Auxiliary Task 2: gender recognition
Hgen_input = FC_Layer( Ffused)
# Feedback incorporating emotion information
Hgen_mid=Add( Hgen_input, Hemo)
Hgen_mid = Conv_Pool_Block( Hgen_mid)
Hgen = Hgen_mid Dgender = Softmax(ReLU(FC_Layer( Hgen)))
# 3. Main Task: depression recognition
Hdep_input= FC_Layer( Ffused)
# Integrate the emotional features from the auxiliary task Hemo Hin_dep=Add( Hdep_input, Hemo)
Hmid_dep=Conv_Pool_Block( Hin_dep)
# Integrate the gender features from the auxiliary task Hgen Hfinal_dep = Concatenate( Hmid_dep, Hgen)
# Final decision
Ddepression = Softmax(ReLU(FC_Layer( Hfinal_dep)))
return Ddepression, Demotion,  Dgender

Throughout the entire process, the CNN module used for processing Mel Spectrograms adopted a lightweight deep structure, consisting of 4 convolutional layers. Each layer was followed by Batch Normalization and ReLU activation functions. The first two layers used 3 ×3 convolution kernels, while the last two layers used 1 ×1 convolution kernels for channel reduction and feature integration. The input channel was 1, and the channel numbers of each layer were set to 32, 64, 128, and 256 respectively. After the 2nd and 4th layers, 2 ×2 max pooling was applied. In the transformer cross-attention module, the number of attention heads was 8, the input dimension was 256, the feedforward network dimension was 1024, and the dropout rate was set to 0.3. The training process used the AdamW optimizer, with an initial learning rate of 1×10−4. The batch size was set to 32, and the number of training iterations was set to 150.

4. Experimental results and discussion

In this section, we comprehensively tested the proposed depression recognition method, 5-Fold Cross-Validation was used for the experimental results, and accuracy and F1-score were selected as the performance evaluation metrics (43). Supposed that tp and tn represent the samples correctly predicted as positive and negative, respectively, fp and fn represent the samples incorrectly predicted as positive and negative, respectively. The formulas for calculating accuracy and F1-score are as shown in Equations 12 and 13:

accuracy=tp+tntp+fn+fp+tn (12)
F1−score=2p*rp+r (13)

where p=tptp+fp, r=tptp+fn, represented the precision and recall of model. Accuracy metric quantified the ratio of accurate predictions generated by the model. F1-score integrated precision and recall, and a higher F1-score signifies the better classification performance.

4.1 1. Hand-crafted feature selection and analysis

In this subsection, the experiment analyzed 75 features highly correlated with depression and emotional valence obtained by mRMR, denoted as Fstat−mRMR. The feature distribution statistics were shown in Table 4.

Table 4.

Feature distribution statistics of Fstat−mRMR.

Category LLDs Fstat−mRMR
Spectrum related MFCC 30
lspFreq 7
logMelFreqBand 10
F0 related F0 3
F0final 2
voicingFinalUnclipped 2
F0finEnv 3
voiced sound 1
unvoiced sound 1
Energy related short-time energy 5
short-time zero crossing rate 2
Formant related b3 1
f2 1
Voice quality related shimmer 2
jitter 3
H1, H1A1c 2

As shown in Table 4, among the 75 features, those associated with spectral characteristics, including MFCC, logMelFreqBand, and lpsFreq, account for 62.7% of the selected features, indicating a strong correlation between spectral features and both emotional valence and depression-related information. Notably, features such as the skewness and kurtosis of the first-order difference of MFCC, as well as upleveltimeX, were also frequently selected. Depressed individuals often exhibit psychomotor retardation: their facial and oral muscles tend to be flaccid, vocal organs move slowly, and speech lacks rapid, forceful state transitions. This manifests as reduced variability in kurtosis and significant shifts in skewness values. UpleveltimeX quantifies the fraction of time spent on rapid articulation within a speech segment. In normal communication, healthy individuals’ speech is marked by frequent phoneme transitions and clear enunciation, leading to frequent, prolonged high-value segments in MFCC first-order differences and thus higher upleveltimeX values. By contrast, the monotonous speech of depressed patients results in lower upleveltimeX.

Fundamental frequency (F0) and energy features capture weakened laryngeal muscle control and reduced lung capacity—consequences of emotional blunting and physical fatigue in depressed patients—directly reflecting the low-arousal state characterized by monotonous intonation (diminished F0 variability) and faint voice (energy attenuation). Voice quality features (e.g., Jitter, Shimmer, and H1) quantify vocal fold vibration instability and breathiness, which are canonical acoustic markers of negative emotional valence (e.g., anxiety, distress). In summary, this feature set accurately maps the abnormal states of depressed patients across emotional dimensions (valence and arousal) and physiological functions. The distribution of features is shown in Figure 3.

Figure 3.

Horizontal bar chart showing types of acoustic features and their counts, grouped by categories: Spectrum, F0, Energy, Formant, and Voice quality. MFCC has the highest count at thirty.

Select hand-crafted features.

4.2. Ablation study for multimodal fusion

Ablation experiments were conducted to validate the effectiveness of transformer-based multimodal fusion. Specifically, four features were compared:

  • FMel−CNN: Only deep features extracted from mel-spectrograms were used;

  • Fstat−mRMR: Only handcrafted statistical features selected via MRMR were adopted;

  • FHybrid−Concat: Deep features and handcrafted features were simply concatenated;

  • Ffused: The Transformer-based multimodal fusion method proposed in this paper was employed.

All extracted features were fed into a fully connected layer followed by a softmax layer for classification. The experimental results are presented in Table 5.

Table 5.

Comparison of different features.

Feature Accuracy F1-Score
FMel−CNN 0.7355 0.7129
Fstat−mRMR 0.7122 0.6851
FHybrid−Concat 0.7681 0.7453
Ffused 0.7950 0.7788

As shown in Table 5, for the single modality, FMel−CNNoutperforms Fstat−mRMR, indicating that deep time-frequency features exhibit stronger representational capacity. FHybrid−Concat outperforms both FMel−CNN and Stat-MRMR, validating the complementarity between deep features and handcrafted features. The transformer fusion feature Ffused achieved the best performance, with an accuracy of 79.5%, which is approximately 2.7% higher than that of simple concatenation. This indicates that the transformer attention mechanism has successfully established an effective mapping between heterogeneous features, which is conducive to enhancing the depressive expression of features.

4.3. Ablation study for G-EIMTNet

This study primarily evaluates the efficacy of incorporating emotional and gender information into depression recognition models. In the experiments, our comparative models include: a single-task depression recognition model (DEP_Indep), a multi-task learning model with only gender recognition as an auxiliary task (MTL-Gen), a multi-task learning model with only emotion recognition as an auxiliary task (MTL-Emo), and our proposed G-EIMTNet. The input feature is Ffused and the experimental results are shown in Table 6. We fixed the main task training until convergence, then gradually introduced the emotion and gender branches, observing the changes in the F1 score of the validation set, and searching for the optimal balance point among the weights of each task. After multiple validation, the optimal values of λ1and λ2 in the total loss Ltotal of G-EIMTNet are 0.65 and 0.35 respectively.

Table 6.

Ablation study for G-EIMTNet.

Model Accuracy F1-Score
DEP-Indep 0.7750 0.7581
MTL-Gen 0.7924 0.7854
MTL-Emo 0.7953 0.7901
G-EIMTNet 0.8367 0.8022

As shown in Table 6, integrating emotion recognition and gender recognition as auxiliary tasks in depression recognition yields consistent performance enhancements across models. The improvement of MTL-Emo proves the strong correlation between emotion states and depression. The performance improvement of MTL-Gen suggests that the learned representation effectively mitigates gender bias, thereby enhancing the discriminative power of features. Notably, compared to DEP-Indep, G-EIMTNet achieves a 6.17% increase in accuracy and a 4.41% improvement in F1-score, indicating the efficacy of embedding emotional and gender information into depression recognition. Emotional states provide auxiliary clues for gender recognition; while the introduction of gender features is equivalent to establishing a physiological baseline for depression detection, enabling the model to effectively distinguish between gender-driven natural acoustic variations and depression-induced pathological acoustic anomalies. This can improve the model’s diagnostic precision.

4.4. t-SNE visualization

In this experiment, t-SNE was employed to reduce the dimensionality of both Mel-CNN features and Transformer-based fusion features to a 2D plane, where scatter plots were generated to visualize and analyze the feature distribution before and after fusion. The corresponding results are presented in Figure 4. As can be clearly seen from Figure 4, after feature fusion, the depressive samples and normal samples exhibit a more distinct clustering effect, with the inter-class distance increasing and the intra-class distance decreasing. This indicates that the features fused by the Transformer demonstrate stronger discriminative power.

Figure 4.

Two scatter plots compare data point separation before and after fusion in a classification task, with blue triangles representing depression and pink crosses representing normal. Before fusion, clusters overlap significantly. After fusion, depression and normal groups are more distinctly separated.

Comparison before and after transformer-based fusion.

4.5. Comparison with the state-of-the-art methods

The purpose of this experiment is to conduct a comparative analysis of the performance of other methods and our proposed method. Our main comparative methods included baseline method proposed by Valstar (40), the CNN fused with LSTM for depression recognition based on Mel-Spectrogram proposed by Ma (44), the emotion-assisted multi-task CNN proposed by Dumpala (45), the gender-assisted multi-task learning proposed by Liu (46), the deep LSTM based on MFCC proposed by Rejaibi (47), spatio-temporal attention (STA) network and eigen evolution pooling(EEP) feature fusion (MAFF) strategy proposed by Niu (24), attention-based acoustic feature fusion network (ABAFnet) proposed by Xu (21), multidimensional convolutional neural networks (MDCF-Net) that incorporate emotional features proposed by Ren (48), and joint networks with tailored Attention (JTA) based on Spectrogram-Based model with Tailored Attention (STA), pre-trained model (WavLM) and Transactive Attention-based feature Fusion module (TAF) proposed by Li (23). The experimental results are showed in Table 6.

As indicated in Table 7, the early DepAudioNet model was constrained by its reliance on a single feature modality, resulting in relatively limited performance. In recent years, Dumpala and Liu have made progress by incorporating emotional and gender information into the depression recognition. Compared to DepAudioNet, their proposed methods achieving 4.18% and 3.49% improvements in accuracy respectively. This validated the predictive value of emotional and gender attributes for depression recognition. With the development of deep learning and attention mechanisms, Rejaibi employed deep LSTM for depression detection, which demonstrated the importance of capturing long-distance temporal dependencies. Niu further introduced STA and EEP, which can capture the dynamic evolution of features over time, achieving an accuracy of 0.7881 and an F1-score of 0.7681. Although Niu and Rejaibi’s method has achieved some success in capturing the temporal dynamics of features, its accuracy and F1-score were both lower than G-EIMTNet. The G-EIMTNet we proposed not only used Transformer multimodal attention to capture the temporal relationship of features and the correlation between deep features and handcrafted features (HSFs), but also incorporates emotion and gender into depression recognition. Compared with Niu’s method, the accuracy of G-EIMTNet had increased by 4.86%.

Table 7.

Results of our method and comparative methods.

Author Methods Year Accuracy F1-score
Valstar et al. (40) HSFs+SVM (baseline method) 2014 0.6779 0.6459
Ma et al. (44) DepAudioNet 2016 0.7232 0.6832
Dumpala et al. (45) Spectrogram + Emotion Task 2021 0.7650 0.7355
Liu et al. (46) MFCC + Gender Task + Attention 2021 0.7581 0.7208
Rejaibi et al. (47) MFCC+LSTM 2022 0.7325 0.7150
Niu et al. (24) STA+EEP 2023 0.7881 0.7681
Xu et al. (21) ABAFnet 2024 0.8032 0.7853
Ren et al. (48) MDCF-Net 2025 0.7754 0.7547
Li et al. (23) JTA(WavLM+Tailored Attention+ Transactive Attention-based feature Fusion) 2025 0.8205 0.7895
G-EIMTNet - 0.8367 0.8022

Both the MDCF-Net proposed by Ren and the ABAFnet proposed by Xu were dedicated to integrating more features such as spectrograms, envelope features, and HSFs. However, ABAFnet implements feature fusion via dynamic weight adjustment, whereas G-EIMTNet employs Transformer-based nonlinear interactive fusion. More importantly, G-EIMTNet introduced a Gender-Emotion Interaction module. In AVEC 2014 dataset, which contain different genders and rich emotional states, simply piling up acoustic features was prone to interference from gender differences.

Compared with the JTA framework proposed by Li, the accuracy of G-EIMTNet was 1.62% higher. The JTA framework leveraged the pre-trained WavLM, whereas the G-EIMTNet achieved an accuracy of 83.67% without relying on large-scale external pre-training data. This was also credited to our fusion of deep features and handcrafted features, as well as the dedicated Gender-Emotion Interaction module proposed in this work.

In conclusion, comparative analyses with state-of-the-art methods demonstrated the effectiveness of the G-EIMTNet proposed in this paper. It also indicates that depression recognition cannot be separated from emotional states, and gender differences can affect the judgment. The visualization of performance comparison is shown in Figures 5 and 6.

Figure 5.

Bar chart comparing accuracy percentages of various models or studies labeled by author, with values ranging from sixty-seven point seven nine percent to eighty-three point six seven percent. Green and red annotations above bars indicate performance differences between methods, highlighting the highest value for “Ours” at eighty-three point six seven percent.

The visualization of accuracy comparison.

Figure 6.

Bar chart comparing F1-scores in percentages for different methods, with values labeled on each bar. The highest score is 80.22 for “Ours”. Green and red annotations indicate numerical increases or decreases between methods.

The visualization of F1-score comparison.

4.6. Discussion

The time complexity of G-EIMTNet is mainly composed of three parts: the feature extraction layer, the Transformer fusion layer, and the multi-task classifier. Among them, the time complexity of the feature extraction layer mainly depends on the convolution operation of CNN on the Mel spectrogram, which is related to the size of the convolution kernel, the number of channels, and the input image. In the Transformer fusion module, the time complexity is O(L2·d), where Lis the sequence length and d is the embedding dimension. Since speech segments are usually divided into short frames, the sequence length L will not be too large. The time complexity of the multi-task classifier is O(Nfc), The three parallel branches mainly consist of fully connected layers and simple convolutional pooling. The feature interaction between tasks only involves vector operations, with extremely low complexity. Therefore, compared to training three separate models for each task, G-EIMTNet significantly reduces the computing time through the parameter sharing mechanism.

5. Conclusions

This work presents the Gender-Emotion Interaction Multi-task Network (G-EIMTNet), a novel framework designed to enhance the robustness of speech-based depression recognition. Our method targets two critical limitations in existing research: first, the lack of effective fusion strategies for heterogeneous acoustic features; second, the oversight of intrinsic interactions among depression, emotion, and gender. In our method, deep Mel-spectral representations from CNN are first fused with MRMR-selected HSFs via a Transformer-based multimodal attention mechanism, which bridges their semantic gap and enables adaptive alignment and enhancement. We further innovatively construct a multi-task interactive framework: unlike traditional methods treating gender or emotion as isolated labels, G-EIMTNet explicitly models their coupling mechanism, with experimental results on the AVEC 2014 dataset validating the method’s effectiveness.

Despite the promising results, this work still has limitations that warrant further improvement. G-EIMTNet introduces multiple loss weights, λ1 and λ2. Finding the optimal weights requires a large number of grid search experiments, and the parameter tuning cost is relatively high. Furthermore, although the transformer excels at local feature fusion, for a disorder like depression, which exhibits long-term behavioral patterns, the model currently tends to analyze short-time speech segments, possibly neglecting more macroscopic trends in intonation changes or long-term psychological fluctuations. Therefore, in subsequent work, we will explore an adaptive weighting strategy for the joint loss function in multi-task learning and simultaneously consider integrating long-term behavioral features into depression recognition to further improve performance.

Funding Statement

The author(s) declared that financial support was received for this work and/or its publication. This work was supported in part by the National Natural Science Foundation of China (Grant No. 61841203), in part by Doctoral Program of Yanyuan Science and technology Innovation Fund (2023BSZX04), in part by Yanyuan Scholar Program, in part by University Faculty Innovation Fund Project(2026A-253).

Footnotes

Edited by: Xuntao Yin, Guizhou Provincial Rehabilitation Hospital, China

Reviewed by: Alwin Poulose, Thiruvananthapuram, India

Chen Zeyang, Beijing Jiaotong University, China

Data availability statement

Publicly available datasets were analyzed in this study. This data can be found here: http://avec2013-db.sspnet.eu/.

Author contributions

YX: Visualization, Writing – original draft, Methodology, Writing – review & editing. RH: Validation, Writing – review & editing. XC: Data curation, Writing – review & editing. PT: Data curation, Writing – review & editing. LC: Visualization, Writing – review & editing.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

References

  • 1.World Health Statistics (2021). Available online at: https://www.who.int/data/gho/publications/world-health-statistics (Accessed May 20, 2021).
  • 2. Oliva V, Fico G, De Prisco M, Gonda X, Rosa AR, Vieta E. Bipolar disorders: an update on critical aspects. Lancet Reg Health Eur. (2024) 48:101135. doi:  10.1016/j.lanepe.2024.101135 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Miner AS, Haque A, Fries JA, Fleming SL, Wilfley DE, Terence Wilson G, et al. Assessing the accuracy of automatic speech recognition for psychotherapy. NPJ Digit Med. (2020) 3:82. doi:  10.1038/s41746-020-0285-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Almaghrabi SA, Clark SR, Baumert M. Bio-acoustic features of depression: A review. BioMed Signal Process Control. (2023) 85:105020. doi:  10.1016/j.bspc.2023.105020 38826717 [DOI] [Google Scholar]
  • 5. Wei Y, Qin S, Liu F, Liu R, Zhou Y, Chen Y, et al. Acoustic-based machine learning approaches for depression detection in Chinese university students. Front Public Health. (2025) 13:1561332. doi:  10.3389/fpubh.2025.1561332 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Deng Q, Luz S, De La Fuente Garcia S. (2025). “ An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech”, in: Interspeech 2025 (ISCA). Rotterdam: ISCA. 5248–52. doi:  10.21437/Interspeech.2025-1968 [DOI] [Google Scholar]
  • 7. Tekindor A, Aydın E. (2024). “ Feature selection improves speech based Parkinson’s disease detection performance”, in: Proceedings of the 17th International Joint Conference on Biomedical Engineering Systems and Technologies (Rome, Italy: SCITEPRESS-Science and Technology Publications; ) 1, 726–32. doi:  10.5220/0012347300003657 [DOI] [Google Scholar]
  • 8. Khan MH, Majid M, Arsalan A, Linguraru MG, Anwar SM. (2025). “ Deep learning based depression detection using speech spectrograms”, in: 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC) (Copenhagen, Denmark: IEEE; ), 1–4. doi:  10.1109/EMBC58623.2025.11253416 [DOI] [PubMed] [Google Scholar]
  • 9. Wang Y, Gu Y, Yin Y, Han Y, Zhang H, Wang S, et al. Multimodal transformer augmented fusion for speech emotion recognition. Front Neurorobot. (2023) 17:1181598. doi:  10.3389/fnbot.2023.1181598 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Teng S, Chai S, Liu J, Tateyama T, Lin L, Chen YW, et al. (2024). “ A sentiment pre-trained text-guided multimodal cross-attention transformer for improved depression detection”, in: 46th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC) (Orlando, FL, USA: IEEE; ), 1–4. doi:  10.1109/EMBC53108.2024.10782904 [DOI] [PubMed] [Google Scholar]
  • 11. Elsayed NM, Vogel AC, Luby JL, Barch DM. Labeling emotional stimuli in early childhood predicts neural and behavioral indicators of emotion regulation in late adolescence. Biol Psychiatry: Cogn Neurosci Neuroimaging. (2021) 6:89–98. doi:  10.1016/j.bpsc.2020.08.018 [DOI] [PubMed] [Google Scholar]
  • 12. Roxo L, Silva M, Perelman J. Gender gap in health service utilisation and outcomes of depression: A cross-country longitudinal analysis of European middle-aged and older adults. Prev Med. (2021) 153:106847. doi:  10.1016/j.ypmed.2021.106847 [DOI] [PubMed] [Google Scholar]
  • 13. Gómez-Zaragozá L, Marín-Morales J, Alcañiz M, Soleymani M. (2025). “ Speech and text foundation models for depression detection: Cross-task and cross-language evaluation”, in: Interspeech 2025 (ISCA). Rotterdam: ISCA. 5253–7. doi:  10.21437/Interspeech.2025-1035 [DOI] [Google Scholar]
  • 14. Zhang Y, Li X, Rong L, Tiwari P. (2021). “ Multi-task learning for jointly detecting depression and emotion”, in: 2021 IEEE International Conference on Bioinformatics and Biomedicine (BIBM). Houston, TX: IEEE, 3142–9. doi:  10.1109/BIBM52615.2021.9669546 [DOI] [Google Scholar]
  • 15. Ali M, Lucasius C, Patel TP, Aitken M, Vorstman J, Szatmari P, et al. Speech as a multimodal digital phenotype for multi-task LLM-based mental health prediction. [Preprint]. (2025). doi:  10.48550/arXiv.2505.23822 [DOI] [Google Scholar]
  • 16. He L, Jiang D, Sahli H. (2015). “ Multimodal depression recognition with dynamic visual and audio cues”, in: 2015 International Conference on Affective Computing and Intelligent Interaction (ACII) (Xi'an, China: IEEE; ), 260–6. doi:  10.1109/ACII.2015.7344581 [DOI] [Google Scholar]
  • 17. Yang L, Jiang D, Sahli H. Integrating deep and shallow models for multi-modal depression analysis—hybrid architectures. IEEE Trans Affect Comput. (2018) 12:239–53. doi:  10.1109/TAFFC.2018.2870398 25079929 [DOI] [Google Scholar]
  • 18. Das AK, Naskar R. A deep learning model for depression detection based on MFCC and CNN generated spectrogram features. BioMed Signal Process Control. (2024) 90:105898. doi:  10.1016/j.bspc.2023.105898 38826717 [DOI] [Google Scholar]
  • 19. Du M, Liu S, Wang T, Zhang W, Ke Y, Chen L, et al. Depression recognition using a proposed speech chain model fusing speech production and perception features. J Affect Disord. (2023) 323:299–308. doi:  10.1016/j.jad.2022.11.060 [DOI] [PubMed] [Google Scholar]
  • 20. Pingping W, Fangfang X, Han L. (2025). “ Enhanced depression detection through optimally weighted spectrogram feature fusion”, in: Proceedings of the 13th International Conference on Computing and Pattern Recognition (ICCPR '24) (New York, NY, USA: Association for Computing Machinery; ), 226–32. doi:  10.1145/3704323.3704375 [DOI] [Google Scholar]
  • 21. Xu X, Wang Y, Wei X, Wang F, Zhang X. Attention-based acoustic feature fusion network for depression detection. Neurocomputing. (2024) 601:128209. doi:  10.1016/j.neucom.2024.128209 38826717 [DOI] [Google Scholar]
  • 22. Zhang X, Li B, Qi G. A novel multimodal depression diagnosis approach utilizing a new hybrid fusion method. BioMed Signal Process Control. (2024) 96:106552. doi:  10.1016/j.bspc.2024.106552 38826717 [DOI] [Google Scholar]
  • 23. Li X-H, Liu Z-T, Liu C-L, Zhong B-L, Chen J, She J, et al. JTA: Joint networks with tailored attention for speech depression detection. Knowledge-Based Syst. (2025) 330:114617. doi:  10.1016/j.knosys.2025.114617 38826717 [DOI] [Google Scholar]
  • 24. Niu M, Tao J, Liu B, Huang J, Lian Z. Multimodal spatiotemporal representation for automatic depression level detection. IEEE Trans Affect Comput. (2023) 14:294–307. doi:  10.1109/TAFFC.2020.3031345 25079929 [DOI] [Google Scholar]
  • 25. Saji AR, P M, Poulose A. (2025). “ Transformer-MIL and AdaBoost fusion for non-invasive depression detection”, in: 2025 International Conference on Robotics and Mechatronics (ICRM) (KOLLAM, India: IEEE; ), 1–6. doi:  10.1109/ICRM66809.2025.11349107 [DOI] [Google Scholar]
  • 26. Chen J, Hu Y, Lai Q, Wang W, Chen J, Liu H, et al. IIFDD: Intra and inter-modal fusion for depression detection with multi-modal information from Internet of Medical Things. Inf Fusion. (2024) 102:102017. doi:  10.1016/j.inffus.2023.102017 38826717 [DOI] [Google Scholar]
  • 27. Li J, Zhang Z, Lang J, Jiang Y, An L, Zou P, et al. (2022). “ Hybrid multimodal feature extraction, mining and fusion for sentiment analysis”, in: Proceedings of the 3rd International on Multimodal Sentiment Analysis Workshop and Challenge (New York, NY, USA: Association for Computing Machinery; ), 81–8. doi:  10.1145/3551876.3554809 [DOI] [Google Scholar]
  • 28. Liu L, Liu L, Wafa H, Tydeman F, Xie W, Wang Y, et al. Diagnostic accuracy of deep learning using speech samples in depression: a systematic review and meta-analysis. J Am Med Inf Assoc JAMIA. (2024) 31:2394–404. doi:  10.1093/jamia/ocae189 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Wang J, Zhang L, Liu T, Pan W, Hu B, Zhu T, et al. Acoustic differences between healthy and depressed people: a cross-situation study. BMC Psychiatry. (2019) 19:300. doi:  10.1186/s12888-019-2300-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Dell'Acqua C, Mologni V, Feraco T, Benvenuti SM. A meta-analysis of the late positive potential for assessing affective processing in depression and depression vulnerability. J Affect Disord. (2026) 45:121539. doi:  10.1016/j.jad.2026.121539 [DOI] [PubMed] [Google Scholar]
  • 31. Xing Y, Liu Z, Chen Q, Li G, Ding Z, Feng L, et al. Depression recognition base on acoustic speech model of multi-task emotional stimulus. BioMed Signal Process Control. (2023) 85:104970. doi:  10.1016/j.bspc.2023.104970 38826717 [DOI] [Google Scholar]
  • 32. Guo W, He Q, Lin Z, Bu X, Wang Z, Li D, et al. Enhancing depression recognition through a mixed expert model by integrating speaker-related and emotion-related features. Sci Rep. (2025) 15:4064. doi:  10.1038/s41598-025-88313-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Hansen L, Zhang YP, Wolf D, Sechidis K, Ladegaard N, Fusaroli R, et al. A generalizable speech emotion recognition model reveals depression and remission. Acta Psychiatrica Scandinavica. (2022) 145:186–99. doi:  10.1111/acps.13388 [DOI] [PubMed] [Google Scholar]
  • 34. Machorro MG, Reichel U, Hecker P, Hammer H, Sagha H, Eyben F, et al. Speech-based depressive mood detection in the presence of multiple sclerosis: A cross-corpus and cross-lingual study. In: Abbas M, Yousef T, Galke L, editors.Proceedings of the 8th International Conference on Natural Language and Speech Processing (ICNLSP-2025). Association for Computational Linguistics, Southern Denmark University, Odense, Denmark: (2025). p. 283–92. Available online at: https://aclanthology.org/2025.icnlsp-1.28/. [Google Scholar]
  • 35. Yang M, El-Attar AA, Chaspari T. Deconstructing demographic bias in speech-based machine learning models for digital health. Front Digital Health. (2024) 6:1351637. doi:  10.3389/fdgth.2024.1351637 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Kim J-W, Yoon H, Oh W, Yoon S, Kim D, Lee D, et al. Domain adversarial training for mitigating gender bias in speech-based mental health detection. [Preprint]. (2025). doi:  10.48550/arXiv.2505.03359 [DOI] [PubMed] [Google Scholar]
  • 37. Wang J, Ravi V, Alwan A. (2023). “ Non-uniform speaker disentanglement for depression detection from raw speech signals”, in: Interspeech, 2023, Dublin: International Speech Communication Association (ISCA). 2343–7. doi:  10.21437/interspeech.2023-2101 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Ravi V, Wang J, Flint J, Alwan A. Enhancing accuracy and privacy in speech-based depression detection through speaker disentanglement. Comput Speech Lang. (2024) 86:101605. doi:  10.1016/j.csl.2023.101605 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Martínez-Nicolás I, Criado D, Gordillo F, Martínez-Sánchez F, Meilán J. Speech analysis for detecting depression in older adults: a systematic review. Front Psychol. (2025) 16:1715538. doi:  10.3389/fpsyg.2025.1715538 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Valstar M, Schuller B, Smith K, Almaev T, Eyben F, Krajewski J, et al. (2014). “ AVEC 2014: 3D dimensional affect and depression recognition challenge”, in: Proceedings of the 4th International Workshop on Audio/Visual Emotion Challenge (Orlando Florida USA: ACM; ), 3–10. doi:  10.1145/2661806.2661807 [DOI] [Google Scholar]
  • 41. Cowie R, Douglas-Cowie Savvidou S, McMahon E, Sawey M, Schröder M. FEELTRACE: an instrument for recording perceived emotion in real time. In: Proc.isca Wsh.on Speech & Emotion. Newcastle, Northern Ireland: International Speech Communication Association (ISCA) (2000). p. 1–6. Available online at: https://api.semanticscholar.org/CorpusID:5962753 (Accessed September 1, 2000). [Google Scholar]
  • 42. Peng HC, Long FH, Ding C. Feature selection based on mutual information: criteria of max-dependency, max-relevance, and min-redundancy. IEEE Transaction Pattern Anal Mach Intell. (2005) 27:1226–38. doi:  10.1109/TPAMI.2005.159 [DOI] [PubMed] [Google Scholar]
  • 43. ShamsEldin T, Gaber S, Ansari S, Elgohary R, Shawky M, Elbahnasawy M, et al. Artificial intelligence for predicting depression anxiety and stress using psychometric data. Sci Rep. (2025) 15:37282. doi:  10.1038/s41598-025-21301-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Ma X, Yang H, Chen Q, Huang D, Wang Y. (2016). “ DepAudioNet: An efficient deep model for audio based depression classification”, in: Proceedings of the 6th International Workshop on Audio/Visual Emotion Challenge (Amsterdam The Netherlands: ACM; ), 35–42. doi:  10.1145/2988257.2988267 [DOI] [Google Scholar]
  • 45. Dumpala SH, Rempel S, Dikaios K, Sajjadian M, Uher R, Oore S, et al. (2021). “ Estimating severity of depression from acoustic features and embeddings of natural speech”, in: ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) (Toronto, Canada: IEEE; ), 7278–82. doi:  10.1109/ICASSP39728.2021.9414129 [DOI] [Google Scholar]
  • 46. Liu Y, Lu X, Shi D, Yuan J, Pan T, An H, et al. (2021). “ Improved depression recognition using attention and multitask learning of gender recognition”, in: 2021 International Conference on Asian Language Processing (IALP) (Singapore, Singapore: IEEE; ), 57–61. doi:  10.1109/IALP54817.2021.9675220 [DOI] [Google Scholar]
  • 47. Rejaibi E, Komaty A, Meriaudeau F, Agrebi S, Othmani A. MFCC-based recurrent neural network for automatic clinical depression recognition and assessment from speech. BioMed Signal Process Control. (2022) 71:103107. doi:  10.1016/j.bspc.2021.103107 38826717 [DOI] [Google Scholar]
  • 48. Ren L, Ling Y, He T, Du J, Xu R, Zhang H, et al. (2025). “ Multidimensional speech feature extraction for depression detection using MDCF-Net”, in: 2025 IEEE International Symposium on Circuits and Systems (ISCAS) (London, United Kingdom: IEEE; ), 1–5. doi:  10.1109/ISCAS56072.2025.11043261 [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Publicly available datasets were analyzed in this study. This data can be found here: http://avec2013-db.sspnet.eu/.


Articles from Frontiers in Psychiatry are provided here courtesy of Frontiers Media SA

RESOURCES