Abstract
Facial emotion recognition (FER) plays a critical role in most applications in human–computer interaction, psychological analysis, and affective computing to make intelligent systems capable of effectively perceiving the emotions of humans. Current approaches lack sufficient strength against problems like poor feature representation, robustness of facial expression variation, and model generalization. In order to counter such shortcomings, this paper presents a new FER model that integrates the SiaCon-DetNet and HySHO algorithm. The most striking novelty of SiaCon-DetNet is its capacity to combine convolutional feature learning with transformer attention mechanisms in order to make strong detection of fine-grained facial features. Also, the suggested framework are based on its intelligent combination of bio-inspired top optimization and deep learning, resulting in an adaptive and efficient emotion detector. Meanwhile, HySHO dynamically adjusts model parameters to enhance learning convergence and reduce computation overhead. This method in the paper presumes an organized working process with the initial step being face region detection by a Siamese convolutional network and feature enhancement by multi-head self-attention in the detection transformer network. Comparative analysis of its performance indicates the new model shows better performance as compared to all other FER methods with up to 99.20% accuracy on JAFFE database and having very short training periods. Emotion-wise correlation and performance testing also validate the reliability of the proposed framework, with precision, recall, and F1-score consistently between 98–99%.
Keywords: Face emotion recognition (FER), Image processing, Deep learning, Optimization and classification
Subject terms: Engineering, Mathematics and computing
Introduction
The human facial expressions serve as essential signals to understand both intentions and emotional states of those we interact with in interpersonal communication. Verbal communication expresses explicit messages by words but nonverbal elements that consist of facial expressions body language and vocal tone together strengthen a person’s understanding during interactions1,2. Nonverbal elements form about two thirds of human communication along with expressions that advance understanding by surpassing verbal messages. The human face provides instantly recognizable communication signals by displaying passion through expressions that reveal natural emotions ranging from happiness to unhappiness to anger along with surprise and fear and disgust throughout nonverbal interactions. These human biological expressions exist as universal signals which different cultures can easily detect thus playing a vital role in social contacts and emotional understanding. Scientific domains such as psychology, neuroscience, artificial intelligence, affective computing and human–computer interaction base their analysis on facial expressions because they represent key elements in their research. Research personnel in perceptual and cognitive sciences study human emotional facial perception together with neural processing that allows face emotion recognition response mechanisms3. The identification of these principles supports the evaluation of social comprehension alongside disorders of mental health and neurological syndromes that affect emotion perception abilities. The development of affective computing systems depends on having facial emotion analysis as a pivotal component because it enables better human–machine interactions. Researchers attempt to create context-sensitive emotional Artificial Intelligence (AI) systems4,5 by implementing emotion recognition which will improve user experiences across virtual assistants and healthcare service and customer service platforms.
Aside from that, facial emotion recognition has also been widely applied in computer animation and entertainment where expressive and natural facial animations must be created so interactive digital content will be effective. Natural facial expression of facial emotion in animated films, virtual reality, or video games maximizes the expressiveness of digital actors so interaction will be optimized to become realistic and immersive6,7. Latest techniques such as deep learning and neural networks have further enhanced the ability of automated synthesis and analysis of facial expressions much more precisely, and it has further extended its influence towards automated emotion analysis, behavior understanding, and sentiment analysis too. As more is researched, face emotion analysis will be at the forefront of an extensive range of applications from psychological science and human–computer interaction to security and surveillance where the detection of emotions would be applied to monitor suspect or distress behavior. Because of its far-reaching breadth of meaning, facial emotion research remains one of the major issues of today, weaving human thinking and AI in the quest to enhance human and technological communication, interaction, and emotional intelligence.
Facial expression recognition is a major technology of intelligent human–machine interaction and combination of human emotional perception and artificial intelligence. Facial expression recognition can make machines sense, understand, and respond to human emotions, hence enhancing human-technology interaction8. Facially based extensive applications of facial expression recognition have rendered facial expression recognition technology as a core technology in medicine, education, game creation, and security. Facial emotion recognition technology has broad application in assistant medicine including patient monitoring, mental status checkup, and pain recognition so that doctors will know the mental and physical states of a patient without relying on verbal information9. It is particularly suitable for such communication disorder patients as autism or speech disability because facial emotion checkup can cause doctors to catch better their sentiments and reactions. Also, facial expression recognition contributes to mental health research by providing automatic stress, depression, and anxiety detection from the facial expression, thus providing an early warning system for mental illness and enhancing patient care using AI-based therapy.
Facial analysis allows computer algorithms to classify learners as engaged, perplexed, irritated, or disengaged and permit the teachers to adjust their pedagogy in accordance with the new data. It is worthwhile especially in e-learning, when teachers have no direct interaction with students and cannot estimate the level of comprehension and emotional attachment to material. Facial expression recognition can be applied to computer games and virtual reality as well since it has the ability to make the game mechanics automatically change based on an individual’s mood. Facial expression recognition is also commonly used in public security networks and surveillance, wherein facial recognition is utilized in surveillance and alerting suspect individuals or potential risk in an affective context. By applying image processing methods, the system identifies characteristic facial features such as eyebrow motion, change in eye shape, and change in mouth shape, which are associated with specific emotions such as happiness, surprise, aversion, and neutrality10. The hence identified features are then processed with the application of deep learning architecture, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), with the aim of classifying expressions with great accuracy. Emotional quantification can help researchers quantify and analyze human emotions scientifically and systematically, and it is applicable in psychology, marketing, and human–computer interaction. As AI technology advances, man-to-machine communication will necessarily become smoother, human beings and machines becoming more and more accustomed to each other naturally and emotionally11. The innovation is proof of the immense importance of face recognition research by facial expressions as, besides making humanity more technologically powerful, it is part of the world’s common good and human and social well-being.
Current models do not possess the capability in handling long-range dependencies, context relationships, and small facial variations across a broad range of diverse populations and weather conditions. Trying to address such matters, deep learning transformer models have been largely characteristic for their capacity in modeling global dependency, sequential data processing effectively, and feature augmentation through self-attention12. As there is more utilization of artificial intelligence in our daily life, from smart teachers and smart assistants to the monitoring and surveillance of one’s mental well-being, the reading of people’s emotions accurately through facial expressions is a need of the times. The proposed system will close the gap that has existed between man-like emotional feeling and machine capabilities through the use of the transformers’ power by gaining deeper and richer facial representation, resulting in propelling human-AI system interaction13,14. Further, the development of such an advanced facial emotion recognition system is also of high social value, in that it will be able to make AI technology more compassionate, increase the mobility of communication-disabled people, and make security stronger by enabling timely detection of abnormal behaviors. As AI keeps growing, transformer models provide a groundbreaking paradigm shift for emotion recognition and therefore are an appealing choice to achieve state-of-the-art results in real-world applications. The primary goal of this work is to leverage the unmatched feature extraction ability of transformers and provide facial emotion recognition systems with greater accuracy, robustness, and adaptability to human emotion uncertainty in changing contexts.
The rest of this paper is organized following the below outline. Section Related Works presents an in-depth introduction to existing face emotion recognition systems by describing the weakness and robustness of standard deep models such as CNNs and RNNs and what has recently become available in terms of transformer models. Section Proposed Methodology presents the proposed approach using a proper and accurate definition for the deep transformer learning-based facial emotion recognition model. Section Results and Discussion presents experimental results for demonstrating the performance of the suggested model on normal datasets and precision, recall, accuracy, and F1-score against other state-of-the-art modern approaches. Section Conclusion concludes the paper by summarizing the entire results, pointing out the outclassing nature of the proposed solution, and revealing future goals to improve emotion recognition accuracy and real-time notification for various cases.
Related works
Gursesli, et al.15 introduce a lightweight convolutional neural network (CNN) for facial emotion recognition, referred to as the Custom Lightweight CNN-based Model (CLCM), from the most popular MobileNetV2 architecture. The reason for designing CLCM is to design a computationally light yet effective model with high accuracy in facial emotion recognition. To make comparative studies of performance of the formulated model, four common public benchmarks FER-2013, RAF-DB, AffectNet, and CK + , which are popular benchmarking datasets of facial expressions used to test ML models, were employed in experimentation with seven human emotions. Two popular light-weight deep learning networks, ShuffleNetV2 and MobileNetV2, also were used in comparative studies on CLCM with their performance results. With a leverage of the best of MobileNetV2 and the addition of personalized optimizations, the authors seek to improve facial emotion recognition in limited environments without a loss of computational efficiency. Ballesteros, et al.16 show research that tests the use of artificial intelligence (AI) and computer vision algorithms for detecting human emotions in video material as viewers view various visual stimuli. The method uses convolutional neural networks (CNNs) to process facial expression attributes and shows emotion can be sensed using deep learning-based training and implementation with software. The study is capable of recognizing many emotions, proving the strength of CNNs in useful facial feature extraction. But the authors note that more training would be required to enhance it, especially in viewing more images and recognizing very similar emotional patterns. The results attest to the effectiveness of emotion recognition techniques based on AI but indicate toward the requirement of additional algorithms and data set enlargements to enable enhanced recognition under varying real-world conditions.
Chowdhury, et al.17 describe face emotion recognition (FER) with convolutional neural networks (CNNs) and Histogram Equalization techniques to enhance the image quality and have high recognition accuracy. The paper compares the impact of histogram equalization, data augmentation, and certain model optimization techniques on various benchmarking datasets like KDEF, CK + , and FER2013. The performance is measured based on a combination of the performance metrics such as accuracy, Area under the Receiver Operating Characteristic Curve (AUC-ROC), AUC for Precision–Recall Curve (AUC-PRC), and Weighted F1-score to measure the model performance in a generalizable manner. The results stress preprocessing operations and hyperparameter tuning in FER performance improvement and ongoing optimization to maximize the use of improved generalizability on novel test sets. Kim, et al.18 propose a novel deep model, Customized Visual Geometry Group-19 (CVGG-19), which is a combination of VGG, Inception-v1, ResNet, and Xception architecture blocks to enhance facial emotion recognition. CVGG-19 tries to balance the accuracy vs. computation trade-off by mixing modules from different mature models. Experimental results show that performance of new CVGG-19 architecture significantly exceeds that of the conventional VGG-19 model with significantly better performance improvement by 59.29%. At the same time, it also decreases computational expenses by 89.5% simultaneously.
Ezquerra, et al.19 propose the use of Facial Emotion Recognition (FER) technology in problem-based learning to overcome the limitation of customary observation and self-reporting methods in assessing the involvement and affective reaction of students. Using a parametric measurement of facial response to frame-by-frame analysis, the study will allow objective, moment-by-moment measurement of students’ affect, interest, and focus of attention during learning. This is a better method compared to that of the past in the regard that it can constantly monitor states of affect without subjecting people to their individual judgments or introspections, which therefore implies that the cognitive and affective processes of learners can be diagnosed as such. In addition, its application in inquiry learning clarifies the processes through which emotions condition learning to enable more sensitive and responsive pedagogy. Kumari and Bhatia20 propose a deep learning-based facial emotion recognition system using saliency maps and advanced preprocessing algorithms to improve the recognition accuracy. Contrast-Limited Adaptive Histogram Equalization (CLAHE)-based Salient Face Detection, Generative Adversarial Network (GAN) image augmentation, bilateral filter, and saliency maps are used in the proposed system to improve the face images before feature extraction. The necessity to combine saliency-guided preprocessing with deep neural models for enhancing the trustworthiness of facial emotion recognition systems is emphasized by this paper and thus it can be a good use case to deploy in affective computing, human–computer interaction, and behavior analysis.
Zupan, et al.21 conduct a systematic review to examine facial and vocal expression of emotion recognition patterns and task and emotion expression feature influence on performance. In order to make a comprehensive analysis, systematic searching was conducted across six leading databases so that the study could cover a wide variety of existing studies on emotion recognition. Through the integration of findings across studies, the review tries to describe the variables on which accuracy and consistency of face and voice emotion recognition depend. The study highlights the degree to which the diverse experimental conditions of task difficulty, emotional intensity, and modality-specificity affect performance on recognition tasks. The study also implies possible differences in face and voice emotion processing and reveals perceptual and cognitive mechanisms of human emotion perception. Guo, et al.22 give technical explanations of machine learning-based methods to facial expression recognition, for instance, some of the key steps of the process like image preprocessing, feature extraction, and image classification. The article describes classical machine learning methods and deep learning methods with precise analysis of models like convolutional neural networks (CNNs), deep belief networks (DBNs), generative adversarial networks (GANs), and recurrent neural networks (RNNs). By comparative examination of these approaches, the authors summarize their strengths and weaknesses on the basis of the fields of accuracy, computational expense, and resistance to change of facial appearance. The survey is an ideal guidebook to researchers and practitioners who want to know the new face of face recognition of expressions, with comments on the merits and demerits of various models and avenues to be explored in the field. Younis, et al.23 introduce Automated Human Emotion Recognition (AHER) as an important area of research in Computer Science with extensive applications in marketing, human–robot interaction, electronic games, and E-learning. AHER is needed for those applications where one needs to be aware of the emotional state of an individual to enable adaptive response. The paper investigates several automatic approaches for emotion recognition using different modalities, i.e., facial expressions, text, speech and biosignals like electroencephalography (EEG), blood volume pulse (BVP) and electrocardiogram (ECG). Each of these signals is processed independently (uni-modal) or together (multi-modal) in order to improve the recognition accuracy. The authors point out that though a lot of research has been conducted, most research has been on laboratory-controlled experiments and specially crafted models with low generalizability in the real world.
However, few studies have addressed the optimization and combination of multi-modal systems, particularly with deep transformer-based models. Existing FER methods also fail to identify vague and subtle emotions that do not clearly belong to the conventional categorical sets like happiness, sadness, or anger. Human emotions are so delicate that they need more sophisticated emotion class models which can measure continuous affective dimensions like valence-arousal models in order to successfully determine why there could be a subtle expression24. There is also another very important lacuna, and that is the unexplainability and lack of interpretability of FER models based on deep learning. Besides, most FER is conducted on static images and not video streams with temporal data. Facial expressions are temporal, and their temporal progression plays an important part in giving adequate context to determine emotion. Application of FER technology to real-world problems also raises concerns regarding user consent, privacy of data, and bias in emotion detection models required to overcome25,26. The majority of the available models are biased towards one segment or another, and this yields discriminatory or wrong predictions. Such bias and ethical issues require greater effort in the area of bias avoidance techniques, federated learning for facilitating privacy-preserving FER, and formulating regulatory guidelines to facilitate responsible AI deployment.
Facial Emotion Recognition (FER) has been a longstanding research problem in classical machine learning and initial deep learning algorithms where much freedom is granted to hand-engineered features extraction. The previous studies employed geometric and appearance features such as Local Binary Patterns (LBP), Histogram of Oriented Gradient (HOG) and Gabor filters to encode displacement of facial muscles and texture changes. These techniques helped to get valuable information to learn the significance of local parts of the face- the eyes, the mouth and eyebrows in distinguishing the emotion27. However, they had less capacity to react to the variation in illumination, pose, occlusion and subject diversity due to their dependency on manual feature engineering. Additionally, these methods lacked hierarchical abstraction of features, and, therefore, were not as appropriate to detect finer emotional details. These limitations prompted the need of automated feature learning algorithms that can capture the dynamics of the multi-dimensional facial features which prompted the transition to the deep learning-based FER frameworks.
Subsequently, research has offered convolutional neural networks (CNNs) as an improvement to the drawbacks of the handcrafted feature-based algorithms because they enable to automatically learn features and in a hierarchical direct manner on raw face images. The CNN-based FER models were more resistant to variability of facial expressions where they have not just managed to capture low levels texture patterns but also high levels semantic representations. Other studies had a couple of researches that discovered the importance of more elaborate architectures and multi-scale convolutional functions to represent local muscle activities and global facial shapes. Despite these developments, CNN-based models were also prone to poor generalization performance due to bias in datasets and overfitting (especially using those with small and limited FER datasets). In addition, the locality constraint of the convolution operation in space also restricted the network to modelling long range dependencies of distant portions of the face which is critical in the perception of complex or compound expression of emotion.
To overcome the space modeling constraints of CNNs, attention mechanisms were suggested in the FER experiments to enhance the feature discrimination in CNNs by attending to the emotionally significant parts of the face. Important information on the manner in which adaptive weighting of spatial features can substantially improve the recognition performance particularly to faint emotions such as fear, disgust and sadness was also provided by the attention based FER models. These studies also established that the attention mechanism allows the model to dynamically focus on significant facially relevant parts rather than focusing all the regions equally. However, most models of attention-based CNNs employed relatively shallow or fixed attention modules, therefore, unable to learn rich inter-regional interactions. Moreover, attention processes were commonly applied to a sequence of feature extraction as a separate feature, and can again be combined with other features to make more coherent hybrid features.
Transformer-based architectures to FER have since been viewed as important to later literature, where it was reported that the mechanisms can be successful in modelling the global contextual relationships using self-attention mechanisms. The same conceptual change was offered by transformer-based FER models, which accounted for facial images as sequences of patches and hence enabled the long-range interactions among facial regions to be learnt. These techniques showed that they were able to handle pose and degree of expression changes and improve the feature contextualization. Nevertheless, the transformer models are typically computationally expensive and require large scale data to train that is not a practical task when using FER. In addition, pure transformer-based approaches are generally incapable of extracting features at low levels, which means that a hybrid network combining convolutional inductive bias with transformer attention must be used, which balances performance.
Proposed methodology
The research proposal suggests a more advanced and effective deep learning facial emotion recognition model based on the Siamese Convolutional Detection Transformer Network (SiaCon-DetNet) and Hybrid SandHawk Optimization (HySHO) with unparalleled accuracy, flexibility, and efficiency. The key contribution of the research paper is that it introduces a completely new designed and optimized structure of the deep learning model capable of extracting efficiently, processing, and classifying the intricate facial expressions with high accuracy irrespective of the occlusions, varying lighting intensities, and micro-expressions. In contrast to most deep learning-deep face emotion recognition models which are susceptible to overfitting, weak ability in feature extraction, and rigidity of parameter tuning, the new method harnesses the potential of transformer-based attention mechanism, Siamese feature extraction, and hybridised optimization so as to resist such vulnerabilities. The fusion model enables the deep learning model to extract proper local and global facial features, rendering it rather robust for emotion classification across various datasets. In addition, HySHO enhances the optimization technique with Sand Cat Swarm Optimization (SCSO) and Northern Goshawk Optimization (NGO) for effective hyperparameter tuning and adaptive emotion variance calculation. The SiaCon-DetNet is a novel advancement in emotion recognition by deep learning that leverages the strengths of Siamese networks for relative feature discovery. Detection Transformers (DETR) are typically used for spatial-attention-based classification. The typical CNNs tend not to maintain context-aware relationships between various facial expressions and thus represent the features in a warped manner. The convolutional layers of Siamese network are also performed as feature extraction modules, and DETR also reduces the obtained features with self-attention processes to allow the model to learn facial landmark contextual and spatial information. This hybrid strategy using deep learning greatly enhances generalization so that SiaCon-DetNet functions well for various demographics, lighting, and occlusions.
To further improve the performance of HySHO method is proposed to optimize model hyperparameters adaptively and conduct adaptive emotion variance calculation. Conventional optimization methods are plagued with premature convergence and enormous computational costs and hence not appropriate for real-time facial emotion recognition. HySHO transcends the challenges by embracing the exploration nature of Northern Goshawk Optimization (NGO) and exploiting the efficiency in exploitation of Sand Cat Swarm Optimization (SCSO). The NGO component ensures that it explores globally using a set of many values for hyperparameters and thereby avoids getting the process of optimization stuck in local minima. Conversely, the block of SCSO is responsible for maximally maximizing search within the potential areas in an attempt to enhance convergence speed as well as better generalization. Hierarchization with hybridization enhances SiaCon-DetNet greatly while maintaining the parameters such as learning rate, dropout rate, transformer attention head, and filter sizes within the convolution dynamically set for achieving better performances.
As shown in Fig. 1, the operational process of the given framework is in a clearly defined pipeline such that facial expressions are recorded, processed, and classified with accuracy. Facial images are preprocessed first where contrast enhancement, histogram equalization, and noise elimination are performed to make the images clear and consistent for input samples. These embeddings are contrasted with one another, allowing the model to acquire fine-grained emotional differences. The fine-grained features output are subsequently fed into a classification head, whereby emotion labels are learned from representations. For further improving classification accuracy and robustness, HySHO undertakes real-time hyperparameter optimization, tuning model parameters like convolutional kernel sizes, attention weights, and learning rates. Moreover, the AEV calculation procedure naturally controls feature extraction processes to ensure that the system can efficiently adapt to different facial expressions among different individuals and datasets.
Fig. 1.

Overview of the proposed work.
The novelty and originality of the suggested framework are based on its intelligent combination of bio-inspired top optimization and deep learning, resulting in an adaptive and efficient emotion detector. Unlike normal deep networks following pre-defined models and strict training protocols, SiaCon-DetNet enables feature representation flexibility through Siamese-based relative learning and transformer-enforced attention. HySHO also introduces a novel hierarchical optimization strategy in that the hyperparameters are learned in real time based on the complexity of the face expression to be solved. The use of AEV calculation also further improves flexibility, and the proposed framework is therefore greatly appropriate for practical application such as mental disease diagnosis, human–computer interaction, and affective computing.
The necessity to introduce a complex hybrid network such as SiaCon-DetNet using HySHO algorithm lies in the fact that the current FER structure has obvious limitations when it comes to processing fine-grained emotional signals, high intra-class variance and the possibility to transfer the results to real-world conditions. Though the conventional CNN-based models are successful in the hierarchical feature learning, models predominantly rely on local receptive fields and would thus not be useful to understand long-range interactions of distant parts of the face that are pertinent in the perception of subtle and compound affective information. Transformer-based attention models, in their turn, are effective at modelling global contexts, and they often lack good inductive biases to low-level feature extraction in faces and do not scale well to themselves. These complementary gaps are bridged by SiaCon-DetNet, which jointly applies Siamese convolutional feature learning in order to robustly make face regions on the basis of detection transformers, which invoke multi-head self-attention to refine inter-regional feature connections to permit more discriminative face representations. However, with this architectural complexity, there is a problem of hyperparameter sensitivity, training stability and computation efficiency. This is the reason why the HySHO algorithm should be presented, which makes sure that the adaptive and bio-inspired optimization is utilized to optimize the model parameters properly, improve quicker convergence, and remove unnecessary computations when training. The joint design between SiaCon-DetNet and HySHO creates a need-driven framework, rather than making it complex to achieve marginal values, the need creates a framework that tackles the issue of feature representation inadequacy, attention inefficiency, and rigidity in optimization in the existing FER systems that gives a balance to address the issues of accuracy, robustness, and efficiency in the emotion recognition tasks.
C. Siamese convolutional detection transformer network (SiaCon-DetNet)
The SiaCon-DetNet stands as a sophisticated deep learning structure which merges Siamese Networks along with Convolutional Neural Networks (CNNs) and Detection Transformers (DETR) for better facial emotion recognition (FER). This method achieves specific performance by analyzing refined spatial characteristics of the face while building connections between facial elements while using self-attention to identify crucial facial areas for emotion detection. The major drawback of transformer-based models is their need for substantial computational power together with large datasets to achieve successful training despite their strong global context awareness. SiaCon-DetNet achieves superior facial emotion recognition performance through its combination of Siamese network architecture with convolutional features and DETR-based attention systems which excerpts the best properties from each framework. Through this combined framework feature extraction performs reliably for small or hidden expressions while handling different kinds of real-world facial characteristics such as obstructed faces and changing illuminations and poses.
SiaCon-DetNet operates through multiple operational layers which optimize both feature extraction and comparison as well as classification procedures. The Siamese network includes two parallel CNN branches which operate on matching facial images while sharing identical weight parameters between identical (same emotion) or different (different emotions) image pairs. Deep spatial features emerge during CNN branch processing due to convolutional layers and batch normalization with ReLU activation functions as the elements that support hierarchical representation learning. The processing progresses from edge and gradient detection in lower layers toward the discrimination of emotional states including happiness and sadness and anger and surprise in higher layers. Feature similarities are computed through Siamese structure by using a contrastive loss or cosine similarity function that optimizes intra-class compactness together with inter-class separability.
The SiaCon-DetNet stands as a sophisticated deep learning structure which merges Siamese Networks along with Convolutional Neural Networks (CNNs) and Detection Transformers (DETR) for better facial emotion recognition (FER). This method achieves specific performance by analyzing refined spatial characteristics of the face while building connections between facial elements while using self-attention to identify crucial facial areas for emotion detection. The major drawback of transformer-based models is their need for substantial computational power together with large datasets to achieve successful training despite their strong global context awareness. SiaCon-DetNet achieves superior facial emotion recognition performance through its combination of Siamese network architecture with convolutional features and DETR-based attention systems which excerpts the best properties from each framework. Through this combined framework feature extraction performs reliably for small or hidden expressions while handling different kinds of real-world facial characteristics such as obstructed faces and changing illuminations and poses. SiaCon-DetNet operates through multiple operational layers which optimize both feature extraction and comparison as well as classification procedures. The Siamese network includes two parallel CNN branches which operate on matching facial images while sharing identical weight parameters between identical (same emotion) or different (different emotions) image pairs. Deep spatial features emerge during CNN branch processing due to convolutional layers and batch normalization with ReLU activation functions as the elements that support hierarchical representation learning. The processing progresses from edge and gradient detection in lower layers toward the discrimination of emotional states including happiness and sadness and anger and surprise in higher layers. Feature similarities are computed through Siamese structure by using a contrastive loss or cosine similarity function that optimizes intra-class compactness together with inter-class separability.
As shown in Fig. 2, SiaCon-DetNet layer-by-layer processing is well-organized and optimized with preservation of computational effectiveness as well as improved feature extraction. Feature detection at low-level edge-level relies on the initial convolutional layers, feature discrimination is improved by the Siamese twin network architecture, and DETR-based transformer layers add spatial consciousness as well as long-range dependencies to enable stronger expression classification. The final layer of classification ensures the model not only learns separable emotion representations but also generalizes to novel facial expressions. Two-class similarity and intra-class difference, two classical difficulties of facial emotion perception, are two of the strongest strengths of SiaCon-DetNet. Most emotions, such as fear and surprise or anger and disgust, have the same facial muscle activations and therefore are hard to differentiate using typical CNNs. Contrastive learning against the Siamese network does, however, assist the model to distinguish between fine variations through learning an effective distance measure between classes of emotion. Second, the DETR module trained using transformers allows the model to place the expressions into context in a manner where even small changes in facial expressions are not capable of misclassifying. It is especially useful for application where there is a requirement of fine-grained emotion detection, e.g., in psychological assessment, human–computer interaction, and affective computing for intelligent systems.
Fig. 2.

Flow of the proposed SiaCon-DetNet model.
The other salient strength of SiaCon-DetNet lies in its computational efficiency as well as ability to scale large emotion datasets. Regular transformer-based models, although computationally efficient, are time and resource hungry and require mammoth training sets. As one attempt to try and close the gap, SiaCon-DetNet saves on the computationally expensive effort by using light-weight convolution backbones, shared-weight Siamese learning, and a limited transformer module judiciously samples on key facial features rather than sweeping across the whole image in bulk. This makes real-time deployment in real-time applications, such as AI-driven surveillance systems, emotion-driven virtual assistants, and learning systems wherein the scope of student engagement needs to be established is key, feasible. The model in SiaCon-DetNet also becomes more robust against variability in the environment and demographics. All previously known FER models up to now are not optimal where age, ethnic, gender, and skin transitions happen and yield biased recognition output. This ensures that the system behaves in a consistent fashion under different sets of users, lighting, and camera resolutions so that it can be implemented in actual, full-scale settings.
After getting the input image, the preprocessing is performed with the use of logarithmic enhancement operation as shown in the following equation:
![]() |
1 |
where,
represents the grayscale intensity of the input image with coordinates
. As a consequence
of this, the affine transformation is applied to compute the Siamese feature
distance based on the following model:
![]() |
2 |
![]() |
3 |
where,
are the feature embedding,
is the
weight value, and
denotes the bias value. Then, the custom distance measure is computed using the
following equation:
![]() |
4 |
where,
is the tunable
sensitivity parameter. Then, the multi-head self attention function in
transformer is estimated with the emotion specific scaling module as shown in
the following equation:
![]() |
5 |
where,
is the query, key and value parameters,
denotes the feature dimension,
represents the scaling factor, and
is the emotion specific feature variance. In addition, the transformer based
emotion query is represented according to the following model:
![]() |
6 |
where,
represents the
non-linear activation function, and
is
the query embedding weight matrix. Furthermore, the attention guided emotion
prediction is performed with the use of laplacian model as mathematically
expressed in below:
![]() |
7 |
where,
represents the
number of attention heads,
denotes the
attention coefficient,
is
the transformed feature, and
is the
learnable parameter. Then, the Siamese contrastive loss function is estimated
with adaptive temperature scaling model according to the following
equation:
![]() |
8 |
where,
is the scaling
parameter, and
denotes the
marginal threshold. In addition to that, the graph based feature fusion is
performed with the use of Laplacian Eigenmaps as mathematically expressed
below:
![]() |
9 |
where,
is
the laplacian matrix. Furthermore, the weight regularization is performed
according to the emotion entropy constraint parameter as represented in
below:
![]() |
10 |
where,
is the
prediction control parameter. Then, the multi-scale feature aggregation is
performed according to the following model:
![]() |
11 |
where,
is the
trainable attention weight value. Moreover, the squeeze excitation mechanism is
applied with the bi-directional feature fusion as shown in below:
![]() |
12 |
Finally, the decision function is computed for determining the accurate emotion as shown in below:
![]() |
13 |
where,
is used
to map the final feature embedding.
D. Hybrid Sandhawk optimization (HySHO) for adaptive emotion variance computation
HySHO represents a new adaptive optimization framework which unites SCSO and NGO to operate effectively for Adaptive Emotion Variance calculations used in facial emotion recognition tasks. The primary objective of this technique exists to improve the precision combined with adjusting abilities and resilient performance of facial emotion recognition (FER) through dynamic deep learning model parameter adjustments. Traditional optimization approaches face difficulties when operating within complex environments which run high-dimensional searches particularly in systems that need real-time adoption for various facial expression changes. The predatory behavior of Northern Goshawks and stealth hunting of Sand Cats forms HySHO as an evolutionary hybrid which allows effective exploration and exploitation balancing in search spaces. By uniting these robust nature-inspired approaches the model becomes able to change its learning parameters automatically which leads to improved real-time operation together with enhanced generalization features for emotion recognition systems. HySHO achieves uniqueness through its hybrid optimization approach which first utilizes NGO to explore the entire space and secondly employs SCSO for local refinement. HySHO avoids premature convergence issues and local optimum traps since it executes aerial scouting for aggressive search and stealthy tracking for adaptive fine-tuning simultaneously. During the first phase of HySHO a large variety of population options are established through NGO’s extended hunting pattern to properly assess model hyperparameter variables like learning rate and weight decay and dropout rates and convolutional filter sizes. ScSO uses adaptative hunting methods in phase two to optimize model performance by concentrating searches on search space regions which cuts down processing expenses while sharpening face feature perception. The hierarchical optimization approach improves efficiency so HySHO becomes more stable, scalable and computationally efficient than traditional methods do.
One of the main benefits of HySHO for facial emotion recognition is dynamic variance computation, whereby the model is dynamically learning and adapting its learning parameters according to variability in facial expressions. Facial emotions are nonlinear and context-sensitive, i.e., an optimization policy set will be incapable of perceiving minute variations in expressions, e.g., blended or micro-expressions. HySHO controls feature extraction layers of deep neural networks closely by repeatedly flipping filter weights and threshold activations in order to provide more and context-sensitive classification. It is highly useful in practical scenarios since illumination, pose, and occlusions are common factors that prevent traditional emotion recognition systems from working perfectly. Compared to other available methods, HySHO is distinct as it takes a hybrid optimization approach that exploits the advantages of two different bio-inspired algorithms. HySHO does not face this drawback since it provides a hierarchical search architecture where NGO provides global convergence and SCSO provides fine-grained hyperparameter regulation and is thus very efficient for adaptive hyperparameter tuning. Additionally, HySHO’s adaptive exploration–exploitation mechanism saves on computations, making it cheaper to apply in real-life uses of affect detection in human–computer interaction, monitoring mental well-being, and affect computing. HySHO’s novelty and robustness position it at the peak of the future facial affect analysis system. Premature convergence problem inherent to other models is the fact that these methods have a tendency of converging too early to poor solutions, without visiting the rest of the search space. HySHO, on the other hand, employs NGO as its global search component, thus the method never stops to discover new hyperparameter configurations without converging. This ability to explore ensures the varied areas of the search space are explored and avoids the optimization procedure from falling into local minima.
Yet another important benefit of HySHO is in its multi-layered optimization strategy, in that it is capable of accommodating effective and adaptive parameter tuning. PSO and GA are traditional methods and static parameter setting-based, thus are ineffective in responding to complicated facial emotion changes. Conversely, HySHO has an adaptive search strategy wherein NGO carries out extensive exploratory searches over hyperparameter spaces and Sand Cat Swarm Optimization (SCSO) obtains best solutions through localized exploitation. The two-stage optimization process enables the system to learn more effectively and operate optimally while ensuring that the deep learning model can capture fine facial emotional changes on various subjects and environmental conditions. Further, the adaptive emotion variance estimation (AEV) that accompanies HySHO provides a secondary adaptability with dynamically adaptive feature extraction processes. This is achieved in order to enable the deep learning model to capture rich expression and subtle emotions more effectively and hence appropriately suited to real-world applications like mental analysis, human–computer interaction, and affective computing.
The other unique strength of HySHO is its noise and occlusion robustness in facial emotion data. One of the primary difficulties in facial emotion recognition is dealing with occlusions (e.g., hands over part of the face), varying lighting, and poor micro-expressions. Most classical optimization methods are powerless to dynamically adapt to these challenging instances and result in decreased recognition performance. HySHO, however, has adaptive fine-tuning techniques that allow the deep learning model to learn invariant and strong features. AEV computation use ensures variations in features due to occlusions and lighting changes are addressed, hence enhancing the capacity of the model to recognize between expressions that are similar. This property makes HySHO especially valuable in real-world usage scenarios where facial affect information is untrustworthy or subject to partial occlusions. Additionally, the hybridization strategy of HySHO improves decision-making and interpretability. The conventional optimisation methods make use of an inbuilt, pre-determined fixed set of heuristics or fitness functions to make decisions, which would not necessarily follow the harmony between the dynamic scenario of facial emotion recognition tasks and these conventional approaches. HySHO has a decision-making strategy that is adaptive, and in it, final classification is implemented by a scheme of confidence-weighted voting so that the final facial expression classification that is optimal and context-appropriate is guaranteed. This implies that HySHO optimizes deep learning parameters efficiently but also maximizes the interpretability of emotion perception outcomes, making it a useful tool in such critical applications like medicine, psychological assessment, and surveillance systems.
Results and discussion
The results section has a detailed discussion of the performance of the proposed model under three popular facial expression datasets, including JAFFE, CK + , and MMI28–31. The datasets include varied facial emotion samples that are necessary for evaluating the robustness of the model in identifying emotions accurately. The JAFFE database contains 213 grayscale facial images of 10 Japanese women with each showing seven states of emotions: Angry, Disgust, Fear, Happy, Neutral, Sad, and Surprised. JAFFE is highly sought after for facial expression analysis because it has high annotations using good-quality emotion labels and grayscale facial images. Even though it is widely used, its small data size presents difficulties to deep learning models needing a huge volume of training data. The CK + (Extended Cohn-Kanade) dataset has 593 video sequences of 123 individuals, seven basic emotions: Angry, Contempt, Disgust, Fear, Happy, Sad, and Surprised. CK + is a highly used facial expression recognition dataset because it has both static images as well as video sequences to enable models to learn temporal changes in facial expressions. The MMI dataset consists of more than 2900 video sequences captured from 25 subjects displaying six universal emotions: Angry, Disgust, Fear, Happy, Sad, and Surprised. MMI is different from JAFFE and CK + in posed and spontaneous facial expressions and thus serves as a good benchmark for evaluating emotion recognition performance under real-world conditions. These datasets combined provide a general reference to compare the efficiency of the proposed method.
The HySHO framework and SiaCon-DetNet proposal is realistic and well-built, and to further justify the proposed theory of performance it is essential that the proposed prototype be put to test with more difficult challenges of FER. Despite the datasets such as JAFFE being quite useful, in case of well-aligned and noise-free environment, it is not a representation of the complexities that are involved in the real world scenarios, involving variations in illumination, pose, occlusion, background clutter, intensity of expressions and demographic variations. The generalization of the proposed model to other appearances of faces and spontaneous displays of emotions might be cautiously checked through the integration of a large amount of in-the-wild data, such as FER2013, RAF-DB, AffectNet, and SFEW among others. The following datasets present real-life problems such as low-resolution images, coverage of face, and disproportionate representation of classes that are significant in evaluating the effectiveness of the advanced features extraction and attention systems. Wide validation would be a long way in confirming the relevancy, applicability and practicability of the suggested FER approach.
Figure 3 shows the sequential processing stages of the proposed facial expression recognition model, exhibiting how an input image is boiled down to its final emotion label. The first stage (Input Image) is the raw face image gathered from either of the above-mentioned datasets (JAFFE, CK + , or MMI). In the second phase (Face Detected), a face detection algorithm within the model is employed to determine and extract the region of the face accurately so that background information that is not required is removed. In the third phase (Pattern Extracted) is feature extraction in which unique patterns of the face like wrinkles, contours, and differences in texture are analyzed through deep feature maps. The fourth step (Grad-CAM Visualization) provides an interpretability feature by showing the most important regions of the face that the model is using to make the prediction, thus better understanding how the model made the decision is provided. Lastly, the predicted emotion label from the extracted features, closing the facial expression recognition pipeline. The step-by-step transformation ensures a proper and readable emotion classification process, making the suggested model more reliable.
Fig. 3.
Sample input and output images.
In Fig. 4 (JAFFE dataset), correlations are noted for the patterns on how expressions such as “Happy” and “Surprised” have some of the same feature attributes, while “Fear” and “Sad” have the same features because of the activations of similar facial muscles. Likewise, Fig. 5 (MMI dataset) notes strong correlations in between emotions such as “Angry” and “Disgust,” since they occupy the similar facial deformities, while “Happy” remains mostly isolated because of its distinctive features of expression. Figure 6 (CK + dataset) sheds more light, especially with the introduction of “Contempt,” which is not included in the other datasets, and has moderate correlation with emotions such as “Disgust” and “Angry” because they share similar facial expressions. These correlation analyses assist in interpreting how the model is recognizing the emotions, indicating probable misclassifications, and fine-tuning feature extraction methods in order to enhance the accuracy of emotion recognition in various datasets.
Fig. 4.

Correlation analysis with respect to different emotions of JAFFE dataset.
Fig. 5.

Correlation analysis with respect to different emotions of MMI dataset.
Fig. 6.

Correlation analysis with respect to different emotions of CK + dataset.
The accuracy curve in Fig. 7a indicates continuous improvement in performance, and this indicates that the model learns distinct facial expression features well in later epochs. Similarly, Fig. 7b is the loss curve, with the decreasing trend of both training and validation loss indicating a sign of correct model convergence and optimization. Consistency in the loss measures towards the later epochs is a promise that the learning process is well regularized without any overfitting.
Fig. 7.

(a). Training and validation accuracy for JAFFE dataset (b). Training and validation loss for JAFFE dataset.
Figure 8a, b indicate the MMI dataset training and validation performance measurements. Accuracy trend in Fig. 8a indicates that the accuracy has some fluctuation before it stabilizes into a linear increase, and this can be attributed to the MMI dataset nature being complex in terms of having illumination changes and occlusions. The model subsequently becomes proficient in being capable of having high accuracy, which also means that it has learned patterns unique to every emotion. The given loss curve in Fig. 8b shows a declining trend, proving effectiveness of optimization. The overall smooth reduction of loss proves effective learning, but slight fluctuations indicate that there can be some emotions in the dataset that have to be fine-tuned by features to classify them optimally.
Fig. 8.

(a). Training and validation accuracy for MMI datase (b). Training and validation loss for MMI dataset.
Figure 9a, b show training and validation accuracy and loss curves for the model trained with CK + dataset. One can observe in Fig. 9a that there is a drastic increase in the accuracy, and the model strongly rapidly adapts with the CK + dataset, where facial expressions of well-annotated images have been used. Again, strong generalization is guaranteed with stability in the training and validation accuracy. In Fig. 9b, loss becomes steadily decreasing, which only favors the strong learning process. The smooth convergence of loss values reflects the good performance of the proposed method in learning emotion-specific features and achieving better results on the CK + dataset compared to JAFFE and MMI.
Fig. 9.

(a). Training and validation accuracy for CK + dataset (b). Training and validation loss for CK + dataset.
Figure 10 is the performance testing of JAFFE dataset, where precision, recall, F1-score, and accuracy are presented for various emotions like Angry, Disgust, Fear, Happy, Neutral, Sad, and Surprised. The result is that all the parameters fall in the same range of 98% to 99%, showing the effectiveness of the proposed method in facial emotion recognition with very minor fluctuation. The results show that the model demonstrates consistent performance when it analyzes different facial expressions from the JAFFE database. The MMI dataset performance results show the six emotions of Angry, Disgust, Fear, Happy, Sad, and Surprised in Fig. 11. The model achieves 98–99% accuracy based on its actual precision recall and F1-score results which meet JAFFE emotional detection standards. The model reaches perfect classification results because it successfully identifies all emotional expressions despite their small between-emotion variations. The method demonstrates effective performance because all emotional features from the MMI corpus show complete convergence across all assessment measures. Figure 12 displays CK + dataset performance through its analysis of Angry, Contempt, Disgust, Fear, and Happy, Sad, and Surprised facial expressions. The method achieves high classification accuracy through its precision and recall and F1-score and accuracy which all reach the 98–99% threshold. The model achieves high accuracy in emotional recognition because it combines two emotions that are similar but distinct into one emotion category. The results show that the model can identify multiple emotions because it tests their performance with different facial expression databases.
Fig. 10.

Performance analysis with respect to different emotions of JAFFE dataset.
Fig. 11.

Performance analysis with respect to different emotions of MMI dataset.
Fig. 12.

Performance analysis with respect to different emotions of CK + dataset.
Figure 13 illustrates a comparative performance of the model under consideration with others of state-of-the-art quality on the basis of accuracy using the JAFFE dataset32. The study assesses traditional methods and deep learning-based methods through its accuracy testing which starts from 83.33 percent and extends to 99.20 percent. The traditional methods show Fine-grained analysis at 83.33% while McFIS and LPTP achieve superior results with their respective scores of 87.60% and 90.20%. The spatial and frequency-domain feature extraction techniques which use SRC + LBP and SRC + Gabor methods show improved performance with 90.30% and 91.21% accuracy results respectively. The advanced learning-based methods LSDP, FER-Net, and Reversible Neural Networks show substantial accuracy improvements with their respective accuracy scores of 92.30%, 95.00%, and 95.90%. The proposed model achieves 99.20% accuracy which surpasses all existing models because it demonstrates superior ability to extract intricate facial emotion features with maximum accuracy (Fig. 14).
Fig. 13.

Comparison with recent state of the model approaches based on accuracy using JAFFE dataset.
Fig. 14.

Comparison with recent models based on test accuracy using CK + dataset.
The study assesses the test accuracy of deep learning methods through their performance on the CK + dataset which includes facial emotion recognition tests. The accuracy of VGG19 reaches 88% and ResNet50 achieves 87% accuracy. Xception achieves higher results than all other methods because it reaches an accuracy of 91% while EfficientNetB0 surpasses this achievement with 93% accuracy. The model presented in this study achieves 99% test accuracy which proves its superiority over all other architectural designs because it uses CG + configuration to extract emotion recognition features from data. The training time of the same models which use the CK + dataset shows computational efficiency through its comparison of different models in Fig. 15. The training time for DenseNet121 lasts 6.38 s, while InceptionV3 requires only 4.25 s to complete its training. VGG16 requires the most time to complete training because it needs 3.48 s, and VGG19 needs 4.48 s to finish its training. ResNet50 requires 4 s to complete training, while Xception needs 3.58 s because both models achieve optimal results between training speed and accuracy. EfficientNetB0 requires 4.46 s to complete its training. The proposed model achieves the shortest training time of 3.08 s, while it achieves the highest accuracy because it shows proficiency in fast training convergence and effective resource optimization.
Fig. 15.

Comparison with recent models based on training time using CK + dataset.
Figure 16 shows test accuracy of other models on JAFFE data, with remarkable differences in how well they can perform. DenseNet121 comes at 88%, followed closely by 82% for InceptionV3 and 90% for VGG16. VGG19, with feature learning capability strength, comes out at 82%, which was just a step lower than that of VGG16. ResNet50 was 87%, and Xception was 78%. EfficientNetB0 performed as equally well as DenseNet121 with 88%. The model outperformed all the other models with a 99.10% test accuracy, which speaks volumes about its prowess in successful learning of complex emotional patterns with the JAFFE dataset. Comparison of training time among models on the JAFFE dataset can be seen in Fig. 17 and gives a feeling of computational proficiency. DenseNet121 trains in 8.7 s, and InceptionV3 trains in 9.1 s, just slightly more. Training is performed in 7.4 s by VGG16 while VGG19 takes a lesser time of 7 s. EfficientNetB0 trains optimally at 6.8 s. The proposed model, however, performs the training time as low as 3.12 s, showing its high efficiency in minimizing the computational cost with the maximum test accuracy.
Fig. 16.

Comparison with recent models based on test accuracy using JAFFE dataset.
Fig. 17.

Comparison with recent models based on training time using JAFFE dataset.
The SiaCon-DetNet system demonstrates superior recognition capabilities together with enhanced class performance when it is compared to HySHO and the leading FER systems available today. The typical accuracy range for traditional CNN-based FER systems on controlled datasets exists between 88 to 94 percent. Recent transformer-based FER models achieve global feature representation improvements which enable them to reach 97 to 98 percent accuracy. The model achieves 99.20 percent accuracy on the JAFFE dataset which represents an important transformation of 1 to 3 percent when compared to existing systems that use transformer-based and attention-based methods. The proposed framework achieves precision, recall, and F1-score results between 98 and 99 percent for all emotion categories. The competing models show severe performance decline because they cannot handle subtle emotions, including fear and disgust which results in F1-scores that drop below 95 percent. The SiaCon-DetNet system together with HySHO maintains consistent class performance throughout all emotions which results in decreased performance variability. The implementation of HySHO enables faster model training by achieving better accuracy than other deep learning and hybrid model approaches which require less time to reach their best performance. The numerical results show that our proposed method provides better overall results than existing FER techniques together with more accurate and consistent emotional recognition capabilities.
Conclusion
This paper presents a holistic facial emotion recognition approach by combining a state-of-the-art deep learning framework with efficient feature extraction and visualization techniques. This paper presents a new facial emotion recognition model using the incorporation of the SiaCon-DetNet and HySHO to realize efficient and precise emotion classification. This paper’s main contributions are made at the architecture level of SiaCon-DetNet that harnesses the capability of feature extraction in convolution networks and attention-driven detection efficacy in transformer architectures for delivering accurate facial expression analysis.
The operational process of the proposed method entails a few major steps: SiaCon-DetNet-based preprocessing and face detection initially, feature extraction, and pattern refinement later. The HySHO then utilizes a hybrid optimization algorithm to optimize network weights for convergence learning. The model is tested comprehensively on the JAFFE, CK + , and MMI databases and exhibits robustness in detecting fine-grained emotional nuances.
Moreover, correlation analysis of different emotions even confirms the strength of feature extracted, further confirming the strength of the model in real usage.
The proposed SiaCon-DetNet with HySHO achieves maximum accuracy of 99.20% on JAFFE and surpasses baseline models like FER-net and Reversible Neural Networks. Comparison with some other newer deep learning models in CK + and JAFFE databases demonstrates the improved performance of the suggested approach with not just the maximum test accuracy but also the least training time of 3.08 s in CK + and 3.12 s in JAFFE. Comparison analysis for emotions also ensures accuracy, recall, and F1-score within 98–99%, attesting to the stability of the model. These outcomes validate that the developed SiaCon-DetNet and HySHO-based facial emotion recognition system is highly accurate, computationally light, and clear and hence maximally suitable to be employed for affective computing, human–computer interaction, and psychological investigation.
Author contributions
MS wrote the manuscript and CM, KT, AA and VD supervised and proofread the manuscript. All authors approved the manuscript.
Data availability
Data used in this article are publicly available in a repository are available at the following URLs: JAFFE-CK 48 https://share.google/JHWjxdu2fpXV9dRrC, MMA FACIAL EXPRESSION- https://share.google/0Vxg30tk67o3XnHqI, CK+ Dataset- https://share.google/VGTl9wE9gDfXuQJgo.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Change history
6/19/2026
The original online version of this Article was revised: The original version of this Article contained an error in Affiliation 4, which was incorrectly given as ‘Department of Artificial Intelligence and Data Science, Rajalakshmi Institute of Technology, Chembarambakkam, Tamil Nadu 600124, India.’ The correct affiliation is listed here: ‘Department of Computer Science and Engineering, SRM Institute of science and Technology, Vadapalani, Chennai, India’. Additionally, the Data availability section was incomplete. It now reads: “Data used in this article are publicly available in a repository are available at the following URLs: JAFFE-CK 48 https://share.google/JHWjxdu2fpXV9dRrC, MMA FACIAL EXPRESSION- https://share.google/0Vxg30tk67o3XnHqI, CK+ Dataset- https://share.google/VGTl9wE9gDfXuQJgo.” The original Article has been corrected.
References
- 1.Zhou, H., Huang, S. & Xu, Y. UA-FER: Uncertainty-aware representation learning for facial expression recognition. Neurocomputing621, 129261 (2025). [Google Scholar]
- 2.Halkiopoulos, C., Gkintoni, E., Aroutzidis, A. & Antonopoulou, H. Advances in neuroimaging and deep learning for emotion detection: A systematic review of cognitive neuroscience and algorithmic innovations. Diagnostics15, 456 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Madni, S. H. H., Pathmanatan, L. A. /L., Faheem, M., Shahzad, H. M. F. & Shah, S. Exploring optimizer efficiency for facial expression recognition with convolutional neural networks. J. Eng.2025, e70060 (2025). [Google Scholar]
- 4.Zhang, F. et al. Towards facial micro-expression detection and classification using modified multimodal ensemble learning approach. Inf. Fusion115, 102735 (2025). [Google Scholar]
- 5.Kumar, A. & Kumar, A. Human emotion recognition using machine learning techniques based on the physiological signal. Biomed. Signal Process. Control100, 107039 (2025). [Google Scholar]
- 6.Akrout, B. Deep facial emotion recognition model using optimal feature extraction and dual-attention residual U-Net classifier. Expert Syst.42, e13314 (2025). [Google Scholar]
- 7.Balachandran, G., Ranjith, S., Chenthil, T. & Jagan, G. Facial expression-based emotion recognition across diverse age groups: A multi-scale vision transformer with contrastive learning approach. J. Comb. Optim.49, 1–39 (2025). [Google Scholar]
- 8.Fatima, N. S. et al. Enhanced facial emotion recognition using vision transformer models. J. Electr. Eng. Technol.10.1007/s42835-024-02118-w (2025). [Google Scholar]
- 9.Di Luzio, F., Rosato, A. & Panella, M. An explainable fast deep neural network for emotion recognition. Biomed. Signal Process. Control100, 107177 (2025). [Google Scholar]
- 10.Kumar, A., Kumar, A. & Gupta, S. Machine learning-driven emotion recognition through facial landmark analysis. SN Comput. Sci.6, 1–10 (2025). [Google Scholar]
- 11.Wu, H. et al. A video-based cognitive emotion recognition method using an active learning algorithm based on complexity and uncertainty. Appl. Sci.15, 462 (2025). [Google Scholar]
- 12.Devi, B. & Preetha, M. M. S. J. Facial emotion recognition using convolutional neural network based krill head optimisation. Expert Syst.42, e13376 (2025). [Google Scholar]
- 13.Li, Q., Liu, Z., Zhang, Z., Wang, Q. & Ma, M. Decoding group emotional dynamics in a web-based collaborative environment: A novel framework utilizing multi-person facial expression recognition. Int. J. Hum. Comput. Interact.41, 3455–3473 (2025). [Google Scholar]
- 14.Tan, Y., Xia, H. & Song, S. Mining label-free consistency regularization for noisy facial expression recognition. Complex Intell. Syst.11, 105 (2025). [Google Scholar]
- 15.Gursesli, M. C. et al. Facial emotion recognition (FER) through custom lightweight CNN model: Performance evaluation in public datasets. IEEE Access10.1109/access.2024.3380847 (2024). [Google Scholar]
- 16.Ballesteros, J. A., Ramírez V, G. M., Moreira, F., Solano, A. & Pelaez, C. A. Facial emotion recognition through artificial intelligence. Front. Comput. Sci.6, 1359471 (2024). [Google Scholar]
- 17.Chowdhury, J. H., Liu, Q. & Ramanna, S. Simple histogram equalization technique improves performance of VGG models on facial emotion recognition datasets. Algorithms17, 238 (2024). [Google Scholar]
- 18.Kim, J. H., Poulose, A. & Han, D. S. Cvgg-19: Customized visual geometry group deep learning architecture for facial emotion recognition. IEEE Access12, 41557–41578 (2024). [Google Scholar]
- 19.Ezquerra, A., Agen, F., Toma, R. B. & Ezquerra-Romano, I. Using facial emotion recognition to research emotional phases in an inquiry-based science activity. Res. Sci. Technol. Educ.43, 62–85 (2025). [Google Scholar]
- 20.Kumari, N. & Bhatia, R. Saliency map and deep learning based efficient facial emotion recognition technique for facial images. Multim. Tools Appl.83, 36841–36864 (2024). [Google Scholar]
- 21.Zupan, B. & Eskritt, M. Facial and vocal emotion recognition in adolescence: A systematic review. Adolesc. Res. Rev.9, 253–277 (2024). [Google Scholar]
- 22.Guo, X., Zhang, Y., Lu, S. & Lu, Z. Facial expression recognition: A review. Multim. Tools Appl.83, 23689–23735 (2024). [Google Scholar]
- 23.Younis, E. M., Mohsen, S., Houssein, E. H. & Ibrahim, O. A. S. Machine learning for human emotion recognition: A comprehensive review. Neural Comput. Appl.36, 8901–8947 (2024). [Google Scholar]
- 24.Jiang, S. et al. CSE-GResNet: A simple and highly efficient network for facial expression recognition. IEEE Trans. Affect. Comput.10.1109/taffc.2025.3535811 (2025). [Google Scholar]
- 25.Song, X. Emotional recognition and feedback of students in English e-learning based on computer vision and face recognition algorithms. Entertain. Comput.52, 100847 (2025). [Google Scholar]
- 26.Leong, S. C., Tang, Y. M., Lai, C. H. & Lee, C. K. M. Facial expression and body gesture emotion recognition: A systematic review on the use of visual data in affective computing. Comput. Sci. Rev.48, 100545 (2023). [Google Scholar]
- 27.Y. H. Yip and J. Hu, Enhanced Multi-scale Hierarchical Network for Micro-expression Recognition, In International Conference on Image and Graphics, 2025, pp. 40–51.
- 28.Kumari, N. & Bhatia, R. Deep learning based efficient emotion recognition technique for facial images. Int. J. Syst. Assur. Eng. Manag.14, 1421–1436 (2024). [Google Scholar]
- 29.Dujaili, M. J. Survey on facial expressions recognition: Databases, features and classification schemes. Multimed. Tools Appl.83, 7457–7478 (2024). [Google Scholar]
- 30.F. Aghabeigi, S. Nazari, and N. Osati Eraghi, An efficient facial emotion recognition using convolutional neural network with local sorting binary pattern and whale optimization algorithm, Int. J. Data Sci. Anal., 2024/07/08 2024.
- 31.Haider, I., Yang, H.-J., Lee, G.-S. & Kim, S.-H. Robust human face emotion classification using Triplet-loss-based deep CNN features and SVM. Sensors23, 4770 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Li, S., Wang, J., Tian, L., Wang, J. & Huang, Y. A fine-grained human facial key feature extraction and fusion method for emotion recognition. Sci. Rep.15, 6153 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
Data used in this article are publicly available in a repository are available at the following URLs: JAFFE-CK 48 https://share.google/JHWjxdu2fpXV9dRrC, MMA FACIAL EXPRESSION- https://share.google/0Vxg30tk67o3XnHqI, CK+ Dataset- https://share.google/VGTl9wE9gDfXuQJgo.














