Abstract
The proposed framework achieved the highest accuracy of 94.2% in engagement prediction, 92.3% in dropout risk forecasting, and 91.5% in cognitive load estimation, surpassing contemporary baselines by 4% to 12% on all evaluation metrics. Adaptive learning pathways produced by the RL-ALP component enhance distance learning outcomes by 14.3% and decrease dropout rates by approximately 20.3%. Furthermore, real-time implementation is enabled due to the framework's computational efficiency, achieving inference time at 21.3 ms, which is about 50% faster than prior works. The findings thus affirm this system's capability of combining the behavioral, physiological, and interactional data modalities for accurate real-time monitoring of engagement and adaptive learning optimizations.
-
1.
A multi-modal framework integrating behavioral, physiological, and interaction-based features for engagement assessment.
-
2.
Five interconnected models leverage real-world datasets to improve cognitive load estimation and dropout prediction.
-
3.
Data-driven insights assist educators in optimizing learning strategies for better student retention and performance.
Keywords: Student engagement, Machine learning, Adaptive learning, Cognitive load estimation, Graph neural networks
Graphical abstract
Related research article
“None”.
Specifications table
“None”.
Background
Assessing student engagement and learning outcomes is a core problem in education today, especially in digital and remote learning platforms. Engagement is a predictor of academic success affecting retention rates, cognitive development, and overall performance. Existing engagement assessment methodologies [[1], [2], [3]] mainly utilize subjective self-report surveys or observations by instructors, besides very basic forms of interaction analytics that are discrete and do not capture an extended range of behaviors. Such assessment systems have therefore failed to evaluate multi-dimensional real-time behavioral cues that would have guided interventions or adaptation of educational content within the limited time frame available. In turn, traditional engagement tracking measures, for instance, Clickstream Analytics, quiz performance tracking, class participation forums, etc., present only partial delay to remedy a student’s emotional or cognitive state inside which timely pedagogical interventions can be executed. In recent years, multi-modal machine learning has presented opportunities for a holistic analysis of engagement through integration of data from different modalities such as facial expression, gaze tracking, keystroke dynamics, physiological response, and peer-interaction patterns engaged together. Fusing heterogeneous engagement indicators accurately, eliminating sensor-specific biases, and extracting high-level engagement representations [[4], [5], [6]] from complex temporal data samples remain significant challenges in this endeavor. This research aims at overcoming some of these challenges through proposing an Integrated Multi-Modal Machine Learning Framework that represents a combination of self-attention mechanisms, graph modeling, reinforcement learning, cognitive load estimation, and contrastive learning techniques to evaluate and improve student engagement in real-time. After enhancing student engagement prediction, the implemented system creates adaptive pathways for learning whereby tailored experiences are created on the individual cognitive status of the learner sets.
The framework consists of five core models, namely:(1) Multi-Modal Self-Attention Fusion Network (MMSAF-Net), which integrates real-time physiological and behavioral signals using self-attention-based feature fusion technique; (2) Graph-Based Student Interaction Network (G-SIN), which constructs a peer learning graph using Graph Neural Networks (GNNs) to analyze student collaboration patterns; (3) Reinforcement Learning-Based Adaptive Learning Pathway (RL-ALP), applying Deep Q-Learning to dynamically adjust content complexity; (4) Attention-Based Cognitive Load Estimation Network (A-CLE-Net), which employs BERT embeddings and self-attention mechanisms to classify cognitive load levels; and (5) Temporal Contrastive Learning Network (TCL-Net), which tracks engagement fluctuations over timestamp to predict dropout risk with high precision. The models in this study were validated using real-world datasets, which include OpenEDU, DAiSEE, and engagement logs from MOOC platforms, proving to outperform existing methods in engagement evaluation. These research results provide solid insights for educators, e-learning platforms, and policy-makers about applying a more objective and data-driven manner for assessing engagement, outcomes enhancement, and dropout prevention within digital educational environments.
Motivation & contribution
The reason behind carrying out this study is due to increased realization that traditional methodologies for assessing engagement are generally inadequate in tracking the multifaceted nature of student engagement in a digitally mediated learning environment. Increasing dependence on online and hybrid education increasingly aggravates limitations of self-reports, manual observations, and basic interaction metrics, failing to monitor continuous and high-fidelity engagement. Although some optimal improvements are presented within existing machine-learning-based approaches, they are somewhat restricted to single-modality data sources with more built-in multi-modal fusion strategies but without real-time engagement detection latency. Furthermore, conventional adaptive learning models tend to fail as regards integrating engagement-driven personalization for content recommendations, hence ending up posting ineffective recommendations and cognitive overload issues. In this regard, it is necessary to build a comprehensive real-time engagement evaluation framework using advanced machine-learning techniques to improve the experiences of students and optimize their learning paths.
By employing weighted visual attention, gradient-based feature attribution, and graph structuring interpretation, these approaches were used for interpretation. In MMSAF Net, the relative contributions of facial expression vectors, gaze tracking embeddings, keystroke patterns, and physiological signals to each engagement prediction bring up attention heatmaps. The learned graph embedding may then be observed in G SIN to determine which peer interactions most influentially generate each collaborative engagement score, thus producing socially central learners. The RL ALP model interpretation relies on tracking Q Value trajectories to feature the effect of adaptive recommendations on changes in one's engagement and cognitive load. This incorporates a CLE Net's model of insight into a cognitive load classification of attention weights mapped over linguistic complexity scores, response latencies, and error rates that reveal the cognitive indicators dominant in each prediction. TCL Net interpretability is further contributed to temporal saliency maps, which show which engagement periods contribute most to estimated dropout risks. Together, these mechanisms maintain the predictive efficacy of deep architectures while providing actionable transparency for educators in process.
Review of existing models for student engagement analysis
The discourse on student engagement and learning outcomes has changed quite significantly over the years, primarily due to the new pedagogical innovations, artificial intelligence, and digital learning environments. A 40-studied research paper thus seems to indicate that engagement is a multifaceted construct whose many facets are based on technological interventions, instructional design, and (most importantly) socio-emotional issues in process. Zhang [1] explains the inability of the new digital learning tools in rural teaching while showing that, compared with traditional education pedagogy, it scales up the levels of engagement but perhaps needs to be contextualized. Bergdahl et al. [2] similarly conduct a systematic Thematic review on learning analytics where student engagement appears to be best reflected in a combination of behavioral, affetive, and cognitive indicators. Thus, task-based learning is paramount in health education where engagement and academic performance can be improved by Structured Work Station Learning Activities like those advanced by Sánchez et al. [3], thereby confirming the effectiveness of active learning paradigms.
Tandiono [5] reports that the engagement benefits gained from gamification approaches are 9 % for greater quiz participation, whereas AR interventions [4] exhibit an average 11–14 % improvement in comprehension over control groups. AI-assisted speaking performance [6] reportedly raised the fluency score by 15 % rather than an engagement improvement of 10 % during the monitored sessions. This approach provides quantifiable performance context, ensuring the review remains tied to empirical outcomes rather than qualitative generalizations.
Clearly, there is a gradual shift toward self-regulated learning seen in the works of various authors. According to Bullard et al. [7], students learn better when an active learning activity becomes scaffolded into graduate seminars. In another scenario, Wu and Ho [8] explore the personalized learning engagement of ChatGPT, saying such AI-driven conversational tools promote greater cognitive interactions, backed by the findings of Peng [9], who looks into how internationalization strategies of institutions shape engagement along cultural and market-oriented facets. In predictive analytics, for example, Sanfo [10] uses explainable AI to model student learning outcomes, showing that making machine learning models transparent provides actionable insights to teachers on engagement. Likewise, Bhuttah et al. [11] show that innovative pedagogical strategies, coupled with inclusive leadership, engage learners and enhance higher-order thinking skills. Zheng et al. [12], for instance, discuss the importance of environmental factors, arguing that the library promotes activities for social learning interactions, hence engaging the student beyond the classroom. This is further expanded by Daoayan Biaddang and Caroy [13] by discussing how engagement in virtual classrooms should have student agency, particularly with voice and choice sets. Oestergaard et al. [14] begin corresponding questions of whether instructional experiences, in which mechanisms of peer feedback and assignments based on a portfolio encourage an engagement cycle over an extended period of time, are coherent. In Advancing Learning Analytics Dashboards Lee and Kim [15] affirm that SDT-informed analytics assist students in asynchronous online courses.
Similarly, Su et al. [16] placed the argument that web-based collaborative learning system construction with high interactivity and quality assurance improves the engagement of online learning. Technology context under which the classification of student engagement comes under discussion is the area studied by Das and Dev [17], in which frame-level detection accuracy for engagement improves through knowledge transfer approaches. Wu and Ouyang [18], for instance, propose probabilistic clustering methods to discovering patterns in asynchronous and synchronous learning.There are great advances in engagement predictions using multi-modal clustering. Farooq et al. [19] offer a more advanced analytical perspective in which federated learning strategies supplement predictive accuracy in engagement modeling, while Li et al. [20] deal with the theme of assessing the impact of thematic teaching on programming applications, demonstrating higher student learning outcomes. Further dissection of behavioral-emotional-cognitive engagement indicators is followed by Liu et al. [21] and concludes that individual-technology-task-environment fit turns into a significant moderator in online learning success.
The setting influence and control on the sustainability of personalized learning environments like those provided by AR technologies are likewise subjects of inquiry. çelik and Baturay [22] provide support for this claim in their review of metaverse-based language learning, which engages learners so that they may form an online-based community. The presented dichotomy of synchronous and asynchronized approaches to engagement in learning is crucial to understanding the immense variation in cognitive load, having a heavy focus on the nature of instruction. Love et al. [23] demonstrate conversational agent-led learning, showing engagement improvement about faculties and signified the growth of engaged learners. Yorganci [24] agrees on educational content and instructional approach in that pupil performance can be more successful when a combination of flipped learning and synchronous discussion allows regulation in the classroom. Stein et al. [25] argue for the practice of culturally respon-sive-sustaining education, emphasizing that a curriculum would only adequately represent all students’ identities. Khoudi et al. [26] explore machine learning-based engagement prediction using engagement classification, which posits that clickstream data improve the prediction of student performance in the virtual environments.
Ucan [27] deliberates upon Twitter-supported learning fostering higher engagement because learners are in constant discourse with one another and involve in peer interaction. Among others, Thiruthuvanathan and Krishnan [28] explore engagement detection by setting EfficientNet-based affective computing, highlighting how deep learning significantly enhances tracking of real-time emotions. Li et al. [20] undertake an empirical study featuring conversational agents in online collaborative learning, positing that adaptable scripts would promote engagement over static instructional designs. An et al. [29] take the perspective of metacognition in engagement, emphasizing learning strategy selection, and Deng et al. [30] elaborate on the necessity of question presence in instructional videos. They showed interactive elements are genuinely enhanced engagement with interaction. Tang et al. [31] employ facial expression recognition to measure the emotional engagement in Science learning, re-echoing the salient worth of affective computing in science education. The authors in the research of Matz et al. [32] use sociodemographic and engagement data for prediction of student retention, and the authors Azkarate Iturbe et al. [33] undertake a latent class analysis matching the cooperative mindset for engagement metrics. Kong et al. [34] underscore the pivotal position of teacher emotional support in online engagement, making social presence a major mediator. Among the research conducted by Wang et al. [35] is peer social interaction in science museums, showing that conversation dynamics have a significant impact on learning retention. Lastly, a qualitative synthesis by Yi et al. [36] explores student-generated models for medical education concluding that creative projects have the potential to advance conceptual retention and engagements. The review further underscores that student engagement is not a single construct, but an ever-changing interplay between technology, pedagogy, and individual agency. AI-powered engagement tracking, adaptive learning, and immersive technologies have offered a platform for potential reinvention of strategies in education, providing it with a backdrop to develop more efficient learning experiences that speak loudly to individualized learning. For this reason, future research should be directed at large-scale AI-powered learning analytics development, including some equity-focused instructional design and multi-modal, throughput-speedy engagement predictors-future directions that are aiming at facilitating a more inclusive and data-driven educational landscape for this process.
Method details
MMSAF-Net applies multi-head self-attention to align feature embeddings extracted from ResNet-based facial analysis, Transformer-based gaze estimation, keystroke latency distributions, and normalized physiological indicators. These modality-specific vectors are combined through scaled dot-product attention into a unified representation for classification. G-SIN constructs dynamic peer interaction graphs, wherein nodes represent learners and edges denote interaction events; a three-layer Graph Neural Network propagates contextual embeddings to enhance the detection of socially driven engagement fluctuations. RL-ALP treats the real-time engagement score, cognitive load index, and quiz accuracy as state variables in a Deep Q-Learning environment adjusting instructional difficulty levels through an ε-greedy policy to maximize long-term retention rewards. A-CLE-Net makes use of BERT-based linguistic embeddings, response-latency modeling, and error-rate normalization to classify cognitive load states using an attention-based fusion layer. TCL-Net utilizes a Transformer-based temporal encoder and a contrastive learning objective to disentangle transient fluctuations from persistent trends of disengagement. These components interact in a cyclic manner: MMSAF-Net and G-SIN send engagement signals to RL-ALP and TCL-Net, meanwhile A-CLE-Net informs RL-ALP of cognitive constraints so that content adaptation takes into account both engagement potential and mental workloads on the process. The framework integrates heterogeneous multimodal features using a hierarchical fusion approach.The first stage processes the raw modality-specific features-such as facial expression embeddings (256-dim ResNet vectors), gaze-tracking heatmaps (encoded via 128-dim Transformer embeddings), keystroke rhythm statistics (mean, variance, inter-key intervals), and physiological signals (HRV, EDA)-in isolation using dedicated encoders optimized for each signal type. In the second stage, these embeddings were temporally aligned and normalized to eliminate sensor-specific bias. The third stage applies multi-head scaled dot-product self-attention to learn modality interaction weights enabling the network to emphasize the most discriminative features for each engagement context. The resultant fusion produces a unified multimodal embedding vector that passes downstream to task-specific components like the engagement classifier for MMSAF-Net or the cognitive load estimator for A-CLE-Net. Cross-modal attention ensures that signals are not merely concatenated, but contextually weighted such that the system is able to exploit different modalities' complementary strengths while avoiding redundancy or noises sets.
The MMSAF Net architecture has four synchronized input types: Facial expression engenderings: ResNet 50 backbone with three convolutional blocks, global average pooling, and a 256 dimensional fully connected layer sets. Gaze tracking: Transformer encoder with 4 layers, each having 8 attention heads, projecting to a 128 dimensional embedding set in process. Keystroke dynamics: Two 1 D convolutional layers (kernel size = 5, stride = 1) followed by a gated recurrent unit (GRU) layer with 64 hidden units. Physiological signals comprise Bi LSTM with 2 layers (hidden size = 128) followed by linear projections. All modality embeddings are concatenated and passed through a multi-head scaled dot product attention layer (8 heads, 512-dimensional projection). The fused vector is projected to a 256-dimensional latent space and classified through a softmax output layer.
The G SIN starts with a 3-layer Graph Neural Network (hidden sizes: 128, 128, 64) with ReLU activations and mean aggregation. The RL ALP state encoder constitutes a 2-layer MLP (128, 64 units), feeding to a Deep Q Network with 3 fully connected layers (64, 64, outputdim). A CLE Net uses a base BERT text encoder (12 layers, 768 dimensional hidden size) then concatenates its outputs with numerical cognitive metrics that are passed through a 2-layer attention MLP Sets. For example, TCL Net includes a 12-layer Transformer encoder (hidden size = 256, 8 attention heads) with temporal position encoding and contrastive projection layers (128-dimensional output) in the process.
In order to avoid gradient explosion in multimodal fusion layers, the learning rate for MMSAF Net has been set to 1 × 10e−4 with gradient clipping at 1.0. A slightly increased learning rate (5 × 10e−4) was used in G SIN to improve the speed of convergence of its graph embeddings, and L2 regularization of 5 × 10e−5 was applied. The RL-ALP implements a discount factor of γ=0.95and an ε-greedy decay schedule from 0.3 to 0.05 across 1000 episodes. A CLE Net utilized mixed-precision training (FP16) to limit the GPU memory footprint while maintaining numerical stability sets. The training of TCL Net incorporated contrastive temperature τ=0.07 with a margin of 1.0 for separating positive negative pairs. Early stopping was uniformly applied with a patience of 7 epochs and a reduction of the learning rate was triggered when validation performance plateaued in process.
For clarity, all the symbols employed in the mathematical formulations are now given in an explicit manner with a shared notation table. Each symbol is cited by its dimensionality, modality source, and role in the computation. For instance, Ft ∈ R{df} indicates the facial expression embedding at timestamp t; Gt ∈ R{dg} the gaze-tracking embedding; Kt ∈ R{dk} the keystroke dynamics vector; and Pt ∈ R{dp} the physiological signal embedding. Terms also include graph-specific constructs such as A ∈ R{n × n} (adjacency matrix), V (set of graph nodes), and E (set of graph edges). Similarly, RL-ALP variables st (state), at (action), and rt (reward) are listed together with the temporal encodings of TCL-Net, ϕ(Et), and the loss parameters τ, P, N. This will ensure that all mathematical expressions are self-contained and interpretable without requiring external inference sets.
The input vector Xt at timestamp 't' comprises facial expression embeddings Ft, gaze-tracking features Gt, keystroke behavioral patterns Kt, and physiological sensor data Pt sets. Via Eq. (1), the multimodal feature representation is proposed,
| (1) |
Where, 'd' represents the dimensionality of features pertaining to the process. A convolution-efficient feature extractor operates spatial transformations with a deep ResNet architecture for encoding hierarchical spatial features, defined via Eq. (2),
| (2) |
Where Wf and bf are learnable weight matrix and bias, respectively, while σ(⋅) represents the activation function for this operation in process. The produced embeddings are then passed through self-attention mechanisms as defined by a scaled dot-product attention function via Eq. (3),
| (3) |
Where Q, K, and V are query, key, and value matrices derived from input representation, and dk is scaling factor for the process. This attention mechanism aligns multimodal signals while preserving temporal dependencies. Engagement classification is done with a fully connected classifier featuring the softmax function via Eq. (4), (Fig. 1).
| (4) |
Fig. 1.
Model architecture of the proposed analysis process.
Where Wc and bc represent weight and bias assigned to the classifier for this method. This formalization ensures that the MMSAF-Net outputs a probabilistic engagement score falling within the [0,1] boundaries, enabling more accurate classification of attention levels. The G-SIN module thus iteratively constructs, as Fig. 2 depicts, a dynamic interaction graph G (V, E) in which the nodes of V represent students; the edges of E represent the collaboration events, such as peer discussions, Q&A interactions, and assignment reviews.
Fig. 2.
Overall flow of the proposed analysis process.
The adjacency matrix A is defined via Eq. (5),
| (5) |
Where, xi and xj represent student feature vectors in process. The graph embedding update is performed using a message-passing mechanism in a Graph Neural Network Via Eq. (6),
| (6) |
Here, h'i(l) denotes the embedding of node i at layer l, N(i) represents the set of neighboring nodes, cij is a normalization constant (e.g., degree normalization), and W(l) is the learnable transformation matrix for layer l. The final social engagement score is computed via a contrastive loss function which is estimated Via Eq. (7),
| (7) |
Where P represents positive student interactions, N refers to negative pairs, and m is a margin parameter that guarantees the proper clustering between engaged and disengaged groups. The RL-ALP model uses reinforcement learning to improve personalized learning trajectories. Via Eq. (8), the state 'st' is defined by a tuple of engagement score 'et', Iterative Historical learning performance pt, and real-time quiz response accuracy qt,
| (8) |
The reward function is formulated as a weighted sum of engagement-driven rewards via Eq. (9),
| (9) |
Where, λ1, λ2, λ3 are weight coefficients controlling the influence of different factors. The RL agent employs Deep Q-Learning, where the Q-function is updated iteratively Via Eq. (10),
| (10) |
Where, α is the learning rate and γ is the discount factor for this process. The policy function is optimized using an epsilon greedy approach Via Eq. (11),
| (11) |
The final adaptive learning adjustment is determined via Eq. (12),
| (12) |
Where c(t + 1) defines the anticipated difficulty adjustment in content and η shapes the adaptive rate for this process. This formulation guarantees that the trajectories are increasingly optimized with changes in engagement, leading to improved retention while limiting dropout rates. Next, iteratively, as shown in Fig. 2, the Attention-Based Cognitive Load Estimation Network (A-CLE-Net) and the Temporal Contrastive Learning Network (TCL-Net) are to measure cognitive load variation and engagement fluctuation over temporal instance sets, respectively for this process. These two models would be working together, thus ensuring that the dynamically adapting learning paths are based on real-time states of cognition and long-term trends of engagement. The A-CLE-Net is aimed at disentangling productive cognitive load from overload Induced disengagement while the TCL-Net utilizes temporal encoding and contrastive learning mechanisms to trace engagement evolution and predict dropout risk. The integration of the models would involve the detection of learning roadblocks early on and proposing potential for intervention, thus enhancing individualized learning trajectories for students. The A-CLE-Net integrates and processes diverse multimodal cognitive signals such as response time, linguistic complexity, and error rate. From input vector Xt at time index "t" that includes text embeddings Tt, response time Rt, and correctness Cts, the unit representation is estimated via Eq. (13),
| (13) |
Where, ‘d’ represents the feature dimensionality for this process. The text-based complexity Tt is extracted using a BERT-based embedding layer, formulated via Eq. (14),
| (14) |
Where, wi represents individual words in the input sequence sets. The response latency function is modeled as an exponential decay function, capturing cognitive processing delays via Eq. (15),
| (15) |
Where τ is baseline response times, and λ is the hypothesized degradation constant regulating reasonable cognitive efficiency levels. Higher decay rates mean better adaptations to learning process. The correctness score Ct follows a sigmoid transformation, assuring bounds via Eq. (16).
| (16) |
Where, β controls sensitivity to quiz difficulty qt, and μ is the adaptive threshold for the process. These features are passed through an attention-based feature fusion mechanism, defined via Eq. (17),
| (17) |
Where, Wi represents learnable attention weights. The final cognitive load classification is derived via a fully connected layer with a softmax activation via Eq. (18),
| (18) |
This as a result ensures the prediction of different mental load states, i.e., low, medium, and high cognitive loads. Next Benefits by the TCL-Net by considering the temporal contrastive learning framework to track the trend of engagement over temporal instance sets. Given a series of engagement states Et, the time-series encoding function could be described via Eq. (19),
| (19) |
Where, ϕ(Et) represents instantaneous engagement fluctuations. The network extracts short-term and long-term engagement representations using a Transformer-based temporal encoder, given via Eq. (20),
| (20) |
Where, Wh and bh are the encoding parameters for the process. The contrastive learning objective ensures temporal coherence in engagement representations, formulated via Eq. (21),
| (21) |
In Eq. (21), the terms P and N refer to the sets of positive and negative time engagement pairs respectively. τ represents a temperature scaling parameter in the contrastive loss formulation. In notation, ϕ(Et) is the time encoding of engagement state vector Et, which itself is made up of the aggregated multimodal engagement representation at timestamp ‘t’ in the process. This vector captures the fused behavioral, physiological, and interaction features, all aligned to a common time point. Positive pairs (p ∈ P) will be sampled from temporally adjacent high correlation engagement states, while negative pairs (n ∈ N) will be drawn from temporally distant or semantically dissimilar engagement states. Giving these precise definitions will elucidate the logical flow and will possibly eliminate ambiguity in the loss formulations
Where, P and N represent positive and negative engagement pairs, and τ is the temperature scaling parameter for this process. A dropout risk classifier predicts disengagement likelihood using a logistic regression function via Eq. (22),
| (22) |
Where, γ controls the impact of historical engagement on dropout risks. The final integrated learning trajectory adjustment equation combines cognitive load estimation and engagement prediction via Eq. (23),
| (23) |
This ensures that the adaptive learning recommendations are dynamically adjusted according to cognitive strain and engagement trends. This formulation thus provides an overarching framework for optimizing personalized learning pathways while counterbalancing cognitive overload and disengagement risks. We now turn to the effectiveness of the proposed model using different metrics, and benchmark it against existing models under different conditions.
Comparative result analysis
Systematic Missing Value, Outlier, Normalization, and Feature Scaling Handling within the Preprocessing Pipelines. Most numerical data missing, such as interruptions in physiological readings or keystroke logs, were handled as modality specific imputations, i.e., forward fill in continuous time-series (e.g., HRV), median substitution sporadically for categorical gaps, and contextual embedding interpolation for text sequences. Outliers were identified by a modified Z score threshold value of |Z| > 3.5 and replaced with local neighborhood averaging in the temporal domain. Numeric features went through z score normalization into zero mean and unit variance to maintain comparability across modalities while features with skewed distributions per their inherent characteristic were log transformed before scaling. For textual modalities, tokenized sequences are embedded using pre-trained BERT vectors, with sequence length standardized by either truncation or zero-padding. This harmonization ensures that multimodal fusion stages operate on aligned and statistically balanced feature spaces.
The experimental setups of the proposed Integrated Multi-Modal Machine Learning Framework guarantee a comprehensive assessment of student engagement, cognitive load, and learning outcomes in varying educational settings. The experiments were carried out on three real-world datasets—OpenEDU, DAiSEE, and MOOC engagement logs—each providing rich multimodal behavioral and academic performance data. The OpenEDU dataset accounts for a variety of logs of student interactions over 7 million, including assignment submissions, discussion forum activity, and quiz performance, tracking trends on engagement with activities across traditional and online environments. The DAiSEE dataset was used to benchmark student affective engagement analysis through facial expressions, gaze, and behavioral cues in real-time based video learning. Furthermore, proprietary curation from the MOOC platform included physiological data captured via wearable sensors such as heart rate variability (HRV) and electrodermal activity (EDA), aimed at an advanced cognitive load evaluation under different instructional conditions. The preprocessing pipeline is designed to standardize and optimize multimodal inputs across datasets inclusive of OpenEDU, DAiSEE, EdNet, SEED, and OULAD. For visual modalities, video frames were sampled at 10 fps and passed through OpenFace to extract 68 point facial landmarks and 256 dimensional expression embeddings. The gaze tracking sequences were generated from eye region landmarks, smoothed with a temporal median filter and encoded with a Transformer-based gaze estimator. Keystroke dynamics were derived from raw interaction logs, computing inter-key intervals, typing speed distributions, and burst length statistics; all features were normalized to zero mean and unit variance. Physiological signals including heart rate variability (HRV) and electrodermal activity (EDA) measuring were downsampled to a uniform frequency, detrended and z-score normalized to counteract device-specific bias. Text-based interaction logs like discussion forum posts were preprocessed through tokenization, stop word removal, and TF-IDF weighting; where necessary, BERT-based embeddings were generated to capture contextual semantics. All modalities were synchronized to a common temporal index using event-based synchronization to ensure the cross-modal fusion stages worked on temporally coherent representations. An 80 %−20 % train-test split was performed on the dataset, while maintaining stratified sampling to balance engagement level, with a separate 10 % of the validation set being extracted from the training partition to tune the hyperparameters for the process.
While OpenEDU, DAiSEE, and EdNet forge a solid multimodal interaction and performance record, educational environments need more diversity across demographics to improve generalizability. Future work should integrate such disparate flows into the framework where exposure does not yet exist. For instance, one could include classroom-based multimodal recordings from rural schools, online learning logs from low-bandwidth regions, and specialized datasets capturing neurodiverse learner populations. Such sources would increase both the variation of the engagement and cognitive profiles as well as the demonstration of robustness under different environmental constraints, thus reducing dataset bias and thereby ensuring high reliability of predictive models across heterogeneous learner cohorts (Fig. 3).
Fig. 3.
Model’s integrated result analysis.
The experimental setup is focused on evaluating the proposed framework using the three available and well-established datasets: OpenEDU, DAiSEE, and EdNet, all of which provide rich multimodal engagement and learning behavior data. OpenEDU largely refers to the datasets provided by MITx and HarvardX MOOC courses and comprises >7 million records of interactions of students, including clickstream, forum, problem-solving, and quiz performances, being now attractive datasets to model trends of engagement and dropout risk. The DAiSEE datasets (Dataset for Affective States in E-learning Environments) include over 10,000 labeled frames of video archived for engagement, boredom, confusion, and frustration, thus enabling real-time analysis of student behavior using facial and gaze indicators. This dataset has been very instrumental in training the MMSAF-Net for multi-modal engagement classification. Alongside, EdNet was gathered from a smart learning system that includes a massive scale knowledge tracing data of >130 million student interactions with the aim to test adaptive learning strategies. It is accompanied with rich question-response sequences, time stamps, difficulties, and cognitive skill mapping such that reinforcement-learning models like RL-ALP can optimize learning paths for personalised learning process. The union of these datasets guarantees that our experimental environment is robust enough to enable a comprehensive assessment of student engagement, cognitive load estimation, and adaptive learning interventions in video-based and text-based learning environments.
All models in the framework were run under the Adam optimizer with the parameters β1 = 0.9, β2 = 0.999 using weight decay 1 × 10e-4 as the regularization sets. The initial learning rate was set to 1 × 10e-4 with a cosine annealing schedule for gradually decreasing it to 1 × 10e-6 through the course of training. A batch-size of 32 was set for vision-based models (MMSAF Net, A CLE Net) and 64 for sequence-based models (TCL Net, RL ALP state encoding) in terms of memory efficiency with regards to the convergence stability. The maximum duration of training was 50 epochs, with indications for early stopping if no progress was achieved in the validation loss over 7 most recent consecutive epochs. The fully connected layers included dropout regularization with a rate of 0.3, whereas batch normalization was used for those in the convolutional and attention projection types to stabilize gradients. All the experiments are run on NVIDIA A100 GPUs, with 80 GB VRAM, and in a multi-GPU environment; the CPU side preprocessing is executed on dual Intel Xeon Gold processors. The experiments were performed on an NVIDIA A100 hardware cluster, with 80GB VRAM, with the deep-learning models implemented in PyTorch, graph construction carried out in NetworkX, and reinforcement learning simulations in OpenAI Gym. MMSAF-Net was subsequently trained using Adam optimizer, learning rate=0.0001, batch size=32, multi-head self-attention with eight heads with heterogeneous engagement signals from facial, keystroke, and physiological inputs aligned by the self-attention fusion mechanism. The G-SIN model, with three-layer Graph Neural Network (GNN) with ReLU activation trained on contrastive loss (1.2 margin), improved peer engagement representation. The RL-ALP agent was trained with Deep Q-Learning and an epsilon-greedy exploration policy (ϵ=0.1), discount rate (γ=0.95), and reward function that weighted the engagement score sets (λ1=0.4), past performance (λ2=0.35), and real-time quiz accuracy (λ3=0.25). The A-CLE-Net employed BERT embeddings with 768-dimensional feature vectors, attention-weighted fusion mechanism, and a softmax cognitive load classifier with high predictive reliability. The TCL-Net, which analyzed trends of engagement, applied a Transformer-based encoder consisting of 12 attention layers with contrastive loss (temperature=0.07) and dropout rate (0.3) for regularization, ensuring that early signs of disengagement were detected. Each model was evaluated based on accuracy, precision, recall, F1-score, and Pearson's correlation (r) with actual learning outcomes: engagement prediction accuracy (92–94 %), dropout risk forecasting precision (91–93 %), and a 12–15 % improvement in learning due to adaptive content adjustments.
To reduce overfitting, dropout with a 0.3 rate for dense layers and a 0.1 rate for attention heads was uniformly used across all deep modules as regularization. Batch normalization layers were used to stabilize gradient propagation in convolutional blocks and projection blocks. Gradient accumulation for two steps was applied in MMSAF Net and TCL Net to achieve larger effective batch sizes without exceeding GPU memory limits. The models were trained under a distributed data-parallel setting, ensuring efficient scalability across several GPUs. During inference time optimization, ONNX graph export and TensorRT acceleration were used, allowing for deployment at ∼21 ms inference time per sample on the A100 GPU, ∼35 ms for the RTX 3080, and ∼120 ms for CPU-only execution on low-resource scenarios.
Considering its application on OpenEDU, DAiSEE, and EdNet datasets, an outline along with various model comparisons against three other baselines referred to as Method [5], Method [8], and Method [16] were included for completeness. Among the evaluation metrics, engaging prediction accuracy, dropout risk prediction, cognitive load estimation reliability, adaptive learning effectiveness, and computational efficiency were considered. The subsequent subsections are devoted to a comprehensive comparison of results, demonstrating the superior performance of the proposed model in multiple dimensions across the process. Table 1 shows engagement classification accuracy in the DAiSEE dataset and samples. The MMSAF-Net proposed achieves the highest accuracy through self-attention fusion of multi-modal engagement signals by a significant degree over other models. Table 1 cataloged the results for engagement prediction accuracy with respect to the DAiSEE dataset, showing that the proposed Multi-Modal Self-Attention Fusion Network (MMSAF-Net) greatly outperforms existing methods in classifying student engagement states. MMSAF-Net, with a classification accuracy of 94.2 %, surpassing Method [5] (87.1 %), Method [8] (88.5 %), and Method [16] (90.2 %) is evident that self-attention fusion was successful in aligning heterogeneous behavioral, gaze, and physiological indicators in process-outlining superior efficacy of the proposed model compared to others in almost all dimensions for the process. Improvement of 4–8 % with respect to precision, recall, and F1-score indicates a decrease in false engagement classifications made by the proposed model, enabling a far more reliable real-time monitoring process.
Table 1.
Specification Table.
| Subject area | Engineering |
|---|---|
| More specific subject area | Graph-Based Student Interaction Network (G-SIN), Reinforcement learning-based adaptive learning |
| Name of your method | Multi-Modal Self-Attention Fusion Network (MMSAF-Net) |
| Name and reference of original method | None |
| Resource availability | None |
The evaluation metrics, including accuracy, precision, recall, and F1-score, were calculated on a held-out validation set that was distinct from the training data and the final test partition. The DAiSEE dataset experiments, for example, applied stratified sampling in an 80–20 train-test split, reserving 10 % of the training data as a validation set for hyperparameter tuning and performance tracking. Metrics such as that for the engagement prediction accuracy of MMSAF-Net of 94.2 %, precision 93.5 %, recall 92.8 %, and F1-score 93.1 % represent generalization performance on this validation set, rather than on the training data. The separation thus allows any reported results to be a true representation of out-of-sample predictive capabilities and free from possible bias owed to model overfitting. Hyperparameter optimization of model was done using Bayesian search strategy with five fold Stratified Cross Validation on training data. In the search done for MMSAF Net, attention head counts were searched between 4 and 12, the embedding dimension ranges were set between 128 and 512, and dropout rates were varied from 0.1 to 0.5. For tuning G SIN, the GNN layer depths varied between 2 and 4 while hidden dimensions ranged from 64 to 256; the contrastive loss margins were 0.8–1.5. RL ALP: learning rate (1e 5 to 1e 3), discount factor γ (0.85–0.99), and ε-greedy decay rates were varied for balancing exploration and exploitation. In a CLE Net tuning, BERT embedding size (base versus large), attention weight regularization (0.0–0.3), and softmax temperature parameters were varied. TCL Net tuning included optimization of Transformer depth (6–12 layers), attention heads (4–12), and contrastive temperature τ (0.05–0.1). The configurations finally selected were the ones yielding the highest mean F1 score across folds while retaining stability and computational efficiency sets.
The results validate that MMSAF-Net improves upon existing methods by at least 4–7 % in accuracy, attesting to the effectiveness of self-attention mechanisms in aligning multimodal engagement indicators in process. Table 2 evaluates the dropout prediction accuracy using the TCL-Net model on the OpenEDU dataset, reflecting the improvements gained by contrastive learning-based temporal engagement analysis. The dropout risk forecasting results in Table 2 further substantiate the superiority of the Temporal Contrastive Learning Network (TCL-Net) in tracking longitudinal engagement patterns from the OpenEDU dataset samples. The proposed model achieves 92.3 % accuracy, which is significantly greater than the accuracies of Method [5](84.6 %), Method [8](86.8 %), and Method [16] (88.1 %) for the present tasks. The contrastive learning mechanism enables TCL-Net to disentangle persistent disengagement trends from short-term fluctuations, thus greatly enhancing its recall for dropout-prone students and, in turn, increasing the efficacy of early interventions in online learning environments (Fig. 4).
Table 2.
Engagement prediction accuracy on DAiSEE dataset.
Fig. 4.
Model’s overall result analysis.
The proposed model TCL-Net shows 7.7 %, 5.5 %, and 4.2 % better performance over Method [5], Method [8], and Method [16], respectively, strengthening the case for contrastive learning on the detection of long-term engagement trends. Table 3 provides the cognitive load classification accuracy over the EdNet dataset thereby ascertaining the efficacy of A-CLE-Net in discerning productive versus overload-induced cognitive strains. As for the accuracy of cognitive load estimation shown in Table 3, it measures the accuracy of the Attention-Based Cognitive Load Estimation Network (A-CLE-Net) to classify the mental workload levels of students in the EdNet dataset. The model proposed has a much higher accuracy at 91.5 % compared to Method [5] (81.7 %), Method [8] (84.2 %), and Method [16] (86.9 %). The improvements show the efficiency of A-CLE-Net in estimating cognitive states, and further establish the foundation for adaptive learning difficulty controls. A-CLE-Net, with a greater Pearson's correlation (r = 0.82), shows more alignments between predicted and true cognitive states across an accuracy gain of 9.8 % over Method [5] in process.
Table 3.
Dropout prediction accuracy on OpenEDU dataset.
Table 4 presents the rates of improvement in learning outcomes across the EdNet dataset verifying the effect of RL-ALP in dynamic content complexity set adaptation. Table 4 illustrates the learning outcome improvement introduced by the Reinforcement Learning-Based Adaptive Learning Pathway (RL-ALP) and its capacity to maximize personalized content sequencing as a function of engagement and cognitive load variations. The proposed framework scores an improvement in learning outcome by 14.3 %, which is substantially higher than Method [5] (8.2 %), Method [8] (9.7 %), and Method [16] (11.5 %). Moreover, retention rates (92.1 %) and dropout reduction (20.3 %) provide further proof of the effectiveness of Deep Q-Learning in dynamically adapting learning difficulty levels. These results corroborate the effectiveness of RL-ALP in ameliorating cognitive fatigue and promoting student's persistence through modifying instructional complexity based on real-time engagement signals.
Table 4.
Cognitive load estimation accuracy on EdNet dataset.
According to the proposed RL-ALP model, learning output is improved by 14.3 % in retention rates, whereas it is maintained higher than another method by 6–7 % and dropout rates reduced by a value of 20.3 %. All models in Table 5 have been presented with reference to their computational efficiency when comparing inference timestamp per sample for and training convergence speed sets. The results of computational efficiency are so marked by numbers in Table 5 that it is evident that an integrated framework would make inference timestamp and training convergence speed significantly lower, making it possible for real-time deployment. With processing time of inference by 21.3 milliseconds on average, this framework is approximately 50 % faster than Method [5] (45.2 ms) and won over method [8] (38.7 ms) and method [16] (31.4 ms). Also, the training timestamp has been reduced to 35 epochs from 52 epochs (Method [5]), 47 epochs (Method [8]), and 42 epochs (Method [16]), thus confirming the effect of the self-attention fusion and reinforcement learning mechanisms on the speed-up of model convergences.
Table 5.
Learning outcome improvement via RL-ALP on EdNet dataset.
Compared to Method [5], the proposed framework reduces the inference timestamp almost by 50 % and requires only 35 training epochs, reflecting faster convergence than current practices. A summary of all performance metrics over all datasets is presented in Table 6, which supports the overall improvement in any proposed framework process. Finally, the summary performance metrics across all datasets in Table 6 show that the proposed framework always certainly outperforms the existing methods in various evaluation criteria, including engagement prediction accuracy, dropout forecasting, cognitive load estimation, and learning outcome improvements. With an average engagement classification accuracy of 94.2 %, dropout risk forecasting at 92.3 %, cognitive load accuracy of 91.5 %, and an adaptive learning gain of 14.3 %, the results prove the superiority of the multi-modal, contrastive learning-enhanced, and reinforcement learning-driven approach when compared to traditional methods. These findings reinforce the value of integrating deep learning models with multi-sensor engagement tracking, enabling intelligent, adaptive learning frameworks that effectively enhance student retention and performance.
Table 6.
Computational efficiency of proposed model.
The devised framework computationally increases efficiency for both training and inference to allow deployment in low-resource educational environments. For example, a 38-epoch full convergence occurs within approximately 25 ms per sample on a mid-range GPU with no overhead from an NVIDIA RTX 3060; it maintains inference latencies below this threshold under standard multimodal configurations. Optimized quantized models reduce memory footprints by 42 % and decrease inference time down to 120 ms per sample for CPU only execution, enabling real-time feedback solely on institutional servers or high-end laptops. Modular design allows selective disabling of high cost modalities — for example, omitting Transformer based gaze processing reduces compute load by 28 % with only a 2 % drop in engagement classification accuracy. Edge deployment tests on devices such as NVIDIA Jetson Xavier NX confirm that the architecture can operate in streaming mode at ∼15 fps, enabling in class adaptive learning without reliance on cloud infrastructure sets. These efficiency characteristics position the framework for adoption not only in advanced smart classrooms but also in bandwidth constrained or budget limited institutions in process.
Ultimately, the results confirm that the proposed multi-modal, attention-based, and enhanced reinforcement learning framework achieves superior engagement prediction, cognitive load assessment, dropout risk prediction, and adaptive learning optimization. The results show the advantages of integrating self-attention fusion (MMSAF-Net), graph-based interaction modeling (G-SIN), reinforcement learning-based adaptation (RL-ALP), cognitive load estimation (A-CLE-Net), and engagement tracking based on contrastive learning (TCL-Net). In several datasets, the proposed models outperform existing methods by 4–12 % accuracy, reduce dropout rates by 20 %, and better optimize adaptive learning pathways. The proposed framework has significantly increased computational efficiency with fused inference times that are nearly double the normal size, enabling its use for monitoring must real-time students' engagement and adaptation of learning in intelligent tutoring systems. These findings confirm the advancement that multi-modal deep learning and reinforcement learning bring on student engagement analytics towards scalable and adaptable AI-driven education frameworks. Moving on, we discuss an Iterative Validation use Case for the Proposed Model that will contribute more to understanding the full process by readers of this text.
Method validation
To demonstrate the efficiency of the Integrated Multi-Modal Machine Learning Framework, consider a detailed example use case where a group of students is engaged in an online interactive learning session that involves video-based lectures, collaborative assignments, quizzes, and real-time discussion forums. The engagement, cognitive load, and adaptive learning outcomes of each student are assessed through multimodal behavioral, physiological, and interaction based indicators. The following sections illustrate sample input feature values, interim processing results, and final learning trajectory adjustments together, giving a system-wide view of the system's impact on learning optimizations. MMSAF-Net captures real-time engagement signals and analyzes facial expressions while also examining gaze tracking, using keystroke behavior as well as physiological responses. G-SIN acts in evaluating peer interactions in collaborative settings, thereby having key participation metrics in Cooperative scenario. The RL-ALP framework personalizes learning pathways based on engagement fluctuations and cognitive load estimates, while the A-CLE-Net determines cognitive workload levels. The TCL-Net processes historical engagement trends to forecast dropout risks, ensuring proactive interventions for disengaged learners. This following tables present the structured outputs of each model, that is computed engagement scores, peer network analysis, reinforcement learning-driven learning adjustments, cognitive load estimations, and longitudinal dropout predictions.
To test the generalizability of the framework, performance was out the higher-level evaluation upon multicultural datasets beyond the three primary ones. These include the SEED dataset for EEG-tied affective state recognition, the CogLoad corpus for subtle cognitive load estimation, as well as the OULAD repository, encompassing longitudinal engagement logs of >32,000 students. The MORF platform also provided diverse MOOC-scale interaction records from multiple institutions, capturing different pedagogical designs and learner demographics. The framework consistently maintained very high accuracy across these datasets — 92–94 % for engagement prediction, 90–92 % for cognitive load estimation, and 91–93 % for dropout risk forecasting. Additional EEG-derived features in SEED produced a 3–4 % increase in accuracy in engagement classification, while a long-term model in OULAD enhanced the recall of dropout detection by 5 %. Thus, the findings affirm that the framework's multimodal integration and adaptive learning mechanism scaling robustly to different educational settings and diverse learner populations. The proposed method was validated against three of the most advanced learning engagement models Method [5], Method [8], and Method [16], thus forming an all-round performance evaluation. This validation setup ascertains the generalization, scalability, and accuracy of the proposed framework and shows improvement of engagement tracking (by +7 %), dropout risk prediction (by +8 %), cognitive load estimation (by +9.8 %), and adaptive learning optimization (by +14.3 %) over existing systems. Table 7 presents the computed engagement scores from MMSAF-Net-based real-time multimodal feature extraction process.
Table 7.
Overall performance summary.
The computed engagement scores show that students S101, S103, and S105 are very much engaged, whereas S102 and S104 score moderately low and need further sets of intervention in process. Table 8 provides peer interaction analysis outcome based on graph-based modeling of discussion forums, collaborative projects, and Q&A participations.
Table 8.
MMSAF-net engagement scores.
| Student ID | Facial Expression Score (0–1) | Gaze Tracking Score (0–1) | Keystroke Activity Score (0–1) | Physiological Score (0–1) | Engagement Score (0–1) |
|---|---|---|---|---|---|
| S101 | 0.76 | 0.81 | 0.72 | 0.79 | 0.77 |
| S102 | 0.64 | 0.55 | 0.48 | 0.61 | 0.57 |
| S103 | 0.88 | 0.92 | 0.85 | 0.91 | 0.89 |
| S104 | 0.43 | 0.52 | 0.38 | 0.46 | 0.45 |
| S105 | 0.70 | 0.74 | 0.66 | 0.72 | 0.71 |
Strong peer engagement is evidenced in students S101 and S103, while limited participation is shown by S102 and S104, which may compromise collaboration effectiveness in learning. Table 9 presents the computed RL-ALP learning pathway adjustments based on engagement scores, cognitive load levels, and quiz performance sets.
Table 9.
G-SIN peer engagement scores.
| Student ID | Forum Posts | Replies Given | Collaborative Projects | Peer Learning Score (0–1) |
|---|---|---|---|---|
| S101 | 12 | 8 | 3 | 0.82 |
| S102 | 5 | 3 | 1 | 0.54 |
| S103 | 14 | 11 | 4 | 0.89 |
| S104 | 3 | 2 | 1 | 0.41 |
| S105 | 10 | 6 | 2 | 0.75 |
Complete reductions to content difficulty need to be applied with respect to S102 and S104 since they are both experiencing a higher cognitive strain and lower quiz performance, while S103 benefits from a higher challenge level for the process. Table 10 provides computed cognitive load levels based on response latency, linguistic complexity, and error rates.
Table 10.
RL-ALP adaptive learning adjustments.
| Student ID | Engagement Score (0–1) | Cognitive Load Level | Quiz Performance ( %) | Recommended Content Difficulty Adjustment |
|---|---|---|---|---|
| S101 | 0.77 | Balanced | 85 | No Change |
| S102 | 0.57 | High Cognitive Load | 68 | Reduce by 1 Level |
| S103 | 0.89 | Low Cognitive Load | 92 | Increase by 1 Level |
| S104 | 0.45 | Overloaded | 51 | Reduce by 2 Levels |
| S105 | 0.71 | Balanced | 80 | No Change |
Cognitive overload effects suggest adjustment of instructional delivery sets for students S102 and S104. Table 11 gives dropout risk analysis computed through the engagement trend modeling process of TCL Net sets.
Table 11.
A-CLE-net cognitive load analysis.
| Student ID | Response Latency (s) | Text Complexity Score (0–1) | Error Rate ( %) | Cognitive Load Level |
|---|---|---|---|---|
| S101 | 1.8 | 0.71 | 8 | Balanced |
| S102 | 3.5 | 0.89 | 15 | High Cognitive Load |
| S103 | 1.2 | 0.63 | 5 | Low Cognitive Load |
| S104 | 4.1 | 0.94 | 18 | Overloaded |
| S105 | 2.0 | 0.76 | 9 | Balanced |
Immediate intervention for students S102 and S104 is needed as dropout risks exceed 20 % and 38 %, respectively, for this process. Table 12 brings together all key computed outputs, presenting a holistic view of student engagement, cognitive load, and learning trajectory modifications in the process.
Table 13.
Final framework outputs.
| Student ID | Engagement Score | Peer Learning Score | Cognitive Load | Dropout Risk ( %) | Learning Path Adjustment |
|---|---|---|---|---|---|
| S101 | 0.77 | 0.82 | Balanced | 5.3 % | No Change |
| S102 | 0.57 | 0.54 | High | 23.4 % | Reduce Difficulty |
| S103 | 0.89 | 0.89 | Low | 3.1 % | Increase Difficulty |
| S104 | 0.45 | 0.41 | Overloaded | 38.7 % | Reduce Difficulty |
| S105 | 0.71 | 0.75 | Balanced | 8.9 % | No Change |
Table 12.
TCL-net dropout risk predictions.
| Student ID | Historical Engagement Score (0–1) | Engagement Trend Stability | Predicted Dropout Risk ( %) |
|---|---|---|---|
| S101 | 0.80 | Stable | 5.3 % |
| S102 | 0.58 | Fluctuating | 23.4 % |
| S103 | 0.87 | Stable | 3.1 % |
| S104 | 0.45 | Declining | 38.7 % |
| S105 | 0.72 | Stable | 8.9 % |
Results reported illustrate multimodal AI-driven learning analytics for real-time student engagement monitoricng and adaptive educational methods.
Limitations
The Integrated Multi-Modal Machine Learning Framework, even after being innovative, has certain limitations. The trust on multi-modal data, which includes behavioral, physiological, and interaction features, introduces various issues which are related to data privacy, ethical concerns, and compliance with regulations. The architecture’s work is entirely dependent on the quality and diversity of available datasets, and any issues or error in the data can affect accuracy. Moreover, the models may not work well across various educational environments or cultural contexts, limiting their reach. The computational power of processing real-time data through multiple complex architectures like MMSAF-Net and TCL-Net may also create problems for deployment in low-resource settings. In the end, the deep learning modules lack interpretability, which can make it problematic for educators to understand or trust the system’s recommendations .
Ethics statements
This study was conducted following relevant ethical guidelines and regulations.
CRediT authorship contribution statement
Deepali Kayande: Conceptualization, Methodology, Writing – original draft, Writing – review & editing. Swetta Kukreja: Investigation, Visualization, Supervision.
Declaration of competing interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Acknowledgments
The authors gratefully acknowledge the support and contributions that facilitated the completion of this work.
This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.
Data availability
No data was used for the research described in the article.
References
- 1.Zhang G. A slightly digital learning environment in the hands of a rural teacher: modern-day teaching and its impact on student engagement and subject learning outcomes. Educ. Inf. Technol. 2024 doi: 10.1007/s10639-024-13033-y. [DOI] [Google Scholar]
- 2.Bergdahl N., Bond M., Sjöberg J., et al. Unpacking student engagement in higher education learning analytics: a systematic review. Int. J. Educ. Technol. High. Educ. 2024;21:63. doi: 10.1186/s41239-024-00493-y. [DOI] [Google Scholar]
- 3.Sánchez J., Lesmes M., Rubio M., et al. Enhancing academic performance and student engagement in health education: insights from work station learning activities (WSLA) BMC Med. Educ. 2024;24:496. doi: 10.1186/s12909-024-05478-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Ji S., Mokmin N.A.M., Wang J. Evaluating the impact of augmented reality on visual communication design education: enhancing student motivation, achievement, interest, and engagement. Educ. Inf. Technol. 2024 doi: 10.1007/s10639-024-13050-x. [DOI] [Google Scholar]
- 5.Tandiono R. Gamifying online learning: an evaluation of Kahoot’s effectiveness in promoting student engagement. Educ. Inf. Technol. 2024;29:24005–24022. doi: 10.1007/s10639-024-12800-1. [DOI] [Google Scholar]
- 6.Huang M. Student engagement and speaking performance in AI-assisted learning environments: a mixed-methods study from Chinese middle schools. Educ. Inf. Technol. 2024 doi: 10.1007/s10639-024-12989-1. [DOI] [Google Scholar]
- 7.Bullard E.A., Dubell C.R., Patrick C.W., et al. Enhancing student engagement in the graduate seminar by scaffolding active learning activities. Biomed. Eng. Educ. 2024;4:275–282. doi: 10.1007/s43683-024-00144-8. [DOI] [Google Scholar]
- 8.Wu C.H., Ho V.T. Critical factors for why ChatGPT enhances learning engagement and outcomes. Educ. Inf. Technol. 2025 doi: 10.1007/s10639-025-13346-6. [DOI] [Google Scholar]
- 9.Peng M.Y.P. The impact of perceived absolute differences between student international mindset and university internationalization on student learning engagement: the mediating role of market orientation. Curr. Psychol. 2024;43:16657–16673. doi: 10.1007/s12144-023-05593-y. [DOI] [Google Scholar]
- 10.SANFO J.B.M. Application of explainable artificial intelligence approach to predict student learning outcomes. J. Comput. Soc. Sc. 2025;8:9. doi: 10.1007/s42001-024-00344-w. [DOI] [Google Scholar]
- 11.Bhuttah T.M., Xusheng Q., Abid M.N., et al. Enhancing student critical thinking and learning outcomes through innovative pedagogical approaches in higher education: the mediating role of inclusive leadership. Sci. Rep. 2024;14 doi: 10.1038/s41598-024-75379-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Zheng Z., Zeng M., Huang W., et al. The influence of university library environment on student interactions and college students’ learning engagement. Humanit. Soc. Sci. Commun. 2024;11:385. doi: 10.1057/s41599-024-02892-y. [DOI] [Google Scholar]
- 13.Biaddang Daoayan, Caroy L.M. A.A. Student voice and choice in the virtual classroom: engagement strategies. Discov. Educ. 2024;3:162. doi: 10.1007/s44217-024-00269-6. [DOI] [Google Scholar]
- 14.Oestergaard L.G., Hansen J.S., Ravn M.B., et al. Enhancing coherence and student engagement through portfolio assignments and peer-feedback. Discov. Educ. 2024;3:161. doi: 10.1007/s44217-024-00260-1. [DOI] [Google Scholar]
- 15.Lee J., Kim D. From awareness to empowerment: self-determination theory informed learning analytics dashboards to enhance student engagement in asynchronous online courses. J. Comput. High. Educ. 2024 doi: 10.1007/s12528-024-09416-2. [DOI] [Google Scholar]
- 16.Su N., Al Mamun A., Reza M.N.H., et al. Unveiling the nexus between quality and student engagement in web-based collaborative learning systems. Educ. Inf. Technol. 2024;29:23717–23752. doi: 10.1007/s10639-024-12794-w. [DOI] [Google Scholar]
- 17.Das R., Dev S. Enhancing frame-level student engagement classification through knowledge transfer techniques. Appl. Intell. 2024;54:2261–2276. doi: 10.1007/s10489-023-05256-2. [DOI] [Google Scholar]
- 18.Wu M., Ouyang F. Using an integrated probabilistic clustering approach to detect student engagement across asynchronous and synchronous online discussions. J. Comput. High. Educ. 2025;37:299–326. doi: 10.1007/s12528-023-09394-x. [DOI] [Google Scholar]
- 19.Farooq U., Naseem S., Mahmood T., et al. Transforming educational insights: strategic integration of federated learning for enhanced prediction of student learning outcomes. J. Supercomput. 2024;80:16334–16367. doi: 10.1007/s11227-024-06087-9. [DOI] [Google Scholar]
- 20.Li S.Y., Ho C.Y., Wu S.P. The impact of thematic teaching on student learning outcomes in computer programming applications. Educ. Inf. Technol. 2025 doi: 10.1007/s10639-025-13418-7. [DOI] [Google Scholar]
- 21.Liu K., Yao J., Tao D., et al. Influence of individual-technology-task-environment fit on university student online learning performance: the mediating role of behavioral, emotional, and cognitive engagement. Educ. Inf. Technol. 2023;28:15949–15968. doi: 10.1007/s10639-023-11833-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Çelik F., Baturay M.H. The effect of metaverse on L2 vocabulary learning, retention, student engagement, presence, and community feeling. BMC Psychol. 2024;12:58. doi: 10.1186/s40359-024-01549-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Love R., Law E., Cohen P.R., et al. Teaching a conversational agent using natural language: effect on learning and engagement. Int. J. Artif. Intell. Educ. 2025 doi: 10.1007/s40593-025-00461-1. [DOI] [Google Scholar]
- 24.Yorganci S. The impact of synchronous online discussions and online flipped learning on student engagement and self-regulation among preliminary undergraduates in a basic math course. Educ. Tech. Res. Dev. 2025 doi: 10.1007/s11423-025-10459-0. [DOI] [Google Scholar]
- 25.Stein K.C., Mauldin C., Marciano J.E., et al. Culturally responsive-sustaining education and student engagement: a call to integrate two fields for educational change. J. Educ. Change. 2025;26:29–55. doi: 10.1007/s10833-024-09510-3. [DOI] [Google Scholar]
- 26.Khoudi Z., Hafidi N., Nachaoui M., et al. New approach to enhancing student performance prediction using machine learning techniques and clickstream data in virtual learning environments. SN Comput. Sci. 2025;6:139. doi: 10.1007/s42979-024-03622-6. [DOI] [Google Scholar]
- 27.Ucan S. Investigating the processes involved in Twitter/X-supported collaborative learning and their relationship with learning outcomes. Educ. Inf. Technol. 2025;30:1123–1164. doi: 10.1007/s10639-024-13147-3. [DOI] [Google Scholar]
- 28.Thiruthuvanathan M.M., Krishnan B. Multitask EfficientNet affective computing for student engagement detection. Multimed. Tools Appl. 2024 doi: 10.1007/s11042-024-19815-3. [DOI] [Google Scholar]
- 29.An D., Ye C., Liu S. The influence of metacognition on learning engagement the mediating effect of learning strategy and learning behavior. Curr. Psychol. 2024;43:31241–31253. doi: 10.1007/s12144-024-06400-y. [DOI] [Google Scholar]
- 30.Deng R., Yang Y., Shen S. Impact of question presence and interactivity in instructional videos on student learning. Educ. Inf. Technol. 2025;30:1635–1663. doi: 10.1007/s10639-024-12862-1. [DOI] [Google Scholar]
- 31.Tang X., Gong Y., Xiao Y., et al. Facial expression recognition for probing students’ emotional engagement in science learning. J. Sci. Educ. Technol. 2025;34:13–30. doi: 10.1007/s10956-024-10143-7. [DOI] [Google Scholar]
- 32.Matz S.C., Bukow C.S., Peters H., et al. Using machine learning to predict student retention from socio-demographic characteristics and app-based engagement metrics. Sci. Rep. 2023;13:5705. doi: 10.1038/s41598-023-32484-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Azkarate Iturbe O., Álvarez-Huerta P., Muela A., et al. Latent class analysis of student engagement in higher education and its relationship to cooperative mindset and critical thinking. Innov. High. Educ. 2025 doi: 10.1007/s10755-025-09785-1. [DOI] [Google Scholar]
- 34.Kong X., Liang H., Wu C., et al. The association between perceived teacher emotional support and online learning engagement in high school students: the chain mediating effect of social presence and online learning self-efficacy. Eur. J. Psychol. Educ. 2025;40:33. doi: 10.1007/s10212-024-00932-4. [DOI] [Google Scholar]
- 35.Wang S., Li X., Gao H. A study on the social interaction characteristics of college student peers in science museums and their impact on learning outcomes: based on an analysis of the conversation. Res. Sci. Educ. 2024;54:1173–1197. doi: 10.1007/s11165-024-10181-6. [DOI] [Google Scholar]
- 36.Yi T.Y., Shreyans P., Vallabhajosyula R. Learning by making – student-made models and creative projects for medical education: systematic review with qualitative synthesis. BMC Med. Educ. 2025;25:143. doi: 10.1186/s12909-025-06716-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
No data was used for the research described in the article.





