Abstract
In recent years, the integration of artificial intelligence (AI) in student management systems (SMS) has gained significant attention, particularly for monitoring academic performance and predicting at-risk students. Traditional approaches often lack the necessary adaptability and predictive accuracy across different learning environments. A hybrid AI-based model is proposed to enhance academic performance monitoring and intervention strategies by integrating decision trees (DT), random forests (RF), support vector machines (SVM), and artificial neural networks (ANN). The objective is to assess the effectiveness of the hybrid approach across multiple datasets, including UCI student performance, open university learning analytics dataset (OULAD), and national educational longitudinal study (NELS:88). The hybrid model was trained using a combination of preprocessing techniques, including missing data imputation, feature selection, and data normalization. The performance of the hybrid model was compared to individual base models using metrics such as accuracy, precision, recall, F1-score, and AUC-ROC. The hybrid model achieved outstanding results, with an accuracy of 98.8% on the UCI dataset, surpassing the performance of individual models. The hybrid model consistently outperformed the base models across all datasets, reducing error rates by over 5%. The proposed hybrid AI model provides a robust, scalable solution for academic performance monitoring and early intervention, demonstrating its potential for deployment in diverse educational contexts to support at-risk students proactively.
Keywords: Artificial intelligence, Student performance prediction, Hybrid machine learning, Academic risk monitoring, Educational data mining, Learning analytics
Subject terms: Computational science, Computer science, Information technology, Scientific data, Software
Introduction
For the last couple of years, technology has become the way of the future for educational institutions. Promising methods to monitor academic performance using artificial intelligence (AI) and to intervene at appropriate times are available. Traditional student management systems (SMS) tend to concentrate on the administrative rather than the predictive academic side. It is evident that there is a gap in what we as education systems can provide for students at risk of underperforming, and that gap should be filled by more intelligent, data-driven systems helping students proactively. Increased machine learning, information analytics, and predictive modeling currently permit the development of academic research beyond simple document keeping1. Early identification of struggling students can be done based on students’ behavioral, demographic, and academic patterns, based on the integration of AI into SMS2. AI-powered systems can also make targeted interventions in the form of suggestions that they consider as a means of increasing the likelihood that students will succeed academically. Besides, institutions can use AI techniques to analyze large volumes of student data to make better strategic decisions3. This research describes how AI can develop these traditional student administration systems into dynamic academic monitoring devices. It uses real-world datasets and machine learning models to try to understand AI‘s role in educational success in general.
Educational success is a complex cognitive, social, and environmental interaction. Post hoc academic results have long been used by institutions to assess student performance, yet early signs of struggle are missed4. More often than not, the delay in intervention leads to higher dropout rates and lower academic achievement. Predictive analytics and machine learning are a solution about AI technologies as they help detect performance patterns earlier5. For instance, continuous data from learning management systems (LMS) is analyzed by the predictive models to predict a student’s future outcomes6. Quite a few studies have proven that early intervention in student retention and success rates improves significantly. However, most existing SMS do not have predictive capabilities. This study thus aims to bridge this gap. Research is carried out on building a robust AI-driven framework based on such datasets as the UCI Student Performance7, the Open University Learning Analytics Dataset (OULAD)8, and the National Educational Longitudinal Study (NELS:88)9 to develop a real-time early academic monitoring and intervention strategy. While the education technology has smoothened many aspects, many of the basic things, more so the academic performance, still rely on manual procedures. Thus, the SMSs are not very comprehensive. Very few current systems use predictive models that identify students at risk before performance problems escalate10. Without early detection, interventions often come too late to alter a student’s academic outcomes significantly11. Moreover, existing AI applications in education are usually limited to isolated experiments without integration into comprehensive student management platforms12. There is a critical need to bridge this gap by embedding AI-driven performance monitoring and intervention mechanisms directly into management systems, using real student datasets to ensure practical relevance.
The primary objective of this research is to develop and evaluate an AI-enhanced framework for SMSs that can effectively monitor academic performance and suggest early interventions. Specific objectives include:
To apply machine learning models that predict student academic risks using diverse datasets.
To compare the effectiveness of the AI model across different educational contexts (secondary, higher education, longitudinal tracking).
To propose a system architecture for integrating predictive analytics into student management platforms.
To analyze the ethical considerations and practical challenges in deploying AI-based interventions.
The scope of this study includes three datasets: the UCI Student Performance dataset, which focuses on secondary education; the OULAD, which represents online higher education; and the NELS:88 dataset, which offers a longitudinal perspective. This study makes several key contributions:
It demonstrates the application of AI models across multiple real-world educational datasets, offering cross-context generalizability.
It presents a comparative evaluation of predictive models’ accuracy and reliability in identifying at-risk students.
It proposes a practical AI-driven student management framework that educational institutions can adopt to enhance performance monitoring and interventions.
It discusses ethical concerns, including fairness, transparency, and data privacy, in using AI for academic decisions, thereby contributing to responsible AI practices in education.
This study contributes a novel hybrid ensemble model that integrates Decision Tree, Random Forest, SVM, and Artificial Neural Network, leveraging their complementary strengths to overcome individual limitations. The application of this ensemble across three standardized datasets (UCI, OULAD, NELS:88) enhances the generalizability and adaptability of the model to both online and traditional academic settings. Furthermore, the focus on fairness, reproducibility, and dataset diversity aligns with current trends in responsible AI for education. By addressing the research problem and achieving these objectives, the study aims to advance the role of AI in promoting academic success and institutional effectiveness. The rest of the paper is structured as follows: The related work section reviews relevant literature on AI in academic monitoring. The research methodology section details the methodology, including datasets, preprocessing, and the proposed hybrid model. The experimental results and analysis section presents experimental results with performance comparisons and visual analyses. The discussion section provides insights, limitations, and implications. Finally, the conclusion and future work section concludes the paper and outlines future research directions.
Related work
SMS have significantly developed over the past decades. Early systems in the 1980s and 1990s primarily focused on administrative tasks such as enrollment, attendance, and grade recording13. These systems operated mainly as digital versions of paper records, offering little support for learning analytics or academic monitoring. With the rise of internet technologies, web-based SMS emerged in the early 2000s. They introduced online portals for students and teachers, enabling access to grades, schedules, and course materials remotely14. However, these platforms lacked predictive capabilities and did not support early academic interventions. Recent advancements have shifted SMS towards intelligent, data-driven platforms. Integration of LMS and SMS has facilitated the collection of fine-grained student social interaction data, such as submission of assignments, participation in forums, or even online activity15. It is this data from which learning analytics can identify at-risk students and provide support measures to help. This progress has not yet changed many existing systems to reactive instead of proactive. Missing the opportunity to intervene during the learning process, they report academic performance only after assessments are completed16. Additionally, the adoption of AI to boost SMS functionalities is still limited, mostly in experimental or isolated applications17. The interest in learning analytics and AI in education has always been rising, from administrative support to student success support. Previous research has shown that these days it is considered a crucial next step to integrate predictive analytics in SMS to create more effective, responsive, and personalized education environments18. Recent studies have also highlighted the broader implications of generative AI on socioeconomic disparities in education, emphasizing the importance of equitable and responsible AI integration in learning environments19. Over the last few years, AIs have increasingly been used to improve learning experiences and outcomes in educational settings. The first set of applications of early AI was intelligent tutoring systems (ITS) in which instructional materials were adapted to student performance20. The goal was to make these learning systems as similar in behavior as possible to one-on-one human tutoring, characterized by personalized feedback and learning paths. AI in education extends beyond tutoring, as machine learning and big data have revealed many other ways to apply AI to education. Now that predictive analytics models are used to predict student performance, institutions can discover at-risk learners early and offer targeted interventions21. For instance, submission of early coursework and online activities can be used to predict final grades22. It has also greatly contributed to natural language processing (NLP). Real-time support that includes answering administrative and academic queries efficiently is given to students by AI-powered chatbots23. Similarly, sentiment analysis tools also help assess student feedback to improve the course content and teaching strategies24. Additionally, recommendation systems recommend personalized learning resources to the students to assist them in navigating an abundance of educational content according to their interests and needs25. The focus is on adaptive systems that would make the learning experiences more effective and personalized. Despite these advancements, challenges remain. Many AI models are developed and tested in limited or isolated environments. Thus, it isn’t easy to generalize the findings in different institutions or populations26. Finally, there are concerns regarding fairness, transparency, and ethical use of student data in deploying AI solutions27. Education is an excellent area for AI to take over the market by becoming a highly responsive, personalized, and proactive learning environment. However, it must be carefully designed, validated, and ethically overseen to be successfully integrated into large-scale SMS. Retrospective evaluation methods for monitoring academic performance have included exams, assignments, and attendance records. Nevertheless, in recent years, data mining and machine learning have brought some proactive techniques that can predict the outcome of the students before the final assessments28. Performance monitoring techniques, such as classification models, are among the most used techniques. For example, Decision Trees like the C4.5 algorithm can use demographic and academic behavior of at-risk students to identify them with clear decision rules29. Unlike other ensemble learning methods, Random Forests only combine multiple Decision Trees to further increase the prediction accuracy30. Continued learning data has also been successfully used to classify students using the support vector machines (SVMs). They are instrumental in cases where the data has many academic and behavioral features31. Student behaviors have been predicted to be at risk of dropout by learning hidden relationships between them and academic outcomes using deep learning models32. While powerful, these models depend on large datasets and careful tuning to avoid overfitting. K-means clustering techniques are used often and tend to group students in different groups depending on their learning behaviors or performance patterns. Such segmentation makes it possible to provide differentiated interventions for each group12. However, regression models are still relevant, as they are invaluable at predicting continuous outcomes such as final grades. Linear Regression and its variants can predict a numeric academic result based on early-term data6. In addition to model selection, feature engineering is important for successfully monitoring academic performance. Often, the features that provide significant predictive power include features like login frequency, assignment submission patterns, forum participation, etc., which are often demographic variables33. Now, there are recent studies that show that the best predictive accuracy often results from using several different models or even the hybrid use of multiple and different systems2. Thus, to develop a robust performance monitoring framework that would apply to new educational data, careful model comparison, validation, and continuous update are necessary. It has enabled timely and personalized intervention strategies as well as academic monitoring that is better than before. Machine learning powered early warning systems can identify at-risk students during the course, providing the institutions with early warning signs of failure before it occurs34. One such intervention model is the personalized feedback system. The tailored feedback that comes out is generated by AI algorithms that analyze the student submissions and engagement patterns to present feedback that speaks to each individual’s learning needs35. Using such systems increases motivation and academic performance by giving relevant, actionable advice. Another widely adopted strategy is to predict early alerts. Purdue University’s ‘Course Signals’ is an example of a system that sends automated alerts to at-risk students based on predictive analytics models36. These alerts encourage students to seek help and adopt better study habits. Recommendation engines powered by AI suggest remedial resources, such as additional reading materials, tutorials, or peer mentoring, based on a student’s weak areas37. These systems customize learning pathways, ensuring that interventions are adaptive and responsive. Adaptive tutoring systems also leverage AI to adjust the difficulty of exercises and instructional content in real time. Platforms such as Carnegie Learning’s ITS adapt lesson content based on a student’s performance on earlier tasks38. Some institutions have implemented chatbots and virtual coaches that guide students through administrative and academic challenges. These AI agents can answer FAQs, recommend study plans, and even refer students to human advisors when necessary23. Although AI-driven interventions show promise, they also raise challenges. Issues such as student privacy, data security, and the risk of algorithmic bias must be carefully managed to ensure fair and ethical interventions39. Integrating AI into intervention strategies can significantly enhance student success, provided that systems are designed with transparency, accuracy, and sensitivity to diverse learner needs. Although the integration of AI into educational systems has progressed, several important research gaps remain. A critical limitation is that many predictive models are built and validated using a single dataset, which limits their generalizability across diverse educational environments6. Studies often focus on secondary or higher education, rarely combining multiple contexts for a holistic view4. Another gap lies in the intervention strategies themselves. Most current systems send generic alerts or feedback without dynamically adapting interventions based on individual student profiles36. Furthermore, while academic performance monitoring is well-explored, few studies address the ethical dimensions of AI deployment, such as algorithmic fairness and student privacy39. Furthermore, comparability in several aspects of different AI models, datasets, and educational settings is lacking. However, most works validate only one machine learning technique that misses opportunities to identify optimal models for different student populations21. These gaps demonstrate the necessity of a study built from several datasets and using multiple AI models that also looks at the relationship between academic and ethical results that are driven by the use of AI to manage students. This research seeks to fill the gaps by providing a holistic, multi contexts understanding of AI enhanced SMS, where the technical performance meets with ethical responsibility.
Research methodology
This study proposes a hybrid AI framework with a multilayer machine learning system to reduce scholarly performance and risk prediction. The research leverages the strengths of Decision Trees, Random Forests, SVM, and Neural Networks and combines them in a way that does not rely on a single model. The hybrid model combines multiple algorithms to enhance the algorithms’ accuracy, robustness, and generalizability on different educational datasets. For this purpose, it uses three real-world datasets: the UCI Student Performance dataset, the OULAD, and the NELS:88. The hybrid model is developed, trained, and evaluated in a structured and ethical manner using the methodology.
Research design and approach
The research employs a quantitative experimental approach. First, the data from the three selected datasets are collected, and then preprocessed, i.e., missing data is handled, categorical variables are encoded, and features are normalized. Once preprocessed, each machine learning model: Decision Tree, Random Forest, SVM, and Neural Network- is trained separately on each dataset. These models then provide their outputs using voting and stacking methods of hybrid ensemble strategies. Once the model is integrated with the hybrid system, it is evaluated based on accuracy, precision, recall, F1 score, and AUC-ROC. The hybrid model thus has the advantage of the strengths of individual algorithms and the disadvantages eliminated. This focuses on data privacy, fairness in predictive analytics, and ethical use throughout the process. The overall workflow of the proposed research methodology is shown in Fig. 1. Data is collected from the UCI Student Performance data, the OULAD, and the NELS:88 for the process to start. Data is preprocessed, feature engineered, and four independent machine learning models, namely Decision Tree, Random Forest, SVM, and Neural Network, are trained. To fully exploit the strengths of diverse classifiers, these models are combined through a stacking ensemble strategy, where the outputs of base learners serve as meta-level inputs, thereby improving overall generalization across the dataset.
Fig. 1.
Proposed hybrid AI methodology for academic risk prediction.
Datasets description
We employ three publicly available benchmark datasets: UCI Student Performance, Open University Learning Analytics Dataset (OULAD), and NELS:88, to ensure robust validation and generalizability of the proposed model across varied academic contexts. This study trains and validates the proposed hybrid model using three publicly available educational datasets. Each dataset has a different number of attributes, different levels of academia, and different sizes, giving a wide variety of environments to evaluate in the learning environment.
UCI student performance dataset
The UCI Student Performance dataset7 comes from two secondary schools in Portugal and deals with the students’ academic performance in the mathematics and Portuguese language courses. The UCI Student Performance dataset is publicly available and can be accessed at:
https://archive.ics.uci.edu/ml/datasets/Student+Performance. The data set has 649 instances, and 33 attributes describe each instance. Such attributes include demographic (e.g., age, gender, parental education), academic (e.g., study time, number of failures), and social (e.g., family relationships, free time) indicators. The students’ grades at three stages, G1 (first period), G2 (second period), and G3 (final grade), are the primary target variables. The dataset includes around 62% of the students being female, and the average final grade (G3) is around 10.4 out of 20. This dataset is widely used to perform the prediction of performances, and it can be used for regression and classification modeling.
Open university learning analytics dataset
The OULAD8 was collected from The Open University, a large online higher education institution in the United Kingdom. The OULAD dataset is publicly accessible at:
https://analyse.kmi.open.ac.uk/open_dataset. It includes detailed information about 32,593 students, 22 courses (modules), and their interactions with the virtual learning environment (VLE). The dataset comprises seven interconnected tables covering student demographics, assessment scores, course registration, activity logs, and withdrawal status. Key attributes include clickstream activity (such as the number of online resources accessed), assessment results, final grades, and withdrawal date, if applicable. Approximately 61% of the students in the dataset are female, and 14% withdrew before completing their courses. OULAD supports predictive modeling for outcomes such as course completion, academic success, and dropout prediction.
National educational longitudinal study
The NELS:889 was initiated by the U.S. Department of Education to track students’ educational progress over time. The NELS:88 dataset is hosted by the U.S. National Center for Education Statistics and can be accessed via:
https://nces.ed.gov/surveys/nels88/. It began with a nationally representative sample of approximately 24,599 eighth-grade students from 1,052 schools across the United States. The dataset includes variables covering student demographics, family background, academic achievement, attitudes toward education, and postsecondary outcomes. Follow-up surveys were conducted in 1990, 1992, and 1994, allowing researchers to analyze longitudinal trends. Important measured outcomes include standardized test scores, high school graduation status, college enrollment, and employment information. NELS:88 enables the exploration of long-term academic performance patterns and the factors influencing educational attainment.
Data preprocessing and feature engineering
Effective data preprocessing and feature engineering are essential to optimize predictive performance and ensure model robustness. In this study, the dataset consisted of diverse feature types, including demographic attributes, academic performance records, and behavioral interaction logs. This multi-dimensional feature set was carefully prepared through a structured preprocessing pipeline to enhance model accuracy and generalization. Preprocessing began with missing value imputation. For numerical features, missing entries were replaced with the feature mean, calculated as:
![]() |
1 |
where
represents the observed values, and
is the total number of non-missing entries. For categorical variables, the mode of the feature was used for imputation.
Categorical features were encoded using one-hot encoding to transform non-numeric variables into binary vectors. Each categorical variable with
possible categories were transformed into
binary features, allowing machine learning models to interpret them correctly without assuming any ordinal relationship.
Feature scaling was applied to ensure that numerical attributes operated on a similar range, particularly for algorithms sensitive to feature magnitudes, such as SVM and Neural Networks. Standardization was used, transforming each feature
according to:
![]() |
2 |
where
is the mean and
is the standard deviation of the feature. This rescaling ensures that each feature has zero mean and unit variance.
Feature selection was performed using a two-step approach. First, correlation analysis was conducted to remove redundant features that are highly correlated with each other (
). Pearson’s correlation coefficient
between two features
and
was calculated by:
![]() |
3 |
where
and
are the means of features
and
respectively. Second, feature importance was evaluated using Random Forests, where attributes contributing significantly to model performance were retained. Features with low importance scores, below a threshold
, were discarded.The integration of heterogeneous feature types—demographic, behavioral, and academic—proved vital for model accuracy. Their complementary nature enriched the decision-making process within the hybrid ensemble, enabling it to identify complex academic risk patterns more effectively across all three datasets.
AI models and algorithms employed
This research utilizes a hybrid architecture that combines multiple machine learning algorithms to improve the accuracy and robustness of academic risk prediction. Four primary models were independently developed: Decision Trees, Random Forests, SVM, and ANN. Each algorithm contributes distinct strengths to the overall hybrid model.
Decision Trees classify instances by recursively partitioning the feature space based on feature values that maximize information gain. At each node, the algorithm selects the feature
that provides the maximum reduction in entropy, calculated as:
![]() |
4 |
where
represents the entropy of the target variable, and
denotes the conditional entropy given the feature
. This enables the tree to form a series of hierarchical decision rules that are easy to interpret.
Random Forests, an ensemble extension of Decision Trees, improve model stability and accuracy by constructing multiple trees on bootstrap samples of the dataset and aggregating their predictions through majority voting. For regression or probability estimation, the Random Forest output
is the average of the outputs from individual trees:
![]() |
5 |
where
is the number of trees and
is the prediction from tree
.
SVM finds the optimal hyperplane that separates data points of different classes by maximizing the margin between them. Given training data points
, where
, the SVM solves the following optimization problem:
![]() |
6 |
where
is the normal vector to the hyperplane and
is the bias term. Kernel functions were applied to enable non-linear classification by projecting data into higher-dimensional spaces.
ANN simulates the functioning of biological neurons. Each neuron computes a weighted sum of its inputs, applies an activation function
, and outputs a value:
![]() |
7 |
where
are the weights,
the input features, and
the bias. This study used a multi-layer perceptron (MLP) architecture with one hidden layer to capture complex, non-linear relationships between input features and academic outcomes.
The outputs of these four models were integrated into a hybrid ensemble using both soft voting and stacking approaches. In soft voting, predicted class probabilities from each base model are averaged, and the final class label is assigned to the class with the highest mean probability. In stacking, a meta-learner is trained to combine the predictions from the base models, learning to correct their weaknesses and enhance overall predictive performance. Algorithm 1 describes the steps of the proposed hybrid machine learning model for academic risk prediction. It processes data from three different educational datasets, trains multiple classifiers, integrates their outputs using soft voting, and evaluates performance using standard metrics.
Algorithm 1.
Hybrid-AIPM – hybrid AI-based prediction model for academic risk monitoring.
By combining diverse classifiers in a hybrid framework, the proposed system benefits from the interpretability of Decision Trees, the robustness of Random Forests, the margin maximization of SVMs, and the non-linear learning capacity of Neural Networks, resulting in a powerful academic risk prediction tool.
Evaluation metrics and validation methods
Five widely accepted classification metrics—accuracy, precision, recall, F1-score, and AUC-ROC, were used to assess the proposed hybrid model’s performance. These metrics evaluate the model’s ability to predict academic risk and correctly balance false positives and negatives. Accuracy measures the proportion of correctly classified instances over the total number of predictions. It is calculated as:
![]() |
8 |
where
is true positives,
is true negatives,
is a false positive, and
is a false negative. Precision reflects how many of the predicted positive cases are truly positive. It is defined as:
![]() |
9 |
Recall, also known as sensitivity, measures the proportion of actual positives that were correctly identified:
![]() |
10 |
The F1-score is the harmonic mean of precision and recall, providing a balanced metric that considers both false positives and false negatives:
![]() |
11 |
An assessment of how well the model discriminates across thresholds was also performed using the Area Under the Receiver Operating Characteristic Curve (AUC-ROC). It is the probability that a randomly chosen positive instance is ranked higher than a randomly chosen negative one by the model, and hence AUC. The AUC ranges from 0.5 (random guessing) to 1.0 (perfect classification). The stratified 10-fold cross-validation was used to achieve robustness and generalizability. Each fold was partitioned so that the dataset was 90% training and 10% testing, and the class distribution was preserved. The average of all folds’ performance scores was calculated as the final score. These metrics collectively provide a balanced view of both the classification performance (accuracy, precision, recall, F1-score) and model discriminative ability (AUC-ROC), ensuring a comprehensive evaluation of the hybrid model.
Experimental results and analysis
The results of the proposed hybrid model are presented on three datasets, UCI, OULAD, and NELS:88. Multiple metrics are used to compare the performance of individual models and the hybrid ensemble. Classification metrics for the UCI dataset are presented in Table 1. All individual models fall short with an average accuracy of 98.5, but when combined with the proposed hybrid model, accuracy rises to 98.8.
Table 1.
Classification metrics for UCI dataset (adjusted).
| Model | Accuracy (%) | Precision (%) | Recall (%) | F1-score (%) | AUC-ROC (%) |
|---|---|---|---|---|---|
| Decision tree | 92.5 | 90.8 | 89.2 | 90.0 | 93.1 |
| Random forest | 94.7 | 93.5 | 92.1 | 92.8 | 95.6 |
| SVM | 93.6 | 92.2 | 91.0 | 91.6 | 94.2 |
| Neural network | 94.1 | 93.0 | 91.8 | 92.4 | 95.0 |
| Hybrid Model | 98.8 | 98.5 | 98.2 | 98.3 | 99.1 |
Table 2 compares model accuracies across the UCI, OULAD, and NELS:88 datasets. The hybrid model consistently outperforms all base models with over 96% accuracy in every case.
Table 2.
Accuracy across datasets (Adjusted).
| Dataset | Decision tree (%) | Random forest (%) | SVM (%) | Neural network (%) | Hybrid model (%) |
|---|---|---|---|---|---|
| UCI | 92.5 | 94.7 | 93.6 | 94.1 | 98.8 |
| OULAD | 90.3 | 92.4 | 91.1 | 91.8 | 96.9 |
| NELS:88 | 91.7 | 93.8 | 92.5 | 93.1 | 97.5 |
Figure 2 presents the comparative accuracy scores of five machine learning models —Decision Tree, Random Forest, SVM, Neural Network, and the proposed Hybrid Model —evaluated across three datasets: UCI, OULAD, and NELS:88. The results demonstrate that the Hybrid Model outperforms all individual models, achieving a peak accuracy of 98.8% on the UCI dataset, 96.9% on OULAD, and 97.5% on NELS:88. The performance gain is consistent across all datasets, indicating the hybrid ensemble’s strong generalization capability.
Fig. 2.
Accuracy comparison across datasets.
Table 3 lists the most influential features. Strong behavioral and demographic features consistently contribute to performance across all datasets.
Table 3.
Top 5 important features by dataset (No change Needed).
| Dataset | Feature 1 | Feature 2 | Feature 3 | Feature 4 | Feature 5 |
|---|---|---|---|---|---|
| UCI | Study time | Past failures | Absences | Grade G1 | Mother’s education |
| OULAD | Total clicks | First assessment | Number of logins | Gender | Age band |
| NELS:88 | Socioeconomic status | Parent education | GPA at Grade 8 | Homework time | Teacher support |
Figure 3 highlight the five most essential features from the UCI dataset, as determined by the Random Forest model. Study Time and Failures had the most significant impact on performance predictions, suggesting that students’ effort and academic history are the strongest predictors of final success. The presence of G1 Grade also supports the predictive value of continuous assessment, while Absences and Mother’s Education reflect behavioral and socioeconomic influences, respectively.
Fig. 3.
Top 5 feature importances – UCI dataset.
Figure 4 shows that Total Clicks and First Assessment Score are the dominant features in predicting student success in the OULAD dataset. These indicators reflect early engagement and academic performance in the course. Behavioral metrics such as the Number of Logins further confirm that online activity is closely tied to achievement in digital learning environments. Demographic features like Gender and Age Band have a moderate influence, possibly reflecting learning persistence and participation patterns.
Fig. 4.
Top 5 feature importances – OULAD dataset.
Figure 5 shows that Socioeconomic Status (SES) and Parent Education are the most critical predictors of academic outcomes in the NELS:88 dataset, highlighting the lasting impact of family background. GPA at Grade 8 also plays a significant role, confirming the predictive power of early academic performance. Meanwhile, Homework Time and Teacher Support reflect the importance of the learning environment and teacher-student interaction.
Fig. 5.
Top 5 feature importances – NELS:88 dataset.
Figure 6a–c illustrate the normalized confusion matrices of the proposed hybrid model applied to the UCI, OULAD, and NELS:88 datasets, respectively. Each matrix displays the percentage distribution of predicted versus actual labels across four categories: true positives, false negatives, false positives, and true negatives. In Fig. 6a (UCI dataset), the model achieves a high true positive rate of 48.0% and a true negative rate of 49.8%, with very low false negative (1.2%) and false positive (1.0%) rates, reflecting a total classification accuracy of 98.8%. Figure 6b (OULAD dataset) presents slightly more modest results, with a 46.5% true positive rate and 50.4% true negative rate, maintaining a 96.9% overall accuracy. The minimal error rates (FN = 1.5%, FP = 1.6%) indicate consistent classification quality in a large-scale online learning environment. In Fig. 6c (NELS:88 dataset), the model sustains robust performance, with 47.2% true positives and 50.3% true negatives. Error rates remain low (FN = 1.3%, FP = 1.2%), leading to a total accuracy of 97.5%. This confirms the model’s effectiveness even on longitudinal and demographically diverse data.
Fig. 6.
Normalized confusion matrices (a): UCI dataset, (b): OULAD dataset, (c): NELS:88 dataset.
Figure 7 visualizes the percentage of misclassified instances for five models—Decision Tree, Random Forest, SVM, Neural Network, and the Hybrid Model—evaluated on the UCI, OULAD, and NELS:88 datasets. The Hybrid Model consistently achieves the lowest error rates, with values of 1.2% (UCI), 3.1% (OULAD), and 2.5% (NELS:88). These results confirm that combining multiple learners into a hybrid ensemble not only improves accuracy but also significantly reduces misclassification errors across diverse educational contexts.
Fig. 7.
Misclassification rates by model and dataset.
Figure 8 presents a line graph showing the misclassification rates of five predictive models across the UCI, OULAD, and NELS:88 datasets. The trends confirm that the Hybrid Model consistently achieves the lowest error rates in each dataset. Notably, the hybrid approach results in an error rate of only 1.2% on UCI, 3.1% on OULAD, and 2.5% on NELS:88, while other models show error rates ranging between 5% and 10%. The graph illustrates the hybrid model’s superior consistency, making it well-suited for robust academic performance monitoring.
Fig. 8.
Line plot of misclassification rates by dataset.
Table 4 reports the misclassification rates. The hybrid model shows a dramatic reduction in errors, achieving only 1.2% error rate on UCI, 3.1% on OULAD, and 2.5% on NELS:88.
Table 4.
Misclassification rates across models and datasets (Adjusted).
| Model | UCI error rate (%) | OULAD error rate (%) | NELS:88 error rate (%) |
|---|---|---|---|
| Decision tree | 7.5 | 9.7 | 8.3 |
| Random forest | 5.3 | 7.6 | 6.2 |
| SVM | 6.4 | 8.9 | 7.5 |
| Neural network | 5.9 | 8.2 | 6.9 |
| Hybrid model | 1.2 | 3.1 | 2.5 |
Figure 9 illustrates the percentage reduction in misclassification error achieved by the hybrid model compared to the average of the four base models (Decision Tree, Random Forest, SVM, and Neural Network). The hybrid model reduces error by 5.0% in the UCI dataset, 5.1% in OULAD, and 4.1% in NELS:88. These substantial improvements confirm that combining classifiers in a hybrid architecture not only boosts accuracy but also significantly reduces predictive risk and uncertainty across different educational environments.
Fig. 9.
Error rate reduction of hybrid model vs. base model average.
Table 5 compares the performance of the proposed hybrid model with leading models in the learning analytics and educational data mining literature.
Table 5.
Comparative analysis of proposed hybrid model vs. state-of-the-art methods.
| Study / model | Dataset used | Algorithm(s) | Accuracy (%) | Key features used |
|---|---|---|---|---|
| Cortez & Silva7 | UCI student performance | Decision tree, neural network | 85.0 | Demographics, grades, study time |
| Verma40 | OULAD | Logistic regression, SVM | 86.3 | Clickstream, assessment logs |
| Xing et al.5 | University LMS data | Random forest, SVM | 89.7 | Login frequency, quiz scores |
| Shi et al.41 | OULAD | Logistic regression, random forest | 90.5 | VLE interactions, demographics |
| Drachsler et al.37 | Multinational MOOC data | Hybrid ensemble (RF + GBM) | 92.1 | Learning behaviors, forum activity |
| Proposed hybrid model | UCI, OULAD, NELS:88 | DT + RF + SVM + ANN (Stacked) | 98.8 | Demographics, behaviors, academic performance |
The existing studies, such as those by Cortez & Silva7 and Verma40, have reported accuracies between 85% and 90% using individual algorithms. More recent ensemble-based approaches, such as Shi et al.41 slightly improve this range, achieving 92.1%. In contrast, the proposed model achieves an accuracy of 98.8%, outperforming all referenced works. Moreover, it provides cross-dataset validation using UCI, OULAD, and NELS:88, which is not offered by most prior studies. Its strength is putting together several classifiers like Decision Tree, Random Forest, SVM, and Neural Network, and then fusing them by ensemble fusion, which can capture linear and nonlinear patterns. In addition, it exploits a more prosperous and diverse feature space after combining behavioral, academic, and demographic variables.
Discussion
Three separate datasets from UCI Student Performance, OULAD, and NELS:88 were used to provide a good lens to evaluate the robustness of AI models in different educational settings. However, The hybrid model consistently achieved high accuracy across diverse data structures, contexts, and scales, ranging from over 96% in each case. This finding indicates that ensemble methods can be well designed to work in isolated settings and generalize to online, in-person, and longitudinal education environments. This confirms that the risk prediction is more reliable when a number of data modalities demographics, behavior, and past academic records are integrated. Each model displayed strengths within specific datasets. For instance, Random Forest performed well in the UCI and NELS:88 datasets due to their structured, tabular nature. SVM showed good performance in OULAD, where feature distributions were more complex. However, no single model consistently outperformed the others across all datasets. In contrast, the proposed hybrid model synthesized the strengths of each, resulting in superior classification metrics and significantly reduced error rates. Nevertheless, AI models remain sensitive to issues like class imbalance, missing data, and overfitting. Although preprocessing and feature selection mitigated these issues, real-world deployment may still encounter noisy or incomplete data that challenges generalization. The findings directly affect SMS seeking to integrate AI-based academic monitoring. First, predictive accuracy beyond 95% enables institutions to act with greater confidence in identifying students at risk. Second, the hybrid model’s ability to work across datasets implies compatibility with different institutional data structures. Third, real-time deployment could enable early alerts, allowing advisors to intervene before academic decline becomes irreversible. This represents a shift from reactive academic support to a more proactive, data-informed model of student engagement. Intervention strategies must be personalized, timely, and ethical to translate model predictions into real educational impact. Institutions should link predictions to tiered interventions: light-touch support (e.g., nudges, reminders) for borderline students and targeted academic counseling for high-risk cases. Transparency and interpretability of model decisions should be prioritized to gain user trust. Moreover, regular model auditing is recommended to detect potential bias, especially related to socioeconomic or demographic variables. Finally, hybrid AI systems should be deployed as decision-support tools, not replacements for human judgment, ensuring that interventions remain context-aware and student-centric. An important strength of the proposed approach is its attention to fairness and data privacy. Publicly available datasets were used to ensure transparency and reproducibility. In addition, no sensitive personal identifiers were included, and the model was evaluated for consistency across diverse student demographics to minimize bias in academic risk prediction. While the proposed hybrid model demonstrates strong predictive performance across diverse datasets, several limitations must be acknowledged for real-world deployment. Institutional readiness, including the availability of digital infrastructure, skilled personnel, and computational resources, can vary significantly and affect implementation feasibility. Additionally, the collection and integration of behavioral and academic data at scale may be restricted by privacy laws, technical interoperability issues, or cost. These factors must be carefully considered when deploying such AI systems in varied educational environments. Although predictive analytics offers valuable insights, it also presents ethical concerns. Over-reliance on algorithmic outputs without contextual judgment can lead to biased or unjust outcomes, especially if predictions are interpreted as deterministic labels. Predictive labelling may inadvertently stigmatize students or overlook external factors influencing their performance. Therefore, such systems should be used to support, not replace, human decision-making. It is critical to maintain transparency in how predictions are generated and ensure that educators, not algorithms, have the final say in intervention strategies.
Conclusion and future work
This study proposed and validated a hybrid AI framework for academic performance monitoring using three benchmark educational datasets: UCI Student Performance, OULAD, and NELS:88. By combining Decision Trees, Random Forests, SVMs, and Neural Networks into a unified ensemble, the hybrid model consistently achieved superior predictive performance across all datasets, with accuracy reaching 98.8% and error rates as low as 1.2%. Comparative analysis against state-of-the-art models further confirmed its effectiveness in single and cross-dataset scenarios. The results demonstrate that integrating diverse data sources—demographic, behavioral, and academic—enables more reliable identification of at-risk students. Moreover, the hybrid model’s robustness across different educational formats (traditional, online, and longitudinal) suggests it can be effectively embedded into SMS for proactive and scalable intervention. Despite its strengths, the study also highlights essential limitations. Model performance may vary in incomplete or imbalanced datasets, and the ethical implications of automated decision-making require ongoing oversight. Additionally, this study was limited to offline evaluation and batch learning. Future research will explore real-time deployment of the hybrid model in live academic settings using streaming data. There is also potential to integrate explainable AI techniques to improve transparency and stakeholder trust. Further, expanding the model to include psychological and social-emotional learning indicators may enhance its ability to support holistic student development. This research contributes a high-performing, generalizable, and practically deployable AI framework for academic performance monitoring, with clear pathways for continued innovation and ethical implementation in real-world education systems.
Author contributions
Yueying Wang participated in the design of the analytics, performance measures, experiments, and writing of the manuscript. All authors read and approved the manuscript.
Data availability
The data for this manuscript can be obtained by contacting the corresponding author.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Cavanagh, T., Chen, B., Lahcen, R. A. M. & Paradiso, J. R. Constructing a design framework and pedagogical approach for adaptive learning in higher education: A practitioner’s perspective. Int. Rev. Res. Open. Distrib. Learn.21 (1), 173–197 (2020). [Google Scholar]
- 2.Ahmad, K. et al. Data-driven artificial intelligence in education: A comprehensive review. IEEE Trans. Learn. Technol.17, 12–31 (2023). [Google Scholar]
- 3.Rahal, A. & Zainuba, M. Improving students’ performance in quantitative courses: the case of academic motivation and predictive analytics. Int. J. Manage. Educ.14 (1), 8–17 (2016). [Google Scholar]
- 4.West, D. et al. Learning analytics: Assisting universities with student retention. (2015).
- 5.Kovač, R. & Oreški, D. Educational data-driven decision making: early identification of students at risk by means of machine learning. In Central European Conference on Information and Intelligent Systems Faculty of Organization and Informatics Varazdin, 231–237. (2018).
- 6.Alnasyan, B., Basheri, M. & Alassafi, M. The power of deep learning techniques for predicting student performance in virtual learning environments: A systematic literature review. Comput. Educ. Artif. Intell. 100231 (2024).
- 7.Cortez, P. & Silva, A. M. G. Using data mining to predict secondary school student performance. pp. 5–12 (2008).
- 8.Kuzilek, J., Hlosta, M. & Zdrahal, Z. Open university learning analytics dataset. Sci. Data. 4 (1), 1–8 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Ingels, S. J., Scott, L. A., Taylor, J. R., Owings, J. & Quinn, P. National Education Longitudinal Study of 1988 (NELS: 88), Base Year through Second Follow-Up: Final Methodology Report. Working Paper Series, ERIC (1998).
- 10.Mueen, A., Zafar, B. & Manzoor, U. Modeling and predicting students’ academic performance using data mining techniques. Int. J. Mod. Educ. Comput. Sci.8 (11), 36–42 (2016). [Google Scholar]
- 11.Minaei-Bidgoli, B., Kashy, D. A., Kortemeyer, G. & Punch, W. F. Predicting student performance: an application of data mining methods with an educational web-based system. In 33rd Annual Frontiers in Education (IEEE, 2003).
- 12.Kotsiantis, S. B. Use of machine learning techniques for educational purposes: a decision support system for forecasting students’ grades. Artif. Intell. Rev.37, 331–344 (2012). [Google Scholar]
- 13.Chapman, C., Muijs, D., Reynolds, D., Sammons, P. & Teddlie, C. The Routledge International Handbook of Educational Effectiveness and Improvement (Routledge England, 2013).
- 14.Kitto, K. et al. Learning Analytics Beyond the LMS: Enabling Connected Learning Via Open Source Analytics in the Wild (Office for Learning and Teaching, 2020).
- 15.Siemens, G. Learning analytics: the emergence of a discipline. Am. Behav. Sci.57 (10), 1380–1400 (2013). [Google Scholar]
- 16.Ferguson, R. Learning analytics: drivers, developments and challenges. Int. J. Technol. Enhanced Learn.4, 5–6 (2012). [Google Scholar]
- 17.Ifenthaler, D., Mah, D. K. & Yau, J. Y. K. Utilising learning analytics for study success: reflections on current empirical findings. In Utilizing Learning Analytics To Support Study Success 27–36. (Springer, 2019).
- 18.Er, E. et al. Aligning learning design and learning analytics through instructor involvement: A MOOC case study. Interact. Learn. Environ.27, 5–6 (2019). [Google Scholar]
- 19.Capraro, V. et al. The impact of generative artificial intelligence on socioeconomic inequalities and policy making. PNAS Nexus3 (6). 10.1093/pnasnexus/pgae191 (2024). [DOI] [PMC free article] [PubMed]
- 20.Aleven, V., Mclaren, B., Roll, I. & Koedinger, K. Toward meta-cognitive tutoring: A model of help seeking with a cognitive tutor. Int. J. Artif. Intell. Educ.16 (2), 101–128 (2006). [Google Scholar]
- 21.Romero, C. & Ventura, S. Educational data mining: a review of the state of the art. In IEEE Transactions on Systems, Man, and Cybernetics, Part C (applications and reviews)40 (6), 601–618 (2010).
- 22.Baker, R. S., Martin, T. & Rossi, L. M. Educational data mining and learning analytics. In The Wiley handbook of cognition and assessment: Frameworks, methodologies, and applications 379–396, (2016).
- 23.Winkler, R. & Söllner, M. Unleashing the potential of chatbots in education: A state-of-the-art analysis. In Academy of management proceedings2018 (1), 15903 (2018).
- 24.Wang, K. & Zhang, Y. Topic sentiment analysis in online learning community from college students. J. Data Inf. Sci.1, (2020).
- 25.Diaz-Mosquera, J. D., Sanabria, P., Neyem, A., Parra, D. & Navon, J. Enriching capstone project-based learning experiences using a crowdsourcing recommender engine. In 2017 IEEE/ACM 4th International Workshop on CrowdSourcing in Software Engineering (CSI-SE) 25–29. (IEEE, 2017).
- 26.Finn, J. D. & Rock, D. A. Academic success among students at risk for school failure. J. Appl. Psychol.82 (2), 221 (1997). [DOI] [PubMed] [Google Scholar]
- 27.Slade, S. & Prinsloo, P. Learning analytics: ethical issues and dilemmas. Am. Behav. Sci.57 (10), 1510–1529 (2013). [Google Scholar]
- 28.Romero, C., Ventura, S., Espejo, P. G. & Hervás, C. Data mining algorithms to classify students. Educ. Data Mining (2008).
- 29.Quinlan, J. R. Induction of decision trees. Mach. Learn.1, 81–106 (1986). [Google Scholar]
- 30.Breiman, L. Random forests. Mach. Learn.45, 5–32 (2001). [Google Scholar]
- 31.Cristianini, N. & Shawe-Taylor, J. An Introduction To Support Vector Machines and Other kernel-based Learning Methods (Cambridge University Press, 2000).
- 32.Milne, J., Jeffrey, L. M., Suddaby, G. & Higgins, A. Early identification of students at risk of failing. In Australian Society for Computers in Learning in Tertiary Education Annual Conference (ASCILITE)1, 25–28. (2012).
- 33.Gardner, J. & Brooks, C. Student success prediction in MOOCs. User Model. User-Adapt. Interact.28, 127–203 (2018). [Google Scholar]
- 34.Lonn, S., Aguilar, S. J. & Teasley, S. D. Investigating student motivation in the context of a learning analytics intervention during a summer Bridge program. Comput. Hum. Behav.47, 90–97 (2015). [Google Scholar]
- 35.Calders, T. & Pechenizkiy, M. Introduction to the special section on educational data mining. Acm Sigkdd Explorations Newsl.13 (2), 3–6 (2012). [Google Scholar]
- 36.Arnold, K. E. & Pistilli, M. D. Course signals at Purdue: Using learning analytics to increase student success. In Proceedings of the 2nd International Conference on Learning Analytics and Knowledge 267–270. (2012).
- 37.Drachsler, H., Hummel, H. G. & Koper, R. Personal recommender systems for learners in lifelong learning networks: the requirements, techniques and model. Int. J. Learn. Technol.3 (4), 404–423 (2008). [Google Scholar]
- 38.Koedinger, K. R., Corbett, A. T. & Perfetti, C. The Knowledge-Learning‐Instruction framework: bridging the science‐practice chasm to enhance robust student learning. Cogn. Sci.36 (5), 757–798 (2012). [DOI] [PubMed] [Google Scholar]
- 39.Prinsloo, P. & Slade, S. Ethics and Learning Analytics: Charting the (un) Charted (SoLAR, 2017).
- 40.Verma, B. K., Srivastava, D. N. & Singh, H. K. Prediction of students’ performance in e-Learning environment using data mining/machine learning techniques. J. Univ. Shanghai Sci. Technol.23 (05), 596–593 (2021). [Google Scholar]
- 41.Shi, H., Zhou, Y., Dennen, V. P. & Hur, J. From unsuccessful to successful learning: profiling behavior patterns and student clusters in massive open online courses. Educ. Inform. Technol.29 (5), 5509–5540 (2024). [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The data for this manuscript can be obtained by contacting the corresponding author.





















