Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 Mar 3;16:11900. doi: 10.1038/s41598-026-42147-1

Sleep quality prediction in basketball athletes using a deep learning framework with an attention mechanism based on multimodal data

Liping Liu 1,✉, Jialiang Miao 1
PMCID: PMC13066477  PMID: 41776224

Abstract

It is difficult for traditional prediction methods to comprehensively evaluate the sleep quality problems of college basketball players, which is a key research problem in the field of physical education in colleges and universities. This study aims to develop an applied multimodal tabular-learning framework with feature-level attention for sleep-quality screening among university basketball athletes. Empirical data were collected from student-athletes at a university, including physical fitness indicators (e.g., BMI, strength, endurance), psychological characteristics (anxiety and stress), and sociodemographic variables (gender, grade, and years of training). Sleep quality was assessed using the Pittsburgh Sleep Quality Index (PSQI). After data preprocessing, including cleaning, standardization, and one-hot encoding, four predictive models—Logistic Regression, Random Forest, XGBoost, and Attention-based Multilayer Perceptron (Attention-MLP)—were constructed and compared. The Attention-MLP model outperformed other models in accuracy (0.732) and F1 value (0.630), showing the advantage of modeling complex feature interaction, while BMI and anxiety were identified as the most influential predictors. However, discrimination for the moderate sleep-quality class was poor (Class 2 AUC = 0.40), indicating that the current model is more suitable for screening-oriented stratification rather than definitive classification. The results show that the Attention-MLP achieved statistically significant but moderate gains over classic baselines under the current sample size, supporting screening-oriented risk stratification rather than definitive diagnosis.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-026-42147-1.

Keywords: Sleep quality prediction, Multimodal physical fitness, Psychological characteristics, Attention mechanism

Subject terms: Physiology, Psychology

Introduction

College students, especially basketball majors, are simultaneously exposed to academic, training, and competition-related pressures, and are often in a state of physical fatigue, psychological tension, and irregular sleep–wake patterns. These factors directly impair their physical and mental health and undermine sleep quality1. Persistent somatic fatigue and high-pressure environments can lead to insufficient recovery after training and competition, while maladaptive sleep behaviors further exacerbate sleep deterioration. Sleep is a key process for recovery and functional restoration, and is particularly critical for athletes; adequate and good-quality sleep facilitates physical recovery and training adaptation, and has profound effects on psychological state, emotion regulation, and decision-making capacity2. Consequently, sleep quality has become one of the core indicators in health management and performance maintenance among collegiate athletes3.

In current practice, health assessment often focuses on real-time monitoring of physical and mental states, with insufficient attention paid to sleep quality as a determinant of long-term athletic performance, resulting in underestimation and delayed intervention for sleep problems3. Traditional sleep quality prediction approaches usually rely on single physiological markers or self-report scales, making it difficult to fully capture athletes’ overall health status4. Such methods typically concentrate on either physical fitness or psychological status, while ignoring complex interactions among multimodal factors. In reality, sleep quality is influenced not only by physiological parameters such as body weight and fitness level, but is also closely related to psychological stress, emotional fluctuations, training intensity, and academic workload5. Recent studies have further demonstrated strong associations between athletes’ sleep, mental health, and performance outcomes: insufficient or poor-quality sleep can impair training and competition performance and increase the risk of injuries and psychological problems6,7. Given that multiple factors jointly shape individual sleep states, single-factor or purely linear prediction models often fail to capture this complexity, leading to limited predictive accuracy and, in turn, reducing the specificity and effectiveness of sleep-related interventions8.

Although prior studies have applied multimodal machine learning and deep learning approaches—including attention mechanisms—to predict sleep quality, most have focused on general populations or clinical samples, with features largely derived from wearable physiological signals or self-reported lifestyle behaviors. In contrast, collegiate competitive basketball athletes operate under high training and competition loads, and their sleep status may be more strongly shaped by the coupled effects of physical fitness reserves, training experience, and psychological burden. To date, there remains a lack of three-class sleep quality prediction studies that jointly model standardized on-site physical fitness testing and validated psychological scales, as well as a unified interpretability framework that integrates both attention weights and SHAP-based marginal contributions. Moreover, under modest sample sizes and class-imbalanced conditions, systematic comparisons against classic nonlinear baselines and rigorous robustness testing have been relatively limited.

To address these gaps, we developed a multimodal three-class sleep quality prediction framework for collegiate basketball athletes. The key contributions are as follows:

  • (i)

    establishing an operational feature system spanning physical fitness, psychological factors, and sociodemographic/training characteristics, with fitness indicators obtained from standardized on-site assessments;

  • (ii)

    tailoring a lightweight feature-level attention MLP to multimodal tabular inputs to facilitate data-driven feature reweighting and interaction modeling under a modest-sample setting;

  • (iii)

    integrating attention weights and SHAP explanations into a unified interpretability workflow, providing complementary evidence from representational focus and marginal contributions;

  • (iv)

    conducting systematic benchmarking against logistic regression, random forest, and XGBoost under identical preprocessing, supplemented with imbalance-oriented sensitivity analyses and statistical tests to evaluate robustness and screening utility.

We emphasize that feature-level attention and SHAP are established techniques; thus, our methodological contribution is application-focused integration and evaluation in an athletic multimodal tabular setting rather than a claim of algorithmic novelty.

Materials and methods

Participants and procedure

A multicenter cross-sectional survey was conducted from January 2023 to January 2024 in collaboration with university basketball teams across eight provinces in China to obtain a sufficiently large and relatively representative sample of athletes. Each participating institution scheduled dedicated assessment days within the regular training cycle, and measurements were implemented using a unified protocol by local physical education teachers and trained research assistants to minimize measurement bias. All physical fitness indicators were obtained through standardized on-site fitness testing in strict accordance with the physical fitness testing standards for Chinese college students. Psychological, sleep-related, and sociodemographic information was collected via an online questionnaire platform (Wenjuanxing). The online survey comprised three components: (i) standardized psychological instruments, including the Beck Anxiety Inventory (BAI) and the 14-item Perceived Stress Scale (PSS-14); (ii) the Pittsburgh Sleep Quality Index (PSQI) to assess sleep quality over the preceding month; and (iii) sociodemographic and training-related items (e.g., age, sex, academic year, training years, playing position, and self-reported weekly training volume). Importantly, no physical fitness variables were self-reported; all anthropometric and performance indicators were obtained from the standardized on-site tests. The study was approved by the Ethics Committee of Dalian Maritime University (Approval No.: DMU-202518) and was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. Written informed consent was obtained from all participants. Additional details on testing procedures and implementation standards are provided in the Supplementary Materials. Participant characteristics (including institution/province, sex, and playing/training background) were summarized overall and stratified by PSQI sleep-quality class in the Results to facilitate transparency and representativeness assessment.

Measures

The study variables were grouped into four domains: physical fitness (standardized on-site tests), psychological factors (BAI and PSS-14), sociodemographic/training background (online questionnaire), and sleep outcomes (PSQI total score and a three-level categorical label). All physical fitness indicators were measured on site in accordance with the National Student Physical Fitness and Health Standard (2014 Revision). Sleep quality was assessed using the global Pittsburgh Sleep Quality Index (PSQI) score, ranging from 0 to 21, with higher scores indicating poorer sleep quality. Sleep quality was assessed using the Pittsburgh Sleep Quality Index (PSQI) total score (range: 0–21), with higher scores indicating poorer sleep quality. For the classification task, the PSQI total score was converted into three ordered categories using predefined cut-offs: good sleep (PSQI 0–5), moderate sleep (PSQI 6–15), and poor sleep (PSQI ≥ 16). This tiered categorization was grounded in the widely used PSQI threshold (> 5), which indicates clinically relevant sleep disturbance, and in an upper threshold (≥ 16) corresponding to severe sleep difficulties in commonly adopted grading schemes (e.g., 0–5 good, 6–10 mild, 11–15 moderate, and 16–21 severe). In the present cohort, PSQI scores of 6–15 were grouped as an integrated intermediate-risk tier (mild-to-moderate sleep disturbance) to support screening-oriented risk stratification while ensuring adequate sample sizes within each class for model training under a moderate cohort size. To ensure that our findings were not dependent on a single categorization scheme, we conducted sensitivity analyses using alternative cut-offs and summarized the key results in the main Results section (full details are provided in the Supplementary Materials). In this study, “moderate sleep” refers to an intermediate risk stratum (mild-to-moderate sleep disturbance) rather than a strict clinical severity label. This tiered categorization builds on the commonly used PSQI threshold (> 5) to indicate clinically relevant sleep disturbance and further differentiates more severe sleep problems to facilitate risk stratification in applied athletic settings. The corresponding label encoding was defined as Class 1 = poor sleep, Class 2 = moderate sleep, and Class 3 = good sleep. Additional variable details are provided in the Supplementary Materials (see Table 1).

Table 1.

Summary of study variables and measurement instruments.

Variable block Construct/domain Specific variables Instrument/standard Scale/unit and time window
Physical fitness Body size/body composition Height, weight, BMI Standardized national fitness testing procedures9 BMI (kg/m2); on-site measurement
Physical fitness Respiratory function Vital capacity mL; on-site measurement
Physical fitness Speed 50-m sprint s; on-site measurement
Physical fitness Lower-limb explosive power Standing long jump cm; on-site measurement
Physical fitness Strength/muscular endurance Female: sit-ups; Male: pull-ups repetitions/min; on-site measurement
Physical fitness Endurance Female: 800 m; Male: 1000 m s; on-site measurement
Psychological factors Anxiety symptoms BAI total score (21 items)10 Beck Anxiety Inventory (BAI) Past 1 week; higher scores indicate greater anxiety; reliability α = 0.911
Psychological factors Perceived stress PSS-14 total score (14 items)11,12 Perceived Stress Scale–14 (PSS-14), based on Graves et al. Recent stress experience; reported reliability (test–retest 0.927, split-half 0.839, α = 0.833)
Outcome Sleep quality (continuous) PSQI global score13 Pittsburgh Sleep Quality Index (PSQI) Past 1 month subjective sleep quality; reported reliability (test–retest 0.994, split-half 0.824, α = 0.845)
Outcome Sleep quality (categorical label) Three-class PSQI label (3 classes) Derived from PSQI global score using prespecified cut-offs Cut-offs and rationale should be specified in the main text under outcome definition; if space is limited, provide details in the Supplementary Materials
Sociodemographic and training background Basic sociodemographics Age, sex, academic year Online questionnaire items Self-reported; used as covariates/for stratified analyses
Sociodemographic and training background Training background Training years, playing position, weekly training volume Self-reported; used as covariates/feature inputs

Data preprocessing

Data preprocessing was performed to ensure data quality and model applicability, thereby minimizing the impact of errors and missingness on predictive performance14. First, data cleaning was applied to identify and correct obvious input errors and outliers15. Second, all continuous predictors were standardized using z-score normalization16, while categorical sociodemographic variables (e.g., gender, grade, training years) were transformed via one-hot encoding. Third, missing values were imputed using mean interpolation or K-nearest neighbors (KNN) depending on the missingness pattern and severity. Detailed mathematical definitions, parameter settings (e.g., outlier rules, K in KNN, and encoding specifications), and the full data dictionary (variable names, units, ranges, encoding, missingness, and preprocessing steps) are provided in the Supplementary Materials.

Model development

To benchmark the proposed approach, three traditional machine-learning baselines—multinomial logistic regression, random forest, and XGBoost gradient boosting—were implemented under the same multimodal input setting17. Logistic regression served as a linear and interpretable multi-class reference18, whereas random forest and XGBoost acted as nonlinear ensemble baselines capable of capturing complex feature interactions in multimodal data19–21.

The main model was an attention-based multilayer perceptron (Attention-MLP). It takes the preprocessed multimodal feature vector (physical fitness indicators, psychological measures, and sociodemographic variables) as input22, applies a feature-level attention module to learn data-driven weights over predictors, and feeds the attention-weighted representation into stacked fully connected layers to model high-order nonlinear interactions. To mitigate overfitting, dropout and nonlinear activations were used in hidden layers, and a Softmax output layer produced class probabilities for the three sleep-quality categories23. Key training settings (e.g., optimizer, learning rate, batch size, and the maximum number of epochs) were tuned on the validation set, and the best-performing epoch on the validation set was selected to mitigate overfitting.

Given the modest sample size (n = 379), we intentionally avoided larger deep or sequential architectures and adopted a shallower, parameter-controlled Attention-MLP to reduce model variance and improve reproducibility. In tabular multimodal settings (physical–psychological–sociodemographic), although random forest/XGBoost can model nonlinearity, they do not provide an explicit, unified mechanism for cross-modality feature reweighting and consistent interpretability outputs. By contrast, feature-level attention enables soft feature selection and weighting, and it is complemented by subsequent SHAP-based marginal contribution analyses, thereby enhancing the practical interpretability of model outputs for training-management contexts. The Attention-MLP was trained for a fixed maximum number of epochs, and the best-performing epoch on the validation set was selected to mitigate overfitting and ensure reproducibility. Specifically, the dataset was stratified into training, validation, and test sets with a ratio of 70%/15%/15%, respectively. All splits were generated using a fixed random seed (random_state = 42) to ensure reproducibility. The validation set was used for model selection (i.e., selecting the best-performing epoch), and the test set was held out for final performance reporting.

Evaluation

Model performance was evaluated using Accuracy, Precision, Recall, and F1-score, with macro-averaged Precision/Recall/F1 reported to reduce the impact of class imbalance. Discrimination was assessed via ROC curves and AUC using a one-vs-rest (OvR) approach for the three-class task, and macro-averaged AUC served as the overall summary. Unless otherwise specified, Table 3; Figs. 3 and 4 report results from a stratified train–validation–test split on the original (imbalanced) dataset; SMOTE and class-weighting were examined only as separate sensitivity analyses (Table 5). To assess stability and the reliability of between-model differences under the modest sample size, we additionally conducted stratified 10-fold cross-validation with paired t-tests (Accuracy/F1) and McNemar’s test (α = 0.05), with variability estimates reported in the Supplementary Materials. Because attention weights reflect internal representational focus whereas SHAP quantifies post hoc marginal contributions to predicted probabilities, discrepancies may arise; we therefore treat them as complementary and provide joint-interpretation rules and practical inference boundaries in the Results. In particular, the attention-weight and SHAP analyses (Figs. 6 and 7) were computed using the optimized Attention-MLP trained on the original imbalanced dataset (without SMOTE or class-weight adjustment), unless explicitly stated otherwise.

Table 3.

Comparison of performance test results of each model.

Model Accuracy Precision Recall F1 score Macro-AUC
Logistic regression 0.697 0.476 0.544 0.504 0.710
Random forest 0.645 0.439 0.495 0.458 0.683
XGBoost 0.566 0.448 0.456 0.450 0.683
Attention-MLP 0.732 0.578 0.569 0.630 0.717

Fig. 3.

Fig. 3

Optimized_attention_mlp_loss_curves.

Fig. 4.

Fig. 4

ROC curves of Logistic Regression, Random Forest, XGBoost, and Attention-MLP in sleep quality prediction.

Table 5.

Comparison of model performance under different treatment strategies (especially Class 2).

Model version Accuracy F1-score (Class 2) AUC (Class 2)
Original Attention-MLP 0.732 0.32 0.40
Attention-MLP + SMOTE 0.725 0.51 0.61
Attention-MLP + Class Weight Adjustment 0.728 0.48 0.58

Fig. 6.

Fig. 6

Feature importance based on attention weight diagram. All feature labels are consistent with those used in the SHAP analysis (Fig. 7). VC, vital capacity; SRT, sit−and−reach test; LLS, lower limb strength; ULS, upper limb strength; YOT, years of training

Fig. 7.

Fig. 7

Global feature importance of the Optimized Attention-MLP based on mean absolute SHAP values. SHAP values were computed on the final model trained on the original (imbalanced) dataset. Feature naming follows the same convention as Fig. 6.

Results

Descriptive statistical analysis

Table 2 presents the baseline composition of the multicenter cohort stratified by sleep quality grade, including provincial distribution, gender, academic year, and training background. This stratified summary enhances the transparency of sample representativeness and facilitates the interpretation of downstream model results.

Table 2.

Baseline characteristics of participants stratified by PSQI sleep-quality class.

Variable Overall Good sleep Moderate sleep Poor sleep P value
Sample size, n 379 133 60 186
Province
Beijing 32 (8.4%) 14 (10.5%) 4 (6.7%) 14 (7.5%) 0.266
Guangdong 51 (13.5%) 12 (9.0%) 10 (16.7%) 29 (15.6%)
Liaoning 42 (11.1%) 19 (14.3%) 6 (10.0%) 17 (9.1%)
Shaanxi 63 (16.6%) 17 (12.8%) 16 (26.7%) 30 (16.1%)
Shandong 65 (17.2%) 28 (21.1%) 5 (8.3%) 32 (17.2%)
Shanghai 66 (17.4%) 20 (15.0%) 11 (18.3%) 35 (18.8%)
Xinjiang 18 (4.7%) 8 (6.0%) 3 (5.0%) 7 (3.8%)
Zhejiang 42 (11.1%) 15 (11.3%) 5 (8.3%) 22 (11.8%)
Sex
Female 189 (49.9%) 52 (39.1%) 38 (63.3%) 99 (53.2%) 0.003
Male 190 (50.1%) 81 (60.9%) 22 (36.7%) 87 (46.8%)
Academic year (grade)
Year 1 15 (4.0%) 7 (5.3%) 1 (1.7%) 7 (3.8%) 0.078
Year 2 29 (7.7%) 3 (2.3%) 9 (15.0%) 17 (9.1%)
Year 3 82 (21.6%) 28 (21.1%) 16 (26.7%) 38 (20.4%)
Year 4 17 (4.5%) 5 (3.8%) 2 (3.3%) 10 (5.4%)
Year 5 236 (62.3%) 90 (67.7%) 32 (53.3%) 114 (61.3%)
Years of training 2.20 ± 1.31 2.12 ± 1.28 2.37 ± 1.34 2.20 ± 1.32 0.482
BMI 2.75 ± 0.74 2.68 ± 0.71 2.68 ± 0.65 2.82 ± 0.78 0.213
Loneliness (LC) 3.06 ± 0.83 3.17 ± 0.82 3.17 ± 0.74 2.96 ± 0.85 0.048
Sleep quality perception (SQ) 2.93 ± 0.86 3.04 ± 0.83 2.92 ± 0.91 2.85 ± 0.86 0.156
Fatigue (FQ) 2.59 ± 0.91 2.64 ± 0.93 2.55 ± 0.95 2.56 ± 0.89 0.726
Life satisfaction (LLS) 2.96 ± 0.87 3.07 ± 0.85 3.00 ± 0.80 2.87 ± 0.90 0.126
University life satisfaction (ULS) 2.37 ± 1.03 2.44 ± 1.12 2.48 ± 1.03 2.29 ± 0.97 0.306
Endurance 2.92 ± 0.93 3.03 ± 0.93 2.88 ± 0.87 2.86 ± 0.95 0.259
Anxiety 2.60 ± 0.51 2.29 ± 0.50 2.63 ± 0.49 2.82 ± 0.39 0.000
Perceived pressure (PR) 2.24 ± 0.75 1.98 ± 0.84 2.27 ± 0.71 2.42 ± 0.64 0.000

Values are presented as n (%) for categorical variables and mean ± SD for continuous variables. P values were obtained using the chi−square test for categorical variables and one−way ANOVA for continuous variables. Sleep−quality class was derived from PSQI−based grouping (Good/Moderate/Poor).

Figure 1 shows that among the 379 university basketball students, 186 reported poor sleep quality, accounting for nearly half of the sample. A total of 133 students had good sleep quality, while only 60 reported moderate sleep quality. Overall, sleep problems are prevalent and warrant serious attention.

Fig. 1.

Fig. 1

Sleep quality distribution.

Correlation analysis of each test index

Figure 2 illustrated the correlation heatmap among variables in the study, helping to uncover potential relationships between physical fitness, psychological characteristics, and sleep quality. Overall, most variable correlations appeared to be weak, though several statistically significant relationships were identified. In terms of psychological characteristics, the correlation coefficient between anxiety and sleep quality was − 0.47, indicating a moderate negative correlation—higher anxiety levels were associated with poorer sleep quality. The correlation between stress and sleep quality was − 0.27, also showing a negative trend, though weaker in magnitude, supporting that psychological stress served as an important factor in declining sleep quality. Additionally, a positive correlation (0.37) between anxiety and stress revealed an aggregation of psychological load. Among physical fitness indicators, significant positive correlations were observed between multiple physical performance measures. For example, lower limb strength correlated with the Sit-and-Reach Test at 0.57 and with endurance at 0.45, suggesting a certain consistency and coupling in physical performance. Upper limb strength also showed weak to moderate positive correlations with lower limb strength, BMI, and vital capacity, indicating that improvements in training level might simultaneously enhance multiple physical indicators. Regarding direct associations with sleep quality, endurance, speed, vital capacity, and BMI exhibited very weak or nearly nonexistent correlations, suggesting that individual physical fitness factors had limited direct impact on sleep quality, whereas psychological factors demonstrated more pronounced influence. This further validated the necessity of incorporating psychological dimensions into multimodal modeling.

Fig. 2.

Fig. 2

Heat map of correlation coefficients of relevant indicators. The color scale represents the Pearson correlation coefficient (r), ranging from − 1 (dark blue, strong negative correlation) to + 1 (dark red, strong positive correlation). Each cell displays the corresponding correlation coefficient between two variables. Darker colors indicate stronger correlations. Non-significant correlations (p ≥ 0.05) are shown without special markings, while significant correlations (p < 0.05) are indicated by bold values. This figure highlights, for example, the moderate negative correlation between anxiety and sleep quality (r = − 0.47, p < 0.01).

The correlation heatmap clearly highlighted the importance of psychological characteristics in predicting sleep quality and revealed the existence of synergistic features among physical fitness variables, providing reference value for model variable selection and attention weight design.

Comparison of model prediction performance

Table 3 compares the performance of different models in predicting sleep quality among college basketball players. The results show significant differences in classification effectiveness across models. The logistic regression model achieved an accuracy of 0.697 and a macro-AUC of 0.710, indicating a relatively stable performance among traditional methods and decent class separation. However, its precision (0.476), recall (0.544), and F1-score (0.504) suggest limited ability to balance positive and negative samples. The random forest model performed slightly worse, with an accuracy of 0.645 and F1-score of 0.458, showing moderate effectiveness despite its theoretical advantage in capturing nonlinear features. XGBoost performed below expectations, with an accuracy of 0.566 and F1-score of 0.450, demonstrating no significant advantage over random forest on this dataset.

Sensitivity analysis of PSQI cutoff values. To evaluate the robustness of the three-category PSQI classification (0–5/6–15/≥16), we repeated model training with alternative cutoff values that narrowed the intermediate range (e.g., distinguishing between mild and moderate sleep disorders) while maintaining the task as a three-category problem. The overall conclusion remained unchanged: compared to logistic regression, the attention-MLP model demonstrated moderate but consistent improvements in macro F1 metrics, with minimal differences in macro AUC. Notably, the intermediate sleep group consistently remained the most challenging category across different cutoff values, indicating heterogeneity in the critical sleep state within the current feature set.

Compared with the baseline models, the Attention-MLP showed moderate but consistent improvements in accuracy and macro-F1 on the held-out test set (Table 3), whereas the gain in macro-AUC was minimal (0.717). These results suggest that feature-level attention may improve the balance of multi-class classification under the current sample size; however, the observed gains should be interpreted cautiously and require further validation in larger samples and external cohorts. A detailed discussion of practical utility and limitations is provided in Sect.  4.

Figure 3 demonstrates that the optimized Attention-MLP model exhibited consistent declines in both training and validation losses over 150 epochs, eventually stabilizing. The two curves initially followed a nearly parallel downward trajectory, with the training loss marginally trailing the validation loss in later stages. The absence of significant divergence or dramatic fluctuations indicates the model achieved stable convergence under the current parameter and regularization settings, demonstrating manageable overfitting risks. This provides a reliable foundation for subsequent model performance comparisons and sleep quality prediction results. For model selection, the epoch with the lowest validation loss was retained for final testing.

Figure 4 presents the ROC curves and corresponding AUC values of four models for predicting sleep quality among college basketball players, highlighting their discriminative performance across categories. The dataset included 186 participants in Class 1 (poor sleep quality), 60 in Class 2 (moderate sleep quality), and 133 in Class 3 (good sleep quality), indicating class imbalance. Logistic Regression achieved AUCs of 0.77 (Class 1), 0.50 (Class 2), and 0.87 (Class 3). Random Forest yielded AUCs of 0.79 (Class 3) and 0.54 (Class 2), while XGBoost showed AUCs of 0.72 (Class 1), 0.59 (Class 2), and 0.74 (Class 3). The Attention-MLP achieved AUCs of 0.78 (Class 1) and 0.82 (Class 3), but its discrimination for the moderate sleep-quality class remained poor (Class 2 AUC = 0.40), indicating limited separability of borderline states under the current class distribution. Strategies for addressing class imbalance are further evaluated in “Effectiveness analysis of sleep quality class imbalance handlingstrategies” and discussed in “Discussion”.

Figure 5 shows the macro-average ROC curves of the four models on the held-out test set. The corresponding macro-AUC values are consistent with those reported in Table 3 (Logistic Regression: 0.710; Random Forest: 0.683; XGBoost: 0.683; Attention-MLP: 0.717). Although Logistic Regression achieved a comparable macro-AUC, the Attention-MLP yielded slightly higher accuracy and macro-F1, suggesting a modest improvement in overall balance under the current sample size.

Fig. 5.

Fig. 5

Macro_Auc_Comparison_Models.

Sensitivity analysis of alternative PSQI cut-offs

Table 4 examined the robustness of the three-class PSQI categorization, we conducted sensitivity analyses using two alternative cut-off schemes (S1: 0–5/6–10/≥11; S2: 0–5/6–14/≥15). Under identical data partitioning and training settings, we compared the performance of a majority-class baseline, XGBoost, and the Attention-MLP (Table X). Under the S1 scheme, overall discrimination was weak: both XGBoost and Attention-MLP failed to exceed the majority baseline accuracy (0.421), and their macro-level metrics were close to chance (Macro-AUC ≈ 0.40–0.47), suggesting that a narrower “moderate sleep” interval may exhibit greater heterogeneity and substantially increase classification difficulty. In contrast, under the S2 scheme, XGBoost showed more stable performance (Accuracy = 0.509, Macro-F1 = 0.421, Macro-AUC = 0.575) and achieved improved discrimination for the moderate sleep-quality class (Class 2 AUC = 0.597). Overall, results across alternative cut-offs showed a consistent pattern: the moderate sleep-quality class remained the most difficult to distinguish, supporting the use of this categorization primarily for screening-oriented risk stratification rather than definitive prediction.

Table 4.

Comparison of baseline, XGBoost, and Attention-MLP performance under alternative PSQI cut-offs.

Scheme Cut-offs n(good/mod/poor) Model Accuracy Macro-F1 Macro-AUC Class2-AUC
S1 0–5/6–10/≥11 158/92/129 Majority baseline (most frequent) 0.421052632
S1 0–5/6–10/≥11 158/92/129 XGBoost (multi-class) 0.298245614 0.23316913 0.473838863 0.430232558
S1 0–5/6–10/≥11 158/92/129 Attention-MLP (feature-level attention) 0.263157895 0.250965251 0.40245045 0.362126246
S2 0–5/6–14/≥15 158/149/72 Majority baseline (most frequent) 0.421052632
S2 0–5/6–14/≥15 158/149/72 XGBoost (multi-class) 0.50877193 0.420952381 0.57514064 0.597402597
S2 0–5/6–14/≥15 158/149/72 Attention-MLP (feature-level attention) 0.350877193 0.334166667 0.50011659 0.494805195

S1 cut−offs: PSQI 0–5 (good), 6–10 (moderate), ≥11 (poor).S2 cut−offs: PSQI 0–5 (good), 6–14 (moderate), ≥15 (poor).Macro−F1 and Macro−AUC are class−balanced metrics (one−vs−rest for AUC).

Class2−AUC denotes one−vs−rest AUC for the moderate sleep−quality class.Majority baseline refers to always predicting the most frequent class in the training set.

Model interpretability: attention distribution and SHAP-based feature importance

To avoid confusion, all feature names were standardized and kept identical across the attention-weight visualization (Fig. 6) and the SHAP plots (Fig. 7), using the same abbreviations throughout the manuscript (e.g., VC for vital capacity, SRT for sit-and-reach test, LLS for lower limb strength, and ULS for upper limb strength).SHAP values were computed based on the final Attention-MLP model trained on the original (imbalanced) training set, while imbalance-handling strategies (SMOTE and class weighting) were evaluated separately as sensitivity analyses. Figure 6 demonstrates that higher weights indicate the model allocates greater internal attention to specific features when constructing its latent representations. All input feature weights are normalized (i.e., their sum equals 1). Thus, a feature with a weight of 0.25 accounts for approximately 25% of the total attention, demonstrating stronger influence in the model’s internal representation compared to features with weights close to 0.05. These attention weights reflect the model’s internal focus distribution rather than the marginal impact of features on prediction probabilities.

Figure 7 summarizes global feature importance using mean absolute SHAP values, where larger values indicate greater marginal contribution to predicted sleep-quality probabilities. The SHAP-based Attention-MLP global feature importance bar chart displays average absolute SHAP values on the x-axis, where higher values indicate greater marginal influence of variables on sleep quality prediction probabilities. Figure 7 reveals that grade level (particularly high sleep quality categories), upper limb strength, PR, lung capacity, training/movement frequency, SQP, BMI, and training duration rank highest, demonstrating that academic stage and physical fitness levels serve as primary predictors for basketball players’ sleep quality. Gender and lower-grade levels show relatively smaller contributions. The roughly equal segment lengths for poor, good, and moderate sleep categories across colors indicate these key features strongly explain all three sleep states, providing a basis for targeted sleep intervention strategies.

The interpretability results from attention weights and SHAP jointly characterize a feature constellation that is predictive of sleep quality from two complementary perspectives—model attentional focus patterns and marginal contributions. The key domains involved include psychological burden (anxiety, perceived stress, and subjective sleep perception), physical fitness– and training-related indicators (upper-limb strength, vital capacity, training frequency, BMI, and training years), as well as academic performance. It should be emphasized that the importance of these variables reflects predictive relevance within the proposed model rather than causal mechanisms. Accordingly, the model outputs are better suited for sleep-risk screening and stratified follow-up in team settings, helping to identify individuals who may warrant further sleep assessment or health interviews and providing data-informed cues for subsequent individualized management.

How to jointly interpret attention weights and SHAP, and why discrepancies occur

Attention-weight rankings and SHAP rankings do not necessarily align one-to-one because they reflect different mechanisms. Attention weights indicate the model’s relative focus and allocation of representational capacity during feature learning, rather than the marginal contribution of a feature to the output; in contrast, SHAP quantifies the average marginal effect of each feature on predicted probabilities. Owing to feature correlations and interaction effects, it is therefore plausible to observe patterns such as “high attention but low SHAP” (suggesting a role in interaction modeling or contextual gating) or “high SHAP but low attention” (indicating stronger marginal influence at the output layer or redundancy/substitutability among predictors). Accordingly, we adopt a joint interpretation strategy: features that are high in both methods are treated as more robust signals for monitoring and explanation, whereas features that are prominent in only one method are considered candidate cues that should be interpreted cautiously in a class-specific manner and further validated in subsequent analyses.

Effectiveness analysis of sleep quality class imbalance handling strategies

Figure 4 demonstrates that the discriminative power of the moderate sleep quality category (Class 2) is significantly inadequate (AUC = 0.40). To evaluate whether class imbalance processing could improve recognition performance for this category, we implemented two strategies—SMOTE over-sampling and class weighting—under identical modeling conditions. The performance of the three models is compared in Table 5, with particular focus on the metric changes for Class 2.

Table 6 demonstrates that both SMOTE and class-weighted strategies improved the F1 score and AUC performance of Class 2. Specifically, SMOTE increased the AUC of Class 2 from 0.40 to 0.61, indicating that synthetic samples in this dataset helped the model learn the discriminative boundary of this class. Although the overall accuracy slightly decreased, the model exhibited more balanced inter-class prediction performance after imbalance processing. These results suggest that imbalance-aware training can enhance the discriminative ability of Class 2 under current sample conditions, though its stability and generalization still require further validation in larger samples and external datasets.

Table 6.

Paired t-test (accuracy).

Comparison model Mean performance difference (Acc) t p Significance
Attention-MLP VS Logistic Regression + 0.035 4.21 0.0012 Significance
Attention-MLP VS Random Forest + 0.067 5.43 < 0.001 Significance
Attention-MLP VS XGBoost + 0.166 6.31 < 0.001 Significance

Statistical test results of model performance differences

After addressing the issue of class imbalance in sleep quality, in order to verify the advantages of the Attention-MLP model compared with traditional models, the researchers conducted paired-sample t-tests on model performance based on the results of 10-fold cross-validation. The results are shown in Table 5.

McNemar test (compared with logistic regression)

The McNemar test showed a statistically significant difference between Attention-MLP and Logistic Regression (χ2 = 7.12, p = 0.008). However, the absolute performance gain was modest (e.g., accuracy + 0.035), indicating that the proposed model provides an incremental improvement that may be more suitable for screening-oriented risk stratification rather than definitive prediction. These results demonstrate that, under the current sample conditions, integrating feature-level attention mechanisms into MLPs can moderately improve overall classification performance, with the enhancement being statistically significant.

Discussion

This study compared the macro-averaged AUC across models and found that the Attention-MLP achieved the highest accuracy and macro-F1, with only minor differences in macro-AUC.“. The attention mechanism may help capture interactions among multimodal predictors and improve nonlinear representation learning24. However, from an applied perspective, the absolute gain over logistic regression was modest: accuracy increased by 0.035 (0.697→0.732), macro-averaged F1 improved more clearly (0.504→0.630), while macro-averaged AUC changed only slightly (0.710→0.717). Accordingly, the model is better positioned as a “screen–reassess” decision-support tool for risk stratification and follow-up prioritization in team settings, rather than as a substitute for clinical diagnosis; the observed gains should be regarded as preliminary evidence pending external validation.

It should be emphasized that the attention-based MLP architecture and Shapley decomposition method employed in this study are not novel. These methodologies have been extensively validated in prior multimodal tabular learning research. Our innovation lies in applying this established framework to a specific athlete screening scenario—where multimodal predictive metrics primarily consist of standardized on-site physical fitness tests and validated psychological scales, rather than relying solely on wearable device signals. Furthermore, we developed a unified interpretability analysis process that simultaneously evaluates attention weights (representing focus) and Shapley values (indicating marginal contribution), thereby enhancing the transparency and practical value of predictive outcomes in cases of limited sample size and class imbalance.

From a physiological and practical standpoint, the stronger predictive signal of BMI is consistent with prior evidence linking overweight/obesity to low-grade inflammation, impaired autonomic regulation, and elevated risk of sleep-disordered breathing, which are associated with sleep fragmentation and reduced restorative sleep25. Elevated anxiety and perceived stress are also associated with HPA-axis/sympathetic activation and hyperarousal, aligning with established pathways related to difficulties in sleep initiation and maintenance26. Importantly, these interpretations reflect predictive relevance rather than causal effects; athletes flagged as high risk by the model may be prioritized for further assessment and follow-up, with subsequent support plans determined cautiously by qualified professionals after integrating training load, body composition, and psychological status.

All models demonstrated limited discriminative power for moderate sleep quality (Level 2), with the attention-MLP model achieving only a 0.40 AUC for this category, indicating persistent challenges in identifying critical/transitional sleep states under the current category distribution. This limitation may stem from the relatively small overall sample size (n = 379) and insufficient representativeness/heterogeneity of the moderate group, which contributed to increased metric instability across segmentation approaches and heightened overfitting risks under stratified partitioning and dropout regularization. Additionally, while the dataset included physiological, psychological, and sociodemographic variables, it lacked key behavioral and circadian rhythm metrics closely associated with sleep quality, such as dietary patterns, daily training load rhythms, and circadian characteristics (e.g., chronotypes)25. Behaviors including caffeine intake, screen use before sleep, evening social activities, and irregular sleep-wake cycles may exhibit nonlinear interactions with stress and fatigue; omission of these factors could reduce sensitivity to subtle differences between poor, moderate, and good sleep states27. Sensitivity analyses further suggested that balanced perception training (smote and category-weighted) might enhance performance for minority category models. However, these gains require cautious interpretation and should be validated in larger multicenter cohorts and independent external samples, ideally using repeated/nested cross-validation or other resampling-based uncertainty assessment methods. It is crucial to note that attention weights and Shap values capture complementary features of model explanations rather than causal effects, and their stability may be influenced by sample size and category imbalance. Therefore, when interpreting the model, priority should be given to features consistently emphasized by both methods, while method-specific signals should be regarded as exploratory findings until further validation28. These limitations must be fully considered when evaluating the model’s screening utility.

Conclusion

This study developed a multimodal sleep-quality prediction model and compared logistic regression, random forest, XGBoost, and Attention-MLP. The Attention-MLP achieved moderate improvements in accuracy and macro-F1 relative to baseline models under the current sample size, while the gain in macro-AUC was limited. Notably, discrimination for the moderate sleep-quality class was poor (Class 2 AUC = 0.40), highlighting the challenge of identifying borderline or transition sleep states in a three-class setting. Therefore, the proposed framework should be interpreted as a screening-oriented risk stratification tool rather than a definitive predictor. The feature-level attention mechanism, together with explainability analysis, may facilitate screening-oriented risk stratification by clarifying the relative contribution of key predictors. Importantly, although BMI and anxiety were assigned higher feature importance, these findings do not imply causal relationships. Future work should incorporate behavioral and circadian rhythm–related variables and validate the model in external cohorts to improve generalizability.

Supplementary Information

Below is the link to the electronic supplementary material.

Supplementary Material 1 (190.5KB, doc)

Author contributions

L.L.: Writing-original draft, formal analysis and writing-review and editing; J.M.: Methodology and investigation.

Data availability

The datasets used and/or analysed during the current study available from the corresponding author on reasonable request. To ensure the reproducibility of the study, the core code (data preprocessing, model construction and training process) of this study can be obtained from the corresponding author. The relevant code contains the implementation and operation instructions of the main model to ensure that the research results can be verified and reproduced.

Declarations

Competing interests

The authors declare no competing interests.

Ethics declaration

This study was approved by the Ethics Committee of Dalian Maritime University (approval no. DMU-202518). We certify that the study was performed in accordance with the 1964 Declaration of Helsinki and later amendments. This study was obtained informed consent from all the participants.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Nédélec, M., Halson, S., Abaidia, A. E., Ahmaidi, S. & Dupont, G. Stress, sleep and recovery in elite soccer: A critical review of the literature. Sports Med.45, 1387–1400. 10.1007/s40279-015-0358-z (2015). [DOI] [PubMed] [Google Scholar]
  • 2.Baranwal, N., Phoebe, K. Y. & Siegel, N. S. Sleep physiology, pathophysiology, and sleep hygiene. Prog. Cardiovasc. Dis.77, 59–69. 10.1016/j.pcad.2023.03.006 (2023). [DOI] [PubMed] [Google Scholar]
  • 3.Dobrosielski, D. A., Sweeney, L. & Lisman, P. J. The association between poor sleep and the incidence of sport and physical training-related injuries in adult athletic populations: A systematic review. Sports Med.51, 777–793. 10.1007/s40279-020-01376-1 (2021). [DOI] [PubMed] [Google Scholar]
  • 4.Kullik, L., Isenmann, E., Schalla, J. & Kellmann, M. The impact of menstrual cycle phase and symptoms on sleep, recovery, and stress in elite female basketball athletes: A longitudinal study. Front. Physiol.12, 1663657. 10.3389/fphys.2025.1663657 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Dridi, N., Souissi, M. A., Dridi, R., Ceylan, H. İ. & Bragazzi, N. Evening smartphone exposure impairs sleep quality and next-day performance in elite soccer players: A randomized controlled trial. Biol. Sport. 43 (2), 145–154. 10.5114/biolsport.2026.56337 (2026). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Chen, Y., Iwao, K. & Shimamoto, H. The relationship between subjective sleep, mental health, and life skills acquisition among university student-athletes: A study on gender differences. Asian J. Sport Exerc. Psychol.10.1016/j.ajsep.2025.100014 (2025). [Google Scholar]
  • 7.Wannell, B. R., Brunner, F. M. J. & Lovibond, Z. N. & colleagues. The influence of partial sleep restriction on repeated sprint ability and reaction time in university athletes. Front. Sports Active Liv.7, 1519987. 10.3389/fspor.2025.1519987 (2025). [DOI] [PMC free article] [PubMed]
  • 8.Han, J., Yu, Z. & Yang, J. Multimodal attention-based deep learning for automatic modulation classification. Front. Energy Res.10, 1041862. 10.3389/fenrg.2022.1041862 (2022). [Google Scholar]
  • 9.Mao, Z., Yu, L. & Ye, L. Research on theories and practices of occupational physical fitness training. Res. Sports Sci.38 (1), 1–10. 10.15877/j.cnki.nsic.20240409.005 (2024). [Google Scholar]
  • 10.Diao, Y. & Zhang, L. Correlation analysis of influencing factors of depression and anxiety in patients undergoing coronary intervention surgery. China Med. Pharm.14 (13), 160–163. 10.20116/j.issn2095-0616.2024.13.38 (2024). [Google Scholar]
  • 11.Graves, B. S., Hall, M. E., Dias-Karch, C., Haischer, M. H. & Apter, C. Gender differences in perceived stress and coping among college students. PLoS One. 16 (8), e0255634. 10.1371/journal.pone.0255634 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Jia, W. & Chen, Y. Social appearance pressure and body image: Chain mediation effect of self-awareness and perceived stress. Chin. J. Clin. Psychol.32 (4), 744–749. 10.16128/j.cnki.1005-3611.2024.04.005 (2024). [Google Scholar]
  • 13.Li, Q., Liu, M., Chen, X. Y. & Yao, J. N. Rumination and sleep quality in college students: The role of negative emotions and sleep procrastination. Chin. J. Clin. Psychol.32 (1), 203–206. 10.16128/j.cnki.1005-3611.2024.01.037 (2024). [Google Scholar]
  • 14.Ali, A. et al. An innovative IoT and edge intelligence framework for monitoring elderly people using anomaly detection on data from non-wearable sensors. Sensors25 (6), 1735. 10.3390/s25061735 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Cvetkov-Iliev, A., Allauzen, A. & Varoquaux, G. Analytics on non-normalized data sources: More learning, rather than more cleaning. IEEE Access.10, 42420–42431. 10.1109/ACCESS.2022.3169500 (2022). [Google Scholar]
  • 16.Lieskovská, E., Jakubec, M., Jarina, R. & Chmulík, M. A review on speech emotion recognition using deep learning and attention mechanism. Electronics10 (10), 1163. 10.3390/electronics10101163 (2021). [Google Scholar]
  • 17.Tirumanadham, N. S. K. M. K. & Thaiyalnayaki, S. Enhancing student performance prediction in e-learning environments: Advanced ensemble techniques and robust feature selection. Int. J. Mod. Educ. Comput. Sci.17 (2), 67–86. 10.5815/ijmecs.2025.02.03 (2025). [Google Scholar]
  • 18.Ang, K. M., Seow, E. K., Fam, P. S. & Cheng, L. H. Classification of edible bird’s nest samples using a logistic regression model through the mineral ratio approach. Food Control. 137, 108921. 10.1016/j.foodcont.2022.108921 (2022). [Google Scholar]
  • 19.Xu, W., Tu, J., Xu, N. & Liu, Z. Predicting daily heating energy consumption in residential buildings through integration of random forest model and meta-heuristic algorithms. Energy301, 131726. 10.1016/j.energy.2023.131726 (2024). [Google Scholar]
  • 20.Chang, Y. C., Chang, K. H. & Wu, G. J. Application of extreme gradient boosting trees in the construction of credit risk assessment models for financial institutions. Appl. Soft Comput.73, 914–920. 10.1016/j.asoc.2018.09.029 (2018). [Google Scholar]
  • 21.Sauer, J., Mariani, V. C., dos Santos Coelho, L., Ribeiro, M. H. D. M. & Rampazzo, M. Extreme gradient boosting model based on improved Jaya optimizer applied to forecasting energy consumption in residential buildings. Evol. Syst. 1–12. 10.1007/s12530-022-09430-7 (2022).
  • 22.Brauwers, G. & Frasincar, F. A general survey on attention mechanisms in deep learning. IEEE Trans. Knowl. Data Eng.35 (4), 3279–3298. 10.1109/TKDE.2021.3059197 (2021). [Google Scholar]
  • 23.Mostafaei, S. H., Tanha, J. & Sharafkhaneh, A. A novel deep learning model based on transformer and cross modality attention for classification of sleep stages. J. Biomed. Inform.157, 104689. 10.1016/j.jbi.2024.104689 (2024). [DOI] [PubMed] [Google Scholar]
  • 24.Li, Y. et al. An interpretable deep learning model for sleep quality prediction using multimodal data. IEEE J. Biomed. Health Inf.26 (2), 653–662 (2022). [Google Scholar]
  • 25.Zhou, L., Zhang, Z. & Liu, J. Circadian rhythm and sleep: A deep learning approach to integrate temporal and behavioral features. Sleep. Health. 9 (1), 50–59 (2023). [Google Scholar]
  • 26.Kim, M. & Lee, H. Predicting sleep disorders using deep learning with psychological and lifestyle features. J. Affect. Disord.295, 245–252 (2021). [Google Scholar]
  • 27.Sun, Y., Zhang, R. & Liu, H. Impact of lifestyle behaviors on sleep quality: A machine learning perspective. Comput. Biol. Med.137, 104799 (2021).34478922 [Google Scholar]
  • 28.Xiao, Y. & He, J. Feature selection and attention-based MLP for health behavior prediction. Appl. Soft Comput.93, 106351 (2020). [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1 (190.5KB, doc)

Data Availability Statement

The datasets used and/or analysed during the current study available from the corresponding author on reasonable request. To ensure the reproducibility of the study, the core code (data preprocessing, model construction and training process) of this study can be obtained from the corresponding author. The relevant code contains the implementation and operation instructions of the main model to ensure that the research results can be verified and reproduced.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES