Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 Jan 22;16:6015. doi: 10.1038/s41598-026-36561-8

Machine learning for individual epigenetic fingerprints as predictors of well-being in young adults

Andrea Caporali 1,2,✉, Alberto Di Domenico 3, Claudio D’Addario 4,5,✉, Francesco de Pasquale 1,6
PMCID: PMC12901966  PMID: 41565854

Abstract

The crisis in youth mental health has intensified, especially after the COVID-19 pandemic. Traditional assessment tools like the Perceived Stress Scale and Highly Sensitive Person (HSP) index provide valuable insights. However, to address the multifaceted nature of mental issues, molecular biomarkers should be integrated with neuropsychological data when modeling these scales, in order to unravel the interplay of genetic, environmental, and psychological factors. This study explores the interaction of these factors using machine learning to model HSP scores in university students. By conducting exhaustive feature selection, a data-driven classification model is trained to provide individual multivariate fingerprints. Despite the limited sample size, the model achieves remarkable accuracy, sensitivity, and precision. The integration of epigenetic features seems crucial, indicating the importance of balancing neuropsychological and genetic influences for accurate modeling. Our findings pave the way for future clinical applications, since the collection of questionnaires and saliva samples might offer accessible avenues for mental health assessment and personalized healthcare.

Supplementary Information

The online version contains supplementary material available at 10.1038/s41598-026-36561-8.

Keywords: Mental health, Machine learning, Epigenetics, Highly sensitive persons, Multivariate fingerprint

Subject terms: Predictive markers, DNA methylation

Introduction

Nowadays, tackling the crisis in youth mental health has become paramount. Exacerbated by the pandemic’s disruptive impact1, the stressors faced by the youths, from academic pressures to overwhelming workplace competition, alter mental well-being2. Stress levels are typically assessed by the Perceived Stress Scale (PSS)3 and the Highly Sensitive Person (HSP) index4. This is a neuropsychological scale evaluating environmental sensitivity, i.e. the individual susceptibility in responding to external stimuli: a highly sensitive person can be easily overwhelmed in a stressful environment and is at higher risk of developing mental issues. Namely, 50% of individuals seeking counselling belong to the 20% of the general population identified as highly sensitive persons5. Assessing the risk of mental issues, e.g. using PSS or HSP, relies on self-reporting or surveys that are economical and easy to access. However, to capture the complex, multifaceted nature of stress, it is important to model these scales by incorporating molecular biomarkers that reveal the interplay of genetic, environmental, and personal factors. This might allow clinicians to identify specific indicators in late adolescents and young adults, and thus to adopt personalized treatments.

Recently, it has been shown that psychological traits such as stress reactivity and environmental sensitivity emerge from Gene × Environment interactions6. Specifically, genetics and epigenetics of dopamine (DAT1) and serotonin (SERT) transporter genes play a major role in personality formation in early young adults7–9. SERT polymorphisms affect amygdala-prefrontal connectivity, influencing emotional regulation and stress reactivity10, while gene–environment interactions involving these genes modulate sensitivity to environmental stressors during this critical developmental period11. Furthermore, the differential susceptibility hypothesis suggests that individuals carrying specific variants of these genes show heightened environmental sensitivity, being more vulnerable to negative environments but also more responsive to positive interventions12. Complementing the genetic evidence, epigenetic regulation has also been documented: increased methylation at DAT 5’UTR has been observed in highly sensitive students reporting high levels of perceived stress as compared to those less sensitive ones9. In this multidimensional framework, Machine Learning (ML) is a valuable tool. ML has been applied across a wide array of disciplines, encompassing epigenetic and psychological research, but studies modeling HSP remain scarce. However, pioneering work by Sadeghzadeh et al. focused on demographic and psychological features and showed a promising ML accuracy13.

Here, for the first time to our knowledge, we combined neuropsychological and epigenetic data using ML to predict HSP scores in university students following the analysis pipeline shown in Fig. 1. We performed an exhaustive feature selection for the classification of HSP. We aimed at: (I) identifying the optimal set of epigenetic and psychological features to explain the HSP trait; (II) training a multivariate, data-driven classification model for HSP; and (III) providing individualized multivariate fingerprints to assess personal mental health status.

Fig. 1.

Fig. 1

Scheme of the developed analysis pipeline.

Materials and methods

Participants

A total of n = 104 volunteers (M = 19, F = 85; Mean age: 20.04 ± 1.77 years) were recruited from various university courses from a single Italian University, namely University of Chieti-Pescara. We note that the gender imbalance in the acquired sample reflects the natural composition of the university courses (primarily psychology) from which participants were recruited. Race and ethnicity data were not collected, consistent with the homogeneous demographic composition of the recruitment context. Subjects with no history of psychiatric disorders, medication usage for psychiatric conditions, or the use of drugs/psychotropic substances, underwent a screening process. They provided informed consent and did not receive any compensation for their involvement. The study was conducted in accordance with the Declaration of Helsinki and approved by the Institutional Review Board of Psychology of the Department of Psychological, Health, and Territorial Sciences at G. D’Annunzio, University of Chieti-Pescara (identification code: 20026; date of approval: 19 February 2021).

Behavioral variables

The HSP index consisted of 12 self-report items assessed on a 7-point Likert scale (from 1 = strongly disagree to 7 = strongly agree)14. This 12-item brief version of the HSP scale developed by Pluess et al.15, was derived from the original 27-item HSP Scale4. The brief version demonstrated strong psychometric properties (Cronbach’s α = 0.77–0.82) and high correlation with the full scale (r > 0.90)15, making it suitable for research contexts where participant burden must be minimized while maintaining measurement validity. HSP total score was obtained by taking the average across all 12 items. Subjects were divided into two classes for supervised classification: the “High” class, subjects with HSP > 4.5, and the “Low/Medium” class, HSP < = 4.5. This 4.5 cutoff threshold for distinguishing high from low/medium sensitivity follows established conventions in HSP research4,14,16, identifying individuals who are particularly sensitive to environmental stimuli. Unfortunately, our sample size did not allow us to consider the three separate classes defined by14: “High” (HSP > 4.5), “Medium” (4 < HSP < 4.5) and “Low” (HSP < 4). In fact, we had only 9 (8.7%) “Low” individuals, 22 “Medium” (21.2%) and 73 “High” (70.2%). As extensively discussed in ML literature17–19, models trained on very small classes suffer from high variance in error estimation and unstable decision boundaries. This issue is even more pronounced in multivariate non-linear classifiers, which typically require larger and more balanced class sizes to train reliably, and become increasingly sensitive to small-n effects in multiclass settings. Although no universal sample-size thresholds exist, it is generally recommended to avoid classes with fewer than a few dozen observations due to unreliable performance estimates17. We therefore merged the Low and Medium groups to obtain a minimum class size of 30, ensuring a more stable binary classification. “Low/Medium” class comprised 31 subjects (average HSP = 4.04 ± 0.41); “High” class comprised 73 subjects (average HSP = 5.37 ± 0.43). Given the gender imbalance in our sample (85 females, 19 males), we first tested whether HSP class membership differed by gender. A Fisher exact test revealed a significant association (p = 0.025): the male subgroup was balanced across classes (10 Low/Medium, 9 High), while females are more represented in the High-HSP group, a pattern consistent with previous literature on environmental sensitivity and stress reactivity4,15,20,21.

As far as it concerns the behavior, we collected nine behavioral predictive features. The Italian 10-item version of the PSS test assessed on a five-point Likert scale was adopted22. The 20-item 5-point Likert scale (0 = not applicable, 5 = always) Internet Addiction Test was considered to take into account a possible dependence of HSP to internet addiction23. Also, the EAT-26, i.e. a psychological self-assessment test to measure symptoms and concerns characteristic of eating disorders, was considered24. Finally, we used the six first-order variables (i.e. Attention, Impulsivity, Self-control, Cognitive complexity, Perseverance, and Cognitive instability) from Barratt Impulsiveness Scale (BIS)25, i.e. a 30-item self-report questionnaire measured on a 4-point scale.

Molecular variables

To obtain the molecular variables we collected saliva samples by spitting approximately 2 ml of saliva into a 15 ml centrifuge tube. To ensure accurate sampling, participants were instructed not to consume food, drugs, and beverages (except water), or use lip products. They were also advised not to smoke or brush their teeth for at least two hours before collection to prevent potential contamination. DNA methylation at DAT1 and SERT gene promoters were analyzed as reported in9. OXTR DNA methylation levels were assessed using the same biological material9, processed as previously described in26. Six CpG sites were considered for DNA methylation analysis for both DAT1 and SERT genes; four CpG sites were considered for the OXTR gene, for a total of 16 sites. To reduce computational costs during the feature selection process, we selected 10 out of 16 positions, chosen according to the highest coefficient of variation (standard deviation relative to the mean) across subjects. Different CpG sites are denoted as “POS. x” in Fig. 2. Three out of ten sites belong to the OXTR gene, three to the DAT1 gene and four to the SERT gene.

Fig. 2.

Fig. 2

Initial set of considered variables (left) and resulting optimal subset for HSP classification model (right).

Feature selection procedure and machine learning algorithm

In this study, we used an iterative Machine Learning approach based on Support Vector Machine (SVM)27. This choice was supported by the current literature (see SI for a discussion and comparison with alternative ML methods). Here, we adopted a polynomial kernel (degree = 2), since this outperformed other kernels such as Gaussian and linear, see the control analyses reported in SI (.S1). The data were standardized, and the kernel scale parameter was determined using MATLAB’s heuristic procedure. To identify the optimal set of features, we conducted an exhaustive feature selection within the Leave-One-Out Cross-Validation (LOOCV) framework, see Fig. 1. This selection evaluates all possible combinations of features by iterating through subsets of increasing size, ensuring that the model is trained and tested with every potential configuration of features, ranging from 1 to 19 variables (9 neuropsychological and 10 epigenetic). Although computationally demanding, this method proved more reliable for feature selection given the complex, non-linear interactions among psychological, neurobiological, and epigenetic features. See the SI for methodological details and comparisons with more common feature selection approaches. To avoid data leakage, the selection process was based solely on the training data in each fold. At each iteration, we computed five metrics: (I) accuracy, i.e. the ratio of correctly predicted instances to the total number of instances; (II) sensitivity, i.e. the ratio of true positive predictions to the total number of actual positive instances; (III) precision, i.e. the ratio of true positive predictions to the total number of positive predictions; (IV) specificity, i.e. the ratio of true negative predictions to the total number of actual negative instances; (V) negative predictive value (NPV), i.e. the ratio of true negative predictions to the total number of predicted negative instances. These measures allowed us to assess the performance through the F1 Score, defined as the harmonic mean of precision and sensitivity for the “High” class or the harmonic mean of NPV and specificity for the “Low/Medium” class. The optimal model corresponded to the highest F1 score on the “Low/Medium” class across all evaluated subsets. We chose the F1-score, rather than accuracy, due to the 70:30 imbalance between classes and the fact that LOOCV cannot be stratified. Thus, each fold naturally reproduced the empirical class distribution. To avoid accuracy inflation under imbalance, the F1-score was used as the only metric guiding model selection, as it provides a more balanced assessment of classifier performance and is recommended for imbalanced classification tasks28.

To evaluate whether the classifier reached a performance plateau, we examined the model’s learning curves, both on training and test set, on progressively larger subsamples of the dataset using LOOCV accuracy29. To this aim, we fitted the test learning curve using a nonlinear saturating function. Consistently with similar approaches reported in the literature30,31, we adopted a saturating exponential model, i.e. Inline graphic, with a, b, and c parameters to be estimated and x the test sample. The asymptote a + b represents the theoretical upper bound of model accuracy, as in29. Following32, we defined stability as the smallest sample size N for which the predicted accuracy of the fitted learning curve was within ± 0.02 of its asymptotic value. This provides a model-based estimate of the sample size at which further increases would yield negligible improvements in performance.

Finally, to investigate the contribution of each predictor to the classification outcome, we computed Shapley Additive Explanations (SHAP) values on the final SVM model33,34. SHAP provides a game-theoretic measure of feature importance by quantifying, for each subject, how much each variable shifts the model’s prediction as compared to the average prediction35. This approach allows a direct interpretation of how psychological and epigenetic variables jointly influence the predicted HSP class and complements the multivariate feature-selection procedure. All analyses were conducted using home-built codes developed in MATLAB (2022, Natick, Massachusetts: The MathWorks Inc.), leveraging parallel computing capabilities. The training and evaluation of all (i.e. 524.287) models required 8.2 h using 18 processors Intel® Core™ i9-10980XE and 256 GB DDR4-2933 SDRAM.

Results

Based on a sample of 104 university students, we collected 20 features encompassing epigenetic and behavioral scores. We labelled our subjects in two classes: “Low/Medium” (HSP ≤ 4.5) and “High” (otherwise). To identify the optimal variables, we adopted an exhaustive feature selection method, following the analysis pipeline represented in Fig. 1: a total of 524.287 classification models were trained and evaluated to provide individual fingerprints.

The initial set of considered variables is reported in Fig. 2 (left). Among them, in the final optimal model, we obtained: “PSS”, “Attention”, “DAT POS. 5”, “DAT POS. 7”, “SERT POS. 3”, and “SERT POS. 6”. Notably, although the procedure was completely data-driven and imposed no prior constraint on the presence of epigenetic data, the method converged towards a combined model encompassing a parsimonious number of epigenetic and psychological variables. Of note, models with a higher number of variables exhibited a lower performance (see Fig. S1). In Fig. 3 we report a radial plot of the six features of the final fingerprint obtained for the two classes. Specifically, we show the Low/Medium (blue) and the High (yellow) class fingerprints computed as the average of the individual fingerprints within each class. For visualization purposes, each feature was normalized by its maximum value across all subjects. We note that these profiles are clearly distinguishable, except for a common space of features shared by the two classes (see Fig. 3 - green area). To assess the quality of our approach, in Fig. 4 we provide the confusion matrix corresponding to our binary classification. As it can be noted, the optimal model achieved an accuracy of 84.6%, a sensitivity of 90.4%, and a precision of 88.0%. Accordingly, this resulted in an F1-score of 89.2% for the “High” class and 73.3% for the “Low/Medium” class. Despite the limited number of subjects considered in this study, the obtained accuracy seems very promising.

Fig. 3.

Fig. 3

Average fingerprints for the “High” (yellow) and “Low/Medium” (blue). The overlap of the two fingerprints is reported in green.

Fig. 4.

Fig. 4

Diagnostics of the proposed approach. The confusion matrix for the optimal model is reported on the left-hand side. Green squares represent correctly classified instances; red ones represent the incorrectly classified. Performance metrics (i.e. accuracy, sensitivity, specificity, precision, and negative predictive value - NPV) are shown.

To assess potential gender bias in model performance, we also computed classification metrics separately for males and females. The classifier achieved comparable accuracy and F1-scores in both groups (Females: Accuracy = 83.5% and average F1-score = 77.1%; Males: Accuracy = 89.5% and average F1-score = 89.4%), indicating stable generalizability despite the imbalance. The lower average F1-score observed in females is explained by the greater imbalance between High and Low/Medium HSP cases within this subgroup, which reduces specificity and consequently lowers the F1-score for the minority class.

To improve model interpretability, we performed a SHAP analysis33, which provides both global and local explanations of feature contributions to the predicted HSP class. We summarized global feature importance using the mean absolute SHAP values. We report the distribution of contributions through a beeswarm plot (see Fig. 5A and B respectively). “DAT POS. 5” emerged as the most influential predictor (mean |SHAP| = 2.13), indicating that variation at this site has the strongest influence on the model’s decision boundary. Psychological variables also contributed substantially: “PSS” (mean |SHAP| = 1.19) and “Attention” (mean |SHAP| = 1.12) ranked as the second and third most important predictors, respectively. Several epigenetic markers also showed non-negligible contributions, including “SERT POS. 3” (|SHAP| = 0.99), “DAT POS. 7” (|SHAP| = 0.79), and “SERT POS. 6” (|SHAP| = 0.63). The beeswarm plot further illustrates the directionality and heterogeneity of feature effects. Higher PSS scores tended to increase the probability of classification into the High-HSP group. A similar, though weaker, trend was observed for “DAT POS. 5”. Other methylation sites showed both positive and negative effects depending on individual values, consistent with the presence of non-linear relationships. Importantly, the spread of SHAP values for each feature indicates that the model did not rely on a single dominant variable but instead integrated information across psychological and epigenetic domains.

Fig. 5.

Fig. 5

SHAP analysis of the classification model. (A) Global feature importance ranked by mean absolute SHAP values, reflecting the average magnitude of each predictor’s contribution to the classification of High vs. Low/Medium HSP individuals. (B) Beeswarm plot showing the distribution of SHAP values across all participants for each feature; each point represents an individual sample, colored according to the corresponding predictor value. Higher SHAP values indicate stronger positive contributions to classification into the “High” HSP class.

Another important aspect that we addressed concerns the model’s learning curve. In Fig. 6, we report the training and test learning curve, along with the exponential model fitted to the test scores. The curve shows a “healthy” learning behavior: as the number of samples increases, the model’s performance on the test data improves substantially. The test curve shows higher variability at the early stages (fewer samples) but progressively stabilizes as more data points are included, reflecting increasingly reliable performance estimates. Additionally, the curve shows a positive slope, indicating that each additional sample contributes to improved generalization to new data. The fitted exponential model (parameter estimates: a = 0.13 ± 0.04, b = 0.81 ± 0.06, c = 0.018 ± 0.004; 95% CI) implies an asymptotic accuracy of approximately 0.94 ± 0.07 (see dashed line in Fig. 6). We defined model stability as the smallest N for which the predicted accuracy lies within ± 2% of the asymptotic value32. According to this criterion, the learning curve begins to stabilize near N ≈ 198, suggesting that additional samples beyond this point would yield only marginal improvements in generalization performance.

Fig. 6.

Fig. 6

Training (blue) and test (black) learning curves of the SVM model, representing model accuracy as a function of the number of training samples. The red line represents a saturating exponential model fitted to the test learning curve. The dashed red line represents the model’s asymptote.

Discussion

In this study, for the first time to our knowledge, a machine learning approach was adopted to model HSP scores and classify university students based on the potential risk of developing a mental issue. In a completely data-driven way, we identified a parsimonious combination of epigenetic and behavioral features as optimal model, namely “PSS”, “Attention”, “DAT POS. 5”, “DAT POS. 7”, “SERT POS. 3”, and “SERT POS. 6”. This is in line with previous studies, providing a deeper insight into the biological significance of these variables. In fact, while the role of genetics on HSP is well documented in the literature36,37, here we observed an important impact of epigenetic features combined with neuropsychological ones. Among them, the presence of “PSS” and “Attention” - identified from the SHAP analysis as the second and third strongest contributors, respectively – aligns well with the existing literature. PSS is widely reported as a reliable measure of perceived stress, see for example3. Individuals with higher PSS often exhibit dysregulated stress response systems, leading to an increased risk for maladaptive outcomes38. Thus, PSS positively correlates with vulnerability to mental health conditions, e.g. anxiety and depression, see2,39. “Attention”, based on the Barratt Impulsiveness Scale, is another critical variable, representing an individual’s cognitive capacity for focus and self-regulation. This measure is closely tied to stress reactivity and mental health disorders, including ADHD and mood dysregulation40. Impairments in attention can exacerbate the impact of stress by reducing an individual’s ability to process and adapt to challenging stimuli. These deficits have been associated with alterations in dopaminergic regulation41. In this context, the epigenetic markers identified in the model (“DAT POS. 5” and “DAT POS. 7”) reflect the key role of the dopamine transporter gene in modulating stress sensitivity. Notably, “DAT POS. 5” exerts the largest influence on the decision boundary. This observation is consistent with the work of Bellia et al.9, who reported that perceived stress and DAT1 methylation at CpG site number 5 are strongly associated with heightened environmental sensitivity in young adults. Moreover, increased methylation at CpG sites within the DAT1 gene, particularly at position 5, has been linked to reduced gene expression and dopaminergic activity, which may lead to increased vulnerability to stress42–44. In fact, DAT1 methylation changes have been observed in response to environmental stressors, suggesting an interaction between genetic predispositions and external influences in shaping stress responses45,46. Also, the presence of serotonin transporter variables in the optimal model, i.e. “SERT POS. 3” and “SERT POS. 6”, can be linked to previous work. In fact, alterations in SERT regulation have been previously implicated in emotional reactivity and sensitivity, particularly in the context of stress and anxiety47. Their presence in our optimal model can be linked to the reported interplay between serotonergic and dopaminergic systems in influencing mental health outcomes48.

Of note, our approach, despite the limited sample size, reached a remarkably high accuracy (close to 90%), sensitivity, and precision. These values align with the performance of typical mental screening tests49, positioning our model as a potential tool for clinical applications. Our metrics are in line with a previous study, where HSP was modelled exclusively with psychological features13. Here, a deep learning model, based on a sample larger than ours (n = 190 vs. n = 104), reached an accuracy of 83.3% (as compared to 84.6%, in our case). This suggests that the impact of the epigenetic features in modeling HSP is crucial and allows us to reach high accuracy even with limited samples. In our case, the sample size does not allow us to adopt a deep learning approach, as in13. This represents a future development of our work.

Regarding the analysis pipeline, our optimal model was obtained after testing exhaustively every possible combination, of any size, of the initial set of genetic and behavioral variables. This step was computationally expensive since we tested more than 5 × 105 models. Interestingly, despite no penalization terms on the model size being included in the cost function, our algorithm converged to a parsimonious six-variable model. This significant reduction in the number of features and thus the model’s complexity enhances its interpretability.

Furthermore, the learning curves depicted in Fig. 6 suggest that the model is robust. Test performance increases steadily with sample size and approaches the asymptotic accuracy predicted by the exponential model (≈ 0.94), indicating that the model continues to benefit from additional data. Importantly, the curve has not yet reached its theoretical plateau: the predicted accuracy remains below the asymptote across the available range, and model stability is achieved only at an estimated sample size of approximately 200 subjects, a number consistent with previous studies17,29. This suggests that performance has not saturated and that additional observations would likely yield further improvements in generalization. With the actual data, the observed gap between the training and test scores is moderate (around 0.10, measured as the average distance for the last 10 points) suggesting a good model generalization ability. In fact, the gradual convergence of training and test accuracy, as the sample size grows, implies a reduction in overfitting. Of note, the proposed procedure, apart from the definition of HSP labels (obtained from standardized values14), was completely data-driven, i.e. no prior knowledge constrained the feature selection. Moreover, overfitting was limited by the parsimony of the optimal model: larger feature subsets systematically yielded lower F1-scores and accuracy (see Fig. S1), further indicating that models with more predictors were more prone to capturing noise rather than meaningful structure. This is consistent with classical bias–variance considerations: in small datasets, adding parameters typically increases variance and reduces generalization18. The convergence toward a compact six-feature model therefore reflects not only empirical optimality but also better control of variance.

In terms of possible clinical applications, we are aware that the proposed model needs to be confirmed on a larger sample. However, differently from other approaches, our approach is not based on the subjects’ compliance in performing a task, e.g. a visual task as in13, but relies on questionnaires and saliva collection. Therefore, analyzing epigenetic data from saliva is remarkably accessible and cost-effective. This opens new avenues to comprehend how our body interacts with the environment and can offer insights before any noticeable symptoms emerge. In fact, in future applications, subjective fingerprints provided by our approach could represent the basis of Artificial Intelligence applications where individual mental-health profiles could be exploited to design personalized treatments.

Limitations and methodological considerations

Several limitations should be acknowledged. First, our sample was not balanced for gender (82% females). Although our results show no difference in the classification accuracies in males and females, gender differences in stress sensitivity and epigenetic regulation have been reported in the literature50,51 and thus in future studies a more balanced sample will be acquired. Further, our sample was drawn from a single Italian university, and this limits the demographic diversity. Regarding the HSP scale, due to the size of our sample, our binary classification may oversimplify the continuous nature of sensitivity. In future studies larger samples might address this aspect. A further limitation concerns external validation, which was not feasible because no independent datasets with comparable psychological and salivary epigenetic measures are currently available. Although LOOCV provides a good internal estimate of performance, it cannot replace true external validation, which remains the gold standard for evaluating reproducibility across populations and data-collection settings. Future studies should therefore address these issues by collecting larger, gender-balanced, demographically diverse, and multi-center samples, enabling external validation of the model, the development of a multi-class or regression-based approaches that capture the full spectrum of HSP variability, and longitudinal analyses to assess causal inferences.

Supplementary Information

Below is the link to the electronic supplementary material.

Supplementary Information (472.1KB, docx)

Author contributions

Claudio D’Addario and Francesco de Pasquale designed research; Andrea Caporali, Claudio D’Addario, Francesco de Pasquale and Alberto Di Domenico performed research; Andrea Caporali and Claudio D’Addario analyzed data; Andrea Caporali, Claudio D’Addario and Francesco de Pasquale wrote the paper. Francesco de Pasquale and Claudio D’Addario contributed equally to this work.

Funding

The research was partially funded by the European Union- Next Generation EU, Mission 4, Component 2, CUP C53D23008470001, acronym EVERYONE for FdP and by the European Union - Next Generation EU - NRRP M6C2 - Investment 2.1 Enhancement and strengthening of biomedical research in the NHS (PROJECT PNRR-MAD-2022-12376693, acronym BOARDING PASS) for AC and CD.

Data availability

The data that support the findings of this study are not openly available due to reasons of sensitivity and are available from the corresponding author upon reasonable request. Data are located in controlled access data storage at University of Teramo.

Declarations

Competing interests

The authors declare no competing interest.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Contributor Information

Andrea Caporali, Email: andrea.caporali@unicam.it.

Claudio D’Addario, Email: cdaddario@unite.it.

References

  • 1.(OSG). O.o.t.S.G., In Protecting Youth Mental Health: The U.S. Surgeon General’s Advisory (Washington (DC), 2021).
  • 2.Cannito, L. et al. The role of stress and cognitive absorption in predicting social network addiction. Brain Sci., 12(5). (2022). [DOI] [PMC free article] [PubMed]
  • 3.Cohen, S., Kamarck, T. & Mermelstein, R. A global measure of perceived stress. J. Health Soc. Behav.24 (4), 385–396 (1983). [PubMed] [Google Scholar]
  • 4.Aron, E. N. & Aron, A. Sensory-processing sensitivity and its relation to introversion and emotionality. J. Pers. Soc. Psychol.73 (2), 345–368 (1997). [DOI] [PubMed] [Google Scholar]
  • 5.Aron, E. N. Psychotherapy and the Highly Sensitive Person: Improving Outcomes for that Minority of People who Are the Majority of Clients (Routledge, 2011).
  • 6.Plomin, R., DeFries, J. C. & Loehlin, J. C. Genotype-environment interaction and correlation in the analysis of human behavior. Psychol. Bull.84 (2), 309–322 (1977). [PubMed] [Google Scholar]
  • 7.Depue, R. A. & Collins, P. F. Neurobiology of the structure of personality: dopamine, facilitation of incentive motivation, and extraversion. Behav. Brain Sci.22 (3), 491–517 (1999). [DOI] [PubMed] [Google Scholar]
  • 8.Ward, R. et al. The role of serotonin in personality inference: Tryptophan depletion impairs the identification of neuroticism in the face. Psychopharmacol. (Berl). 234 (14), 2139–2147 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Bellia, F. et al. Dopamine and serotonin transporter genes regulation in highly sensitive individuals during stressful conditions: A focus on genetics and epigenetics. Biomedicines, 12(9). (2024). [DOI] [PMC free article] [PubMed]
  • 10.Pezawas, L. et al. 5-HTTLPR polymorphism impacts human cingulate-amygdala interactions: a genetic susceptibility mechanism for depression. Nat. Neurosci.8 (6), 828–834 (2005). [DOI] [PubMed] [Google Scholar]
  • 11.Wichers, M. et al. The BDNF Val(66)Met x 5-HTTLPR x child adversity interaction and depressive symptoms: an attempt at replication. Am. J. Med. Genet. B Neuropsychiatr Genet.147B (1), 120–123 (2008). [DOI] [PubMed] [Google Scholar]
  • 12.Homberg, J. R. & Lesch, K. P. Looking on the bright side of serotonin transporter gene variation. Biol. Psychiatry. 69 (6), 513–519 (2011). [DOI] [PubMed] [Google Scholar]
  • 13.Sadeghzadeh, N. et al. SPS vision net: measuring sensory processing sensitivity via an artificial neural network. Cogn. Comput.16 (3), 1379–1392 (2024). [Google Scholar]
  • 14.Lionetti, F. et al. Dandelions, tulips and orchids: evidence for the existence of low-sensitive, medium-sensitive and high-sensitive individuals. Transl. Psychiatry. 8 (1), 24 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Pluess, M. et al. Environmental sensitivity in children: development of the highly sensitive child scale and identification of sensitivity groups. Dev. Psychol.54 (1), 51–70 (2018). [DOI] [PubMed] [Google Scholar]
  • 16.Acevedo, B. P. et al. The highly sensitive brain: an fMRI study of sensory processing sensitivity and response to others’ emotions. Brain Behav.4 (4), 580–594 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Beleites, C. et al. Sample size planning for classification models. Anal. Chim. Acta. 760, 25–33 (2013). [DOI] [PubMed] [Google Scholar]
  • 18.Hastie, T., Tibshirani, R. & Friedman, J. The Elements of Statistical Learning (Springer series in statistics, 2009).
  • 19.van der Ploeg, T., Austin, P. C. & Steyerberg, E. W. Modern modelling techniques are data hungry: a simulation study for predicting dichotomous endpoints. BMC Med. Res. Methodol.14, 137 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Licht, C. L. et al. Serotonin transporter gene (SLC6A4) variation and sensory processing sensitivity-Comparison with other anxiety-related temperamental dimensions. Mol. Genet. Genomic Med.8 (8), e1352 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Lionetti, F. et al. Sensory processing sensitivity and its association with personality traits and affect: A meta-analysis. J. Res. Pers.81, 138–152 (2019). [Google Scholar]
  • 22.Diotaiuti, P., Valente, G. & Mancone, S. Validation study of the Italian version of Temporal focus scale: psychometric properties and convergent validity. BMC Psychol.9 (1), 19 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Young, K. S. Cognitive behavior therapy with internet addicts: treatment outcomes and implications. Cyberpsychol Behav.10 (5), 671–679 (2007). [DOI] [PubMed] [Google Scholar]
  • 24.Garner, D. M. et al. The eating attitudes test: psychometric features and clinical correlates. Psychol. Med.12 (4), 871–878 (1982). [DOI] [PubMed] [Google Scholar]
  • 25.Patton, J. H., Stanford, M. S. & Barratt, E. S. Factor structure of the Barratt impulsiveness scale. J. Clin. Psychol.51 (6), 768–774 (1995). [DOI] [PubMed] [Google Scholar]
  • 26.D’Addario, C. et al. Regulation of Oxytocin receptor gene expression in obsessive-compulsive disorder: a possible role for the microbiota-host epigenetic axis. Clin. Epigenetics. 14 (1), 47 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Cortes, C. & Vapnik, V. Support-vector networks. Mach. Learn.20, 273–297 (1995). [Google Scholar]
  • 28.Erickson, B. J. & Kitamura, F. Magician’s Corner: 9. Performance Metrics for Machine Learning Models p. e200126 (Radiological Society of North America, 2021). [DOI] [PMC free article] [PubMed]
  • 29.Figueroa, R. L. et al. Predicting sample size required for classification performance. BMC Med. Inf. Decis. Mak.12, 8 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Ramsay, C. R. et al. Statistical assessment of the learning curves of health technologies. Health Technol. Assess.5 (12), 1–79 (2001). [DOI] [PubMed] [Google Scholar]
  • 31.Vianna, L. S., Gonçalves, A. L. & Souza, J. A. Analysis of learning curves in predictive modeling using exponential curve fitting with an asymptotic approach. Plos One. 19 (4), e0299811 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Silvey, S. & Liu, J. Sample size requirements for popular classification algorithms in tabular clinical data: empirical study. J. Med. Internet. Res.26, e60231 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Lundberg, S. M. & Lee, S. I. A unified approach to interpreting model predictions. Adv. Neural. Inf. Process. Syst., 30. (2017).
  • 34.Strumbelj, E. & Kononenko, I. An efficient explanation of individual classifications using game theory. J. Mach. Learn. Res.11, 1–18 (2010). [Google Scholar]
  • 35.Ponce-Bobadilla, A. V. et al. Practical guide to SHAP analysis: explaining supervised machine learning model predictions in drug development. Clin. Transl Sci.17 (11), e70056 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Assary, E. et al. Genetic architecture of environmental sensitivity reflects multiple heritable components: a twin study with adolescents. Mol. Psychiatry. 26 (9), 4896–4904 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Chen, C. et al. Contributions of dopamine-related genes and environmental factors to highly sensitive personality: a multi-step neuronal system-level approach. PLoS One. 6 (7), e21636 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Karyotaki, E. et al. Sources of stress and their associations with mental disorders among college students: results of the world health organization world mental health surveys international college student initiative. Front. Psychol.11, 1759 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Lee, E. H. Review of the psychometric evidence of the perceived stress scale. Asian Nurs. Res. (Korean Soc. Nurs. Sci). 6 (4), 121–127 (2012). [DOI] [PubMed] [Google Scholar]
  • 40.Swann, A. C. et al. Two models of impulsivity: relationship to personality traits and psychopathology. Biol. Psychiatry. 51 (12), 988–994 (2002). [DOI] [PubMed] [Google Scholar]
  • 41.Volkow, N. D. et al. Brain dopamine transporter levels in treatment and drug Naive adults with ADHD. Neuroimage34 (3), 1182–1190 (2007). [DOI] [PubMed] [Google Scholar]
  • 42.Chmielowiec, K. et al. DNA methylation of the dopamine transporter DAT1 Gene-Bliss seekers in the light of epigenetics. Int. J. Mol. Sci., 24(6). (2023). [DOI] [PMC free article] [PubMed]
  • 43.Robertson, K. D. DNA methylation and human disease. Nat. Rev. Genet.6 (8), 597–610 (2005). [DOI] [PubMed] [Google Scholar]
  • 44.De Nardi, L. et al. Involvement of DAT1 Gene on Internet Addiction: Cross-Correlations of Methylation Levels in 5’-UTR and 3’-UTR Genotypes, Interact with Impulsivity and Attachment-Driven Quality of Relationships. Int. J. Environ. Res. Public. Health, 17(21). (2020). [DOI] [PMC free article] [PubMed]
  • 45.Annunzi, E. et al. Mild internet use is associated with epigenetic alterations of key neurotransmission genes in salivary DNA of young university students. Sci. Rep.13 (1), 22192 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Catale, C. et al. Early life social stress causes Sex- and Region-Dependent dopaminergic changes that are prevented by Minocycline. Mol. Neurobiol.59 (6), 3913–3932 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Gotlib, I. H. et al. HPA axis reactivity: a mechanism underlying the associations among 5-HTTLPR, stress, and depression. Biol. Psychiatry. 63 (9), 847–851 (2008). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 48.Di Giovanni, G., Matteo, V. D. & Esposito, E. Serotonin-dopamine interaction: experimental evidence and therapeutic relevance. Preface. Prog Brain Res.172, ix (2008). [DOI] [PubMed] [Google Scholar]
  • 49.Richardson, R. et al. Screening for psychological and mental health difficulties in young people who offend: a systematic review and decision model. Health Technol. Assess.19 (1), 1–128 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50.Nugent, B. M. et al. Brain feminization requires active repression of masculinization via DNA methylation. Nat. Neurosci.18 (5), 690–697 (2015). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51.Heim, C. & Binder, E. B. Current research trends in early life stress and depression: review of human studies on sensitive periods, gene-environment interactions, and epigenetics. Exp. Neurol.233 (1), 102–111 (2012). [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Information (472.1KB, docx)

Data Availability Statement

The data that support the findings of this study are not openly available due to reasons of sensitivity and are available from the corresponding author upon reasonable request. Data are located in controlled access data storage at University of Teramo.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES