Skip to main content
BMC Public Health logoLink to BMC Public Health
. 2025 Jul 26;25:2557. doi: 10.1186/s12889-025-23946-z

Exploring stroke risk factors in different genders using Bayesian networks: a cross-sectional study involving a population of 134,382

Liqin Linghu 1,2, Yaxin Huang 2, Lixia Qiu 1,, Xuchun Wang 1, Jia Zhang 3, Lin Ma 2, chenglian Li 2, Lijie Wang 2
PMCID: PMC12297713  PMID: 40713611

Abstract

Background

The exploration of stroke risk factors provides crucial information for healthcare planning and priority setting. This study aims to utilize Bayesian network modeling to explore stroke risk factors in different genders.

Methods

We collected data from 10 cities and 13 counties in Shanxi Province, China, through questionnaire surveys, physical examinations, and laboratory tests. Logistic regression and Bayesian modeling were employed to analyze the risk factors for stroke in different genders. Preliminary analysis of stroke risk factors was conducted using chi-square tests and logistic regression models. Variables that showed statistical significance were included in the construction of the Bayesian model. Bayesian structure learning was achieved using the Max-Min Hill-Climbing algorithm, and parameter learning utilized maximum likelihood estimation.

Results

The study identified both common and gender-specific risk factors for stroke. Common risk factors for both males and females included region, marital status, education level, age, family history of stroke, secondhand smoke exposure, snoring, abnormal blood lipids, hypertension, diabetes, and coronary heart disease. Gender-specific factors were smoking and respiratory pause for males, and alcohol consumption for females. The Bayesian Network (BN) model further revealed structural relationships among these factors, showing that abnormal blood lipids, hypertension, and age were direct risk factors for stroke in males, with snoring, education level, and respiratory pause as indirect factors. For females, direct risk factors included age, hypertension, and secondhand smoke exposure, while snoring was an indirect factor.

Conclusions

Stroke risk factors vary by gender, highlighting the importance of gender-specific prevention and intervention strategies.

Supplementary Information

The online version contains supplementary material available at 10.1186/s12889-025-23946-z.

Keywords: Stroke, Bayesian network, Risk factors, Different genders

Background

Stroke is a major chronic non-communicable disease that poses a serious threat to health. Globally, among adults, stroke is a leading cause of death and disability [1].In the United States, stroke ranks as the fifth leading cause of death, with over 600,000 people experiencing a first stroke each year. It is also a major cause of long-term disability, resulting in an annual cost of 34 billion dollars in the United States [2]. In the European Union, stroke stands as the second most common cause of death, affecting approximately 1.1 million residents annually and causing 440,000 deaths. The estimated cost of stroke-related expenses in 2017 was 45 billion euros, encompassing direct and indirect medical service costs as well as productivity losses. As the population continues to grow and life expectancy increases, it is anticipated that the incidence of stroke events, their long-term consequences, and associated costs will significantly rise [3].

In China, stroke ranks as the third leading cause of death and a primary contributor to Disability-Adjusted Life Years (DALY) in adults. In 2019, there were 28.76 million stroke patients in China, with 3.934 million new cases, resulting in 2.19 million deaths [4, 5].In 2020, approximately 17.8 million Chinese adults experienced a stroke, with 3.4 million cases being first occurrences, and an additional 2.3 million resulting in death. Furthermore, approximately 12.5% of stroke survivors suffered from disabilities [6]. With the ongoing aging population and increasing prevalence of risk factors such as diabetes, hypertension, and hyperlipidemia, coupled with inadequate control measures, the burden of stroke in China continues to rise, posing a significant challenge to the healthcare system.

Due to biological factors such as genetics and hormones [7], and social-cultural factors such as smoking [8], alcohol consumption [9], as well as disease-specific influences like coronary heart disease [10], hypertension [5, 10, 11], diabetes [12], and lipid disorders [1], there are significant differences between men and women in the risk of certain chronic diseases. Research has shown that women have a higher incidence of stroke before the age of seventy, after which men surpass women; however, among the continuously growing elderly population, women again exhibit the highest incidence of stroke. Within the first five years after a stroke, women have a higher probability of recurrent strokes compared to men. These data suggest the presence of underlying biological variations between genders [7]. Furthermore, the early symptoms of stroke are not easily detectable, emphasizing the importance of conducting large-scale epidemiological surveys to identify risk factors and their interactions, regularly monitoring risk factors, and implementing primary and secondary prevention measures for stroke.

In the exploration of stroke risk factors, the commonly employed statistical method is logistic regression. Previous stroke studies [1315], utilized logistic regression to investigate risk factors. These studies indicated a significant association between stroke and factors like hypertension, hyperlipidemia, diabetes, coronary heart disease, and smoking.However, the logistic regression model has some drawbacks [16, 17]. The first issue involves independent variables, as the application of logistic regression is based on the linear assumption, meaning the relationship between independent and dependent variables is considered linear. In real-world scenarios, there might be complex nonlinear relationships among influencing factors that traditional logistic regression cannot capture. Second, when there is high correlation among independent variables, known as multicollinearity [18], the performance of traditional logistic regression may decline. Multicollinearity can lead to model instability, making parameter estimation inaccurate.

In 1987, Thomas Bayes proposed Bayesian theorem, which laid the theoretical foundation for uncertain probability inference. After experiencing a revival in the mid-20th century, and with the advancement of computer technology, Bayesian methods have found widespread applications in the fields of statistics, machine learning, and artificial intelligence. They have become powerful tools for handling uncertainty, modeling complex systems, and inferring unknown parameters [19], Bayesian networks offer several advantages over traditional logistic analysis [19, 20], Firstly, Capture Complex Dependencies: Bayesian networks can capture and quantify complex dependencies among variables. Each node represents a variable, and edges signify dependencies between variables. This makes Bayesian networks more suitable for describing complex systems in the real world, where nonlinear and non-independent relationships among variables exist.Secondly, Probabilistic Inference: Bayesian networks allow for probabilistic inference, meaning the calculation of probability distributions given evidence. This makes them suitable for handling uncertainty and providing more comprehensive information.Thirdly, Graphical Representation: Bayesian networks provide an intuitive graphical representation that clearly illustrates relationships between variables. This enhances the interpretability of the model.These advantages make Bayesian networks a valuable approach for addressing the challenges posed by complex and uncertain systems in various scientific domains.

Bayesian learning involves two stages: structural learning and parameter learning. Structural learning aims to estimate the Directed Acyclic Graph (DAG) of the Bayesian network based on the data. This process involves determining the dependencies between nodes, i.e., defining the structure of the graph. Structural learning includes constraint methods and scoring methods. Constraint methods, reliant on domain knowledge, use statistical tests to identify dependencies between nodes. Scoring methods evaluate the quality of different structures based on specific criteria and select the optimal structure. Constraint methods provide intuitive and interpretable structures but may be subject to subjective biases. Scoring algorithms, flexible and automated, are suitable for large-scale data but face challenges of high computational complexity and overfitting risks.Due to these limitations, hybrid algorithms, combining constraints and scoring methods, have been proposed. These hybrid approaches integrate prior knowledge, constraint tests, and score searches to navigate large search spaces and identify the best structure. The Max-Min Hill-Climbing (MMHC) algorithm is an example of such a hybrid method. Currently, Bayesian networks along with the MMHC algorithm have been applied in the exploration of factors related to chronic obstructive pulmonary disease [20], hyperlipidemia [21], and risk factors for other diseases [19], demonstrating promising performance.

This study aims to establish an intuitive graphical representation through Bayesian modeling, providing a clear depiction of relationships between variables. The goal is to explore the direct and indirect causes of stroke in different genders. By employing Bayesian networks for sequential prediction and introducing probability distributions into the model, effective inferences can be made from limited information. The structured and interpretable model outcomes offer crucial insights for healthcare planning and decision-making in prioritizing settings.

To our knowledge, there has not been a large-scale, gender-stratified investigation into stroke risk factors in Shanxi Province, China. The absence of this information may limit our in-depth understanding of the stroke burden in this region. Therefore, conducting a comprehensive, gender-specific survey of stroke risk factors is of paramount importance. Understanding the differences in stroke risk between men and women through such an investigation is crucial for designing health programs aimed at alleviating the stroke burden in Shanxi Province, China.By delving into the similarities and differences of stroke risk factors between genders, we can provide more targeted guidance for personalized prevention and control strategies. Developing stroke prevention policies that cater to the distinct needs of both men and women would be more practically meaningful.

Methods

Study region and study population

This study was conducted in the northern region of China, specifically in Shanxi Province, involving 13 counties across 10 cities within the province. The inclusion criteria for the study subjects were residents aged 35 to 75 years who had been living in the project area for at least 6 months and were residing in the project area in the 12 months preceding the screening. Residents with mental disorders were excluded. A total of 141,269 individuals participated in the project, and after excluding cases with missing values, the final study cohort comprised 134,382 individuals, including 55,136 males and 79,246 females. All study participants provided informed consent, and the research received approval from the Ethics Committee of Fuwai Hospital.

Data collection

The data were collected through a combination of questionnaire surveys, physical examinations, and laboratory analyses. Trained public health physicians from county-level disease control centers or physicians from community health service centers in township hospitals conducted the questionnaire surveys. The on-site data entry system was developed and designed by the National Cardiovascular Center, and the entire entry process was recorded to ensure quality control. Detailed descriptions of the data collection sources have been previously outlined in the literature [22].

The specific questionnaire content covered demographic information such as name, gender, education level, marital status, etc. Lifestyle information included details on smoking, alcohol consumption, etc. Medical history encompassed respiratory conditions, diabetes, hypertension, cardiovascular history, etc. Additionally, project staff (trained physicians or nurses) conducted blood pressure measurements, height, weight, waist circumference measurements, as well as fasting fingertip blood rapid glucose and lipid tests for screening subjects.

Variable definitions

The method of assessing smoking status involves asking how often you currently smoke. If you smoke occasionally, on most days, or every day, we classify it as smoking. For alcohol consumption, we inquire about the frequency of drinking in the past year. If you drink more than 2 times a week, we consider it as alcohol consumption. The assessment of sleep apnea is done by asking others if they have noticed the frequency of your breathing pauses while sleeping, such as almost every day, 3–4 times a week, 1–2 times a week, or 1–2 times a month, to determine the presence of sleep apnea. Regarding stroke, we inquire whether you have had a stroke or received corresponding treatment. The scope of coronary heart disease includes myocardial infarction, coronary intervention (PCI stent implantation, balloon dilation), coronary artery bypass grafting (CABG), angina, and coronary heart disease diagnosed by a secondary or higher-level hospital. Hyperlipidemia [23] is diagnosed by a secondary or higher-level hospital, taking lipid-lowering drugs, or having total cholesterol (TC) ≥ 5.2mmol/L, triglycerides (TG) ≥ 1.7mmol/L, cholesterol ratio (TC-HDL) ≥ 4.1, low-density lipoprotein cholesterol (LDL) ≥ 3.4mmol/L. Hypertension [24]is diagnosed by a secondary or higher-level hospital, taking antihypertensive drugs, or having systolic blood pressure (SBP) ≥ 130mmHg or diastolic blood pressure (DBP) ≥ 80mmHg. The diagnostic criteria for diabetes include being diagnosed by a secondary or higher-level hospital, taking antidiabetic drugs, or fasting blood glucose > 7.0 mmol/L. Body mass index (BMI) categories include underweight (BMI < 18.5 kg/m2), normal weight (18.5–24.0 kg/m2), overweight (24.0–28.0 kg/m2), and obesity (BMI ≥ 28 kg/m2). Family history is defined as male direct relatives (father, son, and siblings) having a stroke before the age of 55, and female direct relatives (mother, daughter, and siblings) having a stroke before the age of 65. Education level is represented by having a high school degree or higher, and income is based on whether the family annual income exceeds 50,000 RMB. The classification criteria for heart rate and abdominal obesity are: heart rateless than 60 beats/min for bradycardia, 60–100 beats/min for normal, and greater than 100 beats/min for tachycardia. The criteria for abdominal obesity are a waist circumference ≥ 85 cm for males and ≥ 80 cm for females.

Bayesian network

Bayesian Network (BN) is a probabilistic graphical model used to represent the probability dependencies among a set of variables [19, 25, 26]. It consists of a Directed Acyclic Graph (DAG) and Conditional Probability Tables (CPT). Nodes in the directed acyclic graph represent random variables, and edges represent dependencies between variables. In a directed graph, if there is an edge from node A to node B (A → B), node A is the parent node of node B, and node B is the child node of node A. From a probability perspective, the state of the parent node directly influences the state of the current node. The state of the child node depends on its parent nodes, and its state is contingent on the states of its parent nodes. In probabilistic modeling, each node in the Bayesian Network has a Conditional Probability Table (CPT), which describes the probability distribution of the node’s state given the states of its parent nodes. Thus, the structure and parameters of the Bayesian Network jointly determine the joint probability distribution of the entire network.

Max-Min Hill-Climbing algorithm

The MMHC algorithm, short for Max-Min Hill-Climbing, combines the advantages of constraint-based and score-based methods. The algorithm consists of two stages [21]. The first stage is the candidate parent-child node determination phase, which employs a heuristic search algorithm to identify potential parent and child nodes to build the Bayesian network framework. It detects potential parent and child nodes by adding, deleting, and reversing edges within the constraint space using a score-based search method. The second stage is the score-based search phase, utilizing a score-based search method to determine the edges and directions of the network structure. This phase involves a greedy search, starting from a blank graph and iteratively adding, deleting edges, and changing directions to obtain a belief network structure with the highest rating. The MMHC algorithm uses score-based search methods in structural learning, allowing for the exploration of global optimal solutions rather than just local optima.

Statistical analysis

In this study, all variables are categorical, and percentages are used for representation. Firstly, inter-group differences were determined through the chi-square test. Variables with inter-group differences (P < 0.1 considered statistically significant) were included in multivariate stepwise logistic regression analysis. For the multifactor analysis, the standards of αin = 0.05 and αout = 0.10 were used to screen variables to explore the influencing factors of stroke. Variables with statistical significance (P < 0.05 considered statistically significant) would be included in the Bayesian network.

The construction of the Bayesian network structure was achieved using the “mmhc()” function in the “bnlearn” package, and Bayesian parameter learning was performed using the maximum likelihood estimation method. All these analyses were implemented using R 4.3.2 software. Finally, Bayesian network model parameter learning and Bayesian inference were conducted using Netica software (Norsys Software Corp., Vancouver, Canada).

Results

Prevalence estimates

Table 1 presents the demographic information, medical history, family history, and lifestyle characteristics of the study participants and stroke patients. The study included 4,413 stroke patients, with 55.77% being male (2,461 individuals) and 44.23% female (1,952 individuals). Among the stroke group, 68.25% of patients were aged between 60 and 75 years. The educational level of the patients was relatively low, with 80.69% having received no education beyond high school. Approximately 90.12% of patients had an annual income less than 50,000 yuan. Regarding risk behaviors, the majority of stroke patients were exposed to secondhand smoke (89.64%), 27.03% were smokers, and 7.64% consumed alcohol. The prevalence of family history of stroke was 1.20%, and a minority of patients experienced sleep apnea (8.45%), with 38.25% of the population reporting snoring. Among stroke patients, 66.26% had abnormal blood lipids, 91.66% had hypertension, 25.74% had diabetes, and a small proportion had comorbid coronary heart disease (5.89%). Additionally, 68.55% of the population had a BMI indicating overweight or obesity, and 72.20% of patients exhibited abdominal obesity.

Table 1.

Baseline characteristics among stroke and non-stroke groups

Characteristics Categories Non-stroke (N = 129969,%) Stroke
(N = 4413, %)
P
Region Rural 75,800(58.32) 2313(52.41) < 0.001
Urban 54,169(41.68) 2100(47.59)
Gender Male 52,675(40.53) 2461(55.77) < 0.001
Female 77,294(59.47) 1952(44.23)
Age 35–49 40,504(31.16) 254(5.76) < 0.001
50–59 43,829(33.72) 1147(25.99)
60–75 45,636(35.11) 3012(68.25)
Marital status No 6715(5.17) 391(8.86) < 0.001
Yes 123,254(94.83) 4022(91.14)
Education No 97,439(74.97) 3561(80.69) < 0.001
Yes 32,530(25.03) 852(19.31)
Income No 118,009(90.80) 3977(90.12) 0.133
Yes 11,960(9.20) 436(9.88)
Alcohol No 121,553(93.52) 4076(92.36) 0.002
Yes 8416(6.48) 337(7.64)
Smoke No 100,090(77.01) 3220(72.97) < 0.001
Yes 29,879(22.99) 1193(27.03)
Second hand smoking No 23,103(17.78) 457(10.36) < 0.001
Yes 106,866(82.22) 3956(89.64)
Stroke family history No 129,629(99.74) 4360(98.80) < 0.001
Yes 340(0.26) 53(1.20)
Apnea No 122,038(93.90) 4040(91.55) < 0.001
Yes 7931(6.10) 373(8.45)
Snore No 91,624(70.50) 2725(61.75) < 0.001
Yes 38,345(29.50) 1688(38.25)
Heart rate Normal 120,559(92.76) 3986(90.32) < 0.001
Bradycardia 6804(5.24) 341(7.73)
Tachycardia 2606(2.01) 86(1.95)
Dyslipidemia No 59,199(45.55) 1489(33.74) < 0.001
Yes 70,770(54.45) 2924(66.26)
Hypertension No 33,119(25.48) 368(8.34) < 0.001
Yes 96,850(74.52) 4045(91.66)
Diabetes No 111,075(85.46) 3277(74.26) < 0.001
Yes 18,894(14.54) 1136(25.74)
BMI Normal 44,800(34.47) 1346(30.50) < 0.001
Underweight 1582(1.22) 42(0.95)
Overweight 57,650(44.36) 2042(46.27)
Obesity 25,937(19.96) 983(22.28)
Abdominal obesity No 45,526(35.03) 1227(27.80) < 0.001
Yes 84,443(64.97) 3186(72.20)
Coronary heart disease No 127,724(98.27) 4153(94.11) < 0.001
Yes 2245(1.73) 260(5.89)

Univariate analysis

We conducted chi-square tests to explore the differences between the male and female stroke groups and non-stroke groups for each variable. The results showed that in the male group, there were statistically significant differences (P < 0.05) in gender, age, education level, income, marital status, region, abnormal blood lipids, hypertension, diabetes, coronary heart disease, family history of stroke, smoking, alcohol consumption, secondhand smoke exposure, snoring, heart rate, overweight/obesity, and abdominal obesity. In the female group, there was no statistically significant difference in income (P > 0.1), while there were statistically significant differences (P < 0.05) in gender, age, education level, marital status, region, abnormal blood lipids, hypertension, diabetes, coronary heart disease, family history of stroke, smoking, secondhand smoke exposure, snoring, overweight/obesity, and abdominal obesity(Table 2).

Table 2.

Results of a univariate chi-square test stratified by gender

Male Female
Characteristics
Categories
Non-stroke
(N = 52675,%)
Stroke
(N = 2461, %)
Non-stroke
(N = 77294,%)
Stroke
(N = 1952, %)
Region
 Rural 31,138(59.11) 1250(2.37) 44,662(57.78) 1063(54.46)
 Urban 21,537(40.89) 1211(2.30) 32,632(42.22) 889(45.54)
Marital status
 No 1890(3.59) 136(0.26) 4825(6.24) 255(13.06)
 Yes 50,785(96.41) 2325(4.41) 72,469(93.76) 1697(86.94)
Education
 No 37,377(70.96) 1872(3.55) 60,062(77.71) 1689(86.53)
 Yes 15,298(29.04) 589(1.12) 17,232(22.29) 263(13.47)
Income*
 No 47,295(89.79) 2175(4.13) 70,714(91.49) 1802(92.32)
 Yes 5380(10.21) 286(0.54) 6580(8.51) 150(7.68)
Age
 35–49 15,493(29.41) 169(0.32) 25,011(32.36) 85(4.35)
 50–59 17,062(32.39) 632(1.20) 26,767(34.63) 515(26.38)
 60–75 20,120(38.20) 1660(3.15) 25,516(33.01) 1352(69.26)
Smoke
 No 24,012(45.59) 1326(2.52) 76,078(98.43) 1894(97.03)
 Yes 28,663(54.41) 1135(2.15) 1216(1.57) 58(2.97)
Alcohol
 No 44,535(84.55) 2140(4.06) 77,018(99.64) 1936(99.18)
 Yes 8140(15.45) 321(0.61) 276(0.36) 16(0.82)
Second hand smoking
 No 8939(16.97) 245(0.47) 14,164(18.32) 212(10.86)
 Yes 43,736(83.03) 2216(4.21) 63,130(81.68) 1740(89.14)
Stroke family history
 No 52,563(99.79) 2440(4.63) 77,066(99.71) 1920(98.36)
 Yes 112(0.21) 21(0.04) 228(0.29) 32(1.64)
Apnea
 No 48,796(92.64) 2216(4.21) 73,242(94.76) 1824(93.44)
 Yes 3879(7.36) 245(0.47) 4052(5.24) 128(6.56)
Snore
 No 34,425(65.35) 1467(2.79) 57,199(74.00) 1258(64.45)
 Yes 18,250(34.65) 994(1.89) 20,095(26.00) 694(35.55)
Heart rate
 Normal 47,907(90.95) 2170(4.12) 72,652(93.99) 1816(93.03)
 Bradycardia 3912(7.43) 248(0.47) 2892(3.74) 93(4.76)
 Tachycardia 856(1.63) 43(0.08) 1750(2.26) 43(2.20)
Dyslipidemia
 No 27,401(52.02) 969(1.84) 31,798(41.14) 520(26.64)
 Yes 25,274(47.98) 1492(2.83) 45,496(58.86) 1432(73.36)
Hypertension
 No 11,199(21.26) 203(0.39) 21,920(28.36) 165(8.45)
 Yes 41,476(78.74) 2258(4.29) 55,374(71.64) 1787(91.55)
Diabetes
 No 44,682(84.83) 1855(3.52) 66,393(85.90) 1422(72.85)
 Yes 7993(15.17) 606(1.15) 10,901(14.10) 530(27.15)
BMI
 Normal 17,556(33.33) 766(1.45) 27,244(35.25) 580(29.71)
 Underweight 666(1.26) 20(0.04) 916(1.19) 22(1.13)
 Overweight 24,018(45.60) 1162(2.21) 33,632(43.51) 880(45.08)
 Obesity 10,435(19.81) 513(0.97) 15,502(20.06) 470(24.08)
Abdominal obesity
 No 18,573(35.26) 754(1.43) 26,953(34.87) 473(24.23)
 Yes 34,102(64.74) 1707(3.24) 50,341(65.13) 1479(75.77)
Coronary heart disease
 No 51,298(97.39) 2299(4.36) 76,426(98.88) 1854(94.98)
 Yes 1377(2.61) 162(0.31) 868(1.12) 98(5.02)

*In the female group, there were no differences between income groups, p >0.1

Multivariate analysis

The results of the final multivariate logistic regression model are presented in Table 3. A stepwise method (αin = 0.05, αout = 0.10) was employed to analyze the risk factors for stroke. The multivariate analysis indicated that in males, the likelihood of having a stroke was 2.25 times higher in the hypertensive population compared to those without hypertension [(95% confidence interval (CI): 1.94, 2.61)], the 60–75 had a 6.83-fold higher risk of stroke compared to the 35–49 [(95% CI: 5.83, 8.06)], and having a family history of stroke increased the risk 4.22 times compared to those without a family history [(95% CI: 2.51, 6.78)]. In females, the probability of having a stroke was higher in individuals with hypertension (OR: 2.31, 95% CI: 1.97–2.73), in the age group 60–75 compared to 35–49 (OR: 9.89, 95% CI: 7.94-112.47) and 50–59 compared to 35–49 (OR: 4.29, 95% CI: 3.42–5.44), with a family history of stroke (OR: 5.56, 95% CI: 3.69–8.15), and in those with coronary heart disease (OR: 2.50, 95% CI: 1.99–3.10), as shown in Table 3.

Table 3.

Results of multivariate binary logistic regression analysis stratified by gender

Characteristics Male Female
OR(95%CI) OR(95%CI)
Region 1.36(1.25–1.48) 1.15(1.05–1.26)
Marital status 0.71(0.59–0.86) 0.76(0.66–0.88)
Education 0.78(0.71–0.86) 0.81(0.71–0.93)
Age(50–59) 3.22(2.72–3.83) 4.29(3.42–5.44)
Age(60–75) 6.83(5.83–8.06) 9.89(7.94–12.47)
Stroke family history 4.22(2.51–6.78) 5.56(3.69–8.15)
Smoke 1.21(1.11–1.32) -
Alcohol - 2.24(1.28–3.65)
Second hand smoking 1.88(1.64–2.16) 1.79(1.55–2.07)
Apnea 1.16(1.01–1.34) -
Snore 1.25(1.14–1.36) 1.28(1.16–1.41)
Dyslipidemia 1.66(1.52–1.81) 1.37(1.24–1.52)
Hypertension 2.25(1.94–2.61) 2.31(1.97–2.73)
Diabetes 1.38(1.25–1.52) 1.41(1.27–1.56)
Coronary heart disease 1.51(1.27–1.79) 2.50(1.99–3.10)

Bayesian networks model

The Bayesian Network (BN) for men is constructed with 14 nodes and 25 directed edges, providing a clear representation of risk factors, outperforming the Logistic regression model. Directed edges indicate the probabilistic dependencies between connected nodes. The results indicate that dyslipidemia, hypertension, age are direct risk factors for stroke, while snoring, education level, and respiratory pauses are indirect risk factors. Additionally, the model suggests that dyslipidemia, diabetes, and age are direct risk factors for coronary heart disease (Fig. 1).

Fig. 1.

Fig. 1

Bayesian network and prior probability for stroke in men constructed using MMHC

The Bayesian Network (BN) for women is constructed with 13 nodes and 23 directed edges, providing a clear representation of risk factors, outperforming the Logistic regression model. Directed edges indicate the probabilistic dependencies between connected nodes. The results indicate that age, hypertension, secondhand smoke exposure are direct risk factors for stroke, while snoring serves as an indirect risk factor for stroke. Additionally, snoring, hypertension, and secondhand smoke exposure are direct risk factors for dyslipidemia (Fig. 2).

Fig. 2.

Fig. 2

Bayesian network and prior probability for stroke in women constructed using MMHC

Bayesian reasoning

The prior probabilities of variables are shown in Fig. 1. The obtained probability model allows for a quantitative analysis of the impact of these factors on stroke by calculating the conditional probability P(y|xi). From Fig. 1, we can understand that the prior probability of stroke in males is 4.43%. If this individual has dyslipidemia, the probability increases to P(stroke|dyslipidemia) = 5.53% (Fig.S1). If this individual also has hypertension, the probability rises to P(stroke|dyslipidemia, hypertension) = 6.02% (Fig.S2). If a person is between the ages of 60 and 75, the probability reaches P(stroke|dyslipidemia, hypertension, age = 60–75) = 10.8% (Fig.S3).

From Fig. 2, we can understand that the prior probability of stroke in females is 2.48%. If this individual has hypertension, the probability increases to P(stroke|hypertension) = 3.13% (Fig.S4). If this individual is aged between 60 and 75, the probability rises to P(stroke|hypertension, age = 60–75) = 5.49% (Fig.S5). If there is also exposure to second-hand smoke, the probability reaches P(stroke|hypertension, age = 60–75, second-hand smoking) = 5.92% (Fig.S6).

Discussion

This study investigated the prevalence of stroke and predictive factors among adults of different genders in Shanxi Province, China. The stroke prevalence in Shanxi Province was 3.3%, higher than the national average [27]. The stroke prevalence was 4.46% in males and 2.46% in females. Traditional logistic regression analysis revealed that males had a higher risk of stroke, with factors such as abnormal lipid levels, hypertension, diabetes, family history of stroke, coronary heart disease, secondhand smoke exposure, and snoring being identified as risk factors. Urban residents also had a higher risk of stroke, and the risk increased with age.

Additional risk factors for males included respiratory pauses and smoking, while females did not share these factors. Furthermore, the impact of similar risk factors differed between genders. For example, females with a family history of stroke had a 5.56-fold increased risk, whereas males had a 4.22-fold increased risk. Using the MMHC algorithm, Bayesian Networks (BNs) in this study indicated that in males, abnormal lipid levels, hypertension, age were direct risk factors for stroke, while snoring, education level, and respiratory pauses were indirect risk factors. In females, age, hypertension, and secondhand smoke exposure were identified as direct risk factors, while snoring was an indirect risk factor for stroke.

In the study of stroke risk factors, traditional logistic regression methods are typically constructed under the assumption that variables are independent, failing to fully leverage data information and accurately reflect the impact of feature variables on stroke [28]. Traditional logistic regression methods use probabilities to reflect the strength of associations, lacking the ability to comprehensively explain the complex relationships between risk factors [21], and are unable to detect direct or indirect risk factors. Therefore, logistic regression models are not flexible enough in capturing patterns and relationships between data.In contrast, Bayesian Networks (BNs) demonstrate more advantages in building risk factor models compared to logistic regression [29]. Firstly, Bayesian Networks do not require any prior assumptions and have the ability to integrate different variables and analyze their relative importance [28]. Therefore, in recent years, many clinical researchers prefer using Bayesian Networks for quantifying the identification of risk factors in specific pathological diagnoses, prognosis, and supporting medical decision-making in diseases [26, 30].We applied Bayesian Networks (BNs) to the study of stroke risk factors by gender. This not only reveals the risk factors for stroke but also determines their direct and indirect impacts on stroke, providing in-depth insights into the complex network relationships among them. It is noteworthy that when constructing Bayesian Networks, the network becomes more complex with the increase of feature variables, so the construction of Bayesian Networks should be based on the selection of different feature variables. In this study, single-factor chi-square and multi-factor logistic analysis were used to screen variables.

Based on our understanding, our study is the first to apply a Bayesian network to analyze risk factors for stroke based on gender. Compared to traditional logistic regression models, Bayesian networks using the MMHC algorithm have significant advantages in analyzing stroke risk factors. First, the Bayesian network with the MMHC algorithm is a data-driven model constructed on the knowledge base associated with the disease [31], without strict requirements on data distribution. This enables it to better discover potential, less obvious but important data information. This data-driven approach provides a more scientific and comprehensive foundation for the assessment, prediction, and prevention of stroke. Therefore, the application of Bayesian networks allows for a deeper understanding of stroke risk factors, providing more accurate guidance for personalized prevention and control strategies.The second advantage involves the interactions between variables. Logistic regression [31] can only provide risk indications for stroke risk factors. However, when analyzing interactions, logistic regression needs to introduce them into the model through addition or multiplication, which adds complexity and may introduce potential biases. Moreover, logistic regression struggles to clearly illustrate the interactions between variables and is limited in exploring the complex relationships of multiple variables. In contrast, Bayesian networks with the MMHC algorithm allow for an intuitive description of the interconnections between these risk factors through graphical methods and can comprehensively explore their direct and indirect interactions [19].

Stroke is relatively common in the elderly population, and previous research reports indicate that age is an unmodifiable risk factor for stroke, applicable to both men and women [32]. Similar to these study findings, our research reveals a higher incidence of stroke in both men and women in the age range of 60–75 years. The potential mechanisms through which age influences stroke include the natural narrowing and hardening of arteries as individuals age. This change is attributed to alterations resulting from endothelial dysfunction and impaired autoregulation of the brain [33].Additionally, the elderly population often experiences a concomitant state of multiple chronic diseases, such as diabetes, hypertension, atrial fibrillation, as well as coronary artery and peripheral artery diseases. The prevalence of these conditions also increases gradually with age [34].

In this study, hypertension, diabetes, dyslipidemia, coronary heart disease, family history of stroke, secondhand smoke exposure, and snoring were identified as common risk factors for stroke in both men and women. These findings are consistent with previous research results [28, 35]. Hypertension has adverse effects on arteries, leading to atherosclerosis and narrowing, which can result in thrombosis or embolism, triggering a stroke [36]. Elevated blood sugar levels can damage endothelial cells, leading to atherosclerosis and narrowing of arteries, causing vascular damage. Impaired kidney function can increase blood volume, reducing the elasticity of blood vessels [37]. Higher cholesterol levels may increase the inflammation and apoptosis of plaques, making them more prone to rupture. After plaque rupture, platelets and coagulation proteins in the blood aggregate at the site of rupture, forming a clot, ultimately leading to a stroke [36]. Family history may be due to shared lifestyles and habits or genetic factors, such as the association of white matter lesions with vertebrobasilar artery atherosclerotic brain disease, which is an autosomal dominant cerebrovascular disease [38] . Harmful substances in secondhand smoke may have a direct toxic effect on the nervous system, increasing the risk of stroke [39]. Snoring is often accompanied by sleep apnea, where breathing briefly stops during sleep. This can lead to a decrease in blood oxygen levels, increasing the risk of stroke. Snoring may also lead to an increase in inflammation in the body, and inflammation is associated with cardiovascular disease and stroke [40, 41].

The study has several limitations. Firstly, in Bayesian Networks (BNs), directed edges cannot accurately represent causal relationships between nodes and can only express probabilistic dependencies. Secondly, due to the face-to-face survey method used, participants may rely on memory to answer questions, introducing potential reporting or recall biases in estimating the prevalence of various diseases. Additionally, the survey did not collect some important information, including: (a) variables related to women’s characteristics such as menstrual history and reproductive history, making the analysis of risk factors in women potentially incomplete; (b) data on some inflammatory factors, electrocardiogram data, carotid ultrasound, and coronary artery ultrasound; (c) effectiveness data related to dietary factors. Therefore, we were unable to assess the impact of these factors on the risk of stroke. Furthermore, as the study focused on Bayesian Networks using the MMHC algorithm, it did not compare with other hybrid algorithms. Stroke was not differentiated into ischemic and hemorrhagic, and the study did not analyze the differential effects of biochemical indicators such as blood glucose, blood pressure, and cholesterol on stroke in the absence of medication factors. This will be a focus of our future work. Despite these limitations, the findings of this study provide valuable information for the development of health planning and programs aimed at reducing the burden of stroke in Shanxi Province, China.

Conclusions

The identification of risk factors plays a crucial role in disease prevention. Our study results indicate that the Bayesian Network (BN) using the MMHC algorithm can not only reveal the intricate relationships among different gender-specific stroke risk factors but can also effectively predict stroke risks for different genders. This not only provides a scientific basis for the control and treatment of stroke in different genders but also contributes to reducing the incidence of stroke in diverse gender populations. The specific findings of the study are as follows:

The logistic regression model shows that risk factors for stroke in males include: region, marital status, education level, age, family history of stroke, smoking, secondhand smoke exposure, sleep apnea, snoring, dyslipidemia, hypertension, diabetes, and coronary heart disease. Risk factors for stroke in females include: region, marital status, education level, age, family history of stroke, alcohol consumption, secondhand smoke exposure, snoring, dyslipidemia, hypertension, diabetes, and coronary heart disease.

The BN model for stroke in males includes 14 nodes and 25 directed edges. Dyslipidemia, hypertension, age are direct risk factors for stroke, while snoring, education level, and sleep apnea are indirect risk factors.The BN model for stroke in females includes 13 nodes and 23 directed edges. Age, hypertension, secondhand smoke exposure are direct risk factors for stroke, while snoring is an indirect risk factor for stroke.

The Bayesian Network (BN) using the MMHC algorithm can perform probabilistic inference for unknown nodes based on known nodes, flexibly illustrating the interactions of different risk factors for stroke.

Supplementary Information

Supplementary Material 1. (765.1KB, docx)

Acknowledgements

We would like to express our gratitude to all the staff and participants who dedicated their hard work and commitment to this project.

Abbreviations

BN

Bayesian Network

MMHC

Max-Min Hill-Climbing

BMI

Body Mass Index

DAG

Directed Acyclic Graph

CPT

Conditional Probability Table

Authors’ contributions

“LH.LQ. and Q.LX. conceived the idea and designed the study. M.L. L.CL.and W.LJ. collected the data. LH.LQ. H.YX. and W.XC. analyzed the data. LH.LQ. H.YX. and Z.J. drafted the manuscript, and then LH.LQ. and Q.LX. reviewed the manuscript. All authors read and approved the final draft.”

Funding

This research is supported by a grant from the National Natural Science Foundation of China (grant no: 81973155). The funding source played no role in the design of the study, data collection, analysis and interpretation, and manuscript writing process.

Data availability

We declare that the materials described in the manuscript, including all relevant raw data, will be freely available to any scientist wishing to use them for non-commercial purposes, without breaching participant confidentiality. If someone wishes to access the original data, they should contact the corresponding author.

Declarations

Ethics approval and consent to participate

This study was conducted with approval from the Ethics Committee of Fuwai Hospital(Approval number: 2014 − 574). The participants provided their written informed consent in this study.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Baghshomali S, Bushnell C. Reducing stroke in women with risk factor management: blood pressure and cholesterol. Womens Health (Lond). 2014;10(5):535–44. [DOI] [PubMed] [Google Scholar]
  • 2.Khan MM, Roberson S, Reid K, Jordan M, Odoi A. Prevalence and predictors of stroke among individuals with prediabetes and diabetes in Florida. BMC Public Health. 2022;22(1):243. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Wafa HA, Wolfe CDA, Emmett E, Roth GA, Johnson CO, Wang Y. Burden of stroke in Europe: thirty-year projections of incidence, prevalence, deaths, and disability-adjusted life years. Stroke. 2020;51(8):2418–27. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Ma Q, Li R, Wang L, Yin P, Wang Y, Yan C, Ren Y, Qian Z, Vaughn MG, McMillin SE, et al. Temporal trend and attributable risk factors of stroke burden in china, 1990–2019: an analysis for the global burden of disease study 2019. Lancet Public Health. 2021;6(12):e897–906. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Wang YJ, Li ZX, Gu HQ, Zhai Y, Zhou Q, Jiang Y, Zhao XQ, Wang YL, Yang X, Wang CJ, et al. China stroke statistics: an update on the 2019 report from the National center for healthcare quality management in neurological diseases, China National clinical research center for neurological diseases, the Chinese stroke association, National center for chronic and non-communicable disease control and prevention, Chinese center for disease control and prevention and Institute for global neuroscience and stroke collaborations. Stroke Vasc Neurol. 2022;7(5):415–50. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Tu WJ, Wang LD, Special Writing Group of China Stroke Surveillance R. China stroke surveillance report 2021. Mil Med Res. 2023;10(1):33. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Man JJ, Beckman JA, Jaffe IZ. Sex as a biological variable in atherosclerosis. Circ Res. 2020;126(9):1297–319. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Campesi I, Montella A, Sotgiu G, Saderi L, Tonolo G, Seghieri G, Franconi F. Smoking and combined oral contraceptives should be considered as an independent variable in sex and gender-oriented studies. Toxicol Appl Pharmacol. 2022;457: 116321. [DOI] [PubMed] [Google Scholar]
  • 9.Radke AK, Sneddon EA, Frasier RM, Hopf FW. Recent perspectives on sex differences in compulsion-like and binge alcohol drinking. Int J Mol Sci. 2021. 22(7) 10.3390/ijms22073788. [DOI] [PMC free article] [PubMed]
  • 10.Salmantabar P, Abzhandadze T, Viktorisson A, Reinholdsson M, Sunnerhagen KS. Pre-stroke physical inactivity and stroke severity in male and female patients. Front Neurol. 2022;13:831773. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Ekker MS, Verhoeven JI, Schellekens MMI, Boot EM, van Alebeek ME, Brouwers P, Arntz RM, van Dijk GW, Gons RAR, van Uden IWM, et al. Risk factors and causes of ischemic stroke in 1322 young adults. Stroke. 2023;54(2):439–47. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Shepard BD. Sex differences in diabetes and kidney disease: mechanisms and consequences. Am J Physiol-Ren Physiol. 2019;317(2):F456-62. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Xu J, Zhang X, Jin A, Pan Y, Li Z, Meng X, Wang Y. Trends and risk factors associated with stroke recurrence in China, 2007–2018. JAMA Netw Open. 2022;5(6):e2216341. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Tsai CF, Sudlow CLM, Anderson N, Jeng JS. Variations of risk factors for ischemic stroke and its subtypes in Chinese patients in Taiwan. Sci Rep. 2021;11(1):9700. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Dong Y, Cao W, Cheng X, Fang K, Zhang X, Gu Y, Leng B, Dong Q. Risk factors and stroke characteristic in patients with postoperative strokes. J Stroke Cerebrovasc Dis. 2017;26(7):1635–40. [DOI] [PubMed] [Google Scholar]
  • 16.Pan J, Rao H, Zhang X, Li W, Wei Z, Zhang Z, Ren H, Song W, Hou Y, Qiu L. Application of a Tabu search-based bayesian network in identifying factors related to hypertension. Med (Baltim). 2019;98(25):e16058. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Wei Z, Zhang XL, Rao HX, Wang HF, Wang X, Qiu LX. [Using the Tabu-search-algorithm-based bayesian network to analyze the risk factors of coronary heart diseases]. Zhonghua Liu Xing Bing Xue Za Zhi. 2016;37(6):895–9. [DOI] [PubMed] [Google Scholar]
  • 18.Bayman EO, Dexter F. Multicollinearity in logistic regression models. Anesth Analg. 2021;133(2):362–5. [DOI] [PubMed] [Google Scholar]
  • 19.Song W, Qiu L, Qing J, Zhi W, Zha Z, Hu X, Qin Z, Gong H, Li Y. Using bayesian network model with MMHC algorithm to detect risk factors for stroke. Math Biosci Eng. 2022;19(12):13660–74. [DOI] [PubMed] [Google Scholar]
  • 20.Quan D, Ren J, Ren H, Linghu L, Wang X, Li M, Qiao Y, Ren Z, Qiu L. Exploring influencing factors of chronic obstructive pulmonary disease based on elastic net and bayesian network. Sci Rep. 2022;12(1):7563. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Wang X, Pan J, Ren Z, Zhai M, Zhang Z, Ren H, Song W, He Y, Li C, Yang X, et al. Application of a novel hybrid algorithm of bayesian network in the study of hyperlipidemia related factors: a cross-sectional study. BMC Public Health. 2021;21(1):1375. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Lu J, Xuan S, Downing NS, Wu C, Li L, Krumholz HM, Jiang L. Protocol for the China PEACE (patient-centered evaluative assessment of cardiac events) million persons project pilot. BMJ Open. 2016;6(1):e010200. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Li JJ, Zhao SP, Zhao D, Lu GP, Peng DQ, Liu J, Chen ZY, Guo YL, Wu NQ, Yan SK, et al.: 2023 China guidelines for lipid management. J Geriatr Cardiol. 2023;20(9):621–63. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Hypertension NCfCDCMDA, Committee of the Chinese Medical Doctor Association. Chinese society of cardiology CM, association ahcocsme: clinical practice guidelines for the management of hypertension in China. Chin J Cardiol. 2022;50(11):1050–95. [Google Scholar]
  • 25.Hanea AM, Christophersen A, Alday S. Bayesian networks for risk analysis and decision support. Risk Anal. 2022;42(6):1149–54. [DOI] [PubMed] [Google Scholar]
  • 26.Park E, Chang HJ, Nam HS. A bayesian network model for predicting post-stroke outcomes with available risk factors. Front Neurol. 2018;9: 699. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Zhao Y, Hua X, Ren X, Ouyang M, Chen C, Li Y, Yin X, Song P, Chen X, Wu S, et al. Increasing burden of stroke in China: a systematic review and meta-analysis of prevalence, incidence, mortality, and case fatality. Int J Stroke. 2023;18(3):259–67. [DOI] [PubMed] [Google Scholar]
  • 28.Fan ZX, Wang CB, Fang LB, Ma L, Niu TT, Wang ZY, Lu JF, Yuan BY, Liu GZ. Risk factors and a bayesian network model to predict ischemic stroke in patients with dilated cardiomyopathy. Front Neurosci. 2022;16:1043922. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Arora P, Boyne D, Slater JJ, Gupta A, Brenner DR, Druzdzel MJ. Bayesian networks for risk prediction using real-world data: a tool for precision medicine. Value Health. 2019;22(4):439–45. [DOI] [PubMed] [Google Scholar]
  • 30.Agrahari R, Foroushani A, Docking TR, Chang L, Duns G, Hudoba M, Karsan A, Zare H. Applications of Bayesian network models in predicting types of hematological malignancies. Sci Rep. 2018;8(1):6951. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Joo YJ, Kho SY, Kim DK, Park HC. A data-driven bayesian network for probabilistic crash risk assessment of individual driver with traffic violation and crash records. Accid Anal Prev. 2022;176: 106790. [DOI] [PubMed] [Google Scholar]
  • 32.Kelly-Hayes M. Influence of age and health behaviors on stroke risk: lessons from longitudinal studies. J Am Geriatr Soc. 2010;58(Suppl 2):S325–328. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Yousufuddin M, Young N. Aging and ischemic stroke. Aging. 2019;11(9):2542–4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Ferrucci L, Fabbri E. Inflammageing: chronic inflammation in ageing, cardiovascular disease, and frailty. Nat Rev Cardiol. 2018;15(9):505–22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Wang C, Du Z, Ye N, Shi C, Liu S, Geng D, Sun Y. Hyperlipidemia and hypertension have synergistic interaction on ischemic stroke: insights from a general population survey in China. BMC Cardiovasc Disord. 2022;22(1):47. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.Sun L, Clarke R, Bennett D, Guo Y, Walters RG, Hill M, Parish S, Millwood IY, Bian Z, Chen Y, et al. Causal associations of blood lipids with risk of ischemic stroke and intracerebral hemorrhage in Chinese adults. Nat Med. 2019;25(4):569–74. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Denorme F, Portier I, Kosaka Y, Campbell RA. Hyperglycemia exacerbates ischemic stroke outcome independent of platelet glucose uptake. J Thromb Haemost. 2021;19(2):536–46. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Zheng X, Zeng N, Wang A, Zhu Z, Peng H, Zhong C, Xu T, Xu T, Peng Y, Li Q, et al. Family history of stroke and death or vascular events within one year after ischemic stroke. Neurol Res. 2019;41(5):466–72. [DOI] [PubMed] [Google Scholar]
  • 39.Fischer F, Kraemer A. Meta-analysis of the association between second-hand smoke exposure and ischaemic heart diseases, COPD and stroke. BMC Public Health. 2015;15:1202. [DOI] [PMC free article] [PubMed]
  • 40.Li D, Liu D, Wang X, He D. Self-reported habitual snoring and risk of cardiovascular disease and all-cause mortality. Atherosclerosis. 2014;235(1):189–95. [DOI] [PubMed] [Google Scholar]
  • 41.Fan M, Sun D, Zhou T, Heianza Y, Lv J, Li L, Qi L. Sleep patterns, genetic susceptibility, and incident cardiovascular disease: a prospective study of 385 292 UK biobank participants. Eur Heart J. 2020;41(11):1182–9. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1. (765.1KB, docx)

Data Availability Statement

We declare that the materials described in the manuscript, including all relevant raw data, will be freely available to any scientist wishing to use them for non-commercial purposes, without breaching participant confidentiality. If someone wishes to access the original data, they should contact the corresponding author.


Articles from BMC Public Health are provided here courtesy of BMC

RESOURCES