Skip to main content
BMC Medical Research Methodology logoLink to BMC Medical Research Methodology
. 2026 Jul 21;26:213. doi: 10.1186/s12874-026-02944-8

Applications of causal and structural equation modeling in epidemiology: a systematic and critical review

Scholastique Midokpè Merveille Essetcheou 1,✉, Houétchénou Gislain Fortuné Dovonou 1, Souand Peace Gloria Tahi 1, Sèton Calmette Ariane Houetohossou 1, Valère Kolawolé Salako 1, Marcel Tadogbè Donou Hounsode 1, Romain Glèlè Kakaï 1
PMCID: PMC13628753  PMID: 42482173

Abstract

Background

Structural equation modeling (SEM) and causal modeling (CM) are powerful statistical approaches for identifying complex interrelationships among variables. However, their application in epidemiology remains limited and under-documented, especially in infectious disease research, which requires integrated analytical frameworks for effective control.

Methods

To examine how SEM and CM have been applied, their methodological characteristics, and reporting practices, a systematic and critical review was conducted following PRISMA guidelines. The search covered studies published between 1987 and 2025 across PubMed, Scopus, Web of Science, ScienceDirect, SpringerLink, Google Scholar, and the Directory of Open Access Journals. After rigorous screening, 458 articles were thoroughly evaluated.

Results

Most studies focused on neuropsychiatric (32.1%) and chronic (30.1%) conditions, with few addressing infectious diseases (24.0%), primarily malaria, tuberculosis, and HIV, particularly in low-income countries where context-specific evidence is urgently needed to inform targeted interventions. SEM studies predominantly used maximum likelihood estimation (57.7%) and large samples (≥ 200 observations in 85%) with CB-SEM remaining the dominant approach across all sample size categories. In contrast, CM studies showed substantial variability in sample sizes across approaches, ranging from fewer than 100 to over 200 observations (coefficient of variation Inline graphic), with no consistent sample size threshold across methods. Methodological reporting was often incomplete, notably regarding study design (17.4%), measurement validity (15.8%), and model fit criteria (5.2%), reducing transparency and reproducibility.

Conclusion

Overall, broader application of SEM and CM to infectious diseases, combined with improved methodological transparency, could substantially strengthen causal inference and guide evidence-based disease control strategies. Moreover, integrating longitudinal study designs would further enhance the robustness and interpretability of causal findings.

Supplementary Information

The online version contains supplementary material available at https://doi.org/10.1186/s12874-026-02944-8.

Keywords: Structural equation modeling, Causal modeling, Causal inference, Path analysis, Public health

Introduction

Human health is increasingly challenged by a dual burden of disease: chronic non-communicable conditions and infectious diseases. Non-communicable diseases including cardiovascular disorders, cancers, diabetes, and mental health conditions persist alongside endemic, emerging, and re-emerging infections such as malaria, tuberculosis, and Human Immunodeficiency Virus (HIV). Together, these conditions exert substantial strain on healthcare systems, economies, and societies worldwide [1, 2]. This burden is further amplified by globalization, urbanization, climate change, migration, and social inequalities. These forces, shaped by geographical and environmental factors, continue to reshape patterns of disease transmission and exposure [3].

Historically, epidemiological studies have relied on classical statistical approaches such as generalized linear models and correlation analysis, which remain valuable for exploring associations. However, these methods often fall short in observational contexts due to their limited ability to handle confounding factors, selection bias, measurement error, and nonlinear interactions [4, 5]. Moreover, they tend to treat variables in isolation rather than capturing the complex systems in which health determinants interact [6], thereby restricting their capacity to support strong causal inference [7]. As public health challenges become increasingly complex, addressing these methodological limitations has become essential for generating actionable insights.

In response, advanced analytical frameworks have emerged to enhance causal reasoning in epidemiology. Causal inference methods such as Directed Acyclic Graphs (DAGs), the potential outcomes framework, and g-methods allow researchers to formally specify causal hypotheses, design bias-resistant analyses, and improve interpretability [8]. Similarly, Structural Equation Modeling (SEM), long established in the social sciences and increasingly applied in health research, provides a flexible framework for examining latent constructs, mediating relationships, and indirect effects within multivariate systems [9].

SEM and causal modeling (CM) have gained growing prominence across epidemiological applications, from infectious diseases such as malaria [10, 11], tuberculosis [12], and HIV [13], to chronic conditions including hypertension and cancer [14]. These approaches are particularly useful for studying diseases with multifactorial determinants, where environmental, social, and behavioral factors interact in complex ways. By capturing direct and indirect relationships among variables, SEM and CM provide a flexible framework to explore causal mechanisms in diverse epidemiological contexts.

Despite increasing interest in these approaches, the empirical application and methodological rigor of SEM and CM in epidemiological research remain insufficiently documented. Several methodological contributions have highlighted their theoretical strengths and potential for improving causal inference, including guidance on the use of DAGs and confounder selection [15, 16]. Empirical applications of SEM have also expanded across diverse health contexts, including child undernutrition, infectious disease prevention such as COVID-19, and health behavior research [17, 18]. In addition, previous reviews have summarized SEM’s statistical foundations and applications across disciplines [19–21]. However, these works did not specifically examine how SEM and CM are implemented in epidemiology, the disease areas and health outcomes to which they are applied, or the robustness of their methodological practices. This lack of synthesis limits understanding of their real-world utility and constrains their contribution to evidence-based public health.

The present review therefore examines how CM and SEM are applied in epidemiological studies across a range of disease contexts. Specifically, it aims to (i) examine epidemiological studies employing CM or SEM their methodological frameworks, assumptions, and contexts of application; and (ii) evaluate their contributions, limitations, and implications for advancing causal inference in public health.

Overview on structural equation and causal modeling

History of structural equation and causal modeling

In statistics and empirical research, multivariate data analysis has long been used to test hypothetical relationships among variables. These methods are commonly grouped into two generations. The first generation includes classical approaches such as multiple linear regression, logistic regression, and analysis of variance (ANOVA) [22]. Widely applied across disciplines, including epidemiology [23], these methods assume that variables are directly observed and measured without error, and they rely on simplified model structures that limit their ability to represent complex real-world systems [24].

To overcome these limitations, second-generation techniques such as Structural Equation Modeling (SEM) emerged in the early 1970 s [22]. SEM extends general linear modeling by integrating factor analysis and path modeling, allowing simultaneous estimation of measurement and structural models [25]. This makes it particularly suited for analyzing complex systems with interrelated variables and unobserved constructs. Its development was marked by the introduction of the Linear Structural Relationships (LISREL) model and software, through the work of Karl Jöreskog, Ward Keesling, and David Wiley [26–29], which greatly facilitated parameter estimation and model testing. Over subsequent decades, SEM became widely adopted in many fields, including epidemiology, where it supports the study of latent constructs, mediation pathways, and complex interdependencies [14, 30, 31]. A key advantage is its capacity to decompose effects into direct and indirect components, generalizing causal path analysis.

In parallel, causal inference modeling underwent a transformative evolution. Building on the counterfactual framework of Rubin and others in the 1970s [32], Judea Pearl introduced Directed Acyclic Graphs (DAGs) and the do-calculus [33], providing a formal graphical language to specify and test causal assumptions, particularly valuable in observational research where randomization is infeasible. This work bridged counterfactual reasoning and graphical modeling, forming the basis of modern causal theory.

Although SEM and CM developed along separate paths, recent decades have seen growing efforts to integrate their strengths. SEM emphasizes model fit, latent constructs, and covariance structures, while CM prioritizes identification strategies, confounding control, and explicit causal assumptions [34]. Incorporating DAGs into SEM improves causal specification and reduces biases such as collider bias or overadjustment. Their convergence offers epidemiology a powerful toolkit for disentangling the mechanisms linking exposures to outcomes and for guiding effective public health interventions.

Theoretical framework of SEM and CM

Structural equation modeling: theoretical framework and mathematical formulation

Conceptual foundations

SEM is an advanced multivariate statistical approach that combines factor analysis and multiple regression to model complex relationships among variables within a theoretically grounded framework [35, 36]. It distinguishes between observed variables (indicators), directly measured through empirical instruments such as surveys, tests, or sensors, and latent variables (constructs), which are unobservable concepts inferred from the covariation among indicators. This distinction is fundamental in SEM, as it allows researchers to model abstract concepts that cannot be directly measured.

The SEM framework consists of two interconnected sub-models [37]. The measurement model (or outer model) defines the relationships between latent constructs and their observed indicators. These relationships can be specified as reflective (where the latent variable influences its indicators) or formative (where indicators collectively define the latent construct) [38]. The structural model (or inner model) specifies the hypothesized causal relationships among latent constructs, including direct and indirect effects, mediation, and moderation pathways.

SEM is commonly represented using path diagrams, where latent variables are depicted as ovals, observed variables as rectangles, single-headed arrows indicating directional (causal) relationships, and double-headed arrows representing covariances or correlations [39].

Beyond these classical models, SEM supports hybrid specifications such as the Multiple Indicators Multiple Causes (MIMIC) model, where a construct is both reflected by indicators and influenced by external formative inputs [40, 41]. Such models require strong theoretical justification and careful attention to identification issues.

In addition, SEM also distinguishes between exogenous constructs, which are determined outside the model and serve as independent variables, and endogenous constructs, which are influenced by other variables within the model and may themselves act as predictors. This distinction is crucial for establishing causal order and interpreting structural relationships.

The validity of SEM depends on several critical assumptions. First, all specified relationships must be theoretically justified [35, 42]. Second, measurement models must demonstrate reliability and validity [39, 43]. Third, structural relationships are generally assumed to be linear unless explicitly modeled otherwise [44, 45]. When these conditions are met, SEM provides a robust framework for testing complex theoretical models involving latent variables and their interrelationships [46].

Figure 1 illustrates these conceptual foundations through a path diagram combining reflective and formative measurement models, multiple latent constructs, and structural interdependencies.

Fig. 1.

Fig. 1

Conceptual Framework of Structural Equation Modeling. Note: λ: Loadings; γ: Structural coefficients; ϕ: Covariances; δ, ɛ, ζ: Errors; Indicators: y, x

Mathematical formulation of structural equation modeling

Structural Equation Modeling (SEM) combines two core components: the structural model and the measurement model [35, 47].

The structural model is generally formulated as:

graphic file with name d33e762.gif 1

where Inline graphic (m × 1) denotes the endogenous latent variables, Inline graphic (k × 1) the exogenous latent variables, Inline graphic (m × m) captures the relations among endogenous variables (with zeros on the diagonal to avoid self-effects [31]), and Inline graphic (m × k) specifies the influence of exogenous variables on endogenous constructs. The disturbance term Inline graphic represents residual variance, assumed to be multivariate normal with zero mean and uncorrelated with Inline graphic [44].

The measurement model relates latent variables to their observed indicators. In the widely used reflective specification, latent variables generate their measures:

graphic file with name d33e821.gif 2

with Inline graphic (p × 1) and Inline graphic (q × 1) observed indicators for exogenous and endogenous latent variables, respectively. The loading matrices Inline graphic (p × k) and Inline graphic (q × m) quantify the strength of these relationships [48]. Measurement errors Inline graphic and Inline graphic are assumed to be normally distributed with zero mean, uncorrelated with the latent variables and each other [43].

Alternatively, the formative specification models latent variables as composites formed by their indicators [49]:

graphic file with name d33e883.gif 3

where Inline graphic and Inline graphic are weight matrices, and Inline graphic represent unexplained variance. This formulation reverses causal direction, leading to different assumptions regarding causality and error terms.

To ensure model identification and unique parameter estimates, constraints such as fixing one factor loading per latent variable or constraining the variance of exogenous latent variables are typically applied [45, 50].

Under standard SEM assumptions: linearity, multivariate normality, and independent observations, the joint distribution of the observed indicators follows a multivariate normal distribution [51], such that Inline graphic, where the model implied covariance matrix Inline graphic is defined as a function of the model parameters Inline graphic and has the following block structure:

graphic file with name d33e930.gif 4

where Inline graphic is the k × k covariance matrix of the exogenous latent variables, Inline graphic is the m × m covariance matrix of the structural disturbance terms, Inline graphic and Inline graphic denote the covariance matrices associated with the measurement errors Inline graphic and Inline graphic, respectively, Inline graphic is the identity matrix and Inline graphic denotes transpose.

This framework generalizes several analytic approaches. Omitting measurement error yields path analysis [52], while excluding structural relationships reduces to confirmatory factor analysis (CFA) [43].

Based on the specified model structure and under standard assumptions, parameter estimation in SEM is typically performed by minimizing a discrepancy function between the sample covariance matrix Inline graphic and the model-implied covariance matrix Inline graphic [51]. The most commonly used estimation method is maximum likelihood (ML), which assumes multivariate normality and yields consistent and efficient parameter estimates [35]. The estimation problem is typically expressed as: Inline graphic where F denotes a fitting function, often the likelihood-ratio chi-square. The accuracy of parameter estimates is evaluated using asymptotic standard errors, computed from the inverse of the observed information matrix. Alternative estimation techniques, such as generalized least squares (GLS), weighted least squares (WLS), or Bayesian approaches may be used depending on data properties and model assumptions [44, 53].

Causal modeling: theoretical framework and mathematical formulation

Causal modeling provides a rigorous framework for identifying and quantifying causal relationships, moving beyond purely associational analyses. At its core, Structural Causal Models (SCMs) combine graphical representations with structural equations, expressing each endogenous variable Inline graphic as a function of its direct causes Inline graphic and exogenous noise Inline graphic:

graphic file with name d33e1033.gif 5

This mathematical formulation explicitly encodes the data-generating process, distinguishing genuine causal effects from mere correlations [54]. The causal structure is represented through DAGs, where directed edges denote causal relationships and the acyclic property ensures temporal and logical consistency.

Central to causal inference is Pearl’s do-calculus, where the do-operator (Inline graphic) represents an intervention that sets variable X to value x, enabling computation of the post-intervention distribution Inline graphic [33]. Valid causal inference requires three key assumptions: (1) no unmeasured confounding (exchangeability), (2) positivity (0< P(X=x) < 1 for all strata), and (3) correct model specification.

Various estimation approaches address these requirements. Regression adjustment controls for measured confounders, while propensity score methods balance covariate distributions between treatment groups [55]. Marginal structural models handle time-varying confounding [56], and instrumental variables address unmeasured confounding [57]. More recently, doubly robust methods like Targeted Maximum Likelihood Estimation (TMLE) combine machine learning with causal inference [58].

This comprehensive framework supports rigorous causal analysis in epidemiological studies, providing the theoretical foundation for deriving evidence-based interventions from observational data while properly accounting for potential confounding factors. The integration of graphical models, structural equations, and diverse estimation methods offers researchers a flexible yet principled approach to causal discovery and effect estimation.

Methodology

To ensure methodological transparency, replicability, and rigor, this systematic review was conducted in strict compliance with the Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA) 2020 guidelines [59]. The methodology is structured into three interconnected steps, each detailing a key component of the review process: the formulation of research questions, the design and execution of the literature search and the procedures for data extraction and analysis.

Research questions

This systematic review was guided by a set of research questions designed to examine the application of CM and SEM in epidemiology. These questions addressed the objectives of the studies, the methodological approaches used, the disease areas investigated, and the types of data and variables analyzed. They also examined the statistical techniques and software employed, model validation and evaluation practices, as well as the methodological strengths, limitations, and future research directions reported in the literature. The complete list of research questions is presented in Table 1 to facilitate clarity and precision in the review process.

Table 1.

Research questions guiding the systematic review

No. Research questions
RQ1 What are the primary objectives pursued in studies using SEM or CM in epidemiological research?
RQ2 Which types (e.g., standard SEM, path analysis, Mendelian randomization) and approaches (e.g., Covariance-based SEM, Partial Least Squares SEM, Bayesian SEM) of SEM and CM are applied, and what other statistical methods are used alongside these approaches?
RQ3 Which types of diseases and health conditions are most frequently investigated using these approaches, and what are the main determinants examined within each disease or group of diseases?
RQ4 What are the main types of variables analyzed and how are their effects conceptualized?
RQ5 What types of data, data sources, sample sizes, and geographic regions are typically used in SEM and CM studies?
RQ6 What are the most commonly used statistical and estimation techniques and software packages?
RQ7 How are measurement models validated, and what criteria are used for evaluating model fit and reliability?
RQ8 What are the reported methodological strengths and limitations associated with SEM and CM in epidemiology?
RQ9 What future directions and recommendations are proposed to enhance the use of SEM and CM in public health research?

Search strategy

The literature search was conducted across seven major academic databases, including PubMed, which specializes in public health and epidemiological research, as well as Web of Science, Scopus, ScienceDirect, SpringerLink, Google Scholar, and the Directory of Open Access Journals (DOAJ). The search strategy integrated three conceptual categories of keywords covering modeling approaches (“Structural Equation Modeling”, “Structural Equation Modelling”, “Causal Modeling”, “Causal Modelling”, “Causal Inference”, “Path Analysis”, “Latent Variable Models”); epidemiological scope (“Epidemiology”, “Public Health”, “Diseases Surveillance”, “Infectious Diseases”, “Vector-borne Diseases”, “Mosquito-borne Diseases”, “Plasmodium Infections”, “Malaria”, “Noncommunicable Diseases”, “Non-communicable Diseases”, “NCDs”, “Chronic Diseases”, “Cardiovascular Diseases”, “Respiratory Diseases”, “Cancer”, “Diabetes”, “Obesity”); and influencing factors (“Climate Change”, “Environmental Factors”, “Environmental Changes”, “Meteorological Factors”, “Climatic Variables”, “Weather Variability”, “Extreme Weather Events”, “Temperature”, “Rainfall”, “Precipitation”, “Humidity”, “Drought”, “Flooding”, “Seasonal Variation”, “Personal Characteristics”, “Individual Factors”, “Socioeconomic Factors”, “Socioeconomic Status”, “Social Determinants of Health”, “Demographic Factors”, “Health Behavior”, “Hygiene Practices”, “Living Conditions”, “Household Characteristics”, “Education Level”, “Economic Status”, “Occupation”, “Sex”, “Age”). Boolean operators connected these terms systematically (OR within categories and AND between categories), with no date restrictions applied, and the final search was conducted on 28 March 2025.

The retrieved records were then subjected to a rigorous screening and selection process to identify studies meeting the predefined inclusion criteria. Studies were included if they: (i) applied SEM, CM, or related causal inference approaches; (ii) addressed an epidemiological or public health research question; (iii) used primary data or detailed secondary analyses; and (iv) were published as peer-reviewed full-text articles. Conversely, studies were excluded if they applied SEM or CM exclusively outside health-related contexts; were purely theoretical or methodological without empirical application; corresponded to non-empirical publication types (e.g., reviews, book chapters, commentaries, editorials, protocols, or perspectives); or lacked sufficient methodological detail or usable analytical information. Two reviewers independently conducted the selection process, screening titles, abstracts, and full texts at each stage, and resolving any discrepancies through discussion to reach consensus. The updated database search yielded a total of 2,270 records across all sources: PubMed (360), Web of Science (198), Scopus (185), ScienceDirect (354), Google Scholar (915), SpringerLink (210), and DOAJ (48). After removing 459 duplicates using reference management software (Zotero, Mendeley), 1,811 unique articles remained for screening. Titles and abstracts were examined to exclude records lacking explicit reference to SEM or CM within an epidemiological context, while articles mentioning related concepts such as “causal inference” or “mediation analysis” were retained for further evaluation. This step excluded 1021 articles, leaving 790 for full-text assessment. During full-text screening, 332 articles were further excluded: 229 for thematic irrelevance, particularly studies applying SEM or CM exclusively outside health contexts, employing these methods only theoretically, or relying on alternative statistical approaches without meaningful comparison to SEM or CM; 103 for non-empirical studies, including 44 literature reviews, 38 book chapters, 10 commentaries, 6 perspectives, 4 protocols, and 1 editorial. Additional exclusions targeted studies lacking empirical data, sufficient methodological detail, or substantive analytical content. Prior to final exclusion, reference lists from systematic reviews, books, and related articles were examined, but no additional eligible studies were identified through alternative sources such as websites, organizational repositories, or citation tracking. At the conclusion of this process, 458 studies met all inclusion criteria and were retained for in-depth analysis, ensuring a robust and reproducible dataset for subsequent evaluation. The complete selection process is illustrated in the PRISMA flow diagram (Fig. 2).

Fig. 2.

Fig. 2

PRISMA 2020 flow diagram illustrating the study selection process from identification to final inclusion

Data extraction and analysis

Following the selection of the final set of articles, data were systematically extracted, capturing study year, disease type, geographic region, model type (SEM, CM, or both), specific modeling approach, data source type, sample size, model evaluation criteria, and main categories of variables analyzed. Extraction was performed independently by two reviewers using an identical, standardized Microsoft Excel template. To ensure the reliability of data extraction, the inter-reviewer agreement was quantified using Cohen’s kappa statistic (κ = 0.72), which indicated substantial agreement between reviewers [60]. Any discrepancies were discussed and resolved through consensus, ensuring consistency, minimizing potential bias, and maintaining transparency in accordance with PRISMA guidelines.

Analyses were organized into two main categories: (i) descriptive analyses, providing an overview of study characteristics and distributions to address each research question, and (ii) bibliometric analyses, exploring thematic structures and conceptual linkages within the field.

Descriptive analyses examined temporal trends in the use of SEM and CM, distribution of studies by disease type, model type, data source, and geographic coverage, along with additional analyses as needed to address specific research questions. Geographic mapping illustrated the countries of origin of the datasets, highlighting regional and global research dynamics. Outputs were presented in tables and graphical visualizations including histograms, bar graphs, and pie charts produced in R version 4.3.0 (R Core Team, 2023 [61]) using the ggplot2 (v3.5.1) [62] for conventional statistical graphics and tmap for specialized cartographic visualizations. Semantic harmonization techniques were applied to ensure comparability across studies by systematically grouping conceptually similar information into unified analytical categories. For several extracted fields, including study objectives, disease types, causal modeling approaches, determinants, and reported limitations, original descriptions were first recorded verbatim and then iteratively coded into broader categories using an inductive coding approach. For example, study objectives expressed with different wording but addressing similar analytical aims (identifying risk factors, assessing causal pathways, or evaluating intervention effects) were grouped under common objective categories. Likewise, disease such as diabetes, cardiovascular diseases, HIV, and COVID-19 were categorized into broader groups (chronic, infectious diseases, etc.), consistent with the classification framework used in data extraction database, to facilitate synthesis. Similarly, different terminologies describing comparable determinants (socioeconomic status, income level, or education) were harmonized into unified determinant categories. CM techniques were also grouped into standardized model classifications. This iterative coding process was conducted by the review team using a predefined coding framework and consensus discussions to enhance consistency, transparency, and analytical coherence.

Bibliometric analyses mapped keyword co-occurrence networks to identify thematic clusters and explore the field’s conceptual structure [63]. These analyses were conducted using VOSviewer (version 1.6.20).

Finally, research gaps were identified and categorized according to their relevance and implications for future work. Strategic recommendations were formulated based on these findings to guide the advancement of SEM and CM applications in epidemiological research.

Results

Trend in the use of SEM and CM

The studies included in this review were published between 1987 and 2025 (Fig. 3). Overall, the use of SEM and CM in epidemiology has increased over time, albeit with notable fluctuations. SEM appeared sporadically prior to 2010, reflecting its limited early adoption, and subsequently exhibited a marked increase after 2018, rising from 27 publications in 2018 to a peak of 49 publications in 2022, before slightly declining in 2023–2024. These fluctuations may reflect the impact of major global health events, such as the COVID-19 pandemic, which amplified the demand for sophisticated modeling tools to disentangle multifactorial relationships, as well as the growing availability of large, complex datasets requiring multivariate approaches. The slight decline observed after 2022 can also be explained by growing interest in causal inference methods, more thorough methodological scrutiny of SEM applications, or more general disruptions to scientific productivity during the post-pandemic period. By contrast, CM, first introduced in epidemiological research in 2009, has followed a slower yet steady trajectory, reaching 14 publications in 2023. This gradual adoption reflects the growing recognition of causal inference frameworks as essential for generating robust evidence from observational data, although widespread implementation remains constrained by technical demands and ongoing debates regarding their underlying assumptions.

Fig. 3.

Fig. 3

Number of publications using SEM, CM, or both methods

Studies integrating SEM and CM remain exceedingly rare, with only a single instance identified in 2023. This scarcity likely reflects both conceptual and computational challenges inherent to combining two distinct methodological paradigms: SEM, grounded in latent variable modeling, and CM, based on counterfactual frameworks highlighting a significant methodological gap and a promising avenue for future research.

Collectively, these trends underscore a growing yet uneven adoption of advanced modeling approaches in epidemiology. While SEM continues to be favored for testing complex theoretical models, CM is increasingly valued for its explicit causal interpretation. This divergence illustrates persistent methodological fragmentation but also points to opportunities for integrative strategies that could advance causal understanding and shape the methodological landscape of future epidemiological research.

Bibliometric analysis

To identify the dominant orientations of epidemiological research and the thematic structure of studies using SEM and CM, a bibliometric analysis was conducted. Using keyword co-occurrence, this approach illustrates how these methods are applied across public health domains and highlights both established and emerging research fronts. A total of 1,625 author keywords were extracted, of which 101 appeared more than five times. Figure 4 presents the co-occurrence network, where node color indicates thematic clusters, node size reflects frequency, and line thickness denotes the strength of associations. The most frequent keywords included “humans” (158 occurrences), “female” (110), “male” (90), “adult” (69), “middle aged” (62), “COVID-19” (56), “epidemiology” (55), “aged” (49), “cross-sectional studies” (46), and “risk factors” (35). These results highlight the predominance of population-based epidemiological research focusing on demographic characteristics and health determinants. Methodological terms such as “structural equation modeling”, “path analysis”, and “causal inference” also appeared prominently in the network, reflecting the central role of SEM and CM in examining complex causal relationships within public health research. To refine the thematic interpretation, general demographic descriptors were considered alongside more specific health-related topics. The network revealed substantial attention to public health issues such as depression, socioeconomic factors, quality of life, body mass index, and other behavioral and environmental determinants of health. These findings suggest that SEM and CM are widely used to investigate multifactorial pathways linking social, behavioral, and biological determinants in epidemiology.

Fig. 4.

Fig. 4

Visual network map of keywords co-occurrence

Table 2 summarizes the identified clusters, detailing the main keywords, their frequency of occurrence, total link strength (indicating cumulative co-occurrence strength), and thematic overview. Six clusters emerged, delineating major research directions involving SEM and CM in epidemiology. The first cluster (Red) mainly reflects research on maternal and child health. The prominence of terms related to pregnancy, infancy, and early childhood, combined with methodological references to structural equation modeling and path analysis, suggests that many studies rely on SEM to explore early-life health conditions and environmental exposures influencing developmental outcomes. Closely related to population health research, the second cluster (Green) emphasizes demographic characteristics and chronic disease risk factors. The presence of descriptors referring to sex and age groups, together with lifestyle factors such as smoking and broader risk determinants, indicates a strong focus on understanding how demographic and behavioral attributes shape chronic health conditions in epidemiological investigations. A third thematic group (Blue) shifts the focus toward broader epidemiological and social dimensions of health. This cluster brings together concepts related to adult populations, cohort-based research, socioeconomic conditions, and mental health outcomes such as depression. Taken together, these elements point to studies examining how social environments and life-course factors interact to influence population health. The fourth cluster (Yellow) centers on behavioral health among children and adolescents. References to adolescent populations, health behaviors, body mass index, and survey-based research approaches illustrate how SEM has been employed to investigate lifestyle patterns and behavioral determinants affecting health during early stages of life. More recent global health challenges appear prominently in the fifth cluster (Purple), which captures research associated with the COVID-19 pandemic. The combination of terms related to COVID-19, public health, mental health, and pandemics reflects the growing use of SEM and CM to analyze the complex interactions between epidemiological dynamics, behavioral responses, and societal impacts during public health crises. Finally, the sixth cluster (Cyan) highlights studies focusing on health outcomes and inequalities within cross-sectional epidemiological frameworks. Themes related to quality of life, diabetes, anemia, social class, and latent class analysis suggest applications of SEM in investigating chronic conditions, population disparities, and broader indicators of well-being. Taken together, these thematic groupings demonstrate the wide range of contexts in which SEM and CM are currently applied in epidemiological research. Beyond methodological development, these approaches increasingly contribute to the analysis of complex relationships linking demographic characteristics, social determinants, behavioral factors, and emerging public health challenges.

Table 2.

Summary of main keywords and thematic clusters from the updated VOSviewer co-occurrence analysis

Cluster (Color) Main keyword Occ. TLS Theme description
Cluster 1 (Red) humans 158 1285 Maternal and child epidemiology using SEM
pregnancy 18 143
child, preschool 14 145
infant 12 118
structural equation modeling 17 80
path analysis 18 77
Cluster 2 (Green) female 110 1037 Demographic and chronic disease risk factors
male 90 874
middle aged 62 669
aged 49 523
risk factors 35 321
smoking 10 94
Cluster 3 (Blue) adult 69 723 Epidemiology and social determinants of health
epidemiology 55 415
socioeconomic factors 34 285
young adult 31 370
depression 24 159
cohort studies 20 187
Cluster 4 (Yellow) adolescent 33 370 Child behavioral and survey-based health studies
health behavior 22 222
child 17 176
body mass index 15 158
surveys and questionnaires 18 149
feeding behavior 6 75
Cluster 5 (Purple) covid-19 56 220 COVID-19 and public health impacts
public health 29 253
mental health 16 132
sars-cov-2 18 129
pandemics 15 136
mediation analysis 7 61
Cluster 6 (Cyan) cross-sectional studies 46 435 Cross-sectional SEM applications and health outcomes
latent class analysis 12 102
quality of life 8 61
diabetes mellitus 7 62
social class 7 84
anemia 5 36

RQ1: What are the primary objectives pursued in studies using SEM or CM in epidemiological research?

According to the authors, the objectives behind using SEM or CM in epidemiology are diverse. A detailed analysis of the 458 articles included in this review identified four main objective categories: i) analysis of causal relationships and mediation effects, ii) identification of associated factors, iii) theoretical model development and validation, in which SEM or CM is employed to test, refine, or validate conceptual frameworks or hypothesized causal structures, including DAGs and iv) evaluation of intervention or exposure effects (Table 3). Among studies using SEM (n = 378), causal relationship and mediation analysis was the most common objective, underscoring SEM’s established role in exploring complex structures and indirect mechanisms. This was followed by the identification of associated factors, whereas theoretical model development and intervention/exposure effect evaluation were much less frequently addressed. Studies using CM only (n = 79) also prioritized causal relationship and mediation analysis, but placed relatively greater emphasis on theoretical model development and intervention/exposure effect evaluation, while the identification of associated factors was less common. Notably, the only study combining SEM and CM focused on theoretical model development. Critically, this distribution highlights the disproportionate focus on causal and mediation analysis, with limited attention to theoretical model development and intervention assessment. The comparatively higher representation of these objectives in CM studies may reflect the method’s flexibility in addressing complex epidemiological questions; nevertheless, the overall scarcity points to underexplored opportunities for translating SEM and CM findings into practical public health applications. These patterns suggest that future research should broaden its focus beyond causal exploration to include intervention testing and theoretical refinement, thereby enhancing the utility of SEM and CM in epidemiology.

Table 3.

Distribution of studies by model type and objective category

Model type Objective category N Frequency (%)
CM only Causal Relationship and Mediation Analysis 54 68.4
Identification of Associated Factors 6 7.6
Intervention/Exposure Effect Evaluation 8 10.1
Theoretical Model Development and Validation 11 13.9
SEM + CM Theoretical Model Development and Validation 1 100.0
SEM only Causal Relationship and Mediation Analysis 184 48.7
Identification of Associated Factors 154 40.7
Intervention/Exposure Effect Evaluation 13 3.4
Theoretical Model Development and Validation 27 7.1

RQ2: Which types and approaches of SEM and CM are applied, and what other statistical methods are used alongside these approaches?

The studies included in this review employed a diverse array of SEM and CM types in epidemiological research (Table 4). Within SEM applications, standard SEM was by far the most commonly used type, followed by path analysis. More specialised variants such as cross-lagged SEM, multilevel SEM, and multigroup SEM were only rarely employed (Table 4). This pattern suggests that researchers tend to rely on conventional frameworks rather than exploring advanced or context-specific variants that might provide deeper insights. By contrast, CM applications exhibited greater methodological diversity. Matching and weighting approaches were the most frequently applied CM methods, followed by the potential outcomes framework. Other approaches, including Mendelian randomization, graphical causal models, dynamic systems models, and Bayesian methods, were less frequently used. While this variety reflects the adaptability of CM to different epidemiological contexts, the uneven distribution across methods also indicates that several powerful approaches remain underutilised. Notably, both SEM and CM approaches were occasionally combined with complementary techniques such as machine learning methods, spatial analysis, compartmental models, survival models and advanced regression frameworks to enhance performance and inference. However, such integrative practices were relatively uncommon, observed in 45 studies (9.8%) of the reviewed literature.

Table 4.

Distribution of SEM and CM types and SEM approaches across included studies

Model Type/Approach N Frequency (%)
SEM Type Standard SEM 254 67.0
Path Analysis 112 29.6
Cross-lagged SEM 6 1.6
Multilevel SEM 5 1.3
Multigroup SEM 2 0.5
CM Type Matching & Weighting 40 50.0
Potential Outcomes 19 23.8
Graphical Causal Models 6 7.5
Mendelian Randomization 6 7.5
Dynamic Systems 5 6.3
Bayesian Methods 4 5.0
SEM Approach CB-SEM 317 83.6
PLS-SEM 28 7.4
GSEM 18 4.7
BSEM 13 3.4
BGSEM 1 0.3
CB-SEM, PLS-SEM 1 0.3
Piecewise SEM 1 0.3

When examining the specific SEM approaches (Table 4), covariance-based SEM (CB-SEM) clearly dominated the literature, followed by partial least squares SEM (PLS-SEM). Other estimation approaches, such as generalized SEM and Bayesian SEM, were applied less frequently, while hybrid or mixed approaches were extremely rare. Overall, this distribution reflects the predominance of commonly used estimation techniques reported in the SEM literature.

RQ3: Which types of diseases and health conditions are most frequently investigated using these approaches, and what are the main determinants examined within each disease or group of diseases?

Types of diseases and health conditions most frequently investigated using SEM and CM

SEM and CM have been applied across a broad spectrum of disease categories. Across the studies included in this systematic review, a total of 592 disease occurrences were identified and grouped into four major categories: neuropsychiatric disorders (32.1%), chronic diseases (30.1%), infectious diseases (24.0%), and other health-related topics (13.9%). The distribution of these categories is shown in Fig. 5. Overall, neuropsychiatric disorders emerged as the most frequently investigated category, with studies predominantly focusing on mental and emotional disorders, followed by neurocognitive disorders. Regarding chronic diseases, this category represented the second most common group. The most frequently examined conditions included cardiovascular diseases, diabetes, and cancer, followed by respiratory diseases, obesity, and musculoskeletal disorders. Other chronic conditions such as renal diseases, malnutrition, and anaemia were less frequently studied. With respect to infectious diseases, COVID-19 clearly dominated the literature, followed by influenza, HIV, malaria, tuberculosis, dengue, acute diarrhoeal diseases, dental caries, and hepatitis. In contrast, other infectious diseases remained underrepresented, likely due to challenges related to data availability and quality, irregular outbreak patterns, smaller sample sizes, and the inherent complexity of modeling infectious disease transmission dynamics. This underrepresentation highlights a substantial gap in the application of SEM and CM to infectious disease epidemiology, particularly for endemic diseases such as malaria, which continue to impose significant public health burdens. Beyond disease-specific applications, other health-related topics illustrate the methodological versatility of SEM and CM. Studies in this category frequently addressed health behaviors and healthcare utilization, maternal and child health, biomarkers and biological processes, ageing and disability, wellbeing, and other health indicators. These applications demonstrate that SEM and CM are not limited to specific disease outcomes but can also be used to investigate broader health processes and determinants.

Fig. 5.

Fig. 5

Diseases studied using SEM and CM within each health category ((n ≥ 3))

Main determinants examined within each disease or group of diseases

Figure 6 presents the five most frequently examined determinant categories across disease groups. Overall, sociodemographic determinants were the most commonly investigated category in both mental and emotional disorders and neurocognitive disorders (Fig. 6a). In mental and emotional disorders, behavioral determinants ranked second, followed closely by psychological determinants, while clinical and environmental determinants were less frequently represented. A similar pattern was observed for neurocognitive disorders, where behavioral determinants also ranked second, followed by clinical and socioeconomic determinants. These findings underscore the central role of sociodemographic characteristics, such as age, sex, and education, in studies applying CM and SEM to neuropsychiatric outcomes.

Fig. 6.

Fig. 6

Top determinants across disease categories: a neuropsychiatric, b infectious, c other, d chronic

For chronic diseases (Fig. 6b), determinant profiles varied more substantially across conditions. Behavioral determinants were among the most frequently examined categories in cardiovascular diseases, diabetes, musculoskeletal disorders, and obesity, often alongside sociodemographic determinants. By contrast, sociodemographic determinants were most prominent in anaemia, cancer, renal and urinary diseases, and respiratory diseases. Clinical and socioeconomic determinants also featured prominently in several chronic conditions, whereas environmental determinants generally played a more limited role, except in diseases such as respiratory disorders. Malnutrition showed a distinct pattern, with environmental, sociodemographic, and socioeconomic determinants among the leading categories. This distribution is consistent with the established influence of living conditions and socioeconomic disadvantage on nutritional outcomes. The variation observed across chronic diseases likely reflects differences in disease etiology, data availability, and the extent to which certain determinants are considered modifiable in prevention and intervention research.

Infectious diseases (Fig. 6c) likewise displayed heterogeneous determinant profiles. Sociodemographic determinants were the most frequently examined category in several conditions, including COVID-19, influenza, malaria, and tuberculosis. Behavioral determinants were also highly represented and ranked first in HIV, dental caries, and hepatitis, highlighting their importance in transmission and prevention. Environmental determinants were particularly prominent in acute diarrhoea and dengue, where environmental exposures and vector ecology are central to disease dynamics. Clinical and structural determinants were less commonly represented among the leading categories, whereas psychological determinants appeared in several studies on COVID-19, HIV, influenza, and tuberculosis, suggesting growing recognition of psychosocial dimensions in infectious disease epidemiology. These patterns reflect the multifactorial nature of infectious diseases and the fact that the relative importance of determinants depends on transmission pathways, available data, and the specific epidemiological questions being addressed.

Other health-related topics (Fig. 6d) showed a broadly similar pattern, with sociodemographic determinants consistently dominating, followed by behavioral determinants. Clinical, psychological, environmental, and socioeconomic determinants contributed smaller yet relatively comparable proportions. This suggests that social and demographic characteristics remain central even in broader health-related outcomes beyond specific disease categories, although the relative contribution of other determinants varied according to the outcome examined and the conceptual framing of the studies.

Taken together, these results reveal both recurring patterns and important variation in the determinants examined across disease groups. Sociodemographic and behavioral determinants were the most consistently represented categories, reflecting both their epidemiological relevance and their routine availability in datasets. The prominence of other determinants varied by disease group. Psychological determinants were especially visible in neuropsychiatric conditions and some infectious diseases, clinical determinants were more prominent in chronic disease studies, and environmental determinants featured more strongly in conditions such as respiratory diseases, dengue, acute diarrhoeal diseases, and malnutrition, where contextual exposures play a central etiological role. Although socioeconomic determinants appeared across multiple disease groups, they were less consistently represented among the leading categories, suggesting that broader structural influences may still be underexamined in applications of SEM and CM in epidemiology. In addition to the main determinant categories identified here, other factors, including psychosocial, anthropometric, cognitive, genetic, policy-related, biological, and immunological determinants, appeared only sporadically. Broadening the range of determinants incorporated into these models may help better capture the complex, multilevel processes shaping health outcomes and strengthen the explanatory value of epidemiological research.

RQ4: What are the main types of variables analyzed and how are their effects conceptualized?

This review revealed that both qualitative and quantitative variables were examined in studies applying SEM and CM approaches in epidemiology (Fig. 7b), with no major differences in their distribution across the two approaches. Quantitative variables slightly predominated in both SEM (55.6%) and CM (54.8%) studies, while qualitative variables accounted for 44.4% and 45.2%, respectively. This balanced use of variable types indicates an effort to capture both measurable outcomes and categorical factors. Nevertheless, the slight predominance of quantitative variables in SEM may reflect methodological preferences, as latent constructs are more easily operationalized with continuous measures, potentially limiting the exploration of complex categorical phenomena. Across the reviewed studies, five main types of effects were reported (Fig. 7a): direct, indirect, mediation, moderation, and total effects. Direct effects were the most frequently examined in both SEM and CM studies, followed by indirect effects and mediation effects. Moderation effects appeared less frequently overall but were relatively more common in CM studies, whereas total effects were reported only occasionally in both approaches. These patterns suggest that SEM is primarily employed to capture complex indirect and mediation pathways, reflecting its strength in modeling latent constructs and multistep causal mechanisms, whereas CM emphasizes direct and moderation effects, aligning with its utility in testing explicit causal relationships and interactions. The relative underrepresentation of total and moderation effects across both approaches highlights gaps in the current literature and points to opportunities for future studies to systematically explore these effect types, thereby enhancing our understanding of the multifaceted nature of epidemiological relationships.

Fig. 7.

Fig. 7

Distribution of (a) variable effect and (b) variable types by model (SEM vs CM)

RQ5: What types of data, data sources, sample sizes, and geographic regions are typically used in SEM and CM studies?

Types of data, data source, and specific datasets used in SEM and CM studies

The studies included in this review identified four main types of data used in SEM and CM within epidemiology (Fig. 8a): cross-sectional, longitudinal, panel, and time series data. Cross-sectional and longitudinal designs predominated across both approaches, whereas panel and time series data were only marginally represented. Specifically, SEM studies relied overwhelmingly on cross-sectional data, with far fewer studies using longitudinal designs and only rare applications of panel data; time series data were virtually absent. This reliance on cross-sectional designs, while facilitating model estimation and interpretation, provide fewer opportunities to examine temporal ordering and changes over time compared with longitudinal designs. Consequently, the limited use of longitudinal designs constrains the potential for more advanced temporal and dynamic modeling approaches in epidemiological SEM applications. In contrast, CM studies relied more heavily on longitudinal data, which constituted the most common design, followed by cross-sectional data. Panel data were used less frequently, and time series designs remained marginal. This stronger reliance on longitudinal structures is consistent with the emphasis of CM frameworks on temporality and causal pathways. Overall, the relatively limited use of panel and time series data across both approaches suggests that opportunities remain to better exploit data structures capable of capturing dynamic and repeated-measures processes in epidemiological research.

Fig. 8.

Fig. 8

Distribution of (a) data types and (b) data sources used in SEM and CM studies

Regarding data sources, three main categories were identified: primary, secondary, and simulation data (Fig. 8b). SEM studies showed a clear predominance of primary data, whereas secondary data represented a smaller proportion and simulation data were almost absent. In contrast, CM studies displayed a more balanced distribution between primary and secondary data sources, while simulation approaches remained relatively uncommon. Primary data were typically collected through survey questionnaires or direct patient assessments, whereas secondary data originated from national health databases, institutional records, international organisations, and publicly available repositories (see Table 11 in Appendix 3). The heavy reliance on primary data in SEM reflects its historical roots in psychology and social sciences, where survey-based data collection is commonly used. Meanwhile, the greater engagement with secondary and simulation sources in CM highlights its integration with computational and data-driven traditions, offering more scalable applications. Nonetheless, the limited uptake of simulation data across both frameworks signals an underexploited avenue for methodological innovation, particularly for testing counterfactuals and evaluating policy-relevant interventions.

Sample size used accross SEM and CM studies

In terms of sample sizes, substantial variation was observed across SEM and CM studies (Fig. 9). Overall, CB-SEM predominated across all sample size categories within SEM approaches (Fig. 9a). Among studies using large sample sizes (≥ 200 observations), CB-SEM was the most frequent approach, followed by PLS-SEM, GSEM, BSEM, and BGSEM. A similar pattern was observed for studies with marginal sample sizes (100–199 observations), where CB-SEM remained dominant, while Piecewise SEM, BSEM, and PLS-SEM were only marginally represented. Among studies with inadequate sample sizes (<100 observations), CB-SEM also remained the most common approach, followed by GSEM, PLS-SEM, and BSEM.

Fig. 9.

Fig. 9

Sample size distributions: a across SEM approaches and b CM categories

Similarly, sample sizes in CM studies showed considerable variability across approaches (Fig. 9b). Among studies using large sample sizes (≥ 200 observations), Matching & Weighting was the most frequent approach, followed by Potential Outcomes, Mendelian randomization, Graphical Causal Models, Bayesian Methods, and Dynamic Systems. In contrast, studies with marginal sample sizes (100–199 observations) were represented only by Graphical Causal Models and Potential Outcomes. Among studies with inadequate sample sizes (<100 observations), Matching & Weighting and Potential Outcomes were the most frequent approaches, followed by Graphical Causal Models and Dynamic Systems. Overall, these patterns suggest that while some CM approaches tend to rely more on large datasets, others are applied across more diverse sample sizes, highlighting the importance of considering sample size when interpreting causal results.

Geographic regions represented in SEM and CM studies

The geographic distribution of SEM and CM studies, based on study setting, reveals a pronounced concentration in high-income countries (Fig. 10). The United States and China clearly dominate the literature, each contributing more than 51 publications. A second group of moderately represented countries includes Iran (21–50 publications). Countries with moderate research activity include Ethiopia, the United Kingdom, Japan, Brazil, Australia, and India, each contributing between 11 and 20 studies. Moreover, several countries fall into an intermediate category with 6 to 10 studies, including South Africa, Italy, Canada, Malaysia, Indonesia, Norway, the Philippines, Switzerland, France, Spain, and Denmark. In contrast, the majority of African nations (e.g., Benin, Nigeria, Ghana, and Malawi) were underrepresented, with most contributing fewer than six studies and often only one or none. Similar patterns were observed in parts of South and Southeast Asia (e.g., Pakistan, Bangladesh, and Vietnam) and Latin America (e.g., Mexico, Peru, and Chile), where several countries were represented by only a small number of studies. This uneven geographic representation highlights a substantial global imbalance in the application of SEM and CM methodologies, potentially limiting the generalisability of findings to low- and middle-income countries. The paucity of studies from these regions may reflect barriers such as limited access to high-quality datasets, resource constraints, or lower research capacity, underscoring the need for increased methodological dissemination and capacity building to ensure globally relevant epidemiological insights.

Fig. 10.

Fig. 10

Geographic distribution of SEM and CM studies in epidemiology

RQ6: What are the most commonly used software packages, statistical methods, and estimation techniques?

The analysis of the included studies revealed substantial variation in software usage, estimation techniques, and reporting practices between SEM and CM applications in epidemiology. For SEM, Analysis of Moment Structures (AMOS) emerged as the most frequently used software (34.1%), primarily for CB-SEM analyses, followed by Mplus (19.7%), Stata (14.4%), and R (13.1%). In contrast, Statistical Package for the Social Sciences (SPSS), accounted for 9.4% of CB-SEM applications, while Statistical Analysis System (SAS) and LISREL were less frequently used, each representing approximately 3.8% of studies. Other software such as Equations (EQS) and JASP appeared only sporadically. SmartPLS was almost exclusively used for PLS-SEM (86.2%), while BSEM analyses were mainly conducted using Mplus (33.3%) and R (33.3%), followed by AMOS (22.2%) and WinBUGS (11.1%). Less common approaches, such as BGSEM and Piecewise SEM, were implemented exclusively in R. This concentration on a small number of user-friendly platforms reflects not only accessibility and community familiarity, but also the historical influence of software such as LISREL, AMOS, and Mplus in popularising SEM. While this has facilitated uptake, it may have reinforced reliance on conventional tools at the expense of methodological diversity.

For CM studies, Matching and Weighting was the predominant approach, mainly implemented in R (65.6%), followed by SAS (18.8%) and Stata (12.5%). Analyses based on the potential outcomes framework primarily utilised R (73.3%), with SAS and Stata each accounting for 13.3% of studies. Less common CM approaches, including Bayesian methods and Mendelian randomisation, were implemented exclusively in R, while dynamic systems modeling relied mainly on Statistical Parametric Mapping (SPM) and R. Graphical causal models, which include approaches based on DAGs, were primarily implemented in Stata (50%), with additional use of R (25%) and Python (25%). This indicates that R was the most frequently reported software across CM studies.

Regarding estimation techniques (Table 5), SEM analyses were dominated by maximum likelihood, consistent with its standard use in structural models. However, ML relies on assumptions of multivariate normality and large sample sizes, often unmet in epidemiological data, raising concerns about model fit and robustness. Other methods showed moderate diversity, including partial least squares, Weighted Least Squares Mean and Variance adjusted (WLSMV), full information maximum likelihood (FIML), Bayesian estimation, and less frequently used approaches such as WLS, Diagonally Weighted Least Squares (DWLS), Robust Maximum Likelihood (MLR), and Ordinary Least Squares (OLS). A small number of studies (others) employed more specialised techniques, such as Cox regression, GLS, or g-estimation. Notably, 11.6% of SEM studies did not report their estimation method, underscoring limited transparency and reproducibility.

Table 5.

Distribution of estimation methods used in SEM and CM studies (n, %)

Estimation method SEM n (%) CM n (%)
ML 224 (65.3) 10 (29.4)
PLS 32 (9.3) –
WLSMV 23 (6.7) 2 (5.9)
FIML 19 (5.5) –
Bayesian 13 (3.8) 6 (17.6)
WLS 8 (2.3) 2 (5.9)
DWLS 7 (2.0) –
MLR 7 (2.0) –
OLS 2 (0.6) 3 (8.8)
IPW – 2 (5.9)
Others 8 (2.3) 9 (26.5)

CM studies demonstrated greater methodological heterogeneity and limited transparency, with 60.5% failing to specify the estimation approach. Among those reported, ML and Bayesian estimation were most common, followed by OLS, Inverse Probability Weighting (IPW), WLS, and WLSMV. Other approaches including Distributed Lag Non-linear Model (DLNM), g-computation, Inverse Probability of Censoring Weighting (IPCW), log-binomial regression, Poisson log-linear regression, and Vector Autoregression/Vector Error Correction Model (VAR/VECM) were applied only once or twice, indicating fragmented practices and lack of consensus on best methods.

Bootstrapping was employed in 13.4% of SEM studies and 16.46% of CM studies, often in combination with ML or hybrid approaches (e.g., IPW, g-computation). Its limited and inconsistently reported use suggests that opportunities to improve inference under non-normality remain underexploited.

Examining estimation choices by modeling approach (Fig. 11) revealed that each approach tends to rely on specific estimators. Bayesian estimation was used exclusively for BSEM and BGSEM, while ML and PLS were the primary estimators for GSEM and PLS-SEM, respectively. Piecewise SEM mainly applied Cox regression and GLM, whereas CB-SEM showed a dominance of ML. In CM, Bayesian methods were restricted to Bayesian applications, and WLS was applied in Mendelian randomisation studies. Dynamic systems models primarily relied on Bayesian estimation, followed by VAR/VECM. For the potential outcomes framework, ML was most frequently applied, while matching and weighting approaches exhibited varied distributions, with ML, WLSMV, and IPW being the most commonly used. These patterns suggest that estimation choices are largely guided by methodological conventions and the compatibility of specific estimators with the underlying model structures.

Fig. 11.

Fig. 11

Distribution of estimation methods used within each model: a SEM and b CM studies

Overall, findings indicate a dual pattern: reliance on a few historically influential and accessible software platforms, coupled with uneven adoption and underreporting of estimation methods.

RQ7: How are measurement models validated, and what criteria are used for evaluating model fit and reliability?

SEMs consist of two main components: measurement models and structural models. Validating the measurement model is therefore a critical prerequisite for evaluating structural relationships. Across the reviewed studies, measurement model validation primarily relied on Exploratory Factor Analysis (EFA) and Confirmatory Factor Analysis (CFA). Among the 379 SEM articles, 36 conducted an EFA, 171 applied a CFA directly, and 27 combined both approaches. In CM studies, exploratory analyses were also reported, albeit less frequently, with 6 EFA, 7 CFA, and 1 study using both i.e., EFA followed by CFA. The relatively limited application of combined EFA and CFA suggests that some studies may not have fully explored latent structures before confirming them, which could affect construct validity [64, 65]. However, it is equally important to note that proceeding directly with CFA can be entirely appropriate when models are strongly grounded in theory. The more critical concern lies with studies that did not report either EFA or CFA, as this raises questions about the adequacy of construct validation and the robustness of their conclusions.

Beyond factor analysis (EFA and CFA), various reliability and validity criteria were used to assess measurement quality, following methodological guidelines proposed by Hair et al. [39, 66], as summarized in Table 6. Out of the 379 SEM articles, 171 reported at least one criterion. The most frequently reported reliability and validity criteria were Cronbach’s alpha and factor loadings, followed by Composite Reliability (CR) and Average Variance Extracted (AVE). By contrast, the Heterotrait-Monotrait ratio (HTMT) and Variance Inflation Factor (VIF) were less commonly reported, while the Fornell–Larcker criterion, Squared Multiple Correlation (SMC), Intraclass Correlation Coefficient (ICC), internal consistency, and rho-A appeared only rarely. Most reported criteria exceeded established acceptability thresholds (>50%), following recommendations from Hair et al. [66]. However, the infrequent use of certain validity indices highlights heterogeneity in methodological rigor and reporting practices.

Table 6.

Reliability and validity criteria in SEM studies: proportion of studies reporting and meeting recommended thresholds

Reliability criteria Proportion reporting (%) Proportion meeting threshold (%)
Cronbach’s alpha 50.0 88.8
Factor loadings (λ) 48.3 65.1
CR 28.1 76.0
AVE 27.0 75.0
VIF 9.0 87.5
HTMT 5.1 77.8
Fornell-Larcker criterion 2.2 75.0
SMC 1.1 100.0
ICC 0.6 100.0
Internal consistency 0.6 100.0
rho-A 0.6 100.0

Once the measurement models were validated, the overall model quality was assessed using a range of fit indices that capture complementary dimensions of model evaluation. Absolute fit indices such as the Chi-square statistic, Root Mean Square Error of Approximation (RMSEA), Goodness-of-Fit Index (GFI), Adjusted Goodness-of-Fit Index (AGFI), and Standardized Root Mean Square Residual (SRMR) examine how well the hypothesized model reproduces the observed data. Incremental or comparative fit indices, including the Comparative Fit Index (CFI), Tucker-Lewis Index (TLI), Non-Normed Fit Index (NNFI), and Normed Fit Index (NFI), assess the model’s improvement relative to a null or baseline model. Parsimony indices, such as the Akaike Information Criterion (AIC), Bayesian Information Criterion (BIC), Consistent Akaike Information Criterion (CAIC), and Deviance Information Criterion (DIC), penalize excessive model complexity, balancing fit and theoretical simplicity. Finally, explained variance measures like the coefficient of determination (Inline graphic) quantify the proportion of variance in endogenous constructs accounted for by the model. Overall, this review reveals that reporting practices varied considerably across SEM families (Table 7). BSEM most frequently reported the CFI, followed by AIC, GFI, NFI, and RMSEA. CB-SEM studies mainly reported RMSEA and CFI, together with TLI, Chi-square and SRMR. GSEM favored CFI and RMSEA, followed by SRMR, AIC, and BIC, whereas PLS-SEM emphasized Inline graphic, followed by SRMR, NFI, Log-likelihood (LL), and CFI. Piecewise SEM was rarely applied, with only Chi-square reported in one study. These patterns indicate that different SEM approaches prioritize distinct aspects of model evaluation, including global model fit, model parsimony, or explained variance. Less frequently reported indices, such as CAIC, NNFI, AGFI, DIC, Percentile Bootstrap, and Posterior Predictive Checks, further reflect the methodological diversity and the limited standardization in model assessment practices.

Table 7.

Fit indices reported in SEM and CM studies

Approach Fit indices n (%)
BSEM CFI 4 (13.3)
AIC, GFI, NFI, RMSEA 2 (6.7)
CB-SEM RMSEA 236 (21.5)
CFI 230 (20.9)
TLI 129 (11.7)
Chi-square 78 (7.1)
SRMR 78 (7.1)
GSEM CFI, RMSEA 6 (17.1)
SRMR 3 (8.6)
AIC, BIC 2 (5.7)
PIECEWISESEM Chi-square 1 (100)
PLS-SEM RInline graphic 16 (24.6)
SRMR 12 (18.5)
NFI 7 (10.8)
LL 5 (7.7)
CFI 4 (6.2)
CM studies RMSEA 11 (16.2)
CFI 10 (14.7)
AIC, NFI, RInline graphic 6 (8.8)
Chi-square/df 5 (7.4)
GFI, LL, TLI 4 (5.9)
AGFI, BIC, Chi-square, SRMR 3 (4.4)

In CM studies (Table 7), the most commonly reported fit indices were RMSEA and CFI, followed by AIC, NFI, Inline graphic, and the Chi-square/df ratio. Less frequently used indices included LL and a group of traditional SEM measures such as AGFI, BIC, Chi-square, GFI, TLI, and SRMR. Beyond conventional fit indices, other evaluation methods were also applied, notably placebo tests (15.4%), as well as a diverse set of approaches including Bayesian Model Comparison, Posterior Predictive Checks, Balance Diagnostics, Bootstrapping, Permutation Tests, Bonferroni-adjusted Confidence Intervals, Posterior Distributions, Variational Free Energy, DAG-based assessments, and Dirichlet Process Mixture Priors (7.7%). The wide heterogeneity in reporting underscores the lack of consensus and standardised guidelines for evaluating model performance in CM, leading to inconsistent practices across studies. This variability highlights both the methodological flexibility of CM approaches and the need for clearer reporting standards to enhance comparability and reproducibility.

Taken together, the findings indicate that although a wide range of reliability criteria and fit indices is available, most SEM and CM applications rely on a narrow set of conventional measures. While the majority of reported indices satisfied accepted thresholds, the limited use of alternative indicators and the inconsistency in reporting practices raise concerns for transparency, reproducibility, and comparability. These issues highlight the importance of establishing more standardised guidelines for future applications.

RQ8: What are the reported methodological strengths and gaps associated with SEM and CM in epidemiology?

The application of SEM and CM in epidemiology has increasingly been recognised for its ability to rigorously evaluate complex causal relationships and inform public health decision-making. Among the methodological strengths, the most frequently reported was the use of validated instruments and theoretical models (34.5%), encompassing well-established measurement tools, psychometrically robust scales, and theoretical frameworks with confirmed construct validity via CFA and excellent model fit indices such as CFI, TLI, and RMSEA. This reliance on validated frameworks enhances internal validity and ensures that parameter estimates accurately reflect underlying latent constructs. Large and representative samples (16.6%) further strengthen statistical power and improve generalisability, allowing insights to extend across diverse populations and epidemiological contexts. Longitudinal and prospective study designs (11.5%) support stronger causal inference by capturing temporal sequences and trajectories, enabling the assessment of intervention effects or disease progression dynamics. Robust statistical methods (9.5%) including bootstrapping, weighted least squares, and robust maximum likelihood estimation improve resilience to violations of model assumptions and enhance confidence in parameter stability. In addition, innovative methodological approaches (8.8%) illustrate the integration of advanced statistical techniques or novel theoretical frameworks, such as combining latent growth models with SEM for nuanced analysis of time-varying exposures and outcomes. Other strengths reported in the literature include high model fit and explained variance (8.2%), reflecting the ability of SEM-based approaches to capture complex relationships between latent and observed variables. Several studies also highlighted novel and context-specific applications (5.6%), particularly in emerging public health contexts, as well as the explicit use of causal inference techniques (3.6%) and comprehensive adjustment for multiple covariates (1.8%).

Despite these strengths, SEM and CM applications in epidemiology continue to face critical methodological limitations that constrain their reliability and policy relevance. Study design limitations (17.4%) are a primary concern: the pervasive reliance on observational data without randomisation or adequate control groups limits causal inference and leaves analyses vulnerable to residual confounding. In epidemiology, exposures and outcomes are embedded within complex socio-environmental systems that cannot be disentangled in purely observational studies. For example, malaria transmission dynamics are shaped by interactions among climatic factors, human behaviours such as bed-net use, and population mobility, making it particularly challenging to separate causal effects. Such design limitations risk overstating or misrepresenting causal pathways, potentially leading to misguided policy decisions, and highlight the need for quasi-experimental designs, natural experiments, or well-planned longitudinal cohorts to strengthen causal inference.

Closely linked are issues of measurement error and instrument validity (15.8%). Self-reported data, common in epidemiology, are prone to recall and social desirability biases, which can distort latent constructs and bias parameter estimates. In malaria research, self-reported bed-net usage often overestimates adherence, weakening the estimated protective effect and risking ineffective interventions. Similarly, in chronic disease studies, self-reported dietary intake or physical activity is highly susceptible to misreporting, compromising the reliability of SEM constructs. Even validated instruments may lose psychometric robustness across populations, highlighting the importance of rigorous pretesting, cross-validation, and assessment of measurement invariance.

Causal inference challenges (15.6%) further complicate interpretation. Without longitudinal data, distinguishing genuine causal pathways from correlations influenced by unmeasured dynamics is challenging. In HIV research, poor mental health may be modelled as a predictor of treatment non-adherence, yet reverse causality is equally plausible, as deteriorating health may exacerbate psychological distress. Most studies remain cross-sectional, capturing only static snapshots of inherently dynamic phenomena such as disease progression, behaviour change, or intervention effects. This reliance on cross-sectional data limits both causal interpretation and predictive utility, reducing the relevance of findings for proactive public health strategies.

Another recurrent limitation is limited generalisability (12.9%). It arises because many findings are drawn from narrowly defined populations or specific ecological contexts, failing to account for demographic heterogeneity, regional variation, or broader environmental influences. For instance, malaria control strategies modelled on data from one West African region may not translate to areas with different vector species, housing conditions, or health system capacities. A further limitation concerns the lack of longitudinal data (11%), which constrains the ability to examine temporal sequences, delayed effects, or cumulative exposures. Unmeasured confounding (8.9%) further compounds these issues, as latent or unobserved variables can produce biased estimates even in sophisticated models, underscoring the importance of sensitivity analyses, proxy measures, and formal causal frameworks such as directed acyclic graphs or instrumental variable approaches.

Sample size and selection bias (8.3%) further weaken credibility: small, non-representative, or convenience-based samples reduce statistical power, inflate standard errors, and generate unstable estimates. Concerns over model fit and statistical reporting (5.2%) persist, with inconsistent criteria and selective index reporting limiting comparability and transparency. Finally, data access and quality limitations (4.9%) remain major barriers, particularly in low-resource settings. In malaria-endemic regions, incomplete entomological data or irregular health facility reporting can produce substantial gaps that undermine model accuracy and replication, further exacerbating global inequities in research capacity and evidence generation.

Taken together, these findings illustrate the dual nature of SEM and CM in epidemiology: although these methods provide powerful frameworks for integrating complex data and advancing causal understanding, their potential is frequently constrained by design limitations, measurement challenges, temporal ambiguity, generalisability constraints, unmeasured confounding, sampling issues, statistical inconsistencies, and data quality deficiencies. Fully realising the promise of SEM and CM for informing evidence-based public health decision-making requires systematic improvements, including robust longitudinal study designs, rigorous instrument validation, comprehensive confounding control, and transparent reporting practices.

RQ9: What future directions and recommendations are proposed to enhance the use of SEM and CM in public health research?

The reviewed studies outlined several directions to advance SEM and CM applications in epidemiology, though most remain aspirational rather than fully realised. Intervention development and testing (20.9%) emerged as the most frequently cited priority, underscoring the importance of translating analytical insights into concrete programmes or treatments. Yet, the scarcity of illustrative examples highlights a persistent gap between methodological potential and practical implementation. Recommendations for longitudinal and causal studies (18.7%) emphasised the need to strengthen temporal and causal inference; however, their limited use reflects enduring barriers in data availability and the feasibility of long-term study designs, especially in resource-constrained settings. Methodological and measurement improvements (17.5%) were consistently acknowledged as essential, though empirical validation remains scarce. Policy and regulation guidance (13.6%) and population or contextual expansion (10.3%) were also frequently mentioned, but their operationalisation has been inconsistent, with research still concentrated in narrow geographic or demographic contexts. Other directions such as applications to specific public health areas (10.3%), mechanistic and biological studies (5.6%), and multilevel or multidisciplinary approaches (3.1%) remain largely underexplored despite their potential to address complex health challenges. Overall, these findings reveal a critical tension between the theoretical promise of SEM and CM and their translation into practice. Moving the field forward will require not only methodological innovation but also sustained institutional investment and deliberate strategies to embed analytical insights into effective interventions and policy frameworks.

Discussion

Summary of findings

This systematic review provides an integrative synthesis of how SEM and CM have been applied in epidemiological research over the past two decades. By analyzing peer-reviewed studies, we identified key patterns, methodological practices, and thematic focuses in the use of these frameworks. Overall, applications of both SEM and CM have increased substantially since 2010, reflecting their relevance for addressing complex causal questions in public health. Despite this growth, several methodological gaps persist, including limited validation procedures, underutilization of longitudinal data, and uneven disease coverage, particularly for infectious diseases. These gaps shape both the questions researchers can address and the robustness of their findings.

Across the reviewed literature, studies focused on four main analytical themes: examining causal and mediation relationships, identifying determinants of health outcomes, validating theoretical models, and evaluating interventions or exposures [67–69]. Causal and mediation analyses predominated, particularly in SEM studies, reflecting the framework’s ability to test complex multivariate hypotheses and its origins in social sciences [67, 68]. By contrast, theoretical model development and intervention evaluation were underexplored. This suggests that many studies emphasize testing known causal pathways rather than refining conceptual frameworks or translating findings into practice [20, 69]. CM studies showed slightly more focus on theoretical validation and intervention assessment, likely due to their flexibility in handling diverse data types and counterfactual scenarios [70].

Turning to the analytical frameworks themselves, most SEM studies relied on standard approaches such as path analysis and latent variable modeling [71, 72]. Prior reviews confirm that path analysis accounts for a large proportion of SEM applications across journals, largely because of its simplicity and ease of implementation [71, 72]. Advanced SEM techniques, including multilevel, cross-lagged, or BSEM, were seldom applied [73]. CM applications primarily employed regression-based and counterfactual approaches, with limited integration between SEM and CM, highlighting ongoing conceptual and computational challenges [74, 75]. These methodological patterns help explain both the strengths and limitations observed in study outcomes.

The epidemiological focus of studies was uneven. Neuropsychiatric and chronic diseases, particularly mental health and metabolic disorders, were studied most frequently, whereas infectious diseases such as malaria, HIV, and tuberculosis were underrepresented [76–78]. This pattern reflects SEM’s social science origins [20, 45] and the complexity of modeling infectious disease dynamics. The focus on certain diseases also shaped determinant selection. Sociodemographic and behavioral factors were prioritized, while clinical, environmental, and policy-related variables received limited attention [79, 80]. These epidemiological patterns naturally connect to the methodological choices researchers made in study design.

Methodologically, SEM studies mostly used cross-sectional designs with primary data. CM studies drew on a more diverse array of longitudinal, secondary, and simulated datasets [81, 82]. Sample sizes were generally adequate but rarely justified through formal power analyses [68, 83]. SEM software usage concentrated on AMOS, Mplus, and SmartPLS, whereas CM relied primarily on R [47, 84]. Maximum Likelihood estimation was the most common approach, despite limitations for non-normal or small-sample data [85, 86]. Alternative estimation techniques, such as GLS, ULS, or Bayesian procedures, were seldom applied [68, 85]. These choices directly influence both the rigor of measurement validation and the reliability of causal inferences.

Measurement model validation primarily relied on CFA, supported by classical reliability indices including Cronbach’s alpha, CR, and AVE [87–89]. Model fit was most commonly assessed with RMSEA, CFI, and TLI [45, 90], although reporting standards varied. Strengths included the ability to model complex causal structures and latent constructs. Limitations involved insufficient validation, underreporting of estimation procedures, and predominance of cross-sectional designs. These methodological patterns help explain the variability in study reproducibility and interpretability.

Overall, SEM and CM are increasingly applied to model causal and structural relationships in epidemiology. Coverage is expanding across disease domains, determinants, and study designs. Yet methodological practices remain heterogeneous. Certain analytical approaches particularly intervention evaluation, theoretical model refinement, and integration of SEM and CM remain underutilized. This summary naturally leads to a detailed discussion of methodological trends, which are critical for improving rigor, reproducibility, and interpretability in epidemiological research.

Methodological trends

Over the past two decades, methodological practices in SEM and CM have followed a dual trajectory: analytical diversification and persistence of conventional foundations. SEM research has grown steadily since the mid-2000s, driven by outlets such as the Structural Equation Modeling journal [20, 91]. Peaks in 2018–2019 and 2021–2022 corresponded to the emergence of complex datasets and the need for advanced modeling during global health crises such as COVID-19 [92, 93]. In contrast, CM adoption introduced to epidemiology around 2009 has evolved more gradually, reflecting incremental engagement with graphical and counterfactual inference frameworks [33, 94]. Despite this conceptual overlap, integration between SEM and CM remains rare. Embedding DAGs into SEM nonetheless represents a promising step toward hybrid frameworks that enhance causal transparency [74].

Building on these developments, SEM applications still rely heavily on classical path analysis and latent variable modeling. Advanced configurations such as multilevel, longitudinal, cross-lagged, or multigroup SEM remain underused, indicating limited adoption of more complex modeling approaches [71, 72]. Methodological diversification has gradually emerged through three paradigms: covariance-based (CB-SEM), variance-based (PLS-SEM), and Bayesian SEM (BSEM). Each differs in sample requirements, distributional assumptions, and model complexity [68, 89]. In parallel, CM studies follow a distinct yet complementary trajectory, typically employing regression-based estimators, inverse probability weighting, and counterfactual mediation analysis. Recent innovations include Mendelian randomization and Bayesian networks [70, 95]. Together, these trends reveal an uneven methodological landscape, where traditional techniques dominate and computational advances spread slowly.

At the data level, structural differences further distinguish both frameworks. SEM mainly relies on cross-sectional datasets, which capture variables at a single point in time and are commonly used due to their simplicity in estimation, while longitudinal designs were less frequently reported in the included studies. CM, by contrast, more often employs longitudinal or time-series data, enabling explicit causal ordering and feedback analysis [81, 96]. Despite this advantage, the use of simulated data remains rare even though such approaches could enhance robustness and generalizability [82]. This imbalance highlights a persistent tension between feasibility and causal depth.

A similar contrast emerges in variable specification. SEM emphasizes latent constructs and mediation processes, estimating direct, indirect, and total effects. CM focuses instead on direct and moderated effects consistent with counterfactual logic [97, 98]. Quantitative indicators still dominate in both frameworks, yet the inclusion of categorical and ordinal variables is increasing, suggesting a shift toward more context sensitive modeling [89, 99].

In terms of data adequacy, sample size and estimation strategies remain central to model reliability. Most SEM studies meet conventional thresholds (≥200 observations) but seldom provide formal power justifications [83, 100]. When smaller samples are used, bootstrapping or observation-to-parameter ratios (5:1 or 10:1) are applied [45, 101]. Monte Carlo simulations are also increasingly used to assess power under non-normality [100]. Sample requirements vary across paradigms: PLS-SEM accommodates smaller samples, CB-SEM demands larger datasets, and BSEM offers flexibility through prior information [68, 89].

Regarding estimation, ML remains dominant and is often paired with bootstrapping to obtain robust standard errors [68, 85]. Although efficient under multivariate normality, its assumptions are seldom met in epidemiological data [86, 102]. Several alternatives have been proposed. GLS minimizes a weighted residual function and provides consistent and efficient estimates when variances and covariances are correctly specified [85]. Unweighted Least Squares (ULS) removes distributional assumptions by applying an identity weight matrix, offering robustness for ordinal data though with lower efficiency [85]. Asymptotically Distribution-Free (ADF), also known as WLS, further relaxes normality assumptions but requires very large samples [85]. Robust techniques suitable for ordinal or categorical data including DWLS, WLSM, WLSMV, BSEM, and PLS-SEM remain underutilized despite their advantages [68]. CM analyses, by contrast, display broader estimator diversity, encompassing ML, inverse probability weighting, log-binomial and Poisson regression, and time-series models such as VAR and VECM [95].

Software preferences closely reflect these methodological patterns. SEM analyses are most frequently conducted in AMOS, Mplus, SmartPLS, LISREL, and R, each optimized for distinct analytical purposes: LISREL for covariance modeling, AMOS for intuitive interfaces and missing data handling, Equations (EQS) for addressing non-normality, and SmartPLS for variance-based estimation [47, 103, 104]. In contrast, CM studies rely predominantly on R, an open-source environment that supports a wide range of statistical and causal inference tools [84].

Once models are estimated, validation practices vary in rigor. CFA remains dominant, while EFA is rarely combined, and some studies omit factor analysis altogether [68, 105]. Reliability and validity are typically assessed through factor loadings, Cronbach’s alpha, CR, and AVE, adopting thresholds of α>0.70, CR > 0.70, and AVE > 0.50 [88, 89]. Substantial factor loadings support convergent validity; however, their significance depends on sample size, with smaller samples requiring higher loadings (see Appendix 2, Table 10) [39]. Advanced indices such as HTMT, the Fornell–Larcker criterion, and VIF remain underused, leading to inconsistent reporting standards.

Following validation, model adequacy is generally assessed using absolute, incremental, and parsimonious fit indices. Absolute indices (Inline graphic, RMSEA, GFI, SRMR) evaluate congruence between observed and predicted covariance structures. Incremental indices (CFI, TLI/NNFI, NFI, IFI, RFI) assess improvement over a null baseline [45, 90]. Parsimonious indices (PNFI, PCFI, PGFI) account for model simplicity and theoretical parsimony [106]. Bayesian and variance-based frameworks instead rely on BIC, posterior predictive checks, Inline graphic, and SRMR [73, 107]. Although commonly used threshold values are summarized in Table 9, these cutoffs remain informal guidelines, with no universal consensus on acceptable levels [39, 106, 108–110].

Table 9.

Required sample size for significance by factor loading

Factor loading Sample size needed
0.30 350
0.35 250
0.40 200
0.45 150
0.50 120
0.55 100
0.60 85
0.65 70
0.70 60
0.75 50

Overall, these converging trends illustrate a field in methodological transition consolidating established conventions while gradually embracing innovation. Despite advances in analytical sophistication, continued reliance on conventional designs, default estimators, and uneven validation practices shows that methodological evolution, though evident, remains incomplete.

Gaps and challenges

The use of SEM and CM in epidemiology has expanded in recent years. However, important conceptual and methodological gaps persist. Current research remains largely focused on neuropsychiatric and chronic diseases, while infectious diseases such as malaria, HIV, hepatitis, and tuberculosis are underrepresented [76–78]. This imbalance reflects SEM’s social science origins [20, 45] and the complexity of modeling infectious disease dynamics involving environmental, behavioral, and socioeconomic determinants [111, 112]. Limited data availability and research investment in low and middle-income countries further contribute to this gap [113].

Similar patterns emerge in the selection of determinants. Studies predominantly examine sociodemographic and behavioral factors, whereas clinical, environmental, and policy-related variables are rarely addressed [79, 80]. Differences in data sources reinforce this imbalance. SEM mostly uses primary data, while CM integrates more secondary and simulated data [81]. Yet, simulation based and counterfactual designs remain scarce, restricting exploration of hypothetical interventions [82].

Methodologically, insufficient sample size justification and lack of formal power analysis undermine transparency [68]. While SEM studies generally meet recommended thresholds [83], CM applications lack standardized criteria [8, 33]. Estimation practices heavily depend on ML, often used without verifying assumptions of multivariate normality [47, 85], whereas alternative estimators such as WLSMV, GLS, ULS, ADF, etc., remain underused [68].

Measurement model validation also shows inconsistencies. EFA are rarely combined with CFA, and reliability or validity indices (CR, AVE, HTMT) are inconsistently reported [39, 88, 114]. Structural model evaluation often omits fit indices or applies them inconsistently [45, 90], while path analysis remains widely used among SEM applications [71, 72].

Finally, poor methodological reporting regarding estimation techniques, data characteristics, missing data handling, and fit criteria remains widespread [68]. Collectively, these weaknesses including narrow disease focus, determinant imbalance, inadequate validation, and inconsistent reporting constrain the reliability, generalizability, and causal interpretability of SEM and CM in epidemiological research.

Strengths and limitations

This systematic and critical review presents several important strengths. First, it was conducted in accordance with PRISMA guidelines, ensuring a transparent and reproducible approach to study identification, screening, and reporting. The comprehensive search across multiple databases enabled the inclusion of a large and diverse set of studies (n = 458), capturing a wide range of epidemiological applications of SEM and CM. Beyond mapping the literature, this review provides a structured and critical appraisal of methodological patterns, highlighting how these approaches are implemented, the assumptions underlying their use, and recurring methodological gaps. This synthesis therefore contributes to a clearer understanding of the current methodological landscape and offers guidance for future epidemiological research.

Nevertheless, several limitations should be acknowledged. Despite an extensive search strategy, relevant studies may have been missed due to database coverage limitations, publication bias, or the exclusion of grey literature. In addition, although the search strategy aimed to capture epidemiological applications broadly, some disease areas may be unevenly represented depending on indexing practices in bibliographic databases. The substantial heterogeneity in study designs, populations, and analytical frameworks further limited the feasibility of quantitative synthesis and required a narrative and interpretative approach. Moreover, several steps in the review process involved a degree of judgement, particularly in the categorisation of diseases, determinants, and methodological approaches, which may introduce some subjectivity and should be considered when interpreting cross-category comparisons. The review protocol was not prospectively registered prior to study conduct, which may limit transparency regarding a priori methodological decisions. Importantly, no formal methodological quality or risk-of-bias assessment was conducted, meaning that the findings reflect patterns in reporting rather than differences in methodological rigor or robustness. In this context, observed frequencies, including publication patterns, software usage, and estimation methods, should not be interpreted as indicators of methodological superiority, appropriateness, or scientific importance, but solely as descriptive characteristics of the included literature. Finally, incomplete reporting in primary studies may have influenced the assessment of methodological features.

Recommendations for future research

Future research should broaden the application of SEM and CM beyond their predominant focus on neuropsychiatric and chronic diseases to include infectious diseases, particularly in low and middle income countries. Such expansion would capture the complex interactions among environmental, behavioral, and socioeconomic determinants that characterize infectious disease dynamics and strengthen the global relevance of epidemiological modeling.

Methodologically, future studies should adopt longitudinal and simulation based designs, integrate complementary approaches such as machine learning and spatial modeling, and combine SEM and CM to enhance causal inference and explanatory depth. Greater transparency in methodological reporting especially regarding sample size justification, estimation techniques, and model validation is also essential. The use of standardized power analyses and open-source tools would improve reproducibility, while interdisciplinary collaboration between epidemiologists, statisticians, and computational scientists can drive innovation and support evidence-based public health decision-making.

Conclusion

Structural Equation Modeling and Causal Modeling have emerged as essential tools for investigating complex causal relationships in epidemiology. This review systematically and critically assessed their application, emphasizing methodological characteristics, disease focus, and reporting practices. The analysis revealed a predominant focus on neuropsychiatric and chronic conditions, while infectious diseases such as malaria, HIV, and tuberculosis, despite their substantial global burden remain underexplored. Methodologically, most studies relied on large samples, cross-sectional designs, and conventional estimation techniques, yet frequently lacked transparency in reporting measurement validity and model fit, thereby limiting reproducibility and causal interpretability. Nevertheless, the increasing integration of causal inference principles and diversification of analytical approaches reflect a positive methodological evolution toward greater rigor and transparency in public health research. SEM and CM each offer distinct strengths: SEM excels in modeling latent constructs and mediation pathways, whereas CM provides an explicit framework for causal reasoning. Future research should enhance methodological transparency, adopt longitudinal and multilevel designs, and broaden the use of these frameworks to infectious diseases such as malaria, where complex socio-environmental interactions demand integrative analytical perspectives. Strengthening the methodological interface between SEM and CM will advance causal inference, reinforce evidence-based policymaking, and support the development of more effective and context-sensitive public health strategies.

Supplementary Information

Supplementary Material 1 (65.1KB, xlsx)

Acknowledgements

Not applicable.

Abbreviations

ADF

Asymptotically Distribution-Free

AGFI

Adjusted Goodness-of-Fit Index

AIC

Akaike Information Criterion

AMOS

Analysis of Moment Structures

ANOVA

Analysis of Variance

AVE

Average Variance Extracted

BIC

Bayesian Information Criterion

BGSEM

Bayesian Generalised Structural Equation Modeling

BSEM

Bayesian Structural Equation Modeling

CAIC

Consistent Akaike Information Criterion

CFA

Confirmatory Factor Analysis

CFI

Comparative Fit Index

CM

Causal Modeling

CB-SEM

Covariance-based Structural Equation Modeling

CVI/CVR

Content Validity Index / Content Validity Ratio

CR

Composite Reliability

DAGs

Directed Acyclic Graphs

DIC

Deviance Information Criterion

DLNM

Distributed Lag Non-linear Model

DOAJ

Directory of Open Access Journals

DWLS

Diagonally Weighted Least Squares

EFA

Exploratory Factor Analysis

FIML

Full Information Maximum Likelihood

GFI

Goodness-of-Fit Index

GLS

Generalized Least Squares

GSEM

Generalised Structural Equation Modeling

HIV

Human Immunodeficiency Virus

HTMT

Heterotrait–Monotrait Ratio

ICC

Intraclass Correlation Coefficient

IPCW

Inverse Probability of Censoring Weighting

IPW

Inverse Probability Weighting

LISREL

Linear Structural Relationships

ML

Maximum Likelihood

MLR

Robust Maximum Likelihood

MIMIC

Multiple Indicators Multiple Causes

NFI

Normed Fit Index

NNFI

Non-Normed Fit Index

PLS-SEM

Partial Least Squares Structural Equation Modeling

RMSEA

Root Mean Square Error of Approximation

SAS

Statistical Analysis System

SCMs

Structural Causal Models

SEM

Structural Equation Modeling

SMC

Squared Multiple Correlation

SPM

Statistical Parametric Mapping

SPSS

Statistical Package for the Social Sciences

SRMR

Standardized Root Mean Square Residual

TLI

Tucker–Lewis Index

TMLE

Targeted Maximum Likelihood Estimation

ULS

Unweighted Least Squares

VAR

Vector Autoregression

VECM

Vector Error Correction Model

VIF

Variance Inflation Factor

WLS

Weighted Least Squares

WLSMV

Weighted Least Squares Mean and Variance Adjusted

Appendix 1

Fit indices and recommended thresholds

Table 8.

Fit indices and their acceptable threshold levels

Fit index Acceptable threshold Levels
Inline graphic Low value relative to degrees of freedom with an insignificant p value (p > 0.05)
Inline graphic < 3 good; < 5 permissible
RMSEA < 0.07 good; 0.07–0.10 moderate; > 0.10 bad
p value (PCLOSE) > 0.50
GFI > 0.95 good; > 0.90 acceptable
RMR Good models have small RMR
SRMR < 0.09
NFI > 0.90 acceptable; > 0.95 good
NNFI (TLI) > 0.95 good; > 0.80 acceptable
CFI > 0.90 acceptable; > 0.95 good
PNFI > 0.50 acceptable

Appendix 2 

Required sample size for significant factor loadings

Appendix 3 

Data sources used in this review and their access links

Table 10.

Data sources used in this review and their access links (if available)

N Reference Sources Links
1 Yuming Guo et al., 2023 National Hospitalisation And Air Pollution Monitoring
2 Reece AS, Hulse GK (2021) Mendeley Data Repository https://data.mendeley.com/
3 VesnaBarros et al., 2022 World Bank Open Data https://data.worldbank.org/
4 EhsanYaghmae et al., 2023 Real-World (Crwd), Iqvia, Crwd
5 Dzandu MD et al., 2022 GWAS Consortia https://osf.io/
6 Kwan et al., 2024 Kpnc Electronic Health
7 de Bont J et al., 2024 GWAS Consortia
8 Lin et al., 2021 World Bank Open Data https://data.worldbank.org/
9 Shirkou Jaafari et al., 2020 Air Pollution Monitoring, Mortality, Green Space Metrics Derived From Satellite Imagery
10 Khavarian-Garmsir AR et al., 2021 Covid-19 Cases/Deaths: Ac-19 Mobile App (Iran Ministry Of Health); Socioeconomic Variables: Statistical Yearbook Of Tehran, Population And Housing Census, Atlas Of Tehran Metropolis
11 Sydney A. Martinez et al., 2018 GWAS Consortia
12 Vasantha et al., 2021 World Health Organization (WHO) https://www.who.int/data
13 Tiglao NC et al., 2025 GWAS Consortia
14 David Opeoluwa Oyewola et al., 2021 Kaggle https://www.kaggle.com/datasets
15 Reece and Hulse, 2021 GWAS Consortia
16 Mitchell and Knowlton, 2011 Primary (From Ark)
17 Auerbach et al., 2018 GAW20 Simulations
18 Zhang CY, Zhang A, 2019 Us Epa/Environmental
19 Wang Y et al., 2022 Chinese Center for Disease Control and Prevention (HFRS)
20 Svanes C et al., 2022 National Registry
21 Hao Z et al., 2024 Us Epa/Environmental
22 Harford TC et al., 2013 National Registry
23 Kendler & Gardner, 2010 National Registry
24 Bissilimou Rachidatou Orounla et al., 2024 World Health Organization (WHO) https://www.who.int/data
25 Aitken Z et al., 2017 Us Epa/Environmental
26 Bahru, B. A. et al., 2019 Young Lives Cohort Study
27 Friston et al., 2020 Public Health
28 Sokolow SH et al., 2022 World Bank Open Data https://data.worldbank.org/
29 Zhang J et al., 2021 Google Covid-19 Mobility Reports https://www.google.com/covid19/mobility/
30 Rezvani et al., 2022 Pad-Tegecoach Trial Baseline
31 “Egloff et al., 2018” Basel Early Detection Of Psychosis (Fepsy)
32 Choi et al., 2023 National Registry
33 Lee CTC et al., 2021 National Registry
34 Dejan Loncar et al., 2023 Unaids, Global Aids Monitoring, Other International S https://aidsinfo.unaids.org/
35 Ma et al., 2015 Primary (Face-To-Face And Bilingual-Administered Questionnaire)
36 Bandara S et al., 2024 Official Regional (Covid-19 Incidence), Philippine Authority, Regional Meteorological, Wash Coverage
37 Awan N et al., 2020 US National Trauma Data Bank & TBI Model Systems
38 Do DP et al., 2013 Panel Study of Income Dynamics (PSID) https://psidonline.isr.umich.edu/
39 Michel Chavance et al., 2010 FLVS II Longitudinal Study
40 Pala D et al., 2025 NHANES (US National Health and Nutrition Examination Survey) https://www.cdc.gov/nchs/nhanes/
41 Pizon MG et al., 2021 Us Epa/Environmental
42 Prado EL et al., 2019 GWAS Consortia https://gwas.mrcieu.ac.uk/
43 Mao JJ et al., 2023 GWAS Consortia
44 Li H et al., 2022 Tencent Real-Time Epidemic Tracking, Baidu Migration, Baidu Index (Search, Media), Weather Forecast Sites, Global Weather Precision Forecast, National Economic And Social Development Bulletin, National Bureau Of, Online Air Quality Platforms
45 Mahroum N et al., 2018 Online Sources: Google Trends, Twitter, Wikipedia, Google News, Pubmed/Medline, National Surveillance
46 Alang SM, 2014 National Of American Life (Nsal, 2001–2003, Public Use), Icpsr
47 Tang Z et al., 2024 National Registry
48 García Martínez P et al., 2020 National Epidemiologic On Alcohol And Related Conditions (Nesarc), Waves 1 And 2
49 Marb A et al., 2025 Us Epa/Environmental
50 Nguyen PH et al., 2019 National Family Health (Nfhs-4), 2015–16, Representative Sample, N=60,096 Primiparous Women With Births In Last 5 Years
51 Yu HW et al., 2019 Taiwan Longitudinal Study on Aging (TLSA) https://www.bhp.doh.gov.tw/BHPnet/Portal/Them_List.aspx?menu=2300&no=200501250001
52 Singh-Manoux A et al., 2005 Whitehall II Cohort Study https://www.ucl.ac.uk/epidemiology-health-care/research/epidemiology-and-public-health/research/whitehall-ii
53 da Silva A.A.M. et al., 2010 Ribeirão Preto Cohort
54 Kebede SA et al., 2022 GWAS Consortia https://dhsprogram.com/
55 Aamir M et al., 2021 Hubei Province Air-Quality (Hourly/Daily Averages); Tianqihoubao.com https://www.tianqihoubao.com/
56 Bakar et al., 2020 Secondary From Open-Source Ifls5
57 Vasconcelos AGG et al., 1998 GWAS Consortia
58 Musenge E et al., 2013 Agincourt HDSS (Health and Socio-Demographic Surveillance System) https://www.agincourt.co.za/
59 Dotsikas K et al., 2023 GWAS Consortia
60 Davis J.N. et al., 2022 World Bank Open Data https://data.worldbank.org/
61 MacSweeney N et al., 2023 Adolescent Brain Cognitive Development Study (ABCD)
62 Frank LD & Wali B, 2021 Behavioral Risk Factor Surveillance System (BRFSS) https://www.cdc.gov/brfss/
63 Turner J.D. et al., 2020 Literature & Site
64 Li C et al., 2020 National Registry

Appendix 4 

Characteristics of included studies

The characteristics of the studies included in this systematic review were extracted and organized into a structured database. The database includes information on the authors and year of publication, study objectives, disease area, country, sample size, type of structural equation modeling used, causal modeling methods, estimation methods, software used, fit indices, and the main methodological limitations reported in each study. The full database containing the characteristics of the 458 included studies is provided in Additional file 1 (Table S1).

Authors' contributions

S.M.M.E. conceived and designed the study, collected the data, performed the analyses, interpreted the results, and drafted the manuscript. F.D. contributed to the study design and participated in data collection and data extraction. S.P.G.T., S.C.A.H., V.K.S., M.T.D.H., and R.G.K. read and provided comments, suggestions, and critical review to improve the study design, data collection, and the manuscript. R.G.K supervised the article. All authors read and approved the final manuscript.

Funding

This research was funded in whole or in part by Science for Africa Foundation to the Developing Excellence in Leadership, Training and Science in Africa (DELTAS Africa) programme [DEL-22-009] with support from Wellcome Trust and the UK Foreign, Commonwealth & Development Office and is part of the EDCPT2 programme supported by the European Union. For purposes of open access, the author has applied a CC BY public copyright licence to any Author Accepted Manuscript version arising from this submission.

Data availability

The datasets generated and/or analysed during the current study consist of a structured database containing the characteristics of all studies included in the review, namely authors, year of publication, study objectives, disease area, country, sample size, structural equation modeling approaches, causal modeling methods, estimation methods, software used, fit indices, and reported methodological limitations. The full dataset is provided in the supplementary information files.

Declarations

Ethics approval and consent to participate

Not applicable.

Consent for publication

Not applicable.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Fauci AS, Touchette NA, Folkers GK. Emerging infectious diseases: A 10-year perspective from the National Institute of Allergy and Infectious Diseases. Int J Risk Saf Med. 2005;17(3–4):157–67. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.World Health Organization. World Health Statistics 2022: Monitoring Health for the SDGs, Sustainable Development Goals. 2022. https://www.who.int/publications/i/item/9789240051157. Accessed 11 Apr 2025.
  • 3.Arnold KF, Harrison WJ, Heppenstall AJ, Gilthorpe MS. DAG-informed regression modelling, agent-based modelling and microsimulation modelling: A critical comparison of methods for causal inference. Int J Epidemiol. 2019;48(1):243–53. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Greenland S, Pearl J, Robins JM. Causal diagrams for epidemiologic research. Epidemiology. 1999;10(1):37–48. [PubMed] [Google Scholar]
  • 5.Robins JM. Semantics of causal DAG models and the identification of direct and indirect effects. In: Highly Structured Stochastic Systems. Oxford: Oxford University Press; 2003. pp. 70–82.
  • 6.Yang T, Ruan H, Li F. The application of structural equation model approach in epidemiological research. Zhonghua Liu Xing Bing Xue Za Zhi. 2005;26(4):297–300. [PubMed] [Google Scholar]
  • 7.Najera J. A critical review of the field application of a mathematical model of malaria eradication. Bull World Health Organ. 1974;50(5):449. [PMC free article] [PubMed] [Google Scholar]
  • 8.Hernán M, Robins J. Causal Inference: What If. Boca Raton: Chapman & Hall/CRC; 2020. [Google Scholar]
  • 9.Muthén LK, Muthén BO. How to use a Monte Carlo study to decide on sample size and determine power. Struct Equ Model. 2002;9(4):599–620. [Google Scholar]
  • 10.Duo-quan W, Lin-hua T, Heng-hui L, Zhen-cheng G, Xiang Z. Application of structural equation models for elucidating the ecological drivers of Anopheles sinensis in the Three Gorges Reservoir. PLoS ONE. 2013;8(7):e68766. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Leopold SJ, Watson JA, Jeeyapant A, Simpson JA, Phu NH, Hien TT, et al. Investigating causal pathways in severe falciparum malaria: A pooled retrospective analysis of clinical studies. PLoS Med. 2019;16(8):e1002858. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Liu K, Zhang M, Luo D, Zheng Y, Shen Z, Chen B. Influencing Factors of Treatment Outcomes Among Patients with Pulmonary Tuberculosis: A Structural Equation Model Approach. Psychol Res Behav Manag. 2023;16:2989–99. [DOI] [PMC free article] [PubMed]
  • 13.Jiang H, He W, Pan H, Zhong X. A structural equation modeling approach to investigate HIV testing willingness for men who have sex with men in China. AIDS Res Ther. 2023;20(1):64. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Olson K, Hayduk L, Cree M, Cui Y, Quan H, Hanson J, et al. The changing causal foundations of cancer-related symptom clustering during the final month of palliative care: a longitudinal study. BMC Med Res Methodol. 2008;8:1–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Tennant PWG, Murray EJ, Arnold KF, Berrie L, Fox MP, Gadd SC, et al. Use of directed acyclic graphs (DAGs) to identify confounders in applied health research: review and recommendations. Int J Epidemiol. 2021;50(2):620–32. 10.1093/ije/dyaa213. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.VanderWeele TJ. Principles of confounder selection. Eur J Epidemiol. 2019;34(3):211–9. 10.1007/s10654-019-00494-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Endawkie A, Keleb A, Dilnesa T, Tsega Y. Generalized structural equation modeling of direct and indirect determinants of chronic undernutrition among under-five children in Ethiopia: Further analysis of the 2019 mini Ethiopian demographic and health survey. J Health Popul Nutr. 2025;44:73. 10.1186/s41043-025-00792-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Zewdie M, Wolde GM, Tadesse M. Structural equation modeling analysis of health belief model-based determinants of COVID-19 preventive behavior of academic staff: A cross-sectional study. BMC Infect Dis. 2024;24:97. 10.1186/s12879-024-09697-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Sánchez BN, Budtz-Jørgensen E, Ryan LM, Hu H. Structural equation models: a review with applications to environmental epidemiology. J Am Stat Assoc. 2005;100(472):1443–55. [Google Scholar]
  • 20.Amorim LDAF, Fiaccone RL, Santos CAS, Santos TNd, Moraes LTLd, Oliveira NF, et al. Structural equation modeling in epidemiology. Cad Saude Publica. 2010;26(12):2251–2262. [DOI] [PubMed]
  • 21.Cheung MWL, Hong RY. Applications of meta-analytic structural equation modelling in health psychology: Examples, issues, and recommendations. Health Psychol Rev. 2017;11(3):265–79. [DOI] [PubMed] [Google Scholar]
  • 22.Hair Jr JF, Hult GTM, Ringle CM, Sarstedt M, Danks NP, Ray S, et al. Partial least squares structural equation modeling (PLS-SEM) using R: A workbook. Springer; 2021.
  • 23.VanderWeele TJ. Invited commentary: structural equation models and epidemiologic analysis. Am J Epidemiol. 2012;176(7):608–12. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Haenlein M, Kaplan AM. A beginner’s guide to partial least squares analysis. Underst Stat. 2004;3(4):283–97. [Google Scholar]
  • 25.Lee L, Petter S, Fayard D, Robinson S. On the use of partial least squares path modeling in accounting research. Int J Account Inf Syst. 2011;12(4):305–28. [Google Scholar]
  • 26.Jöreskog KG. A general approach to confirmatory maximum likelihood factor analysis. Psychometrika. 1969;34(2):183–202. [Google Scholar]
  • 27.Jöreskog KG. A general method for estimating a linear structural equation system. Psychometrika. 1970;35:443–77.
  • 28.Keesling JW. Maximum likelihood approaches to causal flow analysis. The University of Chicago; 1973.
  • 29.Wiley DE. The identification problem for structural equation models with unmeasured variables. Struct Equ Model Soc Sci. 1973:69–83.
  • 30.Shah S, Novak S, Stapleton LM. Evaluation and comparison of models of metabolic syndrome using confirmatory factor analysis. Eur J Epidemiol. 2006;21(5):343–9. [DOI] [PubMed] [Google Scholar]
  • 31.Silva AA, Mehta Z, O’Callaghan FJ. Duration of breast feeding and cognitive function: Population based cohort study. Eur J Epidemiol. 2006;21:435–41. [DOI] [PubMed] [Google Scholar]
  • 32.Rubin DB. Estimating causal effects of treatments in randomized and nonrandomized studies. J Educ Psychol. 1974;66(5):688–701. [Google Scholar]
  • 33.Pearl J. Causality. Cambridge: Cambridge University Press; 2009. [Google Scholar]
  • 34..Bollen KA, Pearl J. Eight Myths About Causality and Structural Equation Models. In: Morgan SL (Ed.), Handbook of Causal Analysis for Social Research, Springer. Chapter 15, 2013. p. 301–28. 10.1007/978-94-007-6094-3_15. [DOI]
  • 35.Bollen KA. Structural Equations with Latent Variables. Hoboken: Wiley; 1989. [Google Scholar]
  • 36.Hair JF, Black WC, Babin BJ, Anderson RE, Tatham RL, et al. Multivariate data analysis. Cengage learning Hampshire; 2019.
  • 37.Jöreskog KG. Testing a simple structure hypothesis in factor analysis. Psychometrika. 1966;31(2):165–78. [DOI] [PubMed] [Google Scholar]
  • 38.Ravand H, Baghaei P. Partial least squares structural equation modeling with R. Pract Assess Res Eval. 2016;21(11):n11. [Google Scholar]
  • 39.Hair Jnr JF. Black WC, Babin BJ. Multivariate data analysis: Anderson RE; 2010. [Google Scholar]
  • 40.Jöreskog KG. Structural analysis of covariance and correlation matrices. Psychometrika. 1975;40(2):125–31. [Google Scholar]
  • 41.Diamantopoulos A, Winklhofer HM. Formative versus reflective indicators in organizational measure development: A comparison and empirical illustration. Br J Manag. 2008;17(4):263–82. [Google Scholar]
  • 42.Lomax RG. A Beginner’s Guide to Structural Equation Modeling. New York: Psychology Press; 2004.
  • 43.Brown TA. Confirmatory Factor Analysis for Applied Research. New York: Guilford Publications; 2015.
  • 44.Kaplan D. Structural Equation Modeling: Foundations and Extensions, vol. 10. Thousand Oaks: SAGE Publications; 2008. [Google Scholar]
  • 45.Kline RB. Principles and Practice of Structural Equation Modeling. New York: Guilford Publications; 2023.
  • 46.Raykov T, Marcoulides GA. A First Course in Structural Equation Modeling. New York: Routledge; 2012.
  • 47.Hoyle RH. Handbook of Structural Equation Modeling. New York: Guilford Publications; 2014.
  • 48.Jöreskog KG. A general method for analysis of covariance structures. Biometrika. 1970;57(2):239–51. [Google Scholar]
  • 49.Edwards JR. The construct validity of formative measurement. Organ Res Methods. 2007;10(2):264–84. [Google Scholar]
  • 50.Boomsma A. Structural equation modeling: An introduction and review. Sociol Methods Res. 2000;27(3):367–426. [Google Scholar]
  • 51.Jöreskog KG. A general approach to confirmatory maximum likelihood factor analysis. Psychometrika. 1973;38(3):275–300. [Google Scholar]
  • 52.Wright S. The method of path coefficients. Ann Math Stat. 1934;5(3):161–215. [Google Scholar]
  • 53.Lee SY. Structural Equation Modeling: A Bayesian Approach. Hoboken: John Wiley & Sons; 2007. [Google Scholar]
  • 54.Pearl J. Causality: Models, Reasoning, and Inference. New York: Cambridge University Press; 2000. [Google Scholar]
  • 55.Rosenbaum PR, Rubin DB. The central role of the propensity score in observational studies for causal effects. Biometrika. 1983;70(1):41–55. [Google Scholar]
  • 56.Robins JM. Marginal structural models and causal inference in epidemiology. Epidemiology. 2000;11(5):550–60. [DOI] [PubMed] [Google Scholar]
  • 57.Angrist JD, Imbens GW, Rubin DB. Identification of causal effects using instrumental variables. J Am Stat Assoc. 1996;91(434):444–55. [Google Scholar]
  • 58.van der Laan MJ, Rose S. Targeted maximum likelihood estimation: A primer. UCLA Statistics. 2011;4(9):1–91. [Google Scholar]
  • 59.Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, et al. statement: An updated guideline for reporting systematic reviews. BMJ. 2020;2021:372. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics. 1977;33(1):159–74. [PubMed]
  • 61.R Core Team. R: A Language and Environment for Statistical Computing. Vienna; 2023. https://www.R-project.org/.
  • 62.Wickham H. ggplot2: Elegant Graphics for Data Analysis. J Stat Softw. 2016;75(1):1–3. 10.18637/jss.v075.i01. [DOI]
  • 63.Pham X-L, Le TT. Bibliometric analysis and systematic review of research on expert finding: A PRISMA-guided approach. Int Arab J Inf Technol. 2024;21(4):661–74.
  • 64.Prokofieva M, Zarate D, Parker A, Palikara O, Stavropoulos V. Exploratory structural equation modeling: a streamlined step by step approach using the R Project software. BMC Psychiatry. 2023;23(1):546. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 65.Hojat M, LaNoue M. Exploration and confirmation of the latent variable structure of the Jefferson scale of empathy. Int J Med Educ. 2014;5:73. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 66.Hair JF, Black WC, Babin BJ, Anderson RE. Multivariate Data Analysis. 7th ed. Upper Saddle River: Pearson Education; 2014. [Google Scholar]
  • 67.Iacobucci D, Saldanha N, Deng X. A meditation on mediation: Evidence that structural equations models perform better than regressions. J Consum Psychol. 2007;17(2):139–53. [Google Scholar]
  • 68.Fan Y, Chen J, Shirkey G, John R, Wu SR, Park H, et al. Applications of structural equation modeling (SEM) in ecological studies: an updated review. Ecol Process. 2016;5(1):19. [Google Scholar]
  • 69.Mun EY, Von Eye A, White HR. An SEM approach for the evaluation of intervention effects using pre-post-post designs. Struct Equ Modeling. 2009;16(2):315–37. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 70.Hernán MA, Robins JM. Causal Inference: What If. Boca Raton: Chapman & Hall/CRC; 2020. Discusses the use of simulated data and counterfactuals for causal modeling and policy evaluation.
  • 71.MacCallum RC, Austin JT. Applications of structural equation modeling in psychological research. Annu Rev Psychol. 2000;51(1):201–26. [DOI] [PubMed] [Google Scholar]
  • 72.Xiong B, Xia B. Examining the effects of early cost drivers on contingencies. In: Construction Research Congress 2014: Construction in a Global Network. Reston, VA, USA: American Society of Civil Engineers; 2014. p. 1518–27. 10.1061/9780784413517.155. [DOI]
  • 73.Kaplan D, Depaoli S. Bayesian structural equation modeling. In R. H. Hoyle (Ed.), Handbook of structural equation modeling. The Guilford Press; 2012. p. 650–73.
  • 74.Bollen KA, Pearl J. Eight Myths About Causality and Structural Equation Models. In: Handbook of Causal Analysis for Social Research. Springer; 2013. pp. 301–28.
  • 75.Olier I, Zhan Y, Liang X, Volovici V. Causal inference and observational data. BMC Med Res Methodol. 2023;23(1):227. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 76.Grad YH, Miller JC, Lipsitch M. Cholera modeling: challenges to quantitative analysis and predicting the impact of interventions. Epidemiology. 2012;23(4):523–30. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 77.Stannah J, Flores Anato JL, Pickles M, et al. From conceptualising to modelling structural determinants and interventions in HIV transmission dynamics models: a scoping review and methodological framework for evidence-based analyses. BMC Medicine. 2024;22(1):404. [DOI] [PMC free article] [PubMed]
  • 78.Wong S, Flegg JA, Golding N, Kandanaarachchi S. Comparison of new computational methods for spatial modelling of malaria. Malar J. 2023;22:356. 10.1186/s12936-023-04760-7. Highlights sensitivity to modeling assumptions in large-scale geostatistical methods. [DOI] [PMC free article] [PubMed]
  • 79.Tarka P. An overview of structural equation modeling: its beginnings, historical development, usefulness and controversies in the social sciences. Qual Quantity. 2018;52(1):313–54. 10.1007/s11135-017-0469-8. [DOI] [PMC free article] [PubMed]
  • 80.Brady SS, Brubaker L, Fok CS, Gahagan S, Lewis CE, Lewis J, et al. Development of conceptual models to guide public health research, practice, and policy: synthesizing traditional and contemporary paradigms. Health Promot Pract. 2020;21(4):510–24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 81.Liu F, Panagiotakos D. Real-world data: a brief review of the methods, applications, challenges and opportunities. BMC Med Res Methodol. 2022;22(1):287. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 82.Zamanian A, Mareis L, Ahmidi N. Partially Specified Causal Simulations. 2023. arXiv preprint arXiv:2309.10514.
  • 83.Bagozzi RP, Yi Y. Specification, evaluation, and interpretation of structural equation models. J Acad Mark Sci. 2012;40(1):8–34. [Google Scholar]
  • 84.Hu L, Ji J. CIMTx: An R package for causal inference with multiple treatments using observational data. R J. 2022;14(3):213. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 85.Dragan D, Topolšek D. Introduction to structural equation modeling: review, methodology and practical applications. In: The International Conference on Logistics & Sustainable Transport, vol. 6. 2014. pp. 19–21.
  • 86.Hartwell ML, Khojasteh J, Wetherill MS, Croff JM, Wheeler D. Using structural equation modeling to examine the influence of social, behavioral, and nutritional variables on health outcomes based on NHANES data: addressing complex design, nonnormally distributed variables, and missing information. Curr Dev Nutr. 2019;3(5):nzz010. [DOI] [PMC free article] [PubMed]
  • 87.Cronbach LJ. Coefficient alpha and the internal structure of tests. Psychometrika. 1951;16(3):297–334.
  • 88.Fornell C, Larcker DF. Evaluating structural equation models with unobservable variables and measurement error. J Mark Res. 1981;18(1):39–50. [Google Scholar]
  • 89.Hair JF, Hult GTM, Ringle CM, Sarstedt M. A Primer on Partial Least Squares Structural Equation Modeling (PLS-SEM). Thousand Oaks: SAGE Publications; 2014.
  • 90.Hu Lt, Bentler PM. Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Struct Equ Model Multidiscip J. 1999;6(1):1–55.
  • 91.Hershberger SL. The growth of structural equation modeling: 1994–2001. Struct Equ Model. 2003;10(1):35–46. [Google Scholar]
  • 92.Orounla BR, Alaye AE, Salako KV, Agbangba CE, Aheto JMK, Glèlè KR. Direct and Indirect Effects of Environmental and Socio-Economic Factors on COVID-19 in Africa Using Structural Equation Modeling. Stats. 2024;7(3):1051–65. [Google Scholar]
  • 93.Mesquita F. Digital Epidemiology after COVID-19: Impact and Prospects. 2023. arXiv preprint arXiv:2312.04835.
  • 94.Höfler M. Causal Inference Based on Counterfactuals. BMC Med Res Methodol. 2005;5(28):1–12. https://bmcmedresmethodol.biomedcentral.com/articles/10.1186/1471-2288-5-28. [DOI] [PMC free article] [PubMed]
  • 95.Friston KJ, Harrison L, Penny W. Dynamic causal modelling. Neuroimage. 2003;19(4):1273–302. [DOI] [PubMed] [Google Scholar]
  • 96.Gollob HF, Reichardt CS. Taking account of time lags in causal models. Child Dev. 1987;58(1):80–92. [PubMed]
  • 97.Baron RM, Kenny DA. The moderator-mediator variable distinction in social psychological research: Conceptual, strategic, and statistical considerations. J Pers Soc Psychol. 1986;51(6):1173. [DOI] [PubMed] [Google Scholar]
  • 98.Muller D, Judd CM, Yzerbyt VY. When moderation is mediated and mediation is moderated. J Pers Soc Psychol. 2005;89(6):852. [DOI] [PubMed] [Google Scholar]
  • 99.Kline RB. Structural equation modeling. New York: Guilford; 1998. p. 33. [Google Scholar]
  • 100.Schoemann AM, Boulton AJ, Short SD. Determining power and sample size for simple and complex mediation models. Soc Psychol Personal Sci. 2017;8(4):379–86. [Google Scholar]
  • 101.Bentler PM, Chou CP. Practical issues in structural modeling. Sociol Methods Res. 1987;16(1):78–117. [Google Scholar]
  • 102.Lüdtke O, Ulitzsch E, Robitzsch A. A comparison of penalized maximum likelihood estimation and Markov Chain Monte Carlo techniques for estimating confirmatory factor analysis models with small sample sizes. Front Psychol. 2021;12:615162. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 103.Kline RB. Software review: Software programs for structural equation modeling: Amos, EQS, and LISREL. J Psychoeduc Assess. 1998;16(4):343–64. [Google Scholar]
  • 104.Elston RC, Satagopan JM, Sun S, Genetic, Terminology. Statistical Human Genetics: Methods and Protocols. New York: Springer; 2011. pp. 1–9.
  • 105.Marsh HW, Muthén B, Asparouhov T, Lüdtke O, Robitzsch A, Morin AJ, et al. Exploratory structural equation modeling, integrating CFA and EFA: Application to students’ evaluations of university teaching. Struct Equ Modeling. 2009;16(3):439–76. [Google Scholar]
  • 106.Mulaik SA, James LR, Van Alstine J, Bennett N, Lind S, Stilwell CD. Evaluation of goodness-of-fit indices for structural equation models. Psychol Bull. 1989;105(3):430. [Google Scholar]
  • 107.Dash G, Paul J. CB-SEM vs PLS-SEM methods for research in social sciences and technology forecasting. Technol Forecast Soc Chang. 2021;173:121092. [Google Scholar]
  • 108.Byrne BM. Structural Equation Modeling with Mplus: Basic Concepts, Applications, and Programming. New York: Routledge; 2013. [Google Scholar]
  • 109.Kline RB. Principles and Practice of Structural Equation Modeling. 3rd ed. New York: The Guilford Press; 2011. [Google Scholar]
  • 110.McDonald RP, Ho MHR. Principles and practice in reporting structural equation analyses. Psychol Methods. 2002;7(1):64. [DOI] [PubMed] [Google Scholar]
  • 111.Campbell-Lendrum D, Manga L, Bagayoko M, Sommerfeld J. Climate change and vector-borne diseases: what are the implications for public health research and policy? Philos Trans R Soc B Biol Sci. 2015;370(1665):20130552. [DOI] [PMC free article] [PubMed]
  • 112.Parham PE, Waldock J, Christophides GK, Hemming D, Agusto F, Evans KJ, et al. Climate, environmental and socio-economic change: weighing up the balance in vector-borne disease transmission. Philos Trans R Soc B Biol Sci. 2015;370(1665):20130551. [DOI] [PMC free article] [PubMed]
  • 113.Malavige GN, Medigeshi GR, Sasmono RT, Amaratunga C. Funding and collaboration inequalities in infectious disease research–why does it matter?  Trends Microbiol. 2025;33(4):378–81. [DOI] [PubMed]
  • 114.Bagozzi RP, Yi Y. On the evaluation of structural equation models. J Acad Mark Sci. 1988;16(1):74–94. [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1 (65.1KB, xlsx)

Data Availability Statement

The datasets generated and/or analysed during the current study consist of a structured database containing the characteristics of all studies included in the review, namely authors, year of publication, study objectives, disease area, country, sample size, structural equation modeling approaches, causal modeling methods, estimation methods, software used, fit indices, and reported methodological limitations. The full dataset is provided in the supplementary information files.


Articles from BMC Medical Research Methodology are provided here courtesy of BMC

RESOURCES