Skip to main content
Frontiers in Endocrinology logoLink to Frontiers in Endocrinology
. 2026 May 11;17:1649881. doi: 10.3389/fendo.2026.1649881

Global research trends, reporting and handling of missing data in observational studies of type 2 diabetes mellitus with mild cognitive impairment from 2020 to 2025: a systematic review

Yang Liu 1, Maoyi Yang 1, Yifan Liao 1, Zhipeng Hu 1,*
PMCID: PMC13199093  PMID: 42199803

Abstract

Background

Missing data is common in observational studies, and even more so in type 2 diabetes mellitus with mild cognitive impairment(T2DM-MCI), which limits the completion of assessments. We evaluated the extent, current reporting, and handling of missing data, as well as the prevailing research trends in observational studies related to T2DM-MCI.

Methods

A systematic search of PubMed, Embase, and Cochrane Library was conducted from January 2020 to April 2025 to identify observational studies related to T2DM-MCI. Bibliometrics was performed using VOSviewer and CiteSpace to evaluate publishing trends, authors, journals, and keywords. The reporting and handling of missing data were assessed according to the guidelines recommended by STROBE and Sterne et al., with a focus on the recording, causes, mechanisms, processing methods, and sensitivity analysis of missing data. Data analysis was conducted using SPSS 26, and visualization was performed using Origin Pro 2024.

Results

Among the 4,471 screened records, 88 studies (78 in English and 10 in Chinese) were included in this analysis. Among the 78 English articles, the annual publication volume exhibited fluctuations, peaking in 2024. Chinese institutions and authors led in research output. Diabetes, Metabolic Syndrome, and Obesity had the highest publication volume (7, 8.97%). Keyword identified five clusters: 1) resting-state functional magnetic resonance imaging, 2) metabolic disorders, 3) clinical assessment tools, 4) molecular mechanisms, and 5) emerging fields such as the gut microbiome.

Missing data

Only 22.7% (n = 20) of the studies quantified the missing data, with an average of 9.1%. Among studies with missing data (n = 23), 52.2% (n = 12) provided reasons for missing data, primarily citing poor quality of data collection (41.7%) and loss to follow-up (41.7%). Complete case analysis was the predominant method for addressing missing data (93.3%). No study articulated the hypothesized mechanisms underlying the missing data, and only 4.4% (n = 1) performed a sensitivity analysis.

Conclusion

In the domain of T2DM-MCI, research outcomes post-COVID-19 pandemic indicate a rebound, with China maintaining a leading position in scientific research output. However, the reporting of missing data remains ambiguous, and the methods employed to handle such data are insufficient, which may potentially introduce bias.

Systematic Review Registration

https://doi.org/10.17605/OSF.IO/EZDXM.

Keywords: bibliometrics, mild cognitive impairment, missing data, observational study, systematic review, type 2 diabetes

Introduction

Diabetes Mellitus is a severe chronic disease that significantly impacts individuals, families, and society. In 2019, the global prevalence of diabetes was estimated at 9.3% (463 million people), with projections indicating an increase to 10.9% (700 million people) by 2045 (1). As a highly prevalent metabolic disorder, type 2 diabetes mellitus (T2DM) is recognized as a critical risk factor for cognitive dysfunction (2, 3). Large-scale epidemiological studies have demonstrated that the risk of dementia in patients with T2DM is significantly higher than in non-diabetic populations. The underlying pathological mechanisms involve cumulative damage to the central nervous system via multiple pathways, including chronic hyperglycemia, insulin resistance, vascular damage, and neuroinflammation (4, 5).

In the spectrum of cognitive impairments associated with diabetes, mild cognitive impairment (MCI) occupies a significant position. Specifically, type 2 diabetes mellitus with mild cognitive impairment (T2DM-MCI) refers to the clinical condition wherein patients with T2DM experience cognitive decline that exceeds the expected age-related changes, affecting domains such as memory, executive function, attention, or language. However, these patients do not meet the diagnostic criteria for dementia, and their fundamental daily living abilities remain intact (6). It is noteworthy that while MCI is regarded as an intermediate stage between a normal cognitive state and dementia, individuals with MCI face an elevated risk of progressing to dementia (7). Thus, early identification and intervention for MCI are of paramount importance.

In recent years, numerous observational studies have focused on exploring the association between T2DM and MCI, analyzing risk factors, pathological features, and potential biomarkers. These studies provide a crucial basis for clinical risk stratification and prevention strategies. However, they often encounter significant methodological challenges due to inherent selection or recall biases among participants and the suboptimal quality of data collection, particularly concerning missing data. A review of complex, large-scale epidemiological surveys indicates that inadequate reporting and improper handling of missing data remain prevalent (8–14). The results from such studies can introduce substantial selection bias, diminish statistical power, and ultimately distort the true strength and direction of the association between T2DM and MCI (15–17).

It is noteworthy that, despite the considerable impact of missing data on the reliability of conclusions, current research on the models, mechanisms, and processing methods for missing data in observational studies of T2DM-MCI remains underdeveloped. Most existing articles focus on the simplistic application of single statistical methods (18–20) and lack in-depth discussion on systematic evaluation and strategies for handling missing data within the T2DM-MCI domain. Therefore, systematically evaluating the characteristics of missing data in observational studies of T2DM-MCI and elucidating its potential impact mechanisms are essential for enhancing research quality in this area and ensuring the validity of conclusions.

This study aims to address this gap by reviewing the reporting and handling of missing data in observational studies of T2DM-MCI across multiple databases from 2020 to 2025. Furthermore, it seeks to identify current research hotspots through bibliometric and visual analysis, thereby providing a foundation and guidance for the standardized design and data analysis of high-quality observational studies in the future.

Methods

Data sources and search strategy

We conducted a comprehensive search of PubMed, Embase, and Cochrane Library for studies publishedbetween January 1, 2020, and April 24, 2025, to evaluate the globalization trends and the status of missing data processing in observational studies on T2DM-MCI. The search strategy was developed with assistance from the Harvard University Library and experienced medical professionals, incorporating key search terms such as cohort study, cross-sectional study, case-control study, type 2 diabetes, and cognitive impairment. Due to the limitations in reviewing a substantial number of identified records, we restricted our search to publications from 2020 to 2025. The detailed search strategy is provided in Supplementary Material 1.

Study selection

Abstracts of the identified citations were screened for eligibility based on the following criteria: (i) observational studies (cohort studies, cross-sectional studies, case-control studies); (ii) subjects diagnosed with T2DM-MCI. We excluded meta-analyses, randomized controlled trials, animal studies, research protocols, and guidelines. Summaries of meetings were also excluded as they did not provide sufficient information regarding the handling of missing data. Two researchers independently conducted preliminary screenings of titles and abstracts, as well as full-text re-screening. Any disputes were resolved through consultation with the third and fourth researchers.

Data extraction and analysis

We extracted the following components for our analysis: research identity (including title, author, keywords, publication year, and journal), research setting, research design (prospective, retrospective, or cross-sectional), data collection methods, sample size, primary statistical analysis methods, and information regarding missing data. The handling of missing data adheres to the guidelines recommended by STROBE and Sterne et al. (21, 22). Specifically, we applied the following items from the STROBE statement that are directly relevant to missing data: Item 12c (explain how missing data were addressed), Item 12e (describe any sensitivity analyses), Item 13a (report numbers of individuals at each stage of the study, e.g., numbers potentially eligible, examined for eligibility, confirmed eligible, included in the study, completing follow-up, and analyzed), Item 13b (give reasons for non-participation at each stage), Item 14b (indicate the number of participants with missing data for each variable of interest), and Item 17 (report other analyses done, e.g., analyses of subgroups and interactions, and sensitivity analyses).

To ensure transparency and reproducibility, we predefined a standardized coding scheme for data extraction. The extent of missing data was classified as “quantified” (explicit number or percentage reported), “not reported”, or “unclear”. Reasons for missingness were categorized as “poor data collection quality”, “loss to follow-up”, or “other”. Missing data mechanisms were assessed based on the assumptions stated in each study (Missing Completely At Random [MCAR], Missing At Random [MAR], Missing Not At Random [MNAR], or not stated). Handling methods were classified as “complete case analysis”, “multiple imputation”, or “other”. Sensitivity analysis was recorded as “yes” (with description) or “no”. Two researchers independently extracted all data; disagreements were resolved through discussion or consultation with the third and fourth researchers.

Consistent with this framework, we documented the number of missing data points, their causes, the mechanisms behind them, the methods employed to address them, and whether a sensitivity analysis was conducted. In addition, we followed the reporting framework proposed by Sterne et al. for multiple imputation. In instances of multiple imputations, we recorded the variables utilized, the number of imputations performed, the assessment of the imputation process, and the treatment of non-normal or categorical variables. For all variables with incomplete observations, we opted to select the missing covariates or the highest values of the outcome data to avoid redundant calculations. If the extent of missing data for the primary outcome was not clearly stated, we calculated this by determining the difference between the number of participants included in the analysis and the number of registered participants. Based on the reporting and handling of missing data, the data were summarized as frequencies and proportions and analyzed using IBM SPSS Statistics 26.

Bibliometrics and visualization analysis

This study employs Origin Pro 2024 software to analyze trends and proportions of annual publications. Additionally, the data extracted was analyzed and visualized using CiteSpace (version 5.7.R5, 64-bit) and VOSviewer (version 1.6.20). VOSviewer, developed by Waltman et al. in 2009, is a free, Java-based software designed for analyzing large volumes of literature data and presenting it in a map format (23). In this study, VOSviewer is utilized to generate visual charts that identify the highest-yielding journals, authors, and high-frequency keywords. To visualize research results in specific areas through the construction of a literature co-citation network, Professor Chen Chaomei developed CiteSpace (version 5.7.R5), which employs experimental frameworks to explore new concepts and evaluate existing technologies (24). This facilitates a deeper understanding of knowledge domains, research frontiers, and trends, while also aiding in the prediction of future research trajectories. This study leverages CiteSpace to visualize keyword clustering and author clustering.

Results

Literature retrieval and characteristics

Figure 1 illustrates the flowchart of the included studies. A total of 4,471 articles were retrieved, from which 731 duplicate articles were removed, leaving 3,740 abstracts for screening. Among these, 3,631 articles that did not meet the inclusion criteria were excluded. This exclusion comprised 96 meta-analyses, 217 reviews, 29 Mendelian randomized studies, 59 randomized controlled trials, 38 research programs, 7 clinical guidelines, 504 animal experiments, 232 reports, 12 conference proceedings, 2 bibliometric studies, and 2,435 unrelated articles. The remaining 109 articles were read in full text, and 21 were further excluded. Ultimately, a total of 88 studies were included in the review, consisting of 78 articles in English and 10 articles in Chinese. The included studies primarily employed a cross-sectional design (70,79.5%), with single-center studies accounting for 97.7% (n = 86). Data were predominantly collected through questionnaire surveys (72,83.3%). The median sample size was 194 (IQR: 108-321), with 75.0% of the studies having a sample size ranging from 100 to 1,000 cases (Table 1; Figure 2A).

Figure 1.

Flowchart illustrating a systematic review process: 4471 records identified, 731 duplicates excluded, 3740 abstracts screened, multiple exclusion reasons listed, 109 articles assessed, further exclusions detailed, resulting in 88 studies included.

Flowchart of study inclusion process.

Table 1.

Characteristics of included studies.

Description Total (n=88)
Study design, n (%)
Prospective 7 (7.9)
Retrospective 11 (12.5)
Cross-Section 70 (79.6)
Sample size, Median (IQR) 194 (108-321)
Sample size, n (%)
<100 17 (19.3)
100-1000 66 (75.0)
>1000 5 (5.7)
Study site, n (%)
Multisite 2 (2.3)
Single site 86 (97.7)
Method of data collection, n (%)
Administrative data 11 (12.5)
Surveysa 72 (83.3)
Mixed 5 (5.7)

n, number;%, percent.

a

includes clinical report form or any study questionnaire.

Figure 2.

Panel A displays a 3D pie chart showing study types: cross-section at seventy-nine point five percent, cohort study at eleven point four percent, and case-control at nine point one percent, with a color legend. Panel B presents a vertical bar chart of article numbers by year, with values of fifteen for 2020 to 2022, eight in 2023, nineteen in 2024, and six in early 2025, along with a trendline connecting these data points.

(A) Proportion of study types included. (B) A Trend chart of annual publication volume.

We conducted a visual analysis of the final English literature, which encompassed 78 articles from 37 countries and regions, published in 49 journals by 236 institutions and authored by 504 individuals. Since 2020, the annual number of publications has generally shown an unstable trend (Figure 2B). We categorize the annual volume of publications into three distinct stages. From 2020 to 2022, during the rampant spread of COVID-19, the number of publications stabilized at 15 per year, indicating sustained interest from researchers in this field. The publication rate remained relatively stable during this period. In 2023, as the COVID-19 epidemic came under control and people’s lives gradually returned to normal, the number of published articles decreased to 8. This decline suggests that, in the first year following the control of the epidemic, researchers showed reduced interest in this area of study. However, by April 2025, the number of publications had increased, peaking in 2024. This trend indicates that the domain has garnered increased attention since 2024.

Authors and journals

A total of 504 authors conducted observational studies on T2DM-MCI. Table 2 presents the top 10 authors who have published in the past five years, all of whom are from China, including Affiliated Zhongda Hospital of Southeast University, First Affiliated Hospital of Harbin Medical University, and Nanjing Drum Tower Hospital Clinical College of Nanjing Medical University. Among them, the earliest publication dates of the top seven authors were all in 2020, suggesting that they may have begun to concentrate on this field earlier than the last three authors. The collaboration network among the authors (Figure 3A) indicates that ShaoHua Wang, the author with the highest output, maintains a close cooperative relationship with the six authors following him. All of these authors are affiliated with the Affiliated Zhongda Hospital of Southeast University, and their publication dates are primarily concentrated between 2020 and 2023. Bing Zhang and Xin Li also exhibit a strong collaboration, being affiliated with Nanjing Drum Tower Hospital. Most of their publications occurred in 2022, indicating that this group began to focus on this field relatively later. Further analysis of the cooperation network indicates that collaboration within institutions is robust, while inter-agency cooperation remains limited. Consequently, we advocate for enhanced collaboration among authors from different institutions to expand the sample size of observational studies and provide more scientifically robust and convincing evidence for the findings.

Table 2.

Statistics of literature published in the top 10 authors.

Rank Author Articles counts Year
1 Shaohua Wang 9 2020
2 Haoqiang Zhang 7 2020
3 Ke An 7 2020
4 Sai Tian 6 2020
5 JiJing Shi 6 2020
6 Wenwen Zhu 6 2020
7 Wuyou Cao 6 2020
8 Xin Li 5 2022
9 Ziwei Yu 4 2022
10 Bing Zhang 4 2021

Figure 3.

Panel A shows a network visualization illustrating connections among researchers, with color-coded clusters and lines representing collaborative relationships over time; author names appear in red within cluster centers. Panel B presents a density visualization of academic journals, where journal names are distributed within a circular heatmap; areas in brighter yellow indicate higher publication density and names such as “diabetes, metabolic syndrome and obesity: targets and therapy” and “frontiers in aging neuroscience” are more prominent.

(A) A Collaboration network diagram among authors. (B) Density chart of published literature in journals.

Table 3; Figure 3B present the ten most productive journals in the field. The journal with the highest number of published papers is Diabetes, Metabolic Syndrome and Obesity (7 papers, 8.97%), followed by Frontiers in Aging Neuroscience (5 papers, 6.41%), Journal of Alzheimer’s Disease (4 papers, 5.13%), and Journal of Diabetes (3 papers, 3.85%). Among these top ten journals, Frontiers in Aging Neuroscience has the highest impact factor (IF) of 4.1. All listed journals are categorized as Q1, Q2, or Q3, with the Journal of Magnetic Resonance Imaging classified as Q1.

Table 3.

Statistics of literature published in the top 10 journals.

Rank Journal Articles counts Percentage (78) IF Quartile
in category
1 Diabetes, Metabolic Syndrome, and Obesity 7 8.97 2.8 Q3
2 Frontiers in Aging Neuroscience 5 6.41 4.1 Q2
3 Journal of Alzheimer’s Disease 4 5.13 3.4 Q2
4 Journal of Diabetes 3 3.85 3 Q2
5 Frontiers in Neuroscience 3 3.85 3.2 Q2
6 Frontiers in Neurology 3 3.85 2.7 Q2
7 Journal of Magnetic Resonance Imaging 2 2.56 3.1 Q1
8 Journal of Diabetes Investigation 2 2.56 3.1 Q2
9 Journal of Clinical Medicine 2 2.56 3 Q1
10 Frontiers in Nutrition 2 2.56 4 Q2

Keywords

Through the analysis of keywords, we can gain a comprehensive understanding of the general situation and developmental direction of this field. Utilizing the co-occurrence of keywords in VOSviewer software, we identified that, in addition to T2DM and MCI, the most frequently occurring keywords are Alzheimer’s disease (8), followed by insulin resistance (5), dementia (4), and magnetic resonance imaging (4) (Table 4; Figure 4A). We clustered the keywords and constructed a network comprising 119 keywords, resulting in a total of five distinct clusters (Figure 4B). Cluster 1 (yellow) contains 30 keywords, including functional connectivity network and white matter network, which are associated with neuroimaging markers. Cluster 2 (green) comprises 27 keywords, such as insulin resistance, blood glucose fluctuation, lipid abnormality, and uric acid index, which pertain to various metabolic disorders. Cluster 3 (blue) includes 23 keywords, such as the Montreal Cognitive Assessment (MoCA) and neuropsychological tests, primarily serving as clinical assessment tools. Cluster 4 (dark blue) consists of 22 keywords, including oxidative stress, brain-derived neurotrophic factor (BDNF), and other molecular mechanisms. Finally, Cluster 5 (red) contains 17 keywords, primarily focusing on the intestinal microbiome, galectin-3, and other emerging directions. We utilized CiteSpace software to create a clustering map that visualizes the evolution of research hotspots over time (Figure 4C).

Table 4.

List of high-frequency keywords.

Rank Keyword Counts Rank Keyword Counts
1 mild cognitive impairment 57 11 proteomics 2
2 type 2 diabetes 56 12 executive function 2
3 Alzheimer’s disease 8 13 Montreal cognitive assessment 2
4 insulin resistance 5 14 brain-derived
neurotrophic factor
2
5 dementia 4 15 elderly 2
6 magnetic resonance imaging 4 16 insulin 2
7 neuropsychological test 4 17 white matter network 2
8 functional connectivity 3 18 functional
connectivity density
2
9 resting-state functional
magnetic resonance imaging
3 19 neuroimaging 2
10 aging 2 20 cerebral small vessel disease 2

Figure 4.

Panel A shows a clustered network visualization of research keywords such as mild cognitive impairment and type 2 diabetes mellitus using colored nodes and connecting lines. Panel B displays a color-coded cluster map highlighting five main research topics in labeled groups, including resting-state functional magnetic resonance imaging, cognitive impairment, MCI, general diabetes, and type 2 diabetes. Panel C presents a timeline visualization on a black background, mapping the temporal evolution of research keywords across five clusters, with paths connecting related terms and magenta cluster labels.

(A) High-frequency keywords network diagram. (B) Cluster analysis of keywords. (C) Keywords clustering timeline plot.

Burst analysis of keywords

We focused on 15 keywords that exhibited the strongest outbreaks in this field (Figure 5), including Alzheimer’s disease, resting-state functional magnetic resonance imaging, executive function, insulin resistance, dementia, and brain-derived neurotrophic factor. Notably, nomogram, galectin-3, and functional connectivity density have garnered particular attention over the past two years. These keywords signify the current research hotspots and potential future research trends in this area.

Figure 5.

Table listing the top fifteen keywords with the strongest citation bursts from 2020, arranged by strength, year, and duration. Keywords include nomogram, Alzheimer's disease, resting-state functional magnetic resonance imaging, graph theory, and others. Each row provides the keyword, year, citation burst strength, start and end years, and a horizontal bar graph highlighting the period of strongest burst in red for years between two thousand twenty and two thousand twenty-five.

Visualization chart of keywords with the strongest citation bursts.

Report of missing data

To clarify the relationship between the numbers of studies reported in Tables 5, 6, we define the following sub-sets: among all 88 included studies, 23 studies (26.1%) explicitly acknowledged the presence of missing data (either by quantifying it or by stating that some data were missing). Within these 23 studies, 20 (87.0%) reported the exact amount of missing data (mean 9.1%, range 1.23–38.1%), while the remaining 3 studies acknowledged missing data but did not quantify the amount, and 15 studies (65.2%) described a specific method for handling missing data. Notably, among the 20 studies that quantified missing data, 5 (25.0%) did not proceed to describe any analytical strategy for addressing those missing values, indicating a gap between recognition of missing data and appropriate methodological handling. The remaining 65 studies (73.9%) did not mention missing data at all. Among the 23 studies that acknowledged missing data, 14 (60.9%) excluded participants based on missing data. Of these 14 studies, 13 (92.9%) explicitly reported the number of individuals excluded, while 1 study did not provide this information.

Table 5.

Reporting of missing data.

Description n (%)
Reported the amount of missing data (N = 88)
Yes 20 (22.7)
No 65 (73.9)
Unclear 3 (3.4)
Reported reasons for missing data (N = 23)
Yes 12 (52.2)
No 11 (47.8)
Reported number of individuals excluded due to missing data (N = 14)
Yes 13 (92.9)
No 1 (7.1)
Described method used to handle missing data (N = 23)
Yes 15 (65.2)
No 8 (34.8)
Stated the assumptions for missing data methods (N = 15)
Yes 0 (0)
No 15 (100)

Of the 88 included studies, 20 (22.7%) reported the amount of missing data, 65 (73.9%) did not, and 3 (3.4%) acknowledged missing data but did not quantify it. Among the 23 studies with missing data (including the 20 that quantified and the 3 that did not), 12 (52.2%) gave reasons for missingness; among the 14 studies that excluded participants due to missing data, 13 (92.9%) reported the number excluded; among the 23 studies, 15 (65.2%) described a handling method; and among those 15, none stated the assumed missing data mechanism (MCAR, MAR, or MNAR).

Table 6.

Handling of missing data.

Description n (%)
Methods used for dealing with missing data (N = 15)
Complete case analysis 14 (93.3)
Multiple imputation 1 (6.7)
Compared differences between individuals with and without incomplete data (N = 14)
Yes 0 (0.0)
No 14 (100.0)
Performed sensitivity analysis to test robustness of results (N = 23)
Yes 1 (4.4)
No 22 (95.6)
For multiple imputation (N = 1)
Indicated number of imputed datasets
No
Reported variables included in imputation model
No
Described handling of non-normal and categorical variables
Yes
Evaluated multiple imputation analysis
No

Of the 15 studies that handled missing data, 14 (93.3%) used complete case analysis and 1 (6.7%) used multiple imputation. Among the 14 studies that excluded cases, none compared those with versus without missing data. Only 1 of the 23 studies with missing data (4.4%) performed a sensitivity analysis. The single study using multiple imputation did not report the number of imputations, imputation variables, or model evaluation, but did describe handling of non-normal and categorical variables.

“For multiple imputation (N = 1) separately evaluates the reporting quality of the only study that employed multiple imputation, based on the guidelines recommended by Sterne et al. This study did not indicate the number of imputed datasets, did not report the variables included in the imputation model, and did not evaluate the multiple imputation analysis; however, it did describe the handling of non‑normal and categorical variables.”

The report detailing the missing data is presented in Table 5. A total of 77.3% (n = 68) of the studies either did not mention missing data or provided ambiguous information regarding its extent, while only 22.7% (n = 20) quantified the proportion of missing data (mean 9.1%, range 1.23-38.1%). Additionally, 52.2% (n = 12) of the studies offered explanations for the missing data, primarily attributing it to poor record quality (41.7%) and loss to follow-up (41.7%). Most studies (92.9%) reported the number of exclusions due to missing data. Among the 23 studies that documented missing data, the majority (15, 65.2%) described methods for addressing the issue. Unfortunately, none of the studies specified the types of missing mechanisms assumed in their analyses.

Processing of missing data

Table 6 presents the methods employed to address missing data in the included studies. Among the studies that reported their strategies for handling missing data (n = 15), complete case analysis emerged as the most prevalent method, with approximately 93.3% of the studies (n = 14) excluding individual observations of missing data as part of their inclusion criteria, either at the initial stage or during the analysis phase. In the 14 studies that omitted participants based on data integrity, no comparisons were made between missing and non-missing data. Only 4.4% (n = 1) of the studies assessed the robustness of their missing data handling regarding the results. Notably, only one study employed the Multiple Imputation method for missing data processing, which did not adhere to the Sterne guidelines. This study only reported its approach to addressing non-normal distributions and categorical variables, without detailing the number of imputed datasets, the variables included in the imputation model, or the evaluation of the multiple imputation analysis.

Discussion

This study systematically analyzes the global research patterns of T2DM-MCI from 2020 to 2025 using bibliometric methods for the first time. The results indicate that, despite fluctuations in the annual number of publications in this field due to the impact of COVID-19, research interest has significantly rebounded since 2024. This trend suggests an increasing attention from the academic community toward the association between T2DM and cognitive impairment. Chinese scholars hold a dominant position in this field, with high-yield authors and institutions primarily concentrated in the Affiliated Zhongda Hospital of Southeast University. This concentration may be attributed to China’s large population and its status as having the highest number of diabetic patients globally (25). Furthermore, it reflects China’s active contribution to the interdisciplinary study of metabolic and neurodegenerative diseases. However, both international and domestic collaboration networks remain limited, highlighting the need for strengthened cross-agency cooperation to enhance the diversity and universality of research.

Keywords can reflect the research frontiers and trends in a field. We conducted cluster analysis, co-occurrence analysis, and burst analysis on keywords. In general, there are three core directions in the current research field:(i) the application of neuroimaging techniques, such as resting-state functional magnetic resonance imaging, in the early diagnosis of MCI;(ii) the interaction between metabolic disorders, such as insulin resistance and lipid abnormalities, and neuroinflammation; and (iii) the potential effects of the gut microbiome and genetic polymorphisms on cognitive function. Notably, the burst trend of emerging keywords, such as functional connectivity density and galectin-3, suggests that future research may focus on exploring molecular mechanisms and conducting multi-omics integrated analyses.

Overall, there is a significant deficiency in the reporting and handling of missing data in observational studies concerning T2DM-MCI. Common practices include inadequate and ambiguous reporting, the direct exclusion of participants with missing data, and a failure to evaluate the robustness of findings derived from missing data. These issues are also evident in other fields and various studies (26–29).

In some of the included studies, it was not specified whether the data were missing or fully observed, a common oversight in many contemporary reports (30–35). The absence of explanations for missing data can mislead readers and undermine the critical assessment and reproducibility of research outcomes. The average proportion of missing data reported in studies was 9.1%, and improper handling of this data can introduce significant bias (36–40).

We compared them with a previous methodological review by Karahalios et al. (2012), which examined missing data in cohort studies with repeated exposure measures (26). That study reported that 43% of articles reported the amount of missing data, and 66% used complete case analysis. Our findings show even lower reporting rates (only 22.7% quantified missing data) and a much higher reliance on complete case analysis (93.3%), suggesting that the handling of missing data in T2DM-MCI observational studies may be even more problematic than in general cohort studies a decade ago. However, unlike earlier reviews that largely focused on longitudinal designs, our study reveals that even in predominantly cross-sectional designs (79.6% of included studies), the handling of missing data remains suboptimal. This persistence of poor reporting across different study designs underscores a widespread methodological gap that has not improved over the past decade. The predominance of cross-sectional designs (79.6%) in our included studies has important implications for missing data patterns and handling strategies. In cross-sectional studies, missing data typically arise from incomplete questionnaires or failed measurements at a single time point, often driven by participant refusal, item non-response, or administrative errors (14). By contrast, longitudinal studies face additional challenges such as attrition and intermittent missingness over time, which may require more sophisticated methods like mixed-effects models or multiple imputation with time-varying covariates (21). Despite these differences, we observed that even in cross-sectional studies, the vast majority (93.3%) resorted to complete case analysis – a method that assumes missingness is Missing Completely At Random (MCAR). However, in cross-sectional health surveys, missingness is often related to observed characteristics (e.g., older age, lower education, or poorer health status), violating MCAR and potentially biasing results. Therefore, the reliance on complete case analysis in cross-sectional T2DM-MCI studies is particularly concerning. Future research should adopt more robust methods, such as multiple imputation under the Missing At Random (MAR) assumption.

A notable gap identified in our review is the disconnect between acknowledging missing data and providing an analytical solution. Among the 20 studies that quantified missing data, only 15 (75%) described a method for handling it. This means that 5 studies (25% of those with quantified missing data) explicitly reported that data were missing but offered no information on how those missing values were addressed in the statistical analysis. This omission is critical because readers cannot determine whether the reported results are based on a reduced sample (complete case analysis) or some form of imputation. The absence of handling methods, even when missing data are recognized, represents a missed opportunity to enhance transparency and reproducibility. It also increases the risk of selective reporting bias, as authors may implicitly adopt complete case analysis without stating it, thereby understating the potential impact of missing data on their conclusions. Therefore, we urge authors to adhere to the STROBE checklist.

In the included studies, complete case analysis was excessively reliant (93.3%) on data processing, while alternative methods, such as multiple imputation, were infrequently employed. This trend aligns with previous studies (41–45), indicating a low adherence to existing guidelines, such as STROBE, in practice. The application of complete case analysis is predicated on the assumption that the missing data isMCAR, implying that the absence of data is unrelated to both observed and unobserved variables (46). Consequently, the fully observed sample is expected to represent the overall study population adequately. However, the validity of this assumption diminishes as the proportion of missing data increases; significant missing data not only reduces statistical efficiency but also heightens the risk of distortion due to selection bias (47). Interestingly, some studies exclude participants with missing data based on data integrity criteria established during the initial inclusion phase, thereby creating a ‘complete observation’ dataset. Although such practices mitigate the issue of missing data, they may inadvertently introduce biases if systematic differences exist between the complete data group and the group with missing data (48, 49). Unfortunately, the studies included in our analysis did not assess whether differences existed between these two groups when employing complete case analysis. Furthermore, regardless of the missing data mechanism or methodology utilized, the robustness of test results against various alternative hypotheses and methods serves as one of the critical means of evaluating bias (22, 46). However, we observed that sensitivity analyses were conducted in only a limited number of studies (1.1%).

Multiple imputation is a commonly used statistical technique for dealing with missing values in data sets. Based on the MAR assumption, this assumption is considered to be valid in many longitudinal data environments (50). Based on the missing mechanism of existing data and assumptions, it generates possible values to replace missing observations to reproduce multiple complete versions of the original data set, and then combines them into a single result. And the values filled in between multiple data sets are different, reflecting the uncertainty of missing values (51). Unlike complete case analysis, this approach utilizes all available data, thereby minimizing the loss of accuracy and statistical power (21, 52). However, existing guidelines for multiple imputations recommend that researchers clearly describe the procedural elements to facilitate subsequent reviews (21); unfortunately, these details are often inadequately reported in studies. Among the studies reviewed, only one employed multiple imputation methods, and it failed to fully disclose the selected variables, imputation variables, and details regarding model evaluation.

Overall, observational studies on T2DM-MCI exhibit a high proportion of missing data (average 9.1%), along with vague reporting and reliance on a singular processing method.

Limitations

Through the use of visual analysis tools such as CiteSpace and VOSviewer, we have gained insights into the research progress and global trends in the field of T2DM-MCI. However, this study does have several limitations. Firstly, all reports included at the end of this review are observational studies, and their limited number may not adequately reflect the overall trends in the research field. Secondly, due to the existence of synonyms, abbreviations, and full names, keyword clustering may differ from the actual results. Therefore, the interpretation of these results should be approached with caution. Additionally, we did not exclude studies from the same participant cohort, which may lead to duplication or overlap of data and reports. Lastly, some studies are challenging to assess fully due to the vague methodological descriptions, which can obscure the impact of missing data. Future research should aim to expand the database and the coverage of years. Nevertheless, we hope that the data gaps highlighted in this review provide a comprehensive mapping of the current situation in the entire field.

Conclusion

In observational studies concerning T2DM-MCI, the global research output has exhibited fluctuations at various stages, with a notable rebound since 2024. China leads in research volume; however, there is a lack of cross-agency and international collaboration. The primary journals contributing to core output include Metabolic Syndrome and Obesity and various Frontiers series. Currently, trending research directions encompass neuroimaging, metabolic disorders, molecular mechanisms, functional connectivity density, and galectin-3, among others. Concurrently, issues such as inadequate reporting and imprecise handling of missing data are prevalent in this domain. Many studies either fail to address missing data or exclude a significant number of participants with missing information. The methods employed for managing missing data tend to be relatively simplistic and non-standardized, with a majority of studies neglecting to assess the robustness of their results, thereby significantly heightening the risk of bias.

Acknowledgments

We would like to express our gratitude to the Harvard University Library for assisting us in formulating the search strategy.

Funding Statement

The author(s) declared that financial support was received for this work and/or its publication. This work was supported by Project of the National Natural Science Foundation of China for Young Scientists [grant numbers 82305209 and 82505508]; Xinglin Scholar of Chengdu University of Traditional Chinese Medicine [grant numbers MPRC2023012 and MPRC2023017]; Natural Science Foundation of Sichuan Provincial Department of Science and Technology [grant numbers 2023NSFSC1832]; Joint Innovation Fund of Health Commission of Chengdu and Chengdu University of Traditional Chinese Medicine [grant numbers WXLH202403005 and WXLH202403136].

Footnotes

Edited by: Ben Nephew, Worcester Polytechnic Institute, United States

Reviewed by: Sarah Firdausa, Syiah Kuala University, Indonesia

Liliana Letra, University of Coimbra, Portugal

Data availability statement

The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found in the article/Supplementary Material.

Author contributions

YaL: Writing – original draft, Methodology, Data curation, Investigation, Writing – review & editing. MY: Formal analysis, Writing – original draft, Validation. YiL: Visualization, Software, Writing – review & editing. ZH: Conceptualization, Writing – review & editing, Supervision.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fendo.2026.1649881/full#supplementary-material

Supplementary Material 1

Search strategies for databases

DataSheet1.docx (723.7KB, docx)

References

  • 1. Saeedi P, Petersohn I, Salpea P, Malanda B, Karuranga S, Unwin N, et al. Global and regional diabetes prevalence estimates for 2019 and projections for 2030 and 2045: Results from the International Diabetes Federation Diabetes Atlas, 9th edition. Diabetes Res Clin Pract. (2019) 157:107843. doi:  10.1016/j.diabres.2019.107843. PMID: [DOI] [PubMed] [Google Scholar]
  • 2. Biessels GJ, Deary IJ, Ryan CM. Cognition and diabetes: a lifespan perspective. Lancet Neurol. (2008) 7:184–90. doi:  10.1016/s1474-4422(08)70021-8. PMID: [DOI] [PubMed] [Google Scholar]
  • 3. Cukierman T, Gerstein HC, Williamson JD. Cognitive decline and dementia in diabetes–systematic overview of prospective observational studies. Diabetologia. (2005) 48:2460–9. doi:  10.1007/s00125-005-0023-4. PMID: [DOI] [PubMed] [Google Scholar]
  • 4. Li C, Zuo Z, Liu D, Jiang R, Li Y, Li H, et al. Type 2 diabetes mellitus may exacerbate gray matter atrophy in patients with early-onset mild cognitive impairment. Front Neurosci. (2020) 14:856. doi:  10.3389/fnins.2020.00856. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5. Nunes AR, Alves MG, Moreira PI, Oliveira PF, Silva BM. Can tea consumption be a safe and effective therapy against diabetes mellitus-induced neurodegeneration? Curr Neuropharmacol. (2014) 12:475–89. doi:  10.2174/1570159x13666141204220539. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Petersen RC. Clinical practice. Mild cognitive impairment. N Engl J Med. (2011) 364:2227–34. doi:  10.1056/NEJMcp0910237. PMID: [DOI] [PubMed] [Google Scholar]
  • 7. Justin BN, Turek M, Hakim AM. Heart disease as a risk factor for dementia. Clin Epidemiol. (2013) 5:135–45. doi:  10.2147/clep.S30621. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Powney M, Williamson P, Kirkham J, Kolamunnage-Dona R. A review of the handling of missing longitudinal outcome data in clinical trials. Trials. (2014) 15:237. doi:  10.1186/1745-6215-15-237. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Ibrahim F, Tom BD, Scott DL, Prevost AT. A systematic review of randomised controlled trials in rheumatoid arthritis: the reporting and handling of missing data in composite outcomes. Trials. (2016) 17:272. doi:  10.1186/s13063-016-1402-5. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Richter S, Stevenson S, Newman T, Wilson L, Menon DK, Maas AIR, et al. Handling of missing outcome data in traumatic brain injury research: a systematic review. J Neurotrauma. (2019) 36:2743–52. doi:  10.1089/neu.2018.6216. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11. Lieber R, Pandis N, Faggion CM. Reporting and handling of incomplete outcome data in implant dentistry: a survey of randomized clinical trials. J Clin Periodontol. (2020) 47:257–66. doi:  10.1111/jcpe.13222. PMID: [DOI] [PubMed] [Google Scholar]
  • 12. Desai M, Kubo J, Esserman D, Terry MB. The handling of missing data in molecular epidemiology studies. Cancer Epidemiol Biomarkers Prev. (2011) 20:1571–9. doi:  10.1158/1055-9965.Epi-10-1311. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Wood AM, White IR, Thompson SG. Are missing outcome data adequately handled? A review of published randomized controlled trials in major medical journals. Clin Trials. (2004) 1:368–76. doi:  10.1191/1740774504cn032oa. PMID: [DOI] [PubMed] [Google Scholar]
  • 14. Eekhout I, de Boer RM, Twisk JW, de Vet HC, Heymans MW. Missing data: a systematic review of how they are reported and handled. Epidemiology. (2012) 23:729–32. doi:  10.1097/EDE.0b013e3182576cdb. PMID: [DOI] [PubMed] [Google Scholar]
  • 15. Seika P, Klein F, Pelzer U, Pratschke J, Bahra M, Malinka T. Influence of the body mass index on postoperative outcome and long-term survival after pancreatic resections in patients with underlying Malignancy. Hepatobiliary Surg Nutr. (2019) 8:201–10. doi:  10.21037/hbsn.2019.02.05. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Liu Y, Li Y, Bai YP, Fan XX. Association between physical activity and lower risk of lung cancer: a meta-analysis of cohort studies. Front Oncol. (2019) 9:5. doi:  10.3389/fonc.2019.00005. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17. Ayilara OF, Zhang L, Sajobi TT, Sawatzky R, Bohm E, Lix LM. Impact of missing data on bias and precision when estimating change in patient-reported outcomes from a clinical registry. Health Qual Life Outcomes. (2019) 17:106. doi:  10.1186/s12955-019-1181-2. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18. Dalén M, Nielsen S, Ivert T, Holzmann MJ, Sartipy U. Coronary artery bypass grafting in women 50 years or younger. J Am Heart Assoc. (2019) 8:e013211. doi:  10.1161/jaha.119.013211. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19. Batty GD, Deary IJ, Hamer M, Frank P, Bann D. Association of childhood psychomotor coordination with survival up to 6 decades later. JAMA Netw Open. (2020) 3:e204031. doi:  10.1001/jamanetworkopen.2020.4031. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Shiraishi A, Otomo Y, Yoshikawa S, Morishita K, Roberts I, Matsui H. Derivation and validation of an easy-to-compute trauma score that improves prognostication of mortality or the Trauma Rating Index in Age, Glasgow Coma Scale, Respiratory rate and Systolic blood pressure (TRIAGES) score. Crit Care. (2019) 23:365. doi:  10.1186/s13054-019-2636-x. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Sterne JA, White IR, Carlin JB, Spratt M, Royston P, Kenward MG, et al. Multiple imputation for missing data in epidemiological and clinical research: potential and pitfalls. Bmj. (2009) 338:b2393. doi:  10.1136/bmj.b2393. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Vandenbroucke JP, von Elm E, Altman DG, Gøtzsche PC, Mulrow CD, Pocock SJ, et al. Strengthening the reporting of observational studies in epidemiology (STROBE): explanation and elaboration. PloS Med. (2007) 4:e297. doi:  10.1371/journal.pmed.0040297. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. van Eck NJ, Waltman L. Software survey: VOSviewer, a computer program for bibliometric mapping. Scientometrics. (2010) 84:523–38. doi:  10.1007/s11192-009-0146-3. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Chen C. Searching for intellectual turning points: progressive knowledge domain visualization. Proc Natl Acad Sci USA. (2004) 101 Suppl 1:5303–10. doi:  10.1073/pnas.0307513100. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Xu J, Chen F, Liu T, Wang T, Zhang J, Yuan H, et al. Brain functional networks in type 2 diabetes mellitus patients: a resting-state functional MRI study. Front Neurosci. (2019) 13:239. doi:  10.3389/fnins.2019.00239. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Karahalios A, Baglietto L, Carlin JB, English DR, Simpson JA. A review of the reporting and handling of missing data in cohort studies with repeated assessment of exposure measures. BMC Med Res Methodol. (2012) 12:96. doi:  10.1186/1471-2288-12-96. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Rombach I, Knight R, Peckham N, Stokes JR, Cook JA. Current practice in analysing and reporting binary outcome data-a review of randomised controlled trial reports. BMC Med. (2020) 18:147. doi:  10.1186/s12916-020-01598-7. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28. Carroll OU, Morris TP, Keogh RH. How are missing data in covariates handled in observational time-to-event studies in oncology? A systematic review. BMC Med Res Methodol. (2020) 20:134. doi:  10.1186/s12874-020-01018-7. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Hussain JA, Bland M, Langan D, Johnson MJ, Currow DC, White IR. Quality of missing data reporting and handling in palliative care trials demonstrates that further development of the CONSORT statement is required: a systematic review. J Clin Epidemiol. (2017) 88:81–91. doi:  10.1016/j.jclinepi.2017.05.009. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30. Gan S, Dawed AY, Donnelly LA, Nair ATN, Palmer CNA, Mohan V, et al. Efficacy of modern diabetes treatments DPP-4i, SGLT-2i, and GLP-1RA in White and Asian patients with diabetes: a systematic review and meta-analysis of randomized controlled trials. Diabetes Care. (2020) 43:1948–57. doi:  10.2337/dc19-2419. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Tzelnick S, Alkan U, Leshno M, Hwang P, Soudry E. Sinonasal debridement versus no debridement for the postoperative care of patients undergoing endoscopic sinus surgery. Cochrane Database Syst Rev. (2018) 11:Cd011988. doi:  10.1002/14651858.CD011988.pub2. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32. Hooper L, Al-Khudairy L, Abdelhamid AS, Rees K, Brainard JS, Brown TJ, et al. Omega-6 fats for the primary and secondary prevention of cardiovascular disease. Cochrane Database Syst Rev. (2018) 11:Cd011094. doi:  10.1002/14651858.CD011094.pub4. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33. Moraa I, Sturman N, McGuire TM, van Driel ML. Heliox for croup in children. Cochrane Database Syst Rev. (2018) 10:Cd006822. doi:  10.1002/14651858.CD006822.pub5. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34. Keats EC, Haider BA, Tam E, Bhutta ZA. Multiple-micronutrient supplementation for women during pregnancy. Cochrane Database Syst Rev. (2019) 3:Cd004905. doi:  10.1002/14651858.CD004905.pub6. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35. Hernaez R, Lazo M, Bonekamp S, Kamel I, Brancati FL, Guallar E, et al. Diagnostic accuracy and reliability of ultrasonography for the detection of fatty liver: a meta-analysis. Hepatology. (2011) 54:1082–90. doi:  10.1002/hep.24452. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Stone L, Olson B, Mowery A, Krasnow S, Jiang A, Li R, et al. Association between sarcopenia and mortality in patients undergoing surgical excision of head and neck cancer. JAMA Otolaryngol Head Neck Surg. (2019) 145:647–54. doi:  10.1001/jamaoto.2019.1185. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37. Stahlmann K, Kellerhuis B, Reitsma JB, Dendukuri N, Zapf A. Comparison of methods to handle missing values in a continuous index test in a diagnostic accuracy study - a simulation study. BMC Med Res Methodol. (2025) 25:147. doi:  10.1186/s12874-025-02594-2. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38. Razzak MI, Imran M, Xu G. Big data analytics for preventive medicine. Neural Comput Appl. (2020) 32:4417–51. doi:  10.1007/s00521-019-04095-y. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39. Ryu SY, Qian WJ, Camp DG, Smith RD, Tompkins RG, Davis RW, et al. Detecting differential protein expression in large-scale population proteomics. Bioinformatics. (2014) 30:2741–6. doi:  10.1093/bioinformatics/btu341. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40. Mulder T, Kluytmans-van den Bergh MFQ, van Mourik MSM, Romme J, Crolla R, Bonten MJM, et al. A diagnostic algorithm for the surveillance of deep surgical site infections after colorectal surgery. Infect Ctrl Hosp Epidemiol. (2019) 40:574–8. doi:  10.1017/ice.2019.36. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41. Bell ML, Fiero M, Horton NJ, Hsu CH. Handling missing data in RCTs; a review of the top medical journals. BMC Med Res Methodol. (2014) 14:118. doi:  10.1186/1471-2288-14-118. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 42. Lenihan CR, Montez-Rath ME, Scandling JD, Turakhia MP, Winkelmayer WC. Outcomes after kidney transplantation of patients previously diagnosed with atrial fibrillation. Am J Transplant. (2013) 13:1566–75. doi:  10.1111/ajt.12197. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43. Jorge EC, Jorge EN, Botelho M, Farat JG, Virgili G, El Dib R. Monotherapy laser photocoagulation for diabetic macular oedema. Cochrane Database Syst Rev. (2018) 10:Cd010859. doi:  10.1002/14651858.CD010859.pub2. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 44. Bressler SB, Beaulieu WT, Glassman AR, Gross JG, Melia M, Chen E, et al. Panretinal photocoagulation versus ranibizumab for proliferative diabetic retinopathy: factors associated with vision and edema outcomes. Ophthalmology. (2018) 125:1776–83. doi:  10.1016/j.ophtha.2018.04.039. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45. Lee CC, Lee CH, Hong MY, Tang HJ, Ko WC. Timing of appropriate empirical antimicrobial administration and outcome of adults with community-onset bacteremia. Crit Care. (2017) 21:119. doi:  10.1186/s13054-017-1696-z. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46. Ibrahim JG, Chu H, Chen MH. Missing data in clinical studies: issues and methods. J Clin Oncol. (2012) 30:3297–303. doi:  10.1200/jco.2011.38.7589. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47. Graham JW. Missing Data: Analysis and Design. New York: Springer Science+Business Media; (2012). p. 31. [Google Scholar]
  • 48. Shung DL, Au B, Taylor RA, Tay JK, Laursen SB, Stanley AJ, et al. Validation of a machine learning model that outperforms clinical risk scoring systems for upper gastrointestinal bleeding. Gastroenterology. (2020) 158:160–7. doi:  10.1053/j.gastro.2019.09.009. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 49. Okpara C, Edokwe C, Ioannidis G, Papaioannou A, Adachi JD, Thabane L. The reporting and handling of missing data in longitudinal studies of older adults is suboptimal: a methodological survey of geriatric journals. BMC Med Res Methodol. (2022) 22:122. doi:  10.1186/s12874-022-01605-w. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 50. Pedersen AB, Mikkelsen EM, Cronin-Fenton D, Kristensen NR, Pham TM, Pedersen L, et al. Missing data and multiple imputation in clinical epidemiological research. Clin Epidemiol. (2017) 9:157–66. doi:  10.2147/clep.S129785. PMID: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 51. Steyerberg EW, van Veen M. Imputation is beneficial for handling missing data in predictive models. J Clin Epidemiol. (2007) 60:979. doi:  10.1016/j.jclinepi.2007.03.003. PMID: [DOI] [PubMed] [Google Scholar]
  • 52. Patrician PA. Multiple imputation for missing data. Res Nurs Health. (2002) 25:76–84. doi:  10.1002/nur.10015. PMID: [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material 1

Search strategies for databases

DataSheet1.docx (723.7KB, docx)

Data Availability Statement

The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found in the article/Supplementary Material.


Articles from Frontiers in Endocrinology are provided here courtesy of Frontiers Media SA

RESOURCES