Abstract
Multicenter retrospective studies can provide a pragmatic approach to evaluating uncommon pediatric conditions and are less expensive than prospective research. A well-executed retrospective multicenter study, with rigorous study design, systematic data collection, and robust statistical analysis, can produce clinically important and generalizable findings A variety of observational designs can be employed, including cross-sectional, cohort, and case-control studies. Selection bias, ascertainment bias, and confounding are common issues in retrospective research. Key steps include development of a feasible study design, regular contact with site investigators, and detailed data collection and management strategies. Principal investigators must seek to ensure that case ascertainment and data collection are consistent across sites, using manual and/or automated data extraction methods. Operations manuals, training sessions, and regular meetings can be used to ensure data reliability. Ethical considerations include obtaining institutional review board approval and establishing data use agreements. A proactive statistical approach to handling missing data, using techniques like multiple imputation and sensitivity analyses, is necessary. Careful planning, effective collaboration, and embracing technological advancements will enhance the value and accuracy of retrospective multicenter studies. This article discusses important considerations in the performance of a retrospective multicenter study.
INTRODUCTION
While randomized control trials are considered the highest quality in the hierarchy of evidence for studies that directly collect and analyze clinical data,1 their complexity and cost frequently render them impracticable. This may be an even greater challenge in pediatrics given lower patient volumes, the relative infrequency of clinically relevant endpoints (eg, death), and less available funding. With appropriate adherence to a relevant study design, systematic collection of pertinent data, and statistically sound analytic approaches, a well-executed retrospective multicenter study can generate clinically impactful and generalizable findings. Compared with single-center retrospective studies, multicenter retrospective studies can further improve sample size and generalizability but add complexity to data collection and analyses that are important to consider. This review summarizes key steps in the development of a multicenter retrospective study, reviews common pitfalls and challenges, and discusses future steps in their evolution.
PLANNING AND DESIGN
Study Design
Multicenter studies can take the form of case series, retrospective cohort studies, cross-sectional studies, and nested case-control studies. Each is suited to evaluating a specific type of condition and may be subject to specific limitations (Table 1). Investigators should carefully establish study aims to define the objectives or goals of the research and determine a suitable study design. In contrast to quality improvement and implementation science questions (which attempt to measure how effective a method is at changing physician behavior or evidence-based practices, such as the use of an electronic decision support tool to increase guideline-concordant antibiotic use), clinical research questions evaluate diagnosis, treatment, and prevention methods (such as evaluating risk factors for conditions such as pneumonia or sepsis). A thorough literature search should be performed to ensure that the study question is novel and relevant, is able to address limitations of prior research, and considers relevant, previously reported confounders. Developing a schematic, such as a directed acyclic graph, for thinking through the relationship among various factors (eg, confounders, mediators, modifiers) and the exposure and outcome of interest based on prior literature is pivotal to inform an appropriate study design and analytic approach.2 There are additional comprehensive resources available to provide guidance on the design of meaningful clinical research studies.3
TABLE 1.
Study Designs to Consider in Multicenter Retrospective Research
| Type of Study | Definition | Pros | Cons | Level of Evidence1 | Example |
|---|---|---|---|---|---|
| Case series | Description of analysis of patients with a predefined condition | Data collection is limited only to patients with the outcome of interest, decreasing workload | Cannot identify risk factors associated with disease or outcomes | Level 4 | Outcomes of patients with multisystem inflammatory syndrome in children vs acute COVID-1914 |
| Cross-sectional research | Analyze data from a population at a specific point in time to identify patterns and relationships. | Good for descriptive analyses, can study multiple variables, no risk of attrition | Cannot establish causality or be used in risk prediction modeling | Level 2–3 | COVID-19 vaccine acceptance15 |
| Retrospective cohort | Description of a group of individuals who share a common characteristic or experience to determine the relationship between risk factors and outcomes over time. | Useful for studying rare outcomes, can provide temporal relationships, can be used in risk-factor modeling | Differential missingness, data quality issues, limited temporal information, ascertainment bias | Level 3–4 | Bacterial meningitis score16 |
| Nested case control | Cases from an identified cohort and matched to 1–3 controls based on preselected criteria | Fewer cases to review compared to cohort study, helpful for rare conditions | Same as for cohort study, selection of appropriate criteria for matching; more subject to ascertainment bias | Level 4 | Herpes simplex virus infection17 |
Bias
Retrospective studies are susceptible to several types of bias. Selection bias refers to bias that occurs when the study population is not representative of the target population.4 Ascertainment bias can occur when there are systematic differences in the way data are collected or recorded, particularly between cases and noncases. Confounding occurs when the relationship between the exposure and outcome is influenced by a third variable that is associated with both.
Finding Site Coinvestigators
Recruitment of sites requires the identification of coinvestigators at each site who are available, interested, and have appropriate resources. Potential site coinvestigators can be identified by colleagues, mentors, or through research networks. Research networks (eg, Pediatric Research in Inpatient Settings Network, Pediatric Emergency Medicine Clinical Research Committee) may provide benefits such as infrastructure for identifying coinvestigators, feedback from research experts, and connections to funding opportunities.
Collaboration
It is helpful to establish regular contact with investigators to troubleshoot anticipated and unanticipated issues as they arise. Initial meetings can be useful for setting expectations about tasks expected of the site principal investigator (PI), timeline, finances, and authorship. A participating site PI should be cognizant of authorship requirements for journals. Many adhere to the International Committee of Medical Journal Editors, which requires that authors contribute to conception, data acquisition, analysis, or interpretation; draft or critically revise the manuscript; provide final approval; and agree to be accountable for all aspects of the work.5 Group authorship may be used to provide credit for larger groups. Authorship should be established early on. There may be opportunities for the PI to recommend that coinvestigators lead secondary analyses if there are ancillary aims that can be answered with the collected data.
DATA COLLECTION AND MANAGEMENT
Case Ascertainment
A variety of approaches can be used to identify cases. Ideally, investigators should seek and use a previously validated approach. If one does not exist, the next best strategy would be to use an ascertainment approach used in a prior study, and lastly, authors can pilot a validation study at the local study sites. Billing codes are frequently used, though these have limitations, including their primary use for billing (rather than clinical) purposes, missing, and/or incorrectly applied codes. Prior research has attempted to validate diagnosis algorithm codes in children.6 Online repositories, such as PheKB, contain publicly available repositories of electronic algorithms and can also be useful for case selection. If the authors need to develop a novel list of diagnosis codes to identify candidate patients, it may be piloted at one or more institutions to ensure consistency of diagnosis codes for the condition of interest.4 When doing so, a sample of patients identified and not identified using candidate criteria across study sites should be compared against manual medical record review, with subsequent calculations of diagnostic accuracy (sensitivity, specificity, and predictive values). Other approaches for case selection may be tailor-made based using vital signs, chief concerns, administration of a medication, or based on testing results. For multisite studies, it is important to consider that billing practices and clinical practices (eg, medication administration) may differ across sites.
Variable Definitions
A manual of operations that is rigorous and accounts for how variables are defined and where these are found will optimize data reliability. It is important to consider site differences in how each variable may be used clinically (eg, oxygen may be always given to patients on continuous albuterol at one site but not at another, which may be important to consider if trying to define patients who require oxygen), units of measurement for laboratory values, and where data are stored in the electronic health record (EHR). Although it may be tempting to gather additional data while already performing a medical record review “just in case,” care should be taken to avoid gathering unnecessary additional variables that may complicate data collection, extend timelines, and increase costs.
Method of Data Collection
Data can be collected via manual medical record review or automatically through a query of the EHR. Automated queries decrease the time burden and potential for human error, though they may be too complex to use or unavailable at certain sites. It is important to keep this in mind as one is setting up a data collection tool because there are different strategies for optimizing manual medical record review vs an automatic query that must be balanced (eg, ideally includes text validation, field notes, branching logic, and identify fields as required for manual medical record review, but this can be problematic with importing automatic queries). Additionally, as each site likely has differing capabilities of automated data extractions, it can be helpful to create a data dictionary to keep track of how each variable is extracted by the site. To improve the standardization of automatic data extraction, EHR queries, often written in the structured query language programming language, may be shared among partner institutions to decrease time spent in coding, though all queries will need to be adapted to the local EHR environment.
Manual Medical Record Review
Most studies involve some degree of manual medical record review. Training of study personnel can be done through video conferences. Meetings can be used to collectively discuss data issues, share best practices, and address emerging concerns. If the coinvestigator is not directly performing manual medical record review, they may review a sample (eg, 10–20) of medical records, compare answers with the primary medical record reviewer, and identify inefficiencies and inaccuracies in the data collection process. It can also be helpful to implement some degree of overlap in medical record reviews (eg, 5%–10% of medical records) if there is more than 1 investigator at each site to evaluate interrater reliability and improve the trustworthiness of data collection.4 Site audits and data quality checks can be done to identify systematic issues with data entry. Investigators can increase data accuracy in studies requiring manual record abstraction by ensuring abstractors are sufficiently qualified, optimizing communication with abstractors, and providing early and ongoing project-specific training and oversight.7
Data Collection Tool
A robust data collection tool can help ensure that data collection is standardized across study sites. REDCap, a Health Insurance Portability and Accountability Act (HIPAA)–compliant research platform (Vanderbilt University), and similar software has features to optimize data collection. Data collection forms should be designed and trialed to ensure a consistent understanding of study variables and be sufficiently thorough to carry out the study plan. Piloting the data collection tool across all sites can identify issues with how the tool is built (eg, errors in branching logic) and any manual medical record review inefficiencies (eg, reordering variables to better align with the order of medical record review). Tools should be designed to minimize errors in data entry, particularly for variables that are extracted via manual medical record review. Real-time quality control checks can provide alerts to ensure that data are entered as the right type (eg, using text validation to ensure that inputs are entered as numbers, free text, or dates/times with the correct format), values are plausible (eg, minimum and maximum values), and variables are not skipped. PIs can create custom alerts to notify medical record reviewers of inconsistencies across multiple variables (eg, if the date/time of an in-hospital medication administration is after the date/time of discharge) and recheck these potential errors via manual medical record review.
Advances in Technology
Natural language processing is increasingly recognized as a tool to automate data collection. Large language models (LLMs) can be housed within an institutional HIPAA-compliant intranet to ensure that patient-related information is not provided to a third party. Locally hosted LLM can be trained by investigators with clinical notes or relevant literature. These systems may have comparable accuracy to expert-based review for patient medical records reviews, though this performance may degrade with more complex tasks.8,9 LLMs will play a greater role in retrospective research in coming years.
FUNDING APPROACHES
Existing resources may be leveraged when available (eg, a departmental statistician). Grant funding may be helpful for larger projects. Expenses may be related to the time required by study investigators (as a percent allocation of salary support), funds for other study personnel (eg, EHR analysts to pull automatic data, research assistants to perform manual medical record review, research management, data scientists, and statistical support), equipment and supplies, publication fees, and travel costs. Budgets should take into consideration the potential need for duplicating certain medical records to evaluate interrater reliability, the fact that costs likely will differ across sites, and the time for training and checking in with the site PI and other research staff.
DATA ANALYSIS
Missing Data
Common to many retrospective studies is the need for a plan to handle missing data. Multiple imputation is a statistical technique used to handle missing data by creating multiple complete datasets through repeated random sampling to produce estimates that account for the uncertainty caused by the missing data. Health care data are frequently not missing at random, which may lead to systemic bias when using a complete case approach or performing multiple imputation.10 Sensitivity analyses, which consider missing data as being “negative” for dichotomous variables (eg, if there is no reported presence of a specific risk factor, assume this to be negative), a complete case analysis (in which records with missing data for any of the variables used in statistical modeling are removed), and using multiple imputation can be helpful.
Statistical Considerations
Performing a sample size estimation can help determine if the anticipated data collection will be sufficient. The analytic approach used will depend on the underlying type of study design and should be developed in consultation with a statistician or analyst. In observational studies, including cross-sectional, cohort, and case-control designs, unadjusted estimates of the outcome are typically reported, though investigators can also try to minimize the biasing effect of confounding through stratification, matching, or adjustment. These techniques are beyond the scope of this review but highlight the importance of consulting with a statistician or analyst. As children from the same institution are managed more similarly to each other compared with those from other institutions, additional statistical techniques are needed to account for the correlation introduced through this clustering, which can add uncertainty to estimates.11
In general, approaches that provide effect sizes (eg, odds ratios with 95% confidence intervals) are more useful than those that merely convey statistical significance (eg, chi-square tests). Overuse of tests of significance can result in misleading connotations of significance, rely on arbitrary thresholds, do not convey the magnitude of observed differences, and lead to issues with multiple comparisons.12
ETHICAL CONSIDERATIONS
Regulatory
Regulatory approvals and data-sharing agreements can take between 3 to 6 months to establish. IRB approval is required from participating sites. A centralized IRB, which involves one IRB from the PI’s institution rather than an IRB done at each site, allows institutions to enter into an IRB reliance agreement and follow a centralized review process to ensure uniformity across the approval process. PIs with unfunded studies should contact their IRB to determine if this would be an appropriate arrangement for their study.
Data Sharing
Sharing research data across institutions requires the establishment of a data use agreement (DUA). The HIPAA Privacy Rule allows for protected health information use or disclosure without patient authorization if an IRB or privacy board waives the requirement and permits the sharing of a limited dataset via a DUA (Table 2).13 However, because limited datasets may contain identifiable information, they are still personal health information. DUAs are established under the privacy rule and must be completed before data sharing. The agreement should specify the permitted uses and disclosures of the data, identify authorized users or recipients, and prohibit unauthorized use or further disclosure except as allowed by law. It must require recipients to implement safeguards, report unauthorized use or disclosure, and prohibit identifying or contacting individuals.
TABLE 2.
Types of Protected Health Information That Cannot Be Used in a Limited Dataset
| 1. | Names |
| 2. | Geographic subdivisions smaller than a state, except the initial 3 digits of the zip code (with exceptions) |
| 3. | All elements of dates (except year) for dates directly related to an individual (including dates of birth, admission, discharge, death, and ages over 89) |
| 4. | Telephone numbers |
| 5. | Fax numbers |
| 6. | Email addresses |
| 7. | Social Security numbers |
| 8. | Medical record numbers |
| 9. | Health plan beneficiary numbers |
| 10. | Account numbers |
| 11. | Certificate/license numbers |
| 12. | Vehicle identifiers and serial numbers, including license plate numbers |
| 13. | Device identifiers and serial numbers |
| 14. | Web universal resource locators |
| 15. | Internet protocol address numbers |
| 16. | Biometric identifiers, including fingerprints and voiceprints |
| 17. | Full-face photographic images and any comparable images |
| 18. | Any other unique identifying number, characteristic, or code |
Summarized from published guidelines from the National Institutes of Health.13
CONCLUSION
Retrospective multicenter studies offer significant advantages in improving sample size and enhancing the generalizability of findings compared with single-center studies. Effective collaboration across sites, thorough planning, and meticulous data management are crucial to navigating the challenges and pitfalls associated with multicenter research (Table 3). As advancements in technology evolve, the efficiency and accuracy of retrospective data collection and analysis are poised to improve, further enhancing the value of these studies.
TABLE 3.
Take-Home Points
| 1. | Multicenter retrospective studies improve sample size and generalizability compared with single-center studies and can provide impactful and generalizable findings. |
| 2. | Any case ascertainment approach can introduce bias. Ensure a consistent approach towards ascertainment and data collection processes across sites with a rigorous manual of operations. |
| 3. | Several study designs can be employed when using a multicenter retrospective design. Each has its own strengths and limitations. A retrospective cohort study can be used to study rare outcomes but can be onerous. A nested case-control study can decrease the burden of data collection but can introduce bias related to matching. |
| 4. | Single-site Institutional Review Board approval is required at all sites or approval may be obtained via a Central Institutional Review Board with sharing of data performed after entering into a Data Use Agreement. |
| 5. | Address missing data proactively with techniques like multiple imputation and sensitivity analyses. Use models accounting for the clustering of data from different institutions. |
FUNDING:
The authors have no sources of outside funding to disclose.
Footnotes
CONFLICTS OF INTEREST DISCLOSURES: Dr. Ramgopal is supported by the Gerber Foundation (#9940). The authors have no conflicts of interest relevant to this article to disclose.
REFERENCES
- 1.OCEBM Levels of Evidence Working Group. The Oxford 2011 levels of evidence. The Centre for Evidence-Based Medicine. Published 2011. Accessed May 22, 2024. http://www.cebm.net/index.aspx?o=5653 [Google Scholar]
- 2.Digitale JC, Martin JN, Glymour MM. Tutorial on directed acyclic graphs. J Clin Epidemiol. 2022;142:264–267. PubMed doi: 10.1016/j.jclinepi.2021.08.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Hulley SB. Designing Clinical Research. Lippincott Williams & Wilkins; 2007. [Google Scholar]
- 4.Kaji AH, Schriger D, Green S. Looking through the retrospectoscope: reducing bias in emergency medicine chart review studies. Ann Emerg Med. 2014;64(3):292–298. PubMed doi: 10.1016/j.annemergmed.2014.03.025 [DOI] [PubMed] [Google Scholar]
- 5.International Committee of Medical Journal Editors. Recommendations for the conduct, reporting, editing, and publication of scholarly work in medical journals. Published 2024. Accessed May 22, 2024. https://www.icmje.org/icmje-recommendations.pdf [PubMed] [Google Scholar]
- 6.Williams DJ, Shah SS, Myers A, et al. Identifying pediatric community-acquired pneumonia hospitalizations: Accuracy of administrative billing codes. JAMA Pediatr. 2013;167(9):851–858. PubMed doi: 10.1001/jamapediatrics.2013.186 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Zozus MN, Pieper C, Johnson CM, et al. Factors affecting accuracy of data abstracted from medical records. PLoS One. 2015;10(10):e0138649. PubMed doi: 10.1371/journal.pone.0138649 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Ge J, Li M, Delk MB, Lai JC. A comparison of a large language model vs manual chart review for the extraction of data elements from the electronic health record. Gastroenterology. 2024;166(4):707–709.e3. PubMed doi: 10.1053/j.gastro.2023.12.019 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Feng R Optimized large language models accurately identify recurrence of vt after ablation from complex medical notes: Will chart review become obsolete? OpenReview. Published 2023. Accessed July 19, 2024. https://openreview.net/forum?id=n2nPeZ9VJ0 [Google Scholar]
- 10.Austin PC, White IR, Lee DS, van Buuren S. Missing data in clinical research: a tutorial on multiple imputation. Can J Cardiol. 2021;37(9):1322–1331. PubMed doi: 10.1016/j.cjca.2020.11.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Vagenas D, Totsika V. Modelling correlated data: Multilevel models and generalized estimating equations and their use with data from research in developmental disabilities. Res Dev Disabil. 2018;81:1–11. PubMed doi: 10.1016/j.ridd.2018.04.010 [DOI] [PubMed] [Google Scholar]
- 12.Greenland S, Senn SJ, Rothman KJ, et al. Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. Eur J Epidemiol. 2016;31(4):337–350. PubMed doi: 10.1007/s10654-016-0149-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.National Institutes of Health. HIPAA Privacy Rule: Information for researchers. Published 2007. Accessed May 22, 2024. https://privacyruleandresearch.nih.gov/pr_08.asp [Google Scholar]
- 14.Feldstein LR, Tenforde MW, Friedman KG, et al. ; Overcoming COVID-19 Investigators. Characteristics and outcomes of US children and adolescents with multisystem inflammatory syndrome in children (MIS-C) compared with severe acute COVID-19. JAMA. 2021;325(11):1074–1087. PubMed doi: 10.1001/jama.2021.2091 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Baumann BM, Rodriguez RM, DeLaroche AM, et al. Factors associated with parental acceptance of COVID-19 vaccination: A multicenter pediatric emergency department cross-sectional analysis. Ann Emerg Med. 2022;80(2):130–142. PubMed doi: 10.1016/j.annemergmed.2022.01.040 [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Nigrovic LE, Kuppermann N, Macias CG, et al. ; Pediatric Emergency Medicine Collaborative Research Committee of the American Academy of Pediatrics. Clinical prediction rule for identifying children with cerebrospinal fluid pleocytosis at very low risk of bacterial meningitis. JAMA. 2007;297(1):52–60. PubMed doi: 10.1001/jama.297.1.52 [DOI] [PubMed] [Google Scholar]
- 17.Cruz AT, Nigrovic LE, Xie J, et al. Predictors of invasive herpes simplex virus infection in young infants. Pediatrics. 2021;148(3):e2021050052. PubMed doi: 10.1542/peds.2021-050052 [DOI] [PubMed] [Google Scholar]
