Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Feb 1.
Published in final edited form as: Gastroenterology. 2023 Dec 3;166(2):240–247. doi: 10.1053/j.gastro.2023.11.292

A Practical Guide to Evaluating and Using Big Data in Digestive Disease Research

Madeline Alizadeh 1, Natalia Sampaio Moura 2, Alyssa Schledwitz 2, Seema A Patil 2, Hashem El-Serag 3, Jacques Ravel 1, Jean-Pierre Raufman 2,4,5,6
PMCID: PMC10872385  NIHMSID: NIHMS1951572  PMID: 38052336

What is big data?

Across almost all domains of biomedical investigation, the past decade has witnessed an explosive increase in the availability and use of large datasets, so-called big data.1,2 Although the definitions of ‘big data’ vary, the term generally refers to massive amounts of complex information generated by humans or machines (e.g., artificial intelligence) that are too large to store, process, analyze, or interpret using conventional methods. Instead, due to their size and complexity, these large datasets must be processed and analyzed using highly sophisticated computer algorithms and programs.

Big datasets can be large with respect to sample sizes, characteristics, and/or variables. In human studies, sample size units commonly refer to the number of test subjects, clinics, hospitals, or regions, whereas in other experimental settings, units may refer to the number of cells, animals, microorganisms, or genes. Characteristics of big data also vary but commonly include the so-called ‘five V’s’: volume, variety, velocity, value, and veracity. Volume refers to the computing size of datasets (e.g., bytes), variety to the types of data, velocity to the pace at which data are generated and acquired, and value and veracity to the trustworthiness and reliability of the data. Variables are the descriptive quantitative or categorical units (e.g., number, amount) that can be counted or measured in these characteristics. Lastly, big datasets can be structured or unstructured. The former, commonly resulting from well-designed prospective studies, have predefined organizational structure that make the resulting data easier to sort and analyze. In contrast, unstructured datasets, which comprise the most common form of big data, lack predefined organizational structure and are less easily interpreted or analyzed using standard approaches (e.g., medical record entries, static or video imagery, genomic analyses).

When analyzed and interpreted using appropriate tools, these massive volumes of big data provide clinicians, scientists, and others with useful insights that can help them make informed decisions or pursue novel lines of investigation. Analyses of large datasets can be discovery based or non-targeted; parsing such immense datasets with these tools provides a pathway to discovery, for example, yielding the previously unrecognized association of a gene or genes with a particular disorder. Illustrative cases in point include the ability to perform transcriptomic analysis on a single-cell basis or to parse large electronic health records for meaningful associations.3 Large datasets can also be analyzed to pursue specific questions with targeted analysis allowing the efficient examination of rare conditions. The availability, analysis, and application of the resulting immense datasets not only allows, but in fact fosters the expansion of precision personalized medicine or population health and the tailoring of medical advances to unique individual or population characteristics.4 Nonetheless, assessing and applying this vast information correctly can be challenging, particularly for those who lack expertise in the relevant methodologies.

Here, focusing on the increasing application of big data in gastroenterology and hepatology research, we propose practical guidelines to help clinicians and investigators understand and critically appraise such studies, and avoid common pitfalls that can occur when using big data to answer research questions. Terms such as ‘data mining’ underestimate the complexity and rigor of the methods used to conduct these studies. We recommend a systematic approach to utilizing big data, including developing specific research questions and determining the best possible dataset and study design to answer the question.

The increasing use of big data in digestive disease research

Given the nexuses of the gut-brain and liver-brain axes, the extraordinary size and diversity of the gut microbiome and multi-ome, and the interplay of other organs, including the biliary tree and pancreas, obtaining and using big data poses unique challenges in gastroenterology and liver research.5 To understand and apply evidence obtained from such large datasets appropriately, a basic understanding of the techniques used to interrogate big datasets and their correct application, interpretation, and limitations is required.5,6 This knowledge is essential to judge the validity of the results, inferences, and conclusions of scientific studies, to apply information gleaned from large datasets to basic or clinical research, and to translate the findings to scientific investigation or clinical practice. In this guide, we offer practical approaches to the critical assessment of clinical and scientific literature using big data and, for those for whom it may apply, suggestions for devising scientifically rigorous big data studies.

Critically assessing studies that employ large datasets and advanced statistical methods

When reading and interpreting studies based on large datasets, it is important to address three key concerns (Table 1):

  • Was the study adequately powered to address the question(s) asked?

  • Was additional validation carried out to support the conclusions (claims)?

  • Was there an adequate description of the methods and their limitations (e.g., limited population diversity or study subjects predominantly of European or Asian ancestry)?

Table 1.

Tips for assessing research utilizing large datasets.

Concerns Potential resolution
Was the study adequately powered to address the question(s) posed? Address sample size and dimensionality of the dataset:
  • The more features (e.g., microbes, genes, proteins, clinical parameters) being assessed, the more individual samples (e.g., number of stool samples collected) and/or replicates (multiple samples from the same stool sample) needed.

  • Across datatypes, replicates can indicate many things – e.g., with scRNA-seq, multiple cells of each cell type will be in each sample.

Was additional validation required for the data to support the conclusions? If so, was it provided? Examine the stated hypotheses, the expected outcomes, and the clinical or functional significance of the outcomes. Look both at data and analysis quality checks (e.g., sensitivity analyses and model assessment, data visualization). If adequate, assure biological data and controls, with multiple complimentary datatypes, were used to validate mechanistic and/or causal claims.
Are the authors adequately transparent regarding the methods employed and their limitations? Check for reproducibility – are the protocols thorough, with complete information on factors such as sample collection, processing, and data analysis? This should include a high degree of detail. Discussion of limitations should assess the data: type, characteristic, generalizability, and pipeline and imputation methods/limitations. Finally, and perhaps most importantly, publicly available data in an online repository allows for analyses to be reproduced, and even for several sets of data to be compiled, enhancing the overall validity of results.

scRNA-seq, single cell RNA sequencing.

Sample Size Considerations

Sample size refers to the frequency of individuals in human studies or number of cells, animals, or experimental models used in non-clinical studies. In addition to sample size and the expected effect size, power calculations for non-targeted analyses in large datasets need to also consider the dimensionality of the dataset. Dimensionality is a statistical parameter referring to the number of variables in a dataset – in omics datasets, the numbers are often in the 100’s or 1000’s. Variables are the different characteristics or descriptive items captured for each sample in the dataset, and can be either quantitative or categorical, while features are similar in that they are data captured to be analyzed as predictor variables (and are sometimes used interchangeably with variables), but the term more frequently refers to variables that have been processed for use in modeling. Dimensionality reduction is a method by which the number of features being studied can be reduced while retaining the features that contribute most substantially to variance being studied in the dataset; techniques used include principal component analysis, linear discriminant analysis, and t-distributed neighbor embedding. But even then, dimensionality may remain high after using these methods. Because there may be a broad range in the dimensionality of a dataset, sample sizes can vary substantially depending on the omics type studied. For example, single cell RNA sequencing data includes thousands of cells per sample, so even while assessing more than 1,000 genes, biological replicates (i.e., samples using the same conditions) in the single digits may be sufficient. In contrast, microbiome studies assessing taxa using 16S rRNA amplicon sequencing data may result in a few hundred taxa across a cohort, so hundreds of samples may be required. Technical replicates, or repeated measures of the same samples (e.g., sampling a stool specimen multiple times) can also decrease the number of samples required.

On the other end of the spectrum, human genome-wide association studies (GWAS) interrogate thousands of gene loci per individual with anticipated small effect sizes. Thus, multiple potentially causal associations may be observed. Moreover, there is an additional large burden of confounders that is not seen in structured experimental studies − this results in requiring samples from several hundred or likely thousands of individuals to identify significant differences and, to correct for multiple comparisons and reduce false positive results, p-value cutoffs must be more stringent (e.g., p-value cutoff of 5 × 10−8). Even then, loci identification may not always detect a causal variant. For instance, while multiple studies reproducibly identified hundreds of loci associated with Crohn disease (CD), determining the precise location (rather than locus approximation) of risk variants has been difficult because of the expanded sample size requirement and the consequently large number of samples required for multiple loci identification when billions of base pairs are analyzed.7 While complex power calculations are a lesser concern for the analysis of large clinical datasets, the nature of the question(s) posed must still be considered.

Large sample size permits the detection of small differences in rare outcomes, but a potential disadvantage is that statistically significant differences may not be clinically meaningful. One should also not confuse large sample size with the number of outcomes of interest. For example, a study with a sample size of one million subjects that has only 30 outcome events (i.e., a rare cancer) is underpowered. The ability to adjust for potential confounders and effect modifiers is dependent on the number of outcome events, not on the entire underlying sample size. The statistical requirements of causal inference in observational studies may have different power requirements than a clinical trial establishing non-inferiority of a novel treatment. Such variation makes the development of an exact formula for power estimates difficult. Estimates are often derived from a combination of complex simulations and previously published studies using similar datasets.8 Tools such as PERMANOVA, MOPower, and powerSurvEpi, exist to aid in these calculations both for omics and clinical data, and a thorough description of this process should be present in the publication’s methods section.

Study Validation

To ensure the validity and accuracy of inferences and conclusions, the results must support the claims. Authors must explicitly state their a priori hypotheses, and address both the significance and meaningfulness of their findings. Whereas significance describes statistical outcomes, meaningfulness refers to the potential impact of significant results on biological understanding or clinical practice.9,10 Sensitivity analyses, data visualization, and assessment of data characteristics (e.g., distribution, residuals of a model) should be incorporated into analyses. Even if statistically significant, a small difference in outcomes may have little functional or clinical impact. If the authors are clear about these factors and potential limitations, one must then determine if their data answer the questions and supports the claims. For instance, without rigorous validation and supporting biological data, mechanistic assumptions regarding pathogenic effects based solely on transcriptional or compositional changes in the microbiome may be incorrect.5 In experimental studies, did the authors include multiple data types to support their assertions?11 For example, to make mechanistic claims based on a transcriptomic analysis, protein level confirmation with multi-step characterization of the proposed pathway is required. In observational studies, to establish causality for associations between off-label drug use and a clinical outcome, randomized prospective clinical trials should be considered.

Comprehensive Description of the Methodology and Limitations

Recognizing the limitations of big data research may initially appear daunting for those lacking the appropriate analytical or statistical background, but several guidelines may help overcome this obstacle. Verify that the methods section is comprehensive, and that the dataset is available for other investigators to analyze and confirm the reproducibility of results. Seek information in the publication or supplementary materials, including protocols for sample attainment and processing, data generation and data processing, script availability on an open-access repository (e.g., Github), the use of cross-validation for more complex models (i.e., taking a component of the resulting dataset for machine learning and the remainder for retesting to confirm initial results), and that data availability in online data repositories or elsewhere is clearly stated (a new requirement for funding from the National Institutes of Health). This last point is critical for efforts to reproduce results, cross-validation of different cohorts, and meta-analyses. Other important details include relaying the correct versions of software packages and data pre-processing steps12,13 and providing sample collection and processing information, e.g., tissue samples flash frozen immediately after collection may have richer metabolomic profiles, particularly for volatile compounds, than those frozen even a few hours later.5 A highly phenotyped cohort with appropriately selected controls permits a more complete and accurate assessment of confounders that may affect the observed results. Regarding limitations, seek discussion by the authors regarding data types (e.g., the five V’s), pipelines dealing with missing data and/or imputation, and the nature of the data (e.g., sparse data with many zero values versus dense data with many non-zero values), and, importantly, potential sources of bias in data generation.5,14 The guidelines summarized in Figure 1, can help readers evaluate the validity of big data research.

Figure 1. Key considerations to address when assessing big data.

Figure 1.

Studies utilizing big data should be transparent and rigorous regarding the key elements necessary to evaluate the work. The manuscript must provide comprehensive methods and results, describe sample size and power considerations appropriate for the type of study,incorporate some form of mechanistic validation, and the critical elements must be reproducible. While the specifics vary from study to study, each of these points can be assessed by addressing the considerations described in this schematic. Created with Biorender.com.

In big data clinical observational studies (e.g., the Heart and Estrogen/progestin Replacement Study, HERS), where the data were not collected with a specific research question in mind (i.e., unstructured), the completeness and accuracy of information used to define exposures, confounders, and outcomes of interest must be evaluated. For completeness, the reader or investigator should ask whether the data source captures all or most of the study population encounters. The dataset may be large because it contains millions of people, but information on confounders and outcomes may be sparse because of short or intermittent longitudinal follow up, or incomplete recording of diagnoses and laboratory data. In contrast, patients enrolled in a health maintenance organization (HMO) or Veterans Administration (VA) hospital, who are likely to receive all or the bulk of their care within these systems, are also likely to have most of their healthcare utilization captured. The accuracy of information in clinical datasets can be expressed by positive and negative predictive values.

Although most electronic health record-based data include diagnostic codes [e.g., International Classification of Disease (ICD) and Current Procedural Terminology (CPT) codes], their accuracy is a potential limiting factor for studies that use administrative databases and, therefore, should be evaluated prior to using these codes for research. The accuracy of diagnostic codes may vary depending on the disease as well as the database. For example, conditions that become symptomatic rapidly and are readily diagnosed, such as esophageal cancer, are unlikely to remain undiagnosed; therefore, the negative predictive value of codes indicating such conditions is high. To evaluate the accuracy of diagnostic and procedure codes, it is advisable to conduct a survey or chart validation of subjects nested within the study cohort identified in the database.

Crafting research questions that can be addressed using large datasets and omics analyses

Several factors should be considered when designing studies using large datasets. Although big data can provide powerful insights into biology and medicine, its use requires posing appropriate study questions, an informed study design, and an analytical plan that acknowledges big data complexity.15 Computational biologists recognize the steep learning curve for the analysis of large datasets; it is as difficult to learn “what to do” as “what not to do”. Mistakes common to both omics and clinical data analysis include omitting quality checks in the initial assessment of input data, failing to debug (i.e., failing to verify that the codes used are performing their intended functions without unintended consequences), neglecting to consider biological plausibility in framing analyses, and overfitting (i.e. when a model is fitted too closely to training data, thus losing generalizability), and using inappropriate models for data types. While the precise characteristics of these pitfalls are specific to the data type used, pattern recognition can help avoid common errors when designing studies that involve the generation of large datasets (Table 2).

Table 2.

Potential pitfalls and best practices in big data research.

Potential Pitfalls Examples Consequences Best Practices
Data characterization and quality assessment Unidentified missing data
Assuming incorrect data distribution (e.g., assuming normal distribution)
Skewed or a high degree of missing data may impact the ability to answer questions.19
Applying assumptions to inaccurate data distributions corrupt analysis, resulting in misleading results and unsound conclusions.
Develop familiarity with the data type being analyzed. Visualize data several ways to identify patterns and use complementary methods to ensure proper interrogation of key quality metrics.
Step-by-step validation of results Failure to verify the output after data filtration/normalization or not running code against multiple test scenarios Code may accumulate bugs that won’t be recognized if data aren’t verified at each step. Resulting unintended data removal or alteration may greatly impact results. Examine data output and integrity at each step in the analytical pipeline. Tweak codes to identify unintended consequences. Test input data and scenarios that may catch errors.
Biological context Misinterpreting random associations between variables
Not accounting for latent variables in analyses
Performing an increasing number of statistical tests results in a greater likelihood of false positive results if p-values are not corrected accordingly and biological context is not considered.
Signals may be overlooked due to suppressor variables, while confounding may identify an association due to a separate influencing variable affecting both variables of interest.
Review the literature to identify known associations between available data and metadata; apply this knowledge when crafting analyses.20 If an association doesn’t appear plausible, evaluate the data to identify an potential explanation.
Application of analytical methods Incorrect statistical tests for data types, datasets, or small sample sizes.21,22 Failing to verify that all statistical assumptions are met.
Overcomplicating methods or models
Inaccurate and statistically incorrect results can mislead and suggest false relationships by correlating unrelated variables. Violating statistical assumptions can skew the probability of accurate statistics, thus altering false positive and false negative result rates.23
Model overfitting (defined in terms of unfavorable ratio of number of predictor variables to outcome events) can result in false negatives and obscure otherwise strong signals.
Training in statistics and/or bioinformatics is important for working with large datasets. Verify assumptions and requirements for each test used and know the appropriate methods for different datasets and data types.

The initial step in experimental planning is to choose data types that can answer the question(s) posed. Different data types only provide a subset of information that may contain bias. For instance, if one opts to study the gut microbiome, one must choose a method of assessment. 16S rRNA amplicon sequencing offers a relatively inexpensive method that doesn’t require a huge bacterial input and is reliable at high sequencing depth. However, due to their discordance with abundance, one won’t know if the resulting compositional profiles represent living microbes, which microbes are transcriptionally active, or their gene functionality. Additionally, 16S rRNA amplicon sequencing is based on variable regions (e.g., V4 or V3V4) in the 16S rRNA gene, and offers poorer resolution than metagenomic sequencing, meaning species-level differentiation is often limited or isn’t possible.16,17 Shotgun metagenomic sequencing offers insights into functionality and aids strain typing. This may be difficult at adequate depth if DNA contamination from other species is overwhelmingly (e.g., human DNA in biopsy samples) and won’t assess the percentage of living microbes. Another example is the analysis of large clinical datasets – one must consider how data were obtained and their completeness and accuracy. For example, multiple study sites may enroll subjects for a study exploring the extra-intestinal manifestations of inflammatory bowel disease. If some sites inquire about joint pain at every visit while others do this less frequently, this can result in inconsistent recognition of associated spondyloarthritis and make it difficult to make appropriate site-to-site comparisons. Thus, considering and addressing these factors at the experimental planning stage is vital. Given the expansive opportunities for subject and data misclassification, incomplete recording of exposures and their outcomes, and potential confounders, the reader must be convinced that the methods and results are consistent and robust. Given the many different assumptions for accuracy and completeness of disease outcome and exposures, we recommend testing the robustness of findings by performing sensitivity analyses.

For rigorous validation, mechanistic confirmation of informatics- and computationally-derived findings is required. Many data types can help validate correlational and associative discoveries, e.g., genetic, genomic, transcriptomic, metabolomic, proteomic, epigenomic, clinical informatic, medical imaging, and microbiome data.5 The most appropriate model for filling a knowledge gap depends on the question(s) asked (Figure 2). If one is concerned with biomarker discovery, serum metabolomics may yield easily accessible information. If therapeutic development is the goal, pathway elucidation using a combination of transcriptomic, proteomic, and confirmatory bench approaches may be more appropriate. The translational potential of research outcomes must be considered when selecting which data types to employ.

Figure 2. Example of how complementary big data types can be leveraged to address a clinical problem.

Figure 2.

Whereas genetic backgrounds identify persons at higher risk of Crohn disease progression, epigenomic patterns permit longitudinal assessment and dynamic risk stratification. Transcriptomic data may validate genomic and epigenomic findings; used alone, gene expression patterns may not match genetic and epigenetic differences. Metabolomic and proteomic data can identify potential diagnostic markers; an especially useful asset in those with increased genetic susceptibility. Metagenomic analysis of gut microbiome abundance and composition identifies organisms that may contribute to disease progression as well as metabolic pathways that can be integrated longitudinally with metabolomic and transcriptomic data to develop predictive models and validate causality. Clinical informatics can be integrated with other data types to identify novel factors that can be tested in animal and ex vivo models. Imaging analysis can be integrated to augment complex models. US, ultrasound; MRI, magentic resonance imaging. Created with Biorender.com.

Multiple modalities may be required to fill mechanistic gaps and to ensure that consecutive ‘black boxes’ in mechanism chains are validated.5 To provide complementary information, multi-omic data generated using several methods can be analyzed, preferably in conjunction with clinical data. The sum of these data analyzed together offers more information synergistically than would be gleaned from each individual analysis − Figure 2 illustrates an example. To better understand the underlying susceptibility to disease progression and how it may be modulated by other factors, characteristics can be measured using complementary methods and then integrated to obtain actionable information. Identified genetic predispositions to CD progression can be integrated with longitudinal epigenomic and transcriptomic data; for instance, while the underlying genetic risks may be the same for progression to complex disease more generally, the transcriptomic patterns associated with penetrating versus stricturing CD are likely to differ. Metabolomic and proteomic data can be integrated to associate biomarkers with these changes, yielding a potential opportunity to intervene and modulate clinical outcomes. Incorporating deeply phenotyped, data-rich clinical profiles into this work enhances accuracy and specificity and is key to translating the outcomes into precision medicine.

The appropriateness of the methods chosen to analyze study results depends on the proposed goals of the study; nonetheless, to increase the robustness of the findings, using multiple approaches is encouraged.5 To identify the most suitable methods for attaining mechanistic confirmation of associative findings from big data analyses, key principles should be applied consistently. Decide what specific question(s) are being addressed by the proposed study. Without asking specific, focused questions, one runs the risk of amassing data without achieving a clear outcome for analysis. Vague questions result in vague answers and limited novel insights. An ambiguous question might be “What causes progression of CD from inflammatory to complex phenotype?” A more specific question would be, “What hereditary factors increase the likelihood of progressing from an inflammatory to complex CD phenotype?” Since identifying risk loci does not identify how these mutations confer risk, the question can be enhanced by reframing as “What hereditary factors increase the likelihood of progression from an inflammatory to complex phenotype, and how do they do so?”, and further specifying by asking “via what transcriptomic and proteomic mechanisms, and how do these differ between penetrating and stricturing CD?” Thus, one could connect the underlying risk factor to transcribed and translated output that modulates risk. The resulting information could more easily be used to design animal experiments to establish causative relationships by varying levels of a particular transcript and/or protein. An example of a specific clinical question is “Are levels of commonly measured biomarkers (e.g., CRP, fecal calprotectin, and fecal lactoferrin) altered prior to progression of CD from inflammatory to complex phenotype?”

Lastly, it is crucial to account for prior knowledge and bias at each step: when developing the experimental design, processing and analyzing the data, and drawing inferences. No data are truly “raw” − all have undergone some degree of manipulation, whether by the investigators or as a consequence of how they were generated - this can pose problems if it is not acknowledged.18 Such issues regarding data generation, processing, and analysis can yield results that contradict known biology and potentially entrench incorrect conclusions that bias future studies. Recognizing and acknowledging biases introduced by the approaches used to generate data allows investigators to conduct better science and draw more circumspect conclusions. Keeping biological plausibility in mind helps one recognize and address confounders that may mask true signals in a forest of noise.

Conclusions

The increasing availability and use of big data in all forms of research necessitates wider and deeper understanding of its generation, analysis, and interpretation. Given the centrality of the luminal GI tract and liver to multiorgan signal integration and the complex types of omics data originating from digestive organ tissue samples, this is particularly true for those engaged in digestive disease research. While the intricacies of big data use and interpretation may appear initially intimidating and potentially discouraging, applying the guidelines described above when reading the literature and analyzing and developing research studies is likely to instill confidence. For those interested in performing big data studies, adhering to a consistent set of guiding principles can help avoid common pitfalls and yield meaningful results that advance both scientific knowledge and clinical practice.

Acknowledgments

We thank Drs. Michael France and Bing Ma at the Institute for Genome Sciences at the University of Maryland, School of Medicine for reviewing the manuscript and providing helpful insights and advice.

Funding:

M.A. and N.S.M. were supported by an award from the National Institutes of Health, National Institute of Diabetes and Digestive and Kidney Diseases (T32 DK067872-19; J-P Raufman, PI).

Footnotes

Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

Disclosures: The authors have no conflicts to disclose.

References

  • 1.Mallappallil M, Sabu J, Gruessner A & Salifu M A review of big data and medical research. SAGE open medicine 8, 2050312120934839 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Martin-Sanchez F & Verspoor K Big Data in Medicine Is Driving Big Changes. Yearb Med Inform 23, 14–20 (2014). 10.15265/IY-2014-0020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Zeng H et al. Spatially resolved single-cell translatomics at molecular resolution. Science 380, eadd3067 (2023). 10.1126/science.add3067 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Manrai AK, Patel CJ & Ioannidis JPA In the Era of Precision Medicine and Big Data, Who Is Normal? JAMA 319, 1981–1982 (2018). 10.1001/jama.2018.2009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Alizadeh M et al. Big Data in Gastroenterology Research. Int J Mol Sci 24 (2023). 10.3390/ijms24032458 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Davila JA & El-Serag HB GI Epidemiology: databases for epidemiological studies. Aliment Pharmacol Ther 25, 169–176 (2007). 10.1111/j.1365-2036.2006.03207.x [DOI] [PubMed] [Google Scholar]
  • 7.Verstockt B, Smith KGC & Lee JC Genome-wide association studies in Crohn’s disease: Past, present and future. Clinical & Translational Immunology 7, e1001 (2018). https://doi.org: 10.1002/cti2.1001 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Ferdous T et al. The rise to power of the microbiome: power and sample size calculation for microbiome studies. Mucosal Immunol 15, 1060–1070 (2022). 10.1038/s41385-022-00548-1 [DOI] [PubMed] [Google Scholar]
  • 9.Wilkinson M Distinguishing Between Statistical Significance and Practical/Clinical Meaningfulness Using Statistical Inference. Sports Medicine 44, 295–301 (2014). 10.1007/s40279-013-0125-y [DOI] [PubMed] [Google Scholar]
  • 10.LeFort SM The Statistical versus Clinical Significance Debate. Image: the Journal of Nursing Scholarship 25, 57–62 (1993). https://doi.org: 10.1111/j.1547-5069.1993.tb00754.x [DOI] [PubMed] [Google Scholar]
  • 11.Perakakis N, Yazdani A, Karniadakis GE & Mantzoros C Omics, big data and machine learning as tools to propel understanding of biological mechanisms and to discover novel diagnostics and therapeutics. Metabolism 87, A1–A9 (2018). https://doi.org: 10.1016/j.metabol.2018.08.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Papin JA, Mac Gabhann F, Sauro HM, Nickerson D & Rampadarath A Improving reproducibility in computational biology research. PLOS Computational Biology 16, e1007881 (2020). 10.1371/journal.pcbi.1007881 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Sandve GK, Nekrutenko A, Taylor J & Hovig E Ten Simple Rules for Reproducible Computational Research. PLOS Computational Biology 9, e1003285 (2013). 10.1371/journal.pcbi.1003285 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Lay JO, Liyanage R, Borgmann S & Wilkins CL Problems with the “omics”. TrAC Trends in Analytical Chemistry 25, 1046–1056 (2006). https://doi.org: 10.1016/j.trac.2006.10.007 [DOI] [Google Scholar]
  • 15.Kliebenstein DJ Questionomics: Using Big Data to Ask and Answer Big Questions. The Plant Cell 31, 1404–1405 (2019). 10.1105/tpc.19.00344 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.France MT et al. Insight into the ecology of vaginal bacteria through integrative analyses of metagenomic and metatranscriptomic data. Genome Biology 23, 66 (2022). 10.1186/s13059-022-02635-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Franzosa EA et al. Relating the metatranscriptome and metagenome of the human gut. Proceedings of the National Academy of Sciences 111, E2329–E2338 (2014). 10.1073/pnas.1319284111 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Leonelli S The challenges of big data biology. Elife 8 (2019). 10.7554/eLife.47381 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Chicco D Ten quick tips for machine learning in computational biology. BioData Mining 10, 35 (2017). 10.1186/s13040-017-0155-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Greene CS, Tan J, Ung M, Moore JH & Cheng C Big data bioinformatics. J Cell Physiol 229, 1896–1900 (2014). 10.1002/jcp.24662 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Finlayson SG, Beam AL & van Smeden M Machine Learning and Statistics in Clinical Research Articles—Moving Past the False Dichotomy. JAMA Pediatrics (2023). 10.1001/jamapediatrics.2023.0034 [DOI] [PubMed] [Google Scholar]
  • 22.Morgan CJ Use of proper statistical techniques for research studies with small samples. American Journal of Physiology-Lung Cellular and Molecular Physiology 313, L873–L877 (2017). [DOI] [PubMed] [Google Scholar]
  • 23.Nimon K Statistical Assumptions of Substantive Analyses Across the General Linear Model: A Mini-Review. Frontiers in Psychology 3 (2012). 10.3389/fpsyg.2012.00322 [DOI] [PMC free article] [PubMed] [Google Scholar]

RESOURCES