Skip to main content
Transactions of the American Clinical and Climatological Association logoLink to Transactions of the American Clinical and Climatological Association
. 2023;133:56–68.

CLINICAL HETEROGENEITY IN THE AGE OF BIG DATA, ADVANCED ANALYTICS, AND COMPLEXITY THEORY

David Herrington 1,✉, Yue Wang 1
PMCID: PMC10493739  PMID: 37701617

ABSTRACT

Clinical heterogeneity remains a challenge in the practice of medicine and is an underlying motivation for much of biomedical research. Unfortunately, despite an abundance of technologies capable of producing millions of discrete data elements with information about a patient's health status or disease prognosis, our ability to translate those data into meaningful improvements in understanding of clinical heterogeneity is limited. To address this gap, we have applied newer approaches to manifold learning and developed additional and complementary techniques to interrogate and interpret complex, high dimensional omics data. The central premise is that there exist manifolds embedded in high dimensional data that represent fundamental biologic processes that may help address the challenges of clinical heterogeneity. Preliminary evidence from several real-world data sets suggests that these techniques can identify coherent and reproducible manifolds embedded in higher dimensional omics data. Work is currently ongoing to determine the clinical informativeness of these novel data structures.

INTRODUCTION

Clinical heterogeneity, defined as variations in risk factors, clinical manifestations, response to therapy, or prognosis for a given disease, has been a vexing problem for clinicians and a motivation for biomedical investigators throughout history. Clinical judgment alone is often insufficient to navigate the nuances in clinical heterogeneity encountered in the practice of medicine. As early as 400 BC, the Hippocratic Corpus included the famous acknowledgement that for physicians “… experience is fallacious and judgment difficult” (1). At the turn of the last century, William Osler wrote, “Variability is the law of life, and as no two faces are the same, so no two bodies are alike, and no two individuals react alike and behave alike under the abnormal conditions which we know as disease” (2). A contemporaneous review of the biomedical literature, including the Transactions of the American Clinical and Climatological Association, reveals many titles focused on sub-phenotyping, tumor heterogeneity, targeted therapies, and biomarkers to improve disease prediction—all related to the underlying clinical and scientific challenges posed by clinical heterogeneity.

Over time, the medical community has offered various strategies in an attempt to optimize clinical decision making in the face of clinical heterogeneity. For years, the best approach was to rely on “master clinicians” whose intellect, learning, and accumulated experience could discern the proper diagnosis or course of care for their individual patients. Beginning in the 1950s, the use of objectively measurable risk factors to assist in risk assessment emerged, perhaps best exemplified by the work of the Framingham Heart Study investigators (3). However, despite the rapid integration of risk factors into clinical decision making, it remained clear that even multifactor risk scores such as the Framingham Risk Score were insufficient to account fully for heterogeneity in the development of common diseases such as coronary heart disease. Beginning in the 1990s, considerable effort was made to identify additional “novel” risk factors that would complement the information in conventional risk factors with respect to patient risk and/or outcomes. Nevertheless, despite great effort and extensive evaluation of many novel biomarkers, including blood analytes (e.g., C-reactive protein), anatomic markers (e.g., carotid intimal-medial thickness), etc., our ability to predict disease or response to therapy continued to be insufficiently sensitive or specific to meaningfully improve clinical care.

The report by the International Human Genome Sequencing Consortium in 2004 of the sequence of the human genome (4) ushered in a new era in biomedical research by providing an entirely new class of high dimensional data (genomics) to help understand variation in human health and disease. Since then, several technologies have rapidly evolved to permit routine generation of not only genetic variants and even the full genomic sequence but also a wide array of other classes of high dimensional data including transcriptomics, epigenetics, microRNA data, proteomics, metabolomics, high resolution imaging data, etc., in individual patients. Even the medical record has evolved to provide thousands of variables of a different sort for each of our patients. Collectively, these new technical platforms make it possible to characterize individual patients with many million discrete data elements that harbor unique information about their health status or disease prognosis. This wealth of information has created great expectations that we would be able to use these data to provide personalized or precision medical care to each individual patient (5).

Indeed, for selected clinical conditions such as rare monogenetic disorders or certain somatic mutations in tumors, the goals of precision medicine have been partially realized. This is best evidenced by the increasing number of new therapeutic molecular entities approved by the FDA that are specifically designed for patients who meet narrowly defined criteria defined by certain tumor or blood biomarkers (6). However, for the vast majority of common complex clinical conditions such as atherosclerotic vascular disease, hypertension, osteoarthritis, or dementias, our ability to generate high dimensional data on individual patients far exceeds our ability to use those data to dramatically improve clinical decision making or prediction of future outcomes. Many individual “omics” features [e.g., single nucleotide polymorphisms (SNPs), proteins, metabolites, etc.] are statistically associated with clinical traits when studied in large cohorts or populations. However, the prediction performance of these features for common clinical diseases or outcomes is typically limited (7-10). Somewhat surprisingly, the limited prediction performance of individual omics features may not be overcome by creating linear combinations of large numbers of features, each with small effect. For example, combining information from hundreds to millions of SNPs included in a cardiovascular disease polygenic risk score does not dramatically improve prediction performance over existing conventional risk factors (11-13). Likewise, efforts to combine information from multiple omics platforms (e.g., integrative omics, multi-omics, etc.), though extremely valuable from the point of view of molecular mechanisms and pathways, still have failed to produce clinical tools to meaningfully address the problems of clinical heterogeneity.

One of the challenges of mapping omics features to complex clinical traits has to do with redundancy, degeneracy, and feedback—attributes of biologic systems at all levels of organization (14). The fact that distinct molecular structures can produce identical molecular outcomes (redundancy), or that identical molecular structures can produce distinct molecular outcomes in different circumstances (degeneracy), and that most molecular events are subject to positive and/or negative feedback means that even simple molecular pathways are hard to assess with simple, static measurements of the elements of the system. The effects of degeneracy, redundancy, and feedback are even more dramatic when considering hundreds or thousands of molecules that interact within an organic whole. In this setting, the local interactions among individual elements in a biologic system can produce complex adaptive systems with emergent behaviors that are difficult to anticipate based on a simple assessment of the individual elements themselves (15,16). For these reasons, it is naïve to believe that simple linear -models, as are often found in contemporary omics and biomarker -discovery research, are sufficient to map individual omics features to the complex clinical traits that are at the heart of most clinical -decision-making challenges.

A number of analytic frameworks have been explored to overcome the challenges of mapping omics data to complex clinical traits. Among them, manifold learning is emerging as a particularly useful approach. A manifold is a locally smooth, lower dimensional structure embedded in higher dimensional space (see Figure 1). The underlying assumption is that the “state-space” of the object of study (e.g., cells, organisms, humans, or even populations) defined by high dimensional omics data can be summarized by a lower dimensional structure or “manifold.” Furthermore, the location of data points on this manifold are informative about the current state, while the topology of the manifold is informative about possible transitions from one state to another (17). Numerous examples of this approach are applied to RNA-sequencing data to gain deep insights about cellular plasticity during hematopoiesis (18) or tumorigenesis (19). Recently, our lab has been devoted to the development of methods to apply these manifold learning principles to the analysis of large-scale molecular epidemiology data (20-22) with the expectation that these data structures will harbor important information about underlying biologic processes that contribute to clinical heterogeneity.

Fig. 1.

Fig. 1.

Nonlinear manifolds: First described by Bernhard Reimann, manifolds are an example of nonlinear dimensional reduction designed to visualize or describe a lower dimensional data structure embedded in higher dimensional space. A Reimann manifold exploits the concept that the embedded data structure is locally continuous but may be globally nonlinear. In this framework, data points may appear close in conventional Euclidian space, but far apart on the surface of the manifold. Reimann geometry has proven useful in numerous applications including Einstein's theory of general relativity, time series analyses, machine learning, and mathematical finance. Adapted from public domain material on Wikimedia Commons at https://commons.wikimedia.org.

MATERIALS AND METHODS

Plasma lipidomic and metabolomic data were collected from a subset of participants in the Multi-Ethnic Study of Atherosclerosis (MESA) in the United States (23) (n=3,748) and a similar subset of participants from the Airwave Health Monitoring Study, a large civil servant study from the United Kingdom (24) (n=2,997). Both studies generated plasma lipidomic profiles based on 1H NMR spectra acquired at 600 MHz (Bruker DRX600) followed by lipoprotein subclass analysis (Bruker B.I.-LISA platform). The data include 104 variables providing estimates of the lipid and lipoprotein composition for subfractions of the major lipoprotein particles in human plasma.

Similarly, untargeted mass-spec (MS) metabolomics data were acquired from the same plasma samples. Hydrophilic and lipophilic separation using ultra-high performance liquid chromatography, followed by MS analysis in both positive and negative ion mode for lipids and in positive mode for lipid analyses, was performed. Altogether, there were n=4092 MS metabolomic features available for analysis in each cohort.

Unreliable samples and metabolites were removed based on missingness and extreme outliers. Normalization and analytical batch effect correction were performed to harmonize sample means. For the MESA data, Convex Analysis of Mixtures (CAM) (22) was used to decompose the NMR and MS metabolomic data into discrete latent features with optimization based on minimum description length criteria implemented in the CAM software. The optimal number of latent features in the MESA data determined the rank for a similar deconvolution of the corresponding Airwave data. The RV coefficient (25-27) was used to assess the similarity of the latent feature source matrices from the two cohorts. The latent features from the MS data were also used as input to Phate (28) to capture both local and global nonlinear structures and to produce a 3-D projection of the state-space landscape defined by the latent features. In this visualization approach, each individual subject is mapped to a location on a 3-D manifold whose topology is dictated by the information embedded in the CAM latent features.

RESULTS

Initially, we used Convex Analysis of Mixtures (CAM) (22,29) to decompose the nuclear magnetic resonance lipoprotein profiles from the MESA participants into seven (n=7) latent features that represented the data—assuming an overall convex data structure. Independently, we decomposed the corresponding Airwave data into seven (n=7) distinct latent features using the same deconvolution approach. Comparison of the source matrices for the latent features from each population using the RV coefficient (26,27) revealed a high degree of similarity between the two populations (p < 0.0001), suggesting that the independent deconvolutions converged to a common underlying manifold embedded in the plasma lipidomic data space.

Similarly, when examining the high dimensional MS data from MESA and Airwave participants, CAM deconvolution produced 12 distinct latent sources. As with the NMR lipidomic data, the latent sources from MESA and Airwave were highly correlated (p < 0.0001, Figure 2), suggesting that they are the product of a common underlying manifold. Phate 3-D projections of the CAM data yielded qualitatively similar results (28) (Figure 2), suggesting a common state-space landscape that efficiently summarizes a vast amount of metabolomic data as a simple low dimension manifold.

Fig. 2.

Fig. 2.

Examples of Latent Features and Manifold Learning in Metabolomic Data from MESA and Airwave. Panel A: Correlation Between Metabolomic Latent Features in MESA and Airwave. CAM was used to deconvolve the MS metabolomic data into (n=12) latent features in each cohort. Each cell in the heatmap summarizes the correlation between individual metabolomic feature contributions to each latent feature in the two cohorts. The RV coefficient with a permutation test was used to assess the overall similarity of the sources matrices (p < 0.0001). Panel B: MESA MS Metabolomic State-Space. Phate 3-D projection of the metabolomic state-space among 3074 MESA participants derived from MS metabolomics data (Lipid + , Lipid -, HILIC + ). Light blue dots indicate participants with elevated triglyceride cholesterol concentrations. Panel C: Airwave MS Metabolomic State-Space. Phate 3-D projection of the metabolomic state-space among 2997 Airwave participants derived from similar MS metabolomics data (Lipid + , Lipid -, HILIC + ). The similar low rank state-space topologies in MESA and Airwave are consistent with a common underlying metabolomic manifold.

DISCUSSION

The challenges of mapping high dimensional omics data to complex clinical conditions is one of the barriers to realizing the full potential of personalized or precision medicine. Simple linear models of individual omics features, or even aggregations of omics features, are insufficient to describe complex adaptive systems and their underlying biologic processes. The overall goal of the current research is to use state-of-the-art analytic tools to identify manifolds embedded in high dimensional metabolomic data that may be more informative about underlying canonical metabolic processes that are jointly contributing to clinical traits of interest.

Here we explore the use of a novel, unsupervised, deconvolution strategy developed by our team to decompose the metabolomic profiles into distinct latent features and a complementary tool to visualize the resulting data as a 3-D state-space manifold. If the latent features or manifolds detected by these methods truly reflect fundamental biologic processes, we expect to see similar data structures embedded in data from independent cohorts. Accordingly, the initial focus of this work was to explore the reproducibility of the data structures in two separate cohort studies: MESA and Airwave. The source matrices derived from the CAM analysis reflect the contributions of individual metabolomic features to each of the latent features. We used the RV coefficient as a simple metric to assess the similarity of the source matrices derived from the two cohorts and a permutation p-value to evaluate the degree to which the observed similarities might be expected to occur due to chance alone. In both the NMR and MS metabolomic data, the latent features from MESA and Airwave were highly similar. Likewise, the state-space manifold derived from the MS latent features had similar visual topologies.

Encouraged that this approach is identifying reproducible structures embedded in the high dimensional metabolomic data, our current focus is to determine the physiologic significance of these latent features and to determine whether the metabolomic state-space manifold is useful to describe the range of cardiovascular health status evident in the MESA and Airwave cohorts. Our underlying hypothesis is that the latent features that determine the dimensionality of this manifold and the topology of the metabolic state-space landscape will be uniquely informative about aspects of metabolic health and disease that may be more difficult to discern based on measures of individual metabolites or even correlated groups of metabolites. Specifically, we are examining the extent to which these latent features mediate the relationship between known genetic variants associated with coronary heart disease and actual incident cardiovascular events or help to identify patients with impaired fasting glucose who will progress to fully manifest diabetes mellitus.

The deconvolution approach used here is distinctly different from several alternative methods. Classical principal components, and many of their derivatives, use orthogonal projections to account for the greatest amount of variance in the observed data. However, the resulting latent features are often difficult to interpret because of both positive and negative eigenvalues and the orthogonality assumption. Independent Components Analysis (ICA) (30) assumes that the latent features are statistically independent, which is unlikely to be well suited for certain biologic processes that may be synergistic or interact in other ways. Nonnegative matrix factorization (NMF) (31) is perhaps the closest analytic approach to CAM. This approach seeks parsimonious solutions that are nonnegative and is not limited to orthogonal solutions. The limitation of NMF is that the parsimonious solutions, by definition, will not encode contributions from all possible features in the data. CAM relies on the assumption that the multidimensional data reside in a convex hull with vertices that represent extreme values in the embedded latent features (32). Extensive simulation and real-world data have demonstrated the advantages of CAM over many of its competitors (22).

Many limitations should be considered in this work. Among other things, the cohorts themselves are not fully representative of the entirety of adult humanity. There are certainly people with important metabolomic latent features or underlying metabolic processes and exposures that are not reflected in these study participants. More data are needed on a more diverse array of people to determine the generalizability of the latent features described here. Likewise, the metabolomic analytical platforms only sample certain parts of the entire metabolome. Other analytic platforms better optimized for other classes of metabolites will likely yield different latent features and corresponding state-space topologies. Finally, all omics data, and especially metabolomics data, are not immune to confounding, unstable estimates from collinearity, and hidden sources of technical variation.

In summary, these data suggest that advanced analytic approaches are capable of identifying common metabolomic manifolds embedded in high dimensional metabolomics data from separate cohorts. More work is needed to determine if these manifolds are more generalizable across other cohorts and whether they are informative about canonical metabolic processes that underlie some aspect of clinical heterogeneity. Ultimately, the goal is to find new measures derived from metabolomic, or other omic, profiles that will lead to novel strategies for treatment or prevention of common, complex clinical conditions.

ACKNOWLEDGMENTS AND FINANCIAL SUPPORT

This work was funded in part by the National Institutes of Health under Grants HL111362-05A1 and HL133932 and EU COMBI-BIO project (FP7, 305422).

MESA and the MESA SHARe project are conducted and supported by the National Heart, Lung, and Blood Institute (NHLBI) in collaboration with MESA investigators. Support for MESA is provided by contracts 75N92020D00001, HHSN268201500003I, N01-HC-95159, 5N92020D00005, N01-HC-95160, 75N92020D00002, N01-HC-95161, 75N92020D00003, N01- C-95162, 75N92020D00006, N01-HC-95163, 75N92020D00004, N01-HC-95164, 75N92020D00007, N01-HC-95165, N01-HC-95166, N01-HC-95167, N01-HC-95168, N01-HC-95169, UL1-TR-000040, UL1- TR-001079, UL1-TR-001420, UL1-TR-001881, and DK063491.

The Airwave Health Monitoring Study was funded by the Home Office (grant number 780-TETRA; 2003–2018) and is currently funded by the Medical Research Council and the Economic and Social Research Council (MR/R023484/1) with additional support from the National Institute for Health Research (NIHR) Imperial College Biomedical Research Centre.

DISCUSSION

Firestein, San Diego: That was a terrific summary, and I appreciate the great graphics. One of the things that we need to consider is the vagary of the data that come from the electronic health record (EHR) when trying to make clinical correlations. Although some of the data are high quality, experience dictates that ICD-10 coding or other issues with structured data introduce variability. For example, in the Accrual to Clinical Trials platform, which links about 60 or so clinical data warehouses across the country, it's clear that we have to account for missingness and a host of other issues that can have a major impact on clinical phenotype. Could you briefly comment on how you are going to take that into account when making clinical correlations with equally complex data sets?

Herrington, Winston-Salem: Yes, that's a superb and insightful question. You know the famous quote from Shakespeare, you can't make a silk purse from a sow's ear. There's no question the quality of EHR data is a fundamentally important issue for many clinical research efforts. Indeed, these days EHR data are used to define exposures as well as phenotypes or clinical outcomes. Both can be adversely impacted by poor-quality EHR data. One of the motivations of our work is to develop new ways to define exposures or phenotypes and outcomes, based on omics and other high dimensional data that can overcome these limitations in clinical research.

Firestein, San Diego: Thanks very much.

Wilson, Durham, NC: I just want to ask one question about the ethnic and racial diversity of subjects in your data. We've heard about trying to impact especially disadvantaged neighborhoods and patients who may not have been included in the data set and may suffer the possibility of being excluded from your ultimate determinations. Do you have any comments on that?

Herrington, Winston-Salem: I’m happy to say that a lot of the data we use come from the Multi-Ethnic Study of Atherosclerosis which is one of the NHLBI-funded studies that explicitly attempted to achieve greater race and ethnic diversity. Interestingly, although we look routinely to see whether or not there are differences in the configuration of the manifold or the composition of the latent features, we rarely find any. At a more fundamental biologic level, there may be fewer differences than you might expect. This suggests that other factors may be at play beyond the basic canonical elements of molecular biology, such as environmental and social exposures.

Wilson, Durham, NC: Yes, but you risk looking specifically at one demographic and not the entire population. Our country is so different from the European countries with our multiple ethnicities and races as you know. If we rely on data from European countries, we may be just comparing a segment of our population with their population; only those of European descent.

Herrington, Winston-Salem: Excellent point.

Schwartzstein, Boston: I make no claims to understanding the complexity of your analysis but when I work with some of the big data with people at my center, I worry that this feels like pattern recognition on steroids sometimes and we are reducing medicine to this sort of analysis. Is that the future in your view?

Herrington, Winston-Salem: Well, I’m reminded of a quote from Niels Bohr who said the only real science is physics. Everything else is just stamp collecting. I would be the first one to admit that this is very much at the stamp collecting end of the spectrum at the moment. I don't mean to diminish critically important work that goes on to understand each of the individual nodes in these complex networks and the fundamental biologic relationships that exist between them. However, I think we also have to acknowledge that the clinical traits we are most keenly interested in are the product of a vast, complex array of relationships and we need to develop new techniques to look at biologic and clinical questions on that scale as well. This work is an attempt not to look at molecular biology on more of a macroscopic scale, but to do so in a more sophisticated way than just tables of rank-order p-values.

Schwartzstein, Boston: Thank you.

Arnaout, Boston: My understanding is that changes in the transcriptome or the proteome in the same individual can happen very quickly if you move that individual from one environment into another. I wonder if you focus on that in the same individual or a group of individuals whether that will give you faster insight into the rewiring process that makes us all adaptive to our environments.

Herrington, Winston-Salem: I think it is an urgent priority to go beyond static states and look at the dynamics of these systems that you're referring to. As you know, lots of work has been done. We've done work looking at network rewiring, but what we really need to understand is how a network goes from one configuration to another and therein lies another critical part of biology, particularly the behavior of complex adaptive networks that we're still relatively blind to and need to focus more attention on.

Arnaout, Boston: Thank you.

Footnotes

Potential Conflicts of Interest: None Disclosed.

REFERENCES

  • 1. Hippocrates MA: Loeb Classical Library, 1931.
  • 2.Osler W. On the Educational Value of the Medical Society. Boston Med and Surg J. 1903;148(11):275–9. [Google Scholar]
  • 3.Dawber TR, Kannel WB, Revotskie N, Stokes J, 3rd, Kagan A, Gordon T. Some factors associated with the development of coronary heart disease: six years’ -follow-up experience in the Framingham study. Am J Public Health Nations Health. 1959;49(10):1349–56. doi: 10.2105/ajph.49.10.1349. (In Eng.) doi:10.2105/ajph.49.10.1349. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.International Human Genome Sequencing Consortium Finishing the euchromatic sequence of the human genome. Nature. 2004;431(7011):931–45. doi: 10.1038/nature03001. doi: 10.1038/nature03001. PMID: 15496913. [DOI] [PubMed] [Google Scholar]
  • 5.Snyderman R. Personalized health care: from theory to practice. Biotechnol J. 2012;7(8):973–9. doi: 10.1002/biot.201100297. (In Eng.) doi:10.1002/biot.201100297. [DOI] [PubMed] [Google Scholar]
  • 6. Personalized Medicine at FDA: The Scope & Significance of Progress in 2021. Personalized Medicine Coalition; 2021.
  • 7.Goldstein DB. Common genetic variation and human traits. N Engl J Med. 2009;360(17):1696–8. doi: 10.1056/NEJMp0806284. (In Eng.) doi:10.1056/NEJMp0806284. [DOI] [PubMed] [Google Scholar]
  • 8.Ko D, Benson MD, Ngo D, et al. Proteomics profiling and risk of new-onset atrial fibrillation: Framingham Heart Study. J Am Heart Assoc. 2019;8(6):e010976. doi: 10.1161/JAHA.118.010976. (In Eng.) doi:10.1161/jaha.118.010976. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Samani NJ, Erdmann J, Hall AS, et al. Genomewide association analysis of coronary artery disease. N Engl J Med. 2007;357(5):443–53. doi: 10.1056/NEJMoa072366. (In Eng.) doi:10.1056/NEJMoa072366. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Tzoulaki I, Castagné R, Boulangé CL, et al. Serum metabolic signatures of coronary and carotid atherosclerosis and subsequent cardiovascular disease. Eur Heart J. 2019;40(34):2883–96. doi: 10.1093/eurheartj/ehz235. (In Eng.) doi:10.1093/eurheartj/ehz235. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Elliott J, Bodinier B, Bond TA, et al. Predictive accuracy of a polygenic risk score-enhanced prediction model vs a clinical risk score for coronary artery disease. JAMA. 2020;323(7):636–45. doi: 10.1001/jama.2019.22241. (In Eng.) doi:10.1001/jama.2019.22241. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Marston NA, Pirruccello JP, Melloni GEM, et al. Predictive utility of a coronary artery disease polygenic risk score in primary prevention. JAMA Cardiol. 2023;8(2):130–7. doi: 10.1001/jamacardio.2022.4466. doi: 10.1001/jamacardio.2022.4466. PMID: 36576811; PMCID: PMC9857431. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Mosley JD, Gupta DK, Tan J, et al. Predictive accuracy of a polygenic risk score compared with a clinical risk score for incident coronary heart disease. JAMA. 2020;323(7):627–35. doi: 10.1001/jama.2019.21782. (In Eng.). doi:10.1001/jama.2019.21782. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Edelman GM, Gally JA. Degeneracy and complexity in biological systems. Proc Natl Acad Sci U S A. 2001;98(24):13763–8. doi: 10.1073/pnas.231499798. (In Eng.) doi:10.1073/pnas.231499798. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Levin SA. Complex adaptive systems: exploring the known, the unknown and the unknowable. Bull of Am Math Soc. 2003;40(1):3–19. doi:doi.org/10.1090/S0273-0979-02-00965-5. [Google Scholar]
  • 16.Weaver W. Science and complexity. Am Scientist. 1948;36(4):536–44. Available at: http://www.jstor.org/stable/27826254. [PubMed] [Google Scholar]
  • 17.Moon KR, Stanley JS, Burkhardt D, van Dijk D, Wolf G, Krishnaswamy S. Manifold learning-based methods for analyzing single-cell RNA-sequencing data. Cur Opin in Sys Biol. 2018;7:36–46. [Google Scholar]
  • 18.Wolf FA, Hamey FK, Plass M, et al. PAGA: graph abstraction reconciles clustering with trajectory inference through a topology preserving map of single cells. Genome Biology. 2019;20(1):59. doi: 10.1186/s13059-019-1663-x. doi:10.1186/s13059-019-1663-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Burkhardt DB, San Juan BP, Lock JG, Krishnaswamy S, Chaffer CL. Mapping phenotypic plasticity upon the cancer cell state landscape using manifold learning. Cancer Discov. 2022;12(8):1847–59. doi: 10.1158/2159-8290.CD-21-0282. (In Eng.) doi:10.1158/2159-8290.Cd-21-0282. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Chen L, Lu Y, Wu CT, et al. Data-driven detection of subtype-specific differentially expressed genes. Sci Rep. 2021;11(1):332. doi: 10.1038/s41598-020-79704-1. (In Eng.) doi:10.1038/s41598-020-79704-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Chen L, Wu CT, Lin CH, et al. swCAM: estimation of subtype-specific expressions in individual samples with unsupervised sample-wise deconvolution. Bioinformatics. 2022;38(5):1403–10. doi: 10.1093/bioinformatics/btab839. (In Eng.) doi:10.1093/bioinformatics/btab839. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Chen L, Wu CT, Wang N, Herrington DM, Clarke R, Wang Y. debCAM: a bioconductor R package for fully unsupervised deconvolution of complex tissues. Bioinformatics. 2020;36(12):3927–9. doi: 10.1093/bioinformatics/btaa205. (In Eng.) doi:10.1093/bioinformatics/btaa205. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Bild DE, Bluemke DA, Burke GL, et al. Multi-ethnic study of atherosclerosis: objectives and design. Am J Epidemiol. 2002;156(9):871–81. doi: 10.1093/aje/kwf113. (In Eng.) doi:10.1093/aje/kwf113. [DOI] [PubMed] [Google Scholar]
  • 24.Elliott P, Vergnaud AC, Singh D, Neasham D, Spear J, Heard A. The Airwave Health Monitoring Study of police officers and staff in Great Britain: rationale, design and methods. Environ Res. 2014;134:280–5. doi: 10.1016/j.envres.2014.07.025. (In Eng.) doi:10.1016/j.envres.2014.07.025. [DOI] [PubMed] [Google Scholar]
  • 25.Escoufier Y. Le Traitement des Variables Vectorielles. Biometrics. 1973;29(4):751–60. doi:10.2307/2529140. [Google Scholar]
  • 26.Josse J, Pagès J, Husson F. Testing the significance of the RV coefficient. Computational Statistics & Data Analysis. 2008;53(1):82–91. doi: 10.1016/j.csda.2008.06.012. doi: [DOI] [Google Scholar]
  • 27.Kazi-Aoual F, Hitier S, Sabatier R, Lebreton J-D. Refined approximations to permutation tests for multivariate inference. Comp Stat & Data Anal. 1995;20(6):643–56. doi: 10.1016/0167-9473(94)00064-2. doi: [DOI] [Google Scholar]
  • 28.Moon KR, van Dijk D, Wang Z, et al. Visualizing structure and transitions in high-dimensional biological data. Nature Biotechnology. 2019;37(12):1482–92. doi: 10.1038/s41587-019-0336-3. doi:10.1038/s41587-019-0336-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Wang N, Hoffman EP, Chen L, et al. Mathematical modelling of transcriptional heterogeneity identifies novel markers and subpopulations in complex tissues. Sci Rep. 2016;6:18909. doi: 10.1038/srep18909. (In Eng.). doi:10.1038/srep18909. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Jutten C, Hérault J. Blind separation of sources, part I: an adaptive algorithm based on neuromimetic architecture. Signal Process. 1991;24:1–10. [Google Scholar]
  • 31.Lee DD, Seung HS. Learning the parts of objects by non-negative matrix factorization. Nature. 1999;401(6755):788–91. doi: 10.1038/44565. doi:10.1038/44565. [DOI] [PubMed] [Google Scholar]
  • 32.PV DKM. The Simplex Geometry of Graphs. arXiv. 2019;1807:06475. [Google Scholar]

Articles from Transactions of the American Clinical and Climatological Association are provided here courtesy of American Clinical and Climatological Association

RESOURCES