Skip to main content
World Psychiatry logoLink to World Psychiatry
. 2026 Sep 15;25(3):423–424. doi: 10.1002/wps.70091

Is bigger better?

David S Baldwin 1
PMCID: PMC13576921  PMID: 42742587

Psychiatrists sense an enormous opportunity arising from the application into practice of findings from large population datasets. Such “big data” analytic approaches can accommodate the diversity and complexity of mental health problems and have potential to parse the subtle, cumulative effects of multiple adversities in the evolution of psychological distress and mental disorder. However, there is continuing uncertainty about the added value to clinical psychiatry of findings obtained from hypothesis‐light analyses of extensive databases, when compared to insights derived from hypothesis‐driven investigations in smaller samples. Put more simply, it is not yet clear whether “bigger is better”.

Stein et al, in their magisterial review, provide illustrations of this persisting uncertainty. They note, for example, that polygenic risk scores have not proved sufficiently predictive to merit translation into routine practice; that very few biological markers have demonstrated consistent utility in clinical settings; and that individual differences in neuroimaging measures of brain structure and function account for only a small proportion of the variance in mental disorders. Overall, they contend that overenthusiastic support for big data approaches must be tempered with a balanced understanding of their strengths and weaknesses 1 .

Sometimes, “small is beautiful”. Single‐case (“n‐of‐1”) methods based on serial measurement in an individual can be useful for testing theories and evaluating interventions, and have identified beneficial personalized interventions relating to drug and alcohol use, sleep, eating, and treatment adherence. Detailed consideration of individual narratives through incorporating qualitative research methods generates considerable insights into the lived experience of patients with psychiatric illness, to an extent not achievable by merely quantitative approaches 2 . And the finding, mentioned by Stein et al, of little difference in severity, impairment and comorbidity between generalized anxiety disorder of 6‐month duration and generalized anxiety symptoms of shorter durations, emerging from analysis of big data in community surveys involving 17 countries and 85,052 individuals, was established contemporaneously in a single‐centre cohort comprising just 591 participants 3 .

The proponents of using big data approaches to direct strategic decisions in business management cite their beneficial key attributes of volume, variety, velocity and veracity in enhancing impact. But, in clinical management, the availability of databases which are large, multi‐faceted and swiftly accessible may not necessarily inform the care of individual patients impactfully. Patient preference, clinician confidence, and local service availability are all relevant considerations, and insights derived from distantly accumulated datasets which cannot account for sometimes idiosyncratic and parochial concerns may have limited utility in determining treatment choices in clinical practice. Hence, an important challenge to implementing treatment decisions informed by big data approaches is how best to incorporate the views and opinions of patients and to acknowledge the practice‐based experience and internalized tacit “mindlines” of clinicians 4 .

Much of the anticipated success of big data approaches rests on securing widespread access to electronic health records in routine clinical practice. Stein et al provide some persuasive examples of the value of big data approaches centred on the records of treatment‐seeking individuals, in enhancing clinical practice (as examples, improved accuracy in prediction of suicide risk), and in implementation of psychotherapies for common mental health problems. But, as with other areas of medicine, there are persisting doubts about such records in psychiatric services. What is recorded is influenced by patient, clinician, service and societal factors. Individuals and cultures differ in their idioms for expressing distress; patients offer psychological complaints to clinicians in variable, reflexive ways, depending on their perception of support within the encounter; and reporting of symptoms is affected by illness‐related cognitive difficulties and emotion‐processing biases. Furthermore, health professionals differ in their assiduousness when asking questions about pivotal symptoms for accurate diagnosis: as an example, a national audit in England revealed marked variation between and within mental health services in the comprehensiveness of patient assessment 5 .

Other considerations regarding the use of electronic health records in big data analysis include representativeness, comprehensiveness and privacy. Socioeconomically deprived individuals are disadvantaged in gaining access to services, yet are more likely to experience mental health difficulties, and similar problems are seen in the disproportionately low recruitment of individuals from marginalized communities into mental health research. Data analyses based on treatment‐receiving and research‐participating patients may, therefore, have reduced applicability in some populations. Health records tend to focus on measurable processes such as consultations and prescriptions, but may neglect important determinants of mental well‐being, such as childhood adversity, trauma, diet, exercise, joblessness and family breakdown: experience from oncology research indicates that accurate characterization of this influential “exposome” is a substantial challenge 6 . In addition, there are widespread concerns about privacy protection when sharing health‐related data. Although strategies such as “federated learning” (training statistical models whilst keeping data localized) are a potential solution, there is often a “trade‐off” between maintaining privacy and optimizing model performance.

It is frequently stated that patients with psychiatric illness and their attending mental health professionals yearn for more appropriately targeted, individually tailored clinical care, designed to obviate “trial‐and‐error” treatment approaches, to enhance response rates, to alter the trajectory of illness, and to improve other clinical and personal outcomes. This “precision psychiatry” has the commendable goal of getting the right treatment to the right patient at the right time. Its questioning critics contend that, although well‐intentioned, its underlying premises are flawed and bemoan its reductionist approach to mental suffering 7 . Its advocates, by contrast, anticipate that additional insights from big data analysis will enhance understanding of the effects of adversity and the attributes of resilience, increase the validity and accuracy of psychiatric diagnosis, and so transform clinical practice; and they emphasize the pivotal importance of identifying biomarkers derived from such analyses and their subsequent widespread integration into practice 8 .

In Alzheimer's disease, biomarker identification and evaluation has refined diagnosis, modified clinical staging, informed drug development, and contributed to the approval of novel treatments. The chances of similar success in other psychiatric disorders will rest on establishing that any identified biomarkers have optimal sensitivity and specificity, reflect presumed pathophysiological processes, and can predict outcome trajectories accurately. Other pivotal characteristics of future biomarkers should include their acceptability to patients, their ease of incorporation into practice, and their being sufficiently cost‐effective in refining clinical decisions. A biomarker might be adopted widely if it could be used easily and inexpensively.

Stein et al allude to some methodological developments in big data science which may not be widely known among practising psychiatrists. The challenge of linking data from small surveys to those within larger case‐based registers can be met by algorithmic imputation of missing data when creating a “synthetic dataset” which can mimic real data, with sufficient granularity and variability whilst maintaining confidentiality. Electronic health records can permit “emulated target trials”, which overcome formal randomization by comparing observational data from similar groups of patients given differing treatments 9 . Such an approach may permit evaluations of real‐world comparative effectiveness and acceptability, and may also provide supportive data which might underpin subsequent trials of potentially repurposable medicines.

The opportunity provided by big data analytics, if welcomed with curiosity and flexibility but balanced by caution, could lead to a reconceptualization of psychiatric diagnosis, fuller understanding of pathophysiological processes, enhanced characterization of patients, delineation of personalized treatments, and – hopefully – improved clinical and personal outcomes. We can envisage a circular path of discovery, in which novel findings from explorations of big data confer an additional impetus to hypothesis‐driven mechanistic investigations, the insights from which lead in turn to further pre‐specified and pre‐registered explorations in extensive datasets. But it would be essential to ensure that the explanatory models which emerge from such a process are both complex enough to represent reality yet simple enough to prove useful in psychiatric practice.

REFERENCES


Articles from World Psychiatry are provided here courtesy of The World Psychiatric Association

RESOURCES