Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Mar 1.
Published in final edited form as: JACC Heart Fail. 2026 Jan 5;14(5):102851. doi: 10.1016/j.jchf.2025.102851

Artificial Intelligence to Extract Structured Details from Unstructured Medical Records in a Global Heart Failure Trial

Samarra Badrouchi 1, Pablo M Marti-Castellote 1, Alberto Foà 1,2, Brian L Claggett 1, Akshay S Desai 1, Dongchu Xu 1, Jennifer E Ho 3,4, Patrick Ellinor 3,5, Pardeep S Jhund 6, Martina M McGrath 7, Finnian R McCausland 7, John J V McMurray 6, Scott D Solomon 1, Jonathan W Cunningham 1,3
PMCID: PMC12949679  NIHMSID: NIHMS2149699  PMID: 41493411

INTRODUCTION

Global clinical trials collect extensive unstructured medical records, but their narrative format precludes quantitative analysis. Converting these records into structured data could reveal insights into events like heart failure hospitalization (HFH) but is prohibitively labor intensive. Large language models (LLMs) may be able to extract structured presenting features from unstructured trial records. But varied documentation styles, translated text, and scanned or handwritten documents pose challenges to data extraction in clinical trials. Building on previous work applying LLMs for event adjudication1, we developed and validated an LLM prompting workflow to extract 51 variables (symptoms, signs, laboratory and imaging results, and treatments) from HFH medical record dossiers into a structured dataset.

METHODS

We analyzed HFH medical record dossiers submitted for clinical endpoint adjudication in the DELIVER trial2 of dapagliflozin in HFmrEF/HFpEF conducted at 353 sites in 20 countries. Dossiers included discharge summaries, notes, and imaging/laboratory reports; non-English content was translated professionally during the trial. Portable Document Format (PDF) files were converted to text using Apple OCR Vision.3 Local ethics committees approved the study protocol. All patients provided informed consent.

We developed a prompt for OpenAI o1-mini (version 2024‑09‑12), a reasoning-based LLM,4 to extract 51 symptoms, signs, laboratory, imaging, and treatments from the dossiers into a structured format (Figure Panel A). Categorical variables (n=35) were coded as Yes/No/Not Mentioned and numerical variables (n=16) as values or “Not Mentioned”. The prompt was refined iteratively using small subsets of dossiers (n≈5) until false positives were <2 per dossier. Human reviewers were initially blinded to model output, but many discrepancies arose when the reviewer missed information the model had correctly extracted. Thus, for validation, reviewers were exposed to and directly verified the model output.

Central Illustration:

Central Illustration:

LLM-Based Data Extraction Workflow and Performance in Heart Failure Hospitalization Dossiers

Panel A. Excerpt from a mock heart failure hospitalization dossier designed to mimic authentic clinical trial reports. During data extraction, the PDF is converted to plain text via optical character recognition (OCR), parsed by the o1-mini large language model (LLM) into structured JSON key–value pairs for 51 prespecified variables, and imported into a spreadsheet where each variable is assigned “Yes,” “No,” a numeric value, or “Not Mentioned.”

Abbreviations: BNP = B-type natriuretic peptide; JVP = jugular venous pressure; NT-proBNP = N-terminal pro–B-type natriuretic peptide.

Panel B. Forest plot of pooled performance metrics—accuracy, positive predictive value (PPV), and negative predictive value (NPV)—across clinical domains. Each point represents the micro-averaged estimate with cluster-robust 95% confidence intervals accounting for dossier-level clustering.

The LLM was executed within a secure Microsoft Azure environment at Mass General Brigham on March 23, 2025. Output was returned in JavaScript Object Notation (JSON) format. Development and validation followed TRIPOD-LLM guidelines. The prompt, JSON schema, and TRIPOD-LLM checklist are available upon request.

The optimized prompt was evaluated against human review in 125 randomly selected dossiers. Validation set sample size was set to achieve a 95% confidence interval (CI) of +/−5% around anticipated accuracy of 90%, using finite-population correction (N=1,237), based on ~95% accuracy during prompt refinement. Errors were defined as disagreements involving a present finding (“Yes” vs “No/Not Mentioned”); “No” vs “Not Mentioned” discrepancies were quantified separately. Numerical variable correctness required exact value, biomarker subtype, and units; serum creatinine values were harmonized to μmol/L using a fixed conversion factor and rounded to two decimals. Accuracy, positive predictive value (PPV), and negative predictive value (NPV) were calculated for each variable with exact binomial 95% CIs; pooled (overall and domain-specific) metrics used micro-averaging across all variable–dossier observations, with cluster-robust standard errors to account for within-record dependence. To assess inter-reviewer reproducibility, a second physician independently reviewed 25 dossiers.

Following validation, the model was applied to all 1,237 HFHs in DELIVER. Regional variability in documentation rates was assessed using χ2 tests.

RESULTS:

Among 1,237 adjudicated HFHs in 759 patients, 125 randomly selected dossiers were used for validation. The model demonstrated high performance, with pooled accuracy 0.96 (95% CI, 0.95–0.97), PPV 0.94 (95% CI, 0.92–0.95), and NPV 0.97 (95% CI, 0.97–0.98) on human validation.

Pooled PPV and NPV for symptoms, physical signs, and chest-imaging variables were consistently high (PPV ≥0.96; NPV ≥0.95) (Figure Panel B). Accuracy exceeded 90% for all 18 categorical symptoms, signs, and chest-imaging features. Individual variable PPVs ranged from 0.92 (95% CI, 0.75–0.99) to 1.00 (95% CI, 0.93–1.00) for symptoms and from 0.92 (95% CI, 0.75–0.99) to 1.00 (95% CI, 0.85–1.00) for physical signs. The model correctly distinguished “No” from “Not Mentioned” in 91% of absent fields (n=1,450).

For biomarkers (BNP, NT-proBNP, and Troponin-I and -T), categorical assessment for elevated level showed PPV 0.95 (95% CI, 0.89–0.98) and NPV 0.92 (95% CI, 0.86–0.95). Numeric peak natriuretic peptide and troponin values were often extracted without clear subtype designation, yielding higher accuracy for the more common subtypes—0.92 (95% CI, 0.86–0.96) for NT-proBNP and 0.94 (95% CI, 0.88–0.97) for troponin T—while BNP (0.74 [95% CI, 0.65–0.81]) and troponin I (0.86 [95% CI, 0.79–0.92]) were lower due to subtype contamination in 76% the errors. Pooled biomarker NPV of 0.94 (95% CI, 0.92–0.96) reflected infrequent omissions. Numeric first creatinine, including unit conversions, demonstrated 0.94 (95% CI, 0.88–0.97) accuracy. Echocardiography extraction accuracy was ≥0.97 across all 14 variables, with pooled accuracy 0.98 (95% CI, 0.97–0.99), PPV 0.96 (95% CI, 0.91–0.98), and NPV 0.99 (95% CI, 0.97–0.99).

The second physician reviewer agreed with the first in 1,244/1,275 fields (98%). Disagreements reflected subjective interpretation, hard-to-locate information, or page-order issues. On consensus review, Reviewer 1 was correct in 17 disagreements, Reviewer 2 in 13, and 1 remained unresolved. The model was correct in 13, matched a human error in 14, and differed from both in 4.

Applying the model to all 1,237 HFHs in DELIVER generated a comprehensive global dataset (63,083 data points) describing presenting features of HFH. Dyspnea was documented in 92% of HFHs, peripheral edema in 68%, pulmonary crackles in 51%, and pulmonary congestion on imaging in 58%. Intravenous diuretics were used in 84% and oral diuretics increased in 45%. Documentation completeness varied geographically; “Not Mentioned” rates ranged from 25% in North America to 51% in Latin America (p<0.001).

DISCUSSION

In this proof-of-concept study, LLM extraction of presenting features from unstructured HFH dossiers in a global trial achieved accuracy 0.96, PPV 0.94, and NPV 0.97. Scaling the model to all adjudicated HFHs enabled structured characterization of clinical presentation that would otherwise be prohibitively labor-intensive. The results suggest that LLMs can transform routinely collected trial documentation into analyzable data.

This work extends prior studies of LLM-based feature extraction to globally sourced, scanned, and often translated clinical trial documents. The model inferred nuanced findings not explicitly stated (e.g., identifying orthopnea from documentation of sleeping with multiple pillows). Accuracy of greater than 95% for all variable types except biomarkers supports utilization of this tool for post hoc research that would otherwise be impractical. In the future, LLMs may reduce or replace physician review for prospective outcome assessment or trial monitoring, though these higher-stakes use cases will require prospective validation in multiple datasets.

Our study has limitations. Physician reviewers were unblinded, which may have introduced bias. However, in the initial blinded pilot phase, the model often found variables that human reviewers had overlooked, suggesting that a fully blinded validation might underestimate model performance due to human abstraction errors. While OCR conversion may introduce errors, validation directly against the original PDFs ensures that reported performance reflects the combined errors from OCR and LLM extraction. We focused on positively adjudicated HFHs to ensure evaluation against confirmed events; nonetheless, hallucination risk would be better assessed using non-HFH dossiers where findings should be absent. PPV was estimated imprecisely for rarely documented findings like hepatojugular reflex or S3 gallop.

In conclusion, an iteratively refined LLM-based model accurately extracted structured clinical details from unstructured trial documentation, enabling large-scale characterization of HF hospitalization presentations in a global randomized trial. LLM-based data extraction could be extended to unlock quantitative insights from a wide range of narrative trial records.

Clinical Perspective.

What’s New?

  • A large language model (LLM) accurately extracted 51 clinical variables from adjudication dossiers in a global heart failure trial, achieving 96% accuracy against physician review.

  • The LLM processed heterogeneous scanned dossiers from 353 sites across 20 countries, demonstrating the feasibility of large-scale structured data generation in global cardiovascular research.

  • Automated extraction enabled new insights regarding the prevalence of specific presenting features at the time of HF hospitalization.

Clinical Implications

  • Prompt-engineered LLMs can reliably transform unstructured trial records into large-scale structured datasets without the need for laborious manual abstraction.

  • LLM-based extraction could make trials more efficient and deepen understanding of acute HF phenotypes across diverse populations.

  • LLM-based extraction can reduce reliance on manual chart review, streamline case report form population, and improve data quality and monitoring in multinational trials.

Funding:

The DELIVER trial was sponsored by AstraZeneca. Dr. Cunningham is supported by the American Heart Association (23CDA1052151) and the National Heart, Lung, and Blood Institute (1K23HL168163).

Footnotes

Disclosures: Dr Claggett has received personal consulting fees from Alnylam, Bristol Myers Squibb, Cardior, Cardurion, Corvia, CVRx, Eli Lilly, Intellia, Rocket, and has served on a data safety monitoring board for Novo Nordisk. Dr. Desai has received institutional research grants (to Brigham and Women’s Hospital) from Abbott, Alnylam, AstraZeneca, Bayer, Novartis, and Pfizer as well as personal consulting fees from Abbott, Alnylam, AstraZeneca, Bayer, Biofourmis, Boston Scientific, Medtronic, Merck, Novartis, Parexel, Porter Health, Regeneron, River2Renal, Roche, Veristat, Verily, Zydus. Dr. Ho is supported by the NIH (R01 HL168889, R01 HL160003, and K24 HL153669) and has received consulting fees from Eli Lilly. has received sponsored research support from Bayer AG. Dr. Ellinor receives sponsored research support from Bayer AG, Bristol Myers Squibb, Pfizer and Novo Nordisk; he has also served on advisory boards or consulted for Bayer AG. Dr. Jhund reports speakers’ fees from AstraZeneca, Novartis, Alkem Metabolics, ProAdWise Communications, Sun Pharmaceuticals; advisory board fees from AstraZeneca, Boehringer Ingelheim, Novartis; research funding from AstraZeneca, Boehringer Ingelheim, Analog Devices Inc, Roche Diagnostics. PSJ’s employer, the University of Glasgow, has been remunerated for clinical trial work from AstraZeneca, Bayer AG, Novartis and Novo Nordisk, and is a Director of GCTP Ltd. Dr McGrath has received research support from Lexicon paid directly to her institution. Dr Mc Causland has received research fees from Novartis, Lexicon, AstraZeneca, NIH, paid directly to his institution; consulting fees from Aquapass, GSK, and Zydus Therapeutics; expert witness fees from Rubin-Anders Scientific; travel/speakers fees from Bayer and Global Learning Collaborative; his spouse reports consulting fees from Vera Therapeutics and Alexion; serving on DSMB for AstraZeneca. Dr. McMurray reports payments through Glasgow University from work on clinical trials, consulting and grants from: Amgen, AstraZeneca, Bayer, Cardurion, Cytokinetics, GSK and Novartis. Personal consultancy fees from: Alynylam Pharmaceuticals, Amgen, AnaCardio, AstraZeneca, Bayer, Berlin Cures, BMS, Cardurion, Cytokinetics, Ionis Pharmaceuticals, Novartis, Regeneron Pharmaceuticals, River 2 Renal Corp., British Heart Foundation, National Institute for Health – National Heart Lung and Blood Institute (NIH-NHLBI), Boehringer Ingelheim, SQ Innovations, Catalyze Group. Personal lecture fees: Abbott, Alkem Metabolics, Astra Zeneca, Blue Ocean Scientific Solutions Ltd., Boehringer Ingelheim, Canadian Medical and Surgical Knowledge, Emcure Pharmaceuticals Ltd., Eris Lifesciences, European Academy of CME, Hikma Pharmaceuticals, Imagica Health, Intas Pharmaceuticals, J.B. Chemicals & Pharmaceuticals Ltd., Lupin Pharmaceuticals, Medscape/Heart.Org., ProAdWise Communications, Radcliffe Cardiology, Sun Pharmaceuticals, The Corpus, Translation Research Group, Translational Medicine Academy. Data Safety Monitoring Board: WIRB-Copernicus Group Clinical Inc. He is a director of Global Clinical Trial Partners Ltd. Dr. Solomon has received research grants from Alexion, Alnylam, AstraZeneca, Bellerophon, Bayer, BMS, Boston Scientific, Cytokinetics, Edgewise, Eidos, Gossamer, GSK, Ionis, Lilly, MyoKardia, NIH/NHLBI, Novartis, NovoNordisk, Respicardia, Sanofi Pasteur, Theracos, US2.AI and has consulted for Abbott, Action, Akros, Alexion, Alnylam, Amgen, Arena, AstraZeneca, Bayer, Boeringer-Ingelheim, BMS, Cardior, Cardurion, Corvia, Cytokinetics, Daiichi-Sankyo, GSK, Lilly, Merck, Myokardia, Novartis, Roche, Theracos, Quantum Genomics, Janssen, Cardiac Dimensions, Tenaya, Sanofi-Pasteur, Dinaqor, Tremeau, CellProThera, Moderna, American Regent, Sarepta, Lexicon, Anacardio, Akros, Valo. Dr. Cunningham has consulted for Edgewise Therapeutics, Occlutech, Cytokinetics, and us2ai. All other authors have no disclosures.

Data Availability Statement:

Deidentified data underlying the findings described in this manuscript may be obtained following AstraZeneca’s data sharing policy described at https://astrazenecagrouptrials.pharmacm.com/ST/Submission/Disclosure. Participant record dossiers cannot be made available. The complete LLM prompt and JSON output structure used for data extraction are available upon request.

Access to Data Statement:

Drs. Badrouchi, Cunningham, and Solomon had full access to all the data in the study and take responsibility for the integrity of the data and the accuracy of the data analysis.

References

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Deidentified data underlying the findings described in this manuscript may be obtained following AstraZeneca’s data sharing policy described at https://astrazenecagrouptrials.pharmacm.com/ST/Submission/Disclosure. Participant record dossiers cannot be made available. The complete LLM prompt and JSON output structure used for data extraction are available upon request.

Drs. Badrouchi, Cunningham, and Solomon had full access to all the data in the study and take responsibility for the integrity of the data and the accuracy of the data analysis.

RESOURCES