Skip to main content
BMJ Health & Care Informatics logoLink to BMJ Health & Care Informatics
editorial
. 2025 Dec 3;32(1):e101707. doi: 10.1136/bmjhci-2025-101707

Artificial intelligence in clinical risk prediction: promise, performance and the path forward?

Padmanesan Narasimhan 1,, Usman Iqbal 2, Yu-Chuan Li 3
PMCID: PMC12682157  PMID: 41338666

Artificial intelligence (AI) and machine learning are reshaping clinical risk prediction and patient monitoring.1 Two studies show this transformation, highlighting both promise and challenges. Yoshihara et al investigate deep learning for hypertension detection from pharyngeal images in Japanese primary care settings,2 while Watson et al assess transformer-based models for predicting patient deterioration in emergency admissions, comparing them to the widely used National Early Warning Score (NEWS).3 Together, these studies show AI’s expanding diagnostic capabilities, document improvements over traditional methods and reveal hurdles for widespread clinical adoption.

Expanding diagnostic frontiers: from blood pressure cuffs to smartphone cameras

Yoshihara et al address the underdiagnosis of hypertension, a condition affecting over one billion people worldwide and responsible for millions of deaths annually. Traditional diagnosis relies on blood pressure measurement devices, which aren't universally accessible, especially in low-resource settings.2 Their solution uses vascular changes visible in the pharynx due to hypertension, developing a deep learning algorithm that detects hypertension from pharyngeal images.

The results are compelling. Their multi-instance convolutional neural network, trained on over 7700 patients, achieved an area under the receiver operating characteristics curve (AUROC) of 0.922, substantially outperforming models using only demographic data (AUROC 0.887). This approach promises that for telemedicine, where contactless, device-free diagnostics are increasingly valuable. The potential to use smartphone cameras for hypertension screening could democratise early detection, especially in populations with limited healthcare infrastructure. The study addresses the deep learning ‘black box’ problem through Grad-CAM heatmaps that visualise the model’s focus on the posterior pharyngeal wall, where blood vessels are most prominent. Similar AI explainability techniques have previously been used in the dermatological field, where Grad-CAMs increased clinicians’ trust and confidence in AI-driven melanoma diagnosis.4 This reinforces the potential of explainable AI models to enhance clinical adoption in the cardiovascular field; however, its real impact remains underexplored and warrants further research.

Harnessing unstructured data: the hidden value in clinical notes

Watson et al examined whether advanced ML models could improve prediction of clinical deterioration in emergency department (ED) admissions compared with NEWS. Their retrospective study of over 174 000 ED admissions tested both tree-based (LightGBM) and transformer-based (BioClinicalBERT) models, notably incorporating free-text triage notes—unstructured data typically overlooked by traditional scoring systems.3

The results were striking. The best-performing model (BioClinicalBERT with extended tabular data and triage notes) achieved an average precision of 0.92 compared with 0.28 for NEWS (AUROC 0.96 vs 0.66 for NEWS). This suggests that valuable clinical intuition and contextual information captured in triage notes can be effectively harnessed by modern natural language processing techniques to improve risk stratification. Moreover, the transformer-based models showed reduced bias across demographic groups compared with NEWS—an important consideration for equitable care delivery.

Beyond statistical improvements: real-world clinical impact

These findings align with a growing body of literature demonstrating that machine learning (ML) models, particularly those leveraging rich, multimodal data sources, consistently outperform traditional early warning scores across diverse settings and populations. Gradient-boosted machine learning models and deep learning approaches using clinical notes or pager messages have repeatedly shown higher AUROC and sensitivity than NEWS, Modified Early Warning System or proprietary scores like the Epic Deterioration Index. Critically, these improvements extend beyond statistical metrics to meaningful clinical outcomes. Earlier and more accurate identification of high-risk patients enables timely intervention and improved outcomes. Yoshihara et al project a 10% increase in hypertension detection sensitivity, potentially preventing 60 million missed cases globally.2 Similarly, reducing false alarms in acute care settings can alleviate alert fatigue while enhancing both patient safety and clinician efficiency.

Persistent challenges: generalisability and implementation

Despite these advances, significant challenges remain particularly around generalisability, due to variations in data quality, patient populations and healthcare systems. Yoshihara et al acknowledge the ethnic and clinical homogeneity of their Japanese cohort and their use of specialised cameras, raising questions about broader applicability. However, the study by Watson et al emphasises the need for validation across different electronic health record systems and cautions against over-reliance on single-institution training.3

The complexity of advanced AI models also presents implementation hurdles. Although deep learning and transformer models offer superior performance, their opacity can undermine clinician trust and hinder adoption. Both studies attempt to address this through visualisation techniques, Grad-CAM heatmaps and SHAP values for feature attribution, but transparency and generalisability across various clinical settings remain an ongoing challenge that requires continued attention. One prospect to increase trust and reproducibility is the inclusion of transparent methodology through protocol papers. Another way is adherence to standardised reporting guidelines, such as Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis ((TRIPOD (used by Yoshihara et al)) or TRIPOD-AI (used by Watson et al), which is an updated version of the guideline more suitable for reporting an AI-based model.5

Charting the path forward: integration over replacement

The future of clinical risk prediction lies not in wholesale replacement of traditional tools, but in thoughtful integration that harnesses the full spectrum of available data—structured and unstructured, objective and intuitive. Success will depend on rigorous validation protocols, systematic attention to bias and fairness,6 and unwavering commitment to augmenting rather than replacing clinical expertise. Key priorities for the field include ensuring interoperability across healthcare systems, safeguarding data privacy and maintaining appropriate clinician oversight. As healthcare systems increasingly adopt AI-driven tools, establishing robust governance frameworks and validation standards will be essential for realising AI’s promise while maintaining patient safety. The studies by Yoshihara et al and Watson et al2 3 represent important steps forward in this journey, demonstrating that AI can meaningfully improve clinical risk prediction when thoughtfully developed and rigorously evaluated. These studies provide a roadmap for future innovations while highlighting the careful balance required between technological advancement and clinical pragmatism.

Footnotes

Funding: The authors have not declared a specific grant for this research from any funding agency in the public, commercial or not-for-profit sectors.

Patient consent for publication: Not applicable.

Ethics approval: Not applicable.

Provenance and peer review: Commissioned; internally peer reviewed.

References

  • 1.Giddings R, Joseph A, Callender T, et al. Factors influencing clinician and patient interaction with machine learning-based risk prediction models: a systematic review. Lancet Digit Health. 2024;6:e131–44. doi: 10.1016/S2589-7500(23)00241-8. [DOI] [PubMed] [Google Scholar]
  • 2.Yoshihara H, Tsugawa Y, Fukuda M, et al. Detection of hypertension from pharyngeal images using deep learning algorithm in primary care settings in Japan. BMJ Health Care Inform. 2024;31:e100824. doi: 10.1136/bmjhci-2023-100824. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Watson M, Boulitsakis Logothetis S, Green D, et al. Performance of machine learning versus the national early warning score for predicting patient deterioration risk: a single-site study of emergency admissions. BMJ Health Care Inform . 2024;31:e101088. doi: 10.1136/bmjhci-2024-101088. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Chanda T, Hauser K, Hobelsberger S, et al. Dermatologist-like explainable AI enhances trust and confidence in diagnosing melanoma. Nat Commun. 2024;15:524. doi: 10.1038/s41467-023-43095-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi: 10.1136/bmj-2023-078378. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Chen RJ, Wang JJ, Williamson DFK, et al. Algorithmic fairness in artificial intelligence for medicine and healthcare. Nat Biomed Eng. 2023;7:719–42. doi: 10.1038/s41551-023-01056-8. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from BMJ Health & Care Informatics are provided here courtesy of BMJ Publishing Group

RESOURCES