Skip to main content
Wiley Open Access Collection logoLink to Wiley Open Access Collection
. 2026 Jul 28;38(4):e70285. doi: 10.1111/1742-6723.70285

The Sydney Triage to Admission Risk Tool With Artificial Intelligence (START‐AI) to Support Decision Making in Emergency Departments: Model Explainability and Feature Importance Analysis

Michael Dinh 1,2, Elizabeth Corbett 1,2,✉, Thuy Truc Ngo 3, Eliot Salmon 1,2, Saleem Ahmed Khan 4, Farhana Pethani 4,5, Nicholas Moore 6,7, Irena Koprinska 3
PMCID: PMC13416010  PMID: 42522152

ABSTRACT

Objective

Evaluate the importance of specific variables contributing to a recently reported Artificial Intelligence (AI) prediction model called Sydney Triage to Admission Risk Tool with Artificial Intelligence (START‐AI) to predict inpatient admission from the Emergency Department (ED).

Methods

A model explainability analysis was undertaken using single‐centre ED electronic medical record data over 2 years. The START‐AI model, which comprises ensemble machine learning and a transformer‐based algorithm to enhance the original START tool, was re‐run with each feature added sequentially to the model until the full model was complete. Feature importance was calculated using feature permutation or change in area under receiver operator curve (AUROC) and SHapley Additive exPlanations (SHAP) analyses.

Results

The original START tool alone had an AUROC of 0.78 (95% CI 0.77, 0.78) for prediction of inpatient admission, reaching an AUROC of 0.90 (95% CI 0.89, 0.90) after all features of the START‐AI tool were added sequentially. Features associated with a stepwise increase in cumulative AUROC were triage comments, ED case history notes, any blood test result (in particular, lactate and C‐reactive protein), and any CT order. Vital signs did not appear to be associated with stepwise increases in AUROC. Feature importance was highest for the presence of any blood test result, START score, and C‐reactive protein with respect to overall model importance.

Conclusion

A model explainability analysis provided a clearer understanding of the sequential and relative importance of specific features within the START‐AI model, informing how the tool can be further developed and deployed in clinical settings.

Keywords: artificial intelligence, disposition, emergency, explainability

1. Introduction

Artificial Intelligence (AI) has the potential to transform Emergency Departments (ED) over the coming decades through task automation, diagnosis and decision support [1, 2]. In 2016, we devised the Sydney triage to admission risk tool (START) which was a decision support tool designed to estimate the probability of a patient presenting to the ED requiring inpatient admission [3]. The START tool used structured demographic and clinical data only available at triage. Building on the original START tool, the START‐AI model incorporated additional structured data including vital signs, blood test results and CT imaging orders, alongside free text ED clinical notes [4]. A stacked ensemble architecture was implemented, combining a gradient boosting algorithm (XGBoost) with a pre‐trained language model (Clinical BERT). This novel AI tool was designed to support disposition decision making and flag potential clinical deterioration in ED using more complex health data. The START‐AI model had an overall AUROC of 0.88 (95% CI 0.88, 0.89) for prediction of inpatient admission and 0.93 (95% CI 0.93, 0.94) for intensive care unit admission or death.

The present analysis was undertaken to explore how features within START‐AI contributed to overall model performance. The main reasons for this were to: (1) provide a clearer understanding of how the START‐AI model derived predictions, (2) understand which of the twenty‐four features employed contributed most to those predictions, and (3) acknowledge the sequential nature of clinical data acquisition in ED. Clinical information isn't typically available all at once but becomes available in a step‐wise manner throughout a patient's journey in ED. For a complex ensemble AI model like START‐AI, model explainability represents an important initial step in establishing face validity, analogous to reporting a table of adjusted coefficients or parameters for a multivariable regression model. The analysis may also provide insights for how the model could be rationalised with fewer features, which may be particularly helpful when deployed in the clinical environment.

2. Methods

2.1. Design

This was a machine learning analysis using electronic medical record data from an inner‐city tertiary referral hospital in Sydney, Australia, which sees around 85,000 ED patients per annum.

2.2. Variables and Data Sources

Details of data preprocessing, imputation, machine learning algorithms and transformer‐based encoding of clinical free text were described in the START‐AI derivation study [4]. In brief, adults (aged 16 or above) who presented to a single inner city Australian ED from 1 January 2023 to 30 June 2025 were identified from the electronic health record. Demographic data (age, gender), triage characteristics (triage category, presenting problem), clinical notes (ED nursing and medical notes), laboratory results (blood test results) and CT request data were extracted. An ensemble modelling approach was used commencing with START scores derived using logistic regression as originally described [3]. This was then combined with categorised vital signs, blood test results, CT requests and transformer‐encoded clinical free text derived from medical and nursing notes, combined in a stacked extreme gradient boosted decision tree (XGBoost) classification model. The model was trained and tested (validated) using separate randomly allocated datasets in a ratio of 1:1. The outcome variable used for classification or prediction was inpatient admission from ED, excluding patients with an inpatient length of stay of less than 24 h based on our previous sensitivity analysis.

2.3. Outcomes

The outcomes of interest in the present study were:

  1. Sequential model performance for each iteration of the START‐AI model using area under receiver operator curve (AUROC). The sequential nature of the information potentially represents the real‐life ED floor as patient data are progressively obtained over time.

  2. Feature importance as measured by information gain and permutation importance.

  3. Independent feature importance as measured by SHapley Additive exPlanation (SHAP) scores.

2.4. Machine Learning and Statistical Analysis

In the present study, in addition to the above steps, the machine learning approach was applied sequentially on each iteration of the model using the testing (validation) dataset commencing with the START score alone, then adding a variable at each successive iteration as they would normally become available for a typical patient in ED (START score followed by vital signs, clinical notes, blood test results and CT request) until the final model iteration where all variables were included. AUROC values were calculated at each iteration and plotted on a line chart to demonstrate the cumulative change in model performance with each additional feature. A stepwise increase was defined as any increase in AUROC of 0.02 or more associated with a single feature.

For feature importance, features were disaggregated into component categories by creating dummy variables, and trained and tested separately with a single XGBoost model. Permutation importance measured the average reduction in AUROC that occurred if values of a particular feature in a dataset were scrambled randomly, with all other features held constant. These were presented on a simple bar chart to represent the overall feature importance across the whole dataset. Feature importance was supplemented with SHAP scores, based on cooperative gaming theory, providing a model‐agnostic way of representing additive effects of each feature [5]. Calculated SHAP scores for a single hypothetical case example with certain feature characteristics were presented in a waterfall plot, with directionality and magnitude of effect on inpatient probability moving from baseline to final outcome probability, as different feature characteristics were sequentially added to the model. The plot thus provided a schematic representation of what a computation for a single hypothetical case may look like. Confidence intervals were calculated using 1000‐fold bootstrap resampling.

2.5. Computing Resources

Programming was performed using Python Code Version 13.3.6 on Visual Studio Code installed on a standard Windows 10 operated PC (Dell Precision Tower 3620) with NVIDIA 4060 Graphics Processing Unit and CUDA 12.9 installed. All analyses were conducted and data stored on secure network drives within Sydney Local Health District.

2.6. Ethics

Approval was obtained from the Sydney Local Health District Human Research Ethics Committee [Protocol X25‐0116 & 2025/STE01700]. All data were stored and analysed locally. A waiver of consent was granted.

3. Results

A total of 162,915 cases between 1 January 2023 to 30 June 2025 were analysed with overall inpatient admission rate of 22.42% (excluding admissions with inpatient length of stay < 24 h). With the START tool alone, an AUROC of 0.78 (95% CI 0.77, 0.78) for prediction of inpatient admission was demonstrated, reaching an AUROC of 0.90 (95% CI 0.89, 0.90) after all features were added sequentially. Figure 1 shows the change in cumulative AUROC when features were sequentially added and left in the model at each iteration. Features associated with a stepwise increase in cumulative AUROC when added sequentially were triage comments, ED case history notes, any blood test result, in particular lactate and C reactive protein, and any CT order. Vital signs did not appear to be associated with stepwise increases in AUROC. Figure 2 plots feature importance with respect to permutation and information gain, demonstrating the importance of any blood test result, START score, C Reactive Protein and ED medical notes to overall model importance. For instance, the presence or absence of any blood test result changed the AUROC by an average of 0.062 (95% CI 0.061, 0.062). In relation to a high C reactive protein result, a result in the range of 5.0–70 mg/L changed the AUROC by 0.022 (95% CI 0.022, 0.022). Interestingly, the free text variables had a lesser effect on the AUROC, with the ED medical notes contributing 0.015 (95% CI 0.015, 0.015) and triage comments contributing 0.014 (95% CI 0.014, 0.014) to the AUROC.

FIGURE 1.

FIGURE 1

Cumulative area under receiver operator curve (AUROC) for Sydney Triage to admission risk tool with artificial intelligence (START‐AI) model prediction of inpatient admission with sequential addition of each variable at each iteration. GCS = Glasgow coma score.

FIGURE 2.

FIGURE 2

Feature importance for Sydney Triage to admission risk tool with Artificial intelligence (START‐AI) with respect to permutation change in area under receiver operator curve (AUROC).

The SHAP waterfall plot (Figure 3) indicated the sequential effect size and direction of various features for a single hypothetical case example where START score of 8 and other feature characteristics were as shown on the y‐axis. The baseline probability of inpatient admission indicated at the bottom of the plot was 0.21, and the effects of each successive feature added (moving up the y‐axis) on probability of inpatient admission was shown with the final outputted probability of inpatient admission for this case example being 0.095. The SHAP analysis corroborated the relative feature importance of START, blood tests, any CT order, and ED clinical notes in the previous feature importance plot, with additional context of baseline and final probabilities depicted.

FIGURE 3.

FIGURE 3

Shapley additive explanation (SHAP) waterfall plot with change in probability of admission from baseline probability (E(fx)) to final probability (fx) for a hypothetical case example with START score of 8 and varying levels of laboratory results within Sydney Triage to admission risk tool with artificial intelligence (START‐AI) model.

4. Discussion

The original START model remains one of few reported machine learning tools for ED, developed using generalised datasets and prospectively validated in multicentre clinical trials [6]. The advantage of this simple additive risk score was that feature importance in START was readily explainable based on normalised regression coefficients available from the logistic regression model. In contrast, START‐AI used a more complex architecture, combining a gradient boosting model and pre‐trained language model to derive predictions from both structured and free‐text clinical data. These models have an architecture that does not permit directly determining which features most influenced a given prediction and are therefore often characterised as ‘black box’ systems [7, 8]. Explainable AI methods have been developed in the past two to 3 years to address this limitation, enabling researchers and stakeholders to approximate feature importance post hoc, where model structure alone does not make this transparent.

In the present study, we assessed feature importance both independently, assuming all data elements were available simultaneously, and also sequentially, reflecting step‐wise availability of data throughout the patient journey. The analysis revealed a few important and actionable insights. Firstly, seven of the top ten most important features in START‐AI were related to laboratory blood test results. It is known from previous studies that blood tests are important in contributing to both diagnosis, risk stratification, and prognostication for many types of ED presentations including chest pain, abdominal pain, and infection [9]. Using START‐AI to improve disposition decision times and throughput efficiency in ED would necessarily require a focus on earlier decision‐making, such as a model of care that supported early venepuncture in ED, where appropriate, and reduced laboratory turnaround times [10, 11]. Interestingly, feature importance reported in the START‐AI paper identified white cell count as the most important feature compared to the present study, where the presence of any blood test assumed this role. This was likely due to the slight difference in outcome measure, which in this study classified inpatient stays of less than 24 h as a potential discharge (consistent with current ED short stay unit admission practices).

Secondly, even though features related to laboratory testing were important in the additive sense, when analysed sequentially and iteratively across a patient journey, START scores calculated at the point of triage contributed the most to model performance. This allowed contributions of other important features to be compared additively in relation to START scores. For instance, START had an AUROC of 0.78, which was slightly lower than original derivation studies. The presence of blood tests only contributed an additional 0.062 to overall model AUROC over and above START scores, underscoring the primacy of START with respect to disposition and triage.

Thirdly, there were several variables such as vital signs and specific CT orders that contributed very little to overall model performance. This contrasted with other reported models that appeared to rely on vital signs as part of model prediction [12, 13, 14]. One explanation may be that triage categories within START already incorporated clinical assessment and initial vital sign interpretation by triage nurses. Rationalising these features may be considered to improve model efficiency.

There were several acknowledged limitations to this analysis. The present analysis was based on single centre data and likely reflected institutional specific practices and models of care. A further analysis is being planned using multicentre data to validate the START‐AI model. Misclassification by the model could be reflective of unmeasured confounders such as psychosocial factors, family or carer concerns, clinical variation between senior clinicians and other factors that are difficult to capture semantically in electronic medical records. START‐AI is not designed to replace admission or discharge decision‐making by senior ED clinicians. Rather, it is what's termed a ‘human‐in‐the‐loop’ model, designed to rapidly synthesise complex health data to support ED clinicians, who are frequently called upon to make multiple decisions on multiple patients, often with significant time pressures and limited information. Such a decision support tool would not only guide clinicians and reduce risks of unwarranted clinical variation but may be used to guide discussions with patients about ongoing care needs.

Apart from the need for senior clinician confirmation, the present study highlighted a few other considerations for operational deployment. With the large number of features, the model would ideally appear as a flag within electronic patient tracking systems with data feeding into the START‐AI model directly from existing data warehouses. Model recommendations would need to appear sequentially, with probabilities changing in real‐time as data like blood test results and other parameters become available or values change over time. These flags would be actioned by senior ED clinicians and hospital bed managers as appropriate to make decisions about further laboratory testing, streaming, admission to medical assessment units or transfers out of hospital. Finally, further work is being planned to implement and evaluate START‐AI deployed in the clinical environment to determine its impact on ED performance and agreement with clinician‐based predictions.

5. Conclusion

In this model explainability analysis, START scores, blood testing particularly C Reactive Protein, and ED case history notes were the most important features contributing to START‐AI model predictions of inpatient admissions. These insights will inform further model refinement and strategies for operational deployment.

Author Contributions

All authors contributed to data analysis and project design. The manuscript was principally drafted by MD.

Funding

The authors have nothing to report.

Conflicts of Interest

The authors declare no conflicts of interest.

Acknowledgements

The authors have nothing to report.

Data Availability Statement

The data that support the findings of this study are available from NSW Health. Restrictions apply to the availability of these data, which were used under license for this study. Data are available from the author(s) with the permission of NSW Health.

References

  • 1. Petrella R. J., “The AI Future of Emergency Medicine,” Annals of Emergency Medicine 84, no. 2 (2024): 139–153, 10.1016/j.annemergmed.2024.01.031. [DOI] [PubMed] [Google Scholar]
  • 2. Chenais G., Lagarde E., and Gil‐Jardiné C., “Artificial Intelligence in Emergency Medicine: Viewpoint of Current Applications and Foreseeable Opportunities and Challenges,” Journal of Medical Internet Research 25 (2023): e40031, 10.2196/40031. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Dinh M. M., Russell S. B., Bein K. J., et al., “The Sydney Triage to Admission Risk Tool (START) to Predict Emergency Department Disposition: A Derivation and Internal Validation Study Using Retrospective State‐Wide Data From New South Wales, Australia,” BMC Emergency Medicine 16, no. 1 (2016): 46, 10.1186/s12873-016-0111-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Dinh M., Corbett E., Salmon E., et al., “The Sydney Triage to Admission Risk Tool With Artificial Intelligence (START‐AI): Prediction of Inpatient Admission From Emergency Departments Using Ensemble Machine Learning,” Emergency Medicine Australasia 38, no. 2 (2026): e70240, 10.1111/1742-6723.70240. [DOI] [PubMed] [Google Scholar]
  • 5. Song Y., Zhang D., Wang Q., et al., “Prediction Models for Postoperative Delirium in Elderly Patients With Machine‐Learning Algorithms and SHapley Additive exPlanations,” Translational Psychiatry 14, no. 1 (2024): 57, 10.1038/s41398-024-02762-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6. Berendsen Russell S., Seimon R. V., Dixon E., et al., “Applying Sydney Triage to Admission Risk Tool (START) to Improve Patient Flow in Emergency Departments: A Multicentre Randomised, Implementation Study,” BMC Emergency Medicine 24, no. 1 (2024): 39, 10.1186/s12873-024-00956-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7. Ali S., Akhlaq F., Imran A. S., Kastrati Z., Daudpota S. M., and Moosa M., “The Enlightening Role of Explainable Artificial Intelligence in Medical & Healthcare Domains: A Systematic Literature Review,” Computers in Biology and Medicine 166 (2023): 107555, 10.1016/j.compbiomed.2023.107555. [DOI] [PubMed] [Google Scholar]
  • 8. Okada Y., Ning Y., and Ong M. E. H., “Explainable Artificial Intelligence in Emergency Medicine: An Overview,” Clinical and Experimental Emergency Medicine 10, no. 4 (2023): 354–362, 10.15441/ceem.23.145. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Hicks A. J., Carwardine Z. L., Hallworth M. J., and Kilpatrick E. S., “Using Clinical Guidelines to Assess the Potential Value of Laboratory Medicine in Clinical Decision‐Making,” Biochemia Medica 31, no. 1 (2021): 010703, 10.11613/BM.2021.010703. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Hsiao A. L., Santucci K. A., Dziura J., and Baker M. D., “A Randomized Trial to Assess the Efficacy of Point‐Of‐Care Testing in Decreasing Length of Stay in a Pediatric Emergency Department,” Pediatric Emergency Care 23, no. 7 (2007): 457–462, 10.1097/01.pec.0000280506.44924.de. [DOI] [PubMed] [Google Scholar]
  • 11. Dinh M. M., Green T. C., Newsome D., and Bein K. J., “Impact of Technical Assistants for Venepuncture and Intravenous Cannulation on Overall Emergency Department Performance,” Emergency Medicine Australasia 23, no. 6 (2011): 726–731, 10.1111/j.1742-6723.2011.01479.x. [DOI] [PubMed] [Google Scholar]
  • 12. Raita Y., Goto T., Faridi M. K., Brown D. F. M., C. A. Camargo, Jr. , and Hasegawa K., “Emergency Department Triage Prediction of Clinical Outcomes Using Machine Learning Models,” Critical Care 23 (2019): 64, 10.1186/s13054-019-2351-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Piliuk K. and Tomforde S., “Artificial Intelligence in Emergency Medicine. A Systematic Literature Review,” International Journal of Medical Informatics 180 (2023): 105274, 10.1016/j.ijmedinf.2023.105274. [DOI] [PubMed] [Google Scholar]
  • 14. Van Der Haas Y., Roskamp W., Chang‐Willems L. E. M., et al., “Evaluating an AI Decision Support System for the Emergency Department: Retrospective Study,” JMIR Ai 5 (2026): e80448, 10.2196/80448. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The data that support the findings of this study are available from NSW Health. Restrictions apply to the availability of these data, which were used under license for this study. Data are available from the author(s) with the permission of NSW Health.


Articles from Emergency Medicine Australasia are provided here courtesy of Wiley

RESOURCES