Skip to main content
HHS Author Manuscripts logoLink to HHS Author Manuscripts
. Author manuscript; available in PMC: 2026 Mar 11.
Published in final edited form as: AJR Am J Roentgenol. 2025 Jan 29;224(4):e2432341. doi: 10.2214/AJR.24.32341

Use of ChatGPT Large Language Models to Extract Details of Recommendations for Additional Imaging From Free-Text Impressions of Radiology Reports

Kathryn W Li a, Ronilda Lacson a, Jeffrey P Guenette a, Pamela J DiPiro a, Kristine S Burk a, Neena Kapoor a, Fatima Salah a, Ramin Khorasani a
PMCID: PMC12975287  NIHMSID: NIHMS2134994  PMID: 39878409

Abstract

Background:

Automated extraction of actionable details from radiologist recommendations for additional imaging (RAI) would facilitate tracking and timely completion of clinically necessary RAI in an effort to reduce diagnostic error related to delays in diagnosis.

Purpose:

To evaluate if large language models (LLMs) can accurately extract actionable details about RAI in radiology reports, and compare the accuracy of two LLMs, ChatGPT 3.5 and ChatGPT 4.

Materials and Methods:

This institutional review board-approved, retrospective study included diagnostic radiology reports across multiple care settings, modalities, and subspecialties (neuro, abdominal, thoracic, molecular, and musculoskeletal imaging), generated in August 2023 at a large academic quaternary care center, that were identified to contain a RAI by a previously-validated algorithm. Of 231 randomly-selected reports manually confirmed to contain a RAI, 25 reports were used to engineer a prompt that instructed each LLM to extract details about the RAI modality, body part to be imaged, timeframe, and rationale. The remaining 206 reports were used for testing. One 4th year medical student reviewer and one radiologist reviewer from the relevant subspecialty evaluated how well the LLM outputs from testing matched the information in the report. Performance metrics of LLM accuracy by RAI actionable detail (modality, body part, timeframe, and rationale) were compared using Chi-square and McNemar tests.

Results:

Reviewers reached 100% consensus on evaluation of the LLM outputs after discussion. Overall, ChatGPT 4 was more accurate than ChatGPT 3.5 in extracting RAI modality (94.2% [194/206 responses] vs. 85.4% [176/206] respectively, p<0.001), body part (86.9% [179/206] vs. 77.2% [159/206] respectively, p=0.004), and timeframe (99.0% [204/206] vs. 95.6% [197/206] respectively, p=0.02). ChatGPT 4 and ChatGPT 3.5 both achieved 91.7% (189/206 responses) accuracy at identifying the RAI rationale.

Conclusion:

Tested LLMs had excellent accuracy in extracting RAI actionable details from free-text radiology reports, with ChatGPT 4 outperforming ChatGPT 3.5.

Summary Statement:

Large language models can accurately extract actionable details about recommendations for additional imaging in free-text radiology reports, representing an innovative method to facilitate completion of clinically necessary recommendations.

Introduction

Recommendations for additional imaging (RAI) are included in more than 10% of radiology reports (13) with substantial RAI rate variation by modality, radiology subspecialty, and practice site(36). Although delays or failure to perform clinically necessary recommendations can lead to patient harm(7), adherence rates remain low ranging between 29–77%(810) due to a variety of factors (9,11). Actionable RAIs, defined using a validated RAI taxonomy(12,13) include RAI modality, clinical rationale, and timeframe, are more likely to be performed in a timely fashion (11). While most RAIs in radiology reports are not actionable (5,13,14), substantial improvements can be achieved with IT-enabled initiatives(5,12,13,15). However, such IT tools are not widely available.

Automated tracking and safety net programs (4,9,10,16) have also been reported to improve RAI adherence. To be effective, such programs will require large scale extraction of RAIs and their actionable details (i.e, modality, body part, timeframe and rationale) from radiology reports(5,9,16). Newly developed large language models (LLMs) allow users to implement state-of-the-art, multi-billion parameter natural language processing (NLP) models with little fine tuning.(17) These models can assist in a wider variety of tasks compared to previous NLP models. Some LLMs have been tested in a variety of medical settings, including extraction of oncologic progression information from free-text radiology reports.(18,19)

In this study we evaluated whether we could prompt LLMs to accurately extract actionable details about RAI in radiology reports with the ultimate goal of improving timely completion of clinical necessary RAIs. We secondarily compared the accuracy of two recently developed LLM-based chatbots, ChatGPT 3.5 and ChatGPT 4, similarly prompted for this task.

Materials and Methods

Study Setting and Population

This retrospective, Health Insurance Portability and Accountability Act-compliant study was conducted at an academic quaternary care center. It was Institutional Review Board-approved with waiver of informed consent.

To evaluate if the LLMs could accurately extract actionable details about RAI from free-text radiology reports, reports across five imaging subspecialties (abdominal, molecular, musculoskeletal [MSK], neuro, and thoracic), and multiple care settings (inpatient, outpatient, emergency), generated at the study site in August 2023 were identified. A previously-validated NLP algorithm(20) was used to identify the subset of reports containing a RAI. As that algorithm has not been validated on breast and obstetric reports, those reports were excluded. A randomly-selected subset of 250 reports was then manually reviewed (by a 4th year medical student reviewer and a radiology fellow) to ensure they contained a RAI, based on previously-established criteria.(20) The 231 reports confirmed to contain at least one RAI constituted the final study population and were divided into 2 sets: 25 for prompt engineering and 206 reports for testing.

Study Design

Per study design (Figure 1), all study reports were manually deidentified and only the Impression portion of the reports was used; patient data were not collected. The reports reserved for prompt engineering were used to generate instructions for the LLMs to extract the imaging modality, body part, timeframe, and rationale for each RAI (hereinafter “RAI actionable details”) (Figure 2). Each Impression section from the testing reports was pasted into the prompt and run once through ChatGPT 3.5 and ChatGPT 4 chatbots. The chat was cleared prior to each run and no follow-up questions were asked.

Figure 1: Study design flowchart.

Figure 1:

All reports in the test set (n=206) were run through both large language models (ChatGPT 3.5 and ChatGPT 4). Reports were separated into five subspecialties for secondary review: abdominal imaging (n=61), molecular imaging (n=23), musculoskeletal imaging (n=53), neuroimaging (n=22), and thoracic imaging (n=47).

Figure 2: Prompt to Extract Recommendation for Additional Imaging Details.

Figure 2:

Report text was inserted where the braces are in the prompt, in between the quotation marks. This prompt was used for both large language models (ChatGPT3.5 and ChatGPT4).

Sample Annotation

The medical student reviewer assessed each report in the test sample to determine the presence of each RAI actionable detail, and the relevant radiology subspecialty. Reports with multiple RAI options for a single finding were considered to have a single RAI with multiplicity, as previously defined,(12) while reports with multiple RAI for multiple unique findings were considered to have more than 1 RAI. To ensure accuracy, a radiologist reviewed the RAI actionable details and subspecialty designation for 40 randomly-selected annotated reports.

Manual Evaluation of LLM Outputs

The medical student reviewer and one expert reviewer in the subspecialty of the original report (an abdominal radiologist, a thoracic and oncologic radiologist, an MSK radiologist and a neuroradiologist) evaluated each detail on the LLM outputs vs. the original report. For modality, body part, and timeframe, reviewers determined if the LLM response was correct or incorrect. For rationale, reviewers rated how well the answer matched the report on a Likert scale (Table S1). Likert scores were binarized; 1 or 2 indicating an incorrect response and 3–5 indicating a correct response. Examples of the evaluation guidelines are in Table S2. Reviewers were also instructed to flag responses with confabulation.(21) When the evaluations of the two reviewers differed, or a reviewer indicated uncertainty about a response, consensus was reached through discussion. If consensus was not reached, a third radiologist reviewer made the final evaluation.

LLM Performance Analysis

The LLM results were categorized into true positive, true negative, false positive, and false negative results. True positives were defined as responses where the LLM correctly identified all instances of a specific detail. True negatives were defined as responses where the LLM correctly identified that a specific detail was not included in the original report. False positives were defined as responses where: the LLM identified a detail when none was given in the original report, or the LLM returned an incorrect detail vs. the one stated in the original report (e.g., the report indicated “follow-up with CT” and the LLM indicated the modality is “MRI”). False negatives were defined as responses where: the LLM indicated that there was no detail included in the report despite one being present, or the LLM was unable to identify all instances of a specific detail. Sensitivity and specificity were calculated for both LLMs.

Statistical Analysis

A minimum test sample size of 204 was calculated based on prevalence of reports containing the modality (88%), body part (80%), timeframe (16%), and rationale (96%) for the RAI in the prompt generation subset; confidence of 95%; and precision of 0.055. ChatGPT 3.5 and ChatGPT 4 accuracy was compared using pairwise McNemar tests. Comparison of RAI actionable detail prevalence and assessment of LLM performance variability across different details and subspecialties were done initially using Chi-squared tests of independence followed by post-hoc pairwise Chi-squared comparisons. These compared the RAI actionable detail or subspecialty with the highest prevalence or accuracy to each remaining RAI actionable detail or subspecialty. Prevalence of each RAI actionable detail was reported on a per-RAI basis because reports could contain multiple RAI. Other study metrics were calculated on a per-report basis, as individual reports were run through the LLMs. Statistical significance was set to p<0.05. All statistical analysis was done in Python 3.8.18, using the sci-kit learn module(22) and the statsmodels package.(23)

Results

Study Population

The NLP algorithm identified 3,687 of 59,622 (6.2%) reports generated in August 2023 as containing at least one RAI. Nineteen of the 250 randomly-selected reports classified by the algorithm to contain RAI were excluded on manual review after they were found not to contain a RAI. Among the resulting 206 testing reports, there were 61 abdominal imaging reports, 23 molecular imaging reports, 53 MSK imaging reports, 22 neuroimaging reports, and 47 thoracic imaging reports.

Prevalence of RAI Actionable Details

Ten of the 206 test set radiology reports contained multiple RAI for a total of 216 RAI, 63 in abdominal imaging reports, 26 in molecular imaging reports, 55 in MSK imaging reports, 22 in neuroimaging reports, and 50 in thoracic imaging reports.

Across the test set reports, 172/216 (79.6%) of RAI included information about the modality, 202/216 (93.5%) RAI had information about the body part for imaging, 43/216 (19.9%) RAI contained timeframe information, and 214/216 (99.1%) RAI included the rationale for the recommendation. RAI actionable detail prevalence by subspecialty is shown in Table 1. There were significant differences between the prevalences of different RAI actionable details in the combined test set (p<0.001), and in each of the subspecialties (p<0.001 for all subspecialties). Of note, the proportion of RAI with information about the timeframe was significantly less than the proportion of RAI with information about the rationale across all datasets. Further pairwise posthoc analyses are shown in Table 1.

Table 1:

Characteristics of Recommendations for Additional Imaging (RAI)

RAI Actionable Detail Number of RAI with reported actionable detail P-value
All RAI (n=216)*

100 (%)
Modality 172 (79.6) p<0.001**
Body part 202 (93.5) p=0.005 **
Timeframe 43 (19.9) p<0.001**
Rationale 214 (99.1) Reference
Abdominal Imaging RAI (n=63)

29.2 (%)
Modality 57 (90.5) p=0.12
Body part 63 (100) p>0.99
Timeframe 11 (17.4) p<0.001**
Rationale 62 (98.4) Reference
Molecular Imaging RAI (n=26)

12 (%)
Modality 9 (34.6) p<0.001**
Body part 26 (100) N/A
Timeframe 0 (0) p<0.001**
Rationale 26 (100) Reference
Musculoskeletal Imaging RAI (n=55)*

25.5 (%)
Modality 49 (89.1) p=0.04**
Body part 47 (85.5) p=0.005**
Timeframe 10 (18.2) p<0.001**
Rationale 55 (100) Reference
Neuroimaging RAI (n=22)

10.2 (%)
Modality 18 (81.8) p=0.12
Body part 20 (90.9) p=0.47
Timeframe 3 (13.6) p<0.001**
Rationale 22 (100) Reference
Thoracic Imaging RAI (n=50)*

23.1 (%)
Modality 39 (78) p=0.006**
Body part 47 (94) p=0.61
Timeframe 19 (38) p<0.001**
Rationale 49 (98) Reference

P-values calculated using pairwise chi-squared tests between the prevalence of given detail and that of the detail labelled “Reference”

*

A radiology report could contain more than one RAI

**

Denotes a statistically significant difference

Inter-rater agreement

A total of 73 of 412 LLM outputs required consensus discussion, and a third reviewer was required to reach consensus in 6 cases.

LLM Performance

Table 2 summarizes the performance of the two LLMs across the entire test set. Using each radiology report as the unit of analysis, ChatGPT 4 was more accurate than ChatGPT 3.5 at extracting the modality (94.2% [194/206 responses] vs. 85.4% [176/206], p<0.001), the body part (86.9% [179/206] vs. 77.2% [159/206], p=0.004), and the timeframe (99.0% [204/206] vs. 95.6% [197/206], p=0.02). ChatGPT 4 and ChatGPT 3.5 performed similarly in extracting the RAI rationale: both LLMs achieved 91.7% (189/206 responses) accuracy, however, ChatGPT4 had a sensitivity of 97.9% compared to a sensitivity of 95.5% by ChatGPT 3.5. Neither LLM produced any responses with confabulation. Both versions of ChatGPT had statistically significant differences in accuracy at extracting different RAI actionable details (p<0.001 for ChatGPT 3.5, p=0.002 for ChatGPT 4). Pairwise post hoc analyses are shown in Table S3.

Table 2:

Comparison of Performance by Large Language Model (ChatGPT Version 3.5 vs. Version 4)

ChatGPT 3.5 ChatGPT 4
RAI Actionable Detail Sensitivity Specificity Accuracy Sensitivity Specificity Accuracy GPT 3.5 vs. GPT 4*
Modality 87.3 78.0 85.4 95.2 90.0 94.2 p<0.001**
Body Part 77.4 75.0 77.2 88.2 73.7 86.9 p=0.004**
Time Frame 86.0 98.2 95.6 95.3 100.0 99.0 p=0.016**
Rationale 95.5 0 91.7 97.9 0 91.7 p>0.99

Notes:

RAI = recommendation for additional imaging

Data are presented as percentages unless otherwise stated.

*

McNemar test for association between GPT 3.5 and 4 responses

**

Denotes a statistically significant difference

Performance of both LLMs in individual subspecialties is shown in Figure 3. ChatGPT 4 and ChatGPT 3.5 performed similarly across the subspecialties at extracting each RAI actionable detail, except in thoracic imaging where ChatGPT 4 was more accurate than ChatGPT 3.5 at identifying the RAI modality (97.9% [46/47 responses] vs. 76.6% [36/47], p=0.002). ChatGPT 4 performance did not differ significantly across all subspecialties for each RAI actionable detail (p=0.16 for modality, p=0.17 for body part, p=0.45 for timeframe), except for the rationale (p=0.04). Pairwise post hoc analyses for accuracy at extracting the rationale between subspecialties are shown in Table S4.

Figure 3: Performance of the large language models (ChatGPT 3.5 and ChatGPT 4) by Subspecialty.

Figure 3:

Accuracy of each large language model at extracting all recommendation for additional imaging details is shown for: (A) abdominal imaging reports (n=61), (B) molecular imaging reports (n=23), (C) musculoskeletal (MSK) imaging reports (n=53), (D) neuroimaging reports (n=22), and (E) thoracic imaging reports (n=47).

* denotes a statistically significant difference (p<0.05).

Discussion

We found that two recently released LLMs, ChatGPT 3.5 and ChatGPT 4, have excellent accuracy at extracting RAI actionable details from free-text radiology reports. ChatGPT 4 outperformed ChatGPT 3.5 at identifying the modality (94.2% vs. 85.4%), body part (86.9% vs. 77.2%), and timeframe (95.6% vs. 99.0%), while both LLMs were 91.7% accurate at identifying the rationale.

A few automated tools have been developed to extract RAI actionable details from radiology reports. Mabotuwana et al used a combination of regular expressions and the National Center for Biomedical Ontology annotation service to extract modality, timeframe, and body part information.(24) This method assumes that reports use similar language and only extracts anatomy defined by SNOMED-CT ontology, whereas LLMs are not constrained by the same phrasing requirements. Lau et al leveraged deep learning NLP systems to extract the modality, timeframe, and rationale for the RAI.(25) However, both methods developed by Lau and Mabotuwana require and only use information from the sentence containing the RAI. In our study, we found that information about RAI actionable details can sometimes be inferred from other parts of the report (e.g., “Multiple hypoechoic avascular liver lesions, not seen on prior studies…A contrast-enhanced CT or MRI scan is suggested for further assessment.” – the body part to be imaged is the liver, which is inferred from the prior sentence). By including the entire Impression section of the report, our method can detect additional details that other algorithms would likely miss. Moreover, our method extracts all RAI actionable details (modality, body part, timeframe, and rationale) simultaneously.

Accurate, RAI actionable detail extraction tools enable semi-automated or automated tracking of RAI adherence. Structured reporting systems have been shown to increase RAI adherence, however the published methods generally rely on radiologist’s voluntary usage of a structured template, whereas our method is suitable for the free-text entry that most radiologists use, and is robust to radiologist deviations from structure and natural inter-radiologist variation.(12,14,15) RAI actionable detail extraction tools could also be used to improve the inclusion of these details in RAI. For example, personalized radiologist feedback reports, which have been shown to elicit small but significant changes in radiologist behaviors,(26) could be quickly generated with data from the LLM extraction tool. Potentially better, integrating our RAI actionable detail extraction model or similar models within radiology report generation software may allow for real-time notifications, reminders and decision support, prompting radiologists to include missing RAI actionable attributes as the report is being generated. Finally, integration of these models in the electronic health record could be used to automatically and accurately generate a clinical imaging examination order to elicit agreement/disagreement of the referring provider with the RAI and streamline the ordering of the clinically necessary RAI (e.g., as defined by agreement of the ordering provider with the radiologist’s recommendation)(5,9,15,16,27).

This study has several limitations. It was conducted at a single site, limiting generalizability, although reports from multiple subspecialties were included. Prompts were only run through the LLM once, without the opportunity to further prompt the LLM if the initial answer was incorrect. We felt this approach may best simulate future applications of this technology. While we used a chatbot for this initial proof-of-concept study, future work could utilize the full application programing interface, which offers more fine-tuning properties and upscaling potential.(28) The accuracy of each RAI was evaluated independently, which may impact overall performance in reports with more than one RAI (10/206 reports in our study [4.9%]). We limited the report text to the Impression section of radiology reports as typically RAI are only found in this section,(29) and to optimize computing costs and model accuracy. However, in discussion with radiologist reviewers, we discovered that RAI actionable details could occasionally be inferred from other parts of the report. For example, a radiologist could recommend “continued short term imaging follow-up” for a brain imaging finding where the implied modality and body part for the RAI would be the same as the initial report. This information would be included in the whole report text, but not the Impression section. Finally, although LLM outputs were reviewed by radiologists, evaluation by referring providers (e.g., primary care physicians, surgeons) should also be considered, as these providers will ultimately need to order the appropriate additional test based on the report recommendation.

In conclusion, in this study, LLMs accurately identified the modality, body part, timeframe, and rationale of RAI made in free-text radiology reports. Further studies would be needed to assess whether integrating this technology with initiatives for real-time monitoring of RAI actionability, and for generation of clinically necessary RAI orders, will improve patient safety and quality of care by reducing diagnostic errors and patient harm.

Supplementary Material

Supplementary Material

Table S1: Likert Scale assessment of Recommendation for Additional Imaging Rationale extracted by Large Language Models (LLMs)

Table S2: Evaluation Guidelines for Large Language Model (LLM) Outputs

Table S3: Posthoc analysis of differences in accuracy between recommendation for additional imaging (RAI) details

Table S4: Posthoc analysis of differences in accuracy between subspecialties

Key Results:

  1. In this retrospective study, two large language models, ChatGPT 3.5 and ChatGPT 4, were prompted to extract actionable details about recommendations for additional imaging in 206 free-text radiology reports.

  2. ChatGPT 4 outperformed ChatGPT 3.5 in extracting modality (94.2% vs. 85.4%; p<0.001), body part (86.9% vs. 77.2%; p=0.004), and timeframe (99.0% vs. 95.6%; p=0.016) for the recommended imaging. Models performed similarly (91.7%) at extracting the rationale for the recommendation (p>0.99).

Funding:

Ronilda Lacson MD, PhD, Jeffrey P. Guenette MD, MPH, Pamela J. DiPiro MD, Kristine S. Burk MD, Neena Kapoor MD, Fatima Salah MD, and Ramin Khorasani MD, MPH were all partly supported by the Agency for Healthcare Research and Quality R18 HS029348.

Abbreviations:

LLMs

large language models

NLP

natural language processing

RAI

recommendations for additional imaging

Footnotes

Disclosures: None

References

  • 1.Carrodeguas E, Lacson R, Swanson W, Khorasani R. Use of Machine Learning to Identify Follow-Up Recommendations in Radiology Reports. Journal of the American College of Radiology. 2018;0(0). doi: 10.1016/j.jacr.2018.10.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Sistrom CL, Dreyer KJ, Dang PP, et al. Recommendations for additional imaging in radiology reports: multifactorial analysis of 5.9 million examinations. Radiology. 2009;253(2):453–461. doi: 10.1148/radiol.2532090200. [DOI] [PubMed] [Google Scholar]
  • 3.Cochon LR, Kapoor N, Carrodeguas E, et al. Variation in Follow-up Imaging Recommendations in Radiology Reports: Patient, Modality, and Radiologist Predictors. Radiology. 2019;291(3):700–707. doi: 10.1148/radiol.2019182826. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Mabotuwana T, Hall CS, Hombal V, et al. Automated Tracking of Follow-Up Imaging Recommendations. American Journal of Roentgenology. 2019;212(6):1287–1294. doi: 10.2214/AJR.18.20586. [DOI] [PubMed] [Google Scholar]
  • 5.DeSimone AK, Kapoor N, Lacson R, et al. Impact of an Automated Closed-Loop Communication and Tracking Tool on the Rate of Recommendations for Additional Imaging in Thoracic Radiology Reports. Journal of the American College of Radiology. Elsevier; 2023;20(8):781–788. doi: 10.1016/j.jacr.2023.05.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Guenette JP, Lynch E, Abbasi N, et al. Recommendations for Additional Imaging on Head and Neck Imaging Examinations: Interradiologist Variation and Associated Factors. American Journal of Roentgenology. 2024;AJR.23.30511. doi: 10.2214/AJR.23.30511. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Committee on Diagnostic Error in Health Care, Board on Health Care Services, Institute of Medicine, The National Academies of Sciences, Engineering, and Medicine. Improving Diagnosis in Health Care. Balogh EP, Miller BT, Ball JR, editors. Washington (DC): National Academies Press (US); 2015. http://www.ncbi.nlm.nih.gov/books/NBK338596/. [PubMed] [Google Scholar]
  • 8.Mabotuwana T, Hombal V, Dalal S, Hall CS, Gunn M. Determining Adherence to Follow-up Imaging Recommendations. J Am Coll Radiol. 2018;15(3 Pt A):422–428. doi: 10.1016/j.jacr.2017.11.022. [DOI] [PubMed] [Google Scholar]
  • 9.Kapoor N, Lynch EA, Lacson R, et al. Predictors of Completion of Clinically Necessary Radiologist-Recommended Follow-Up Imaging: Assessment Using an Automated Closed-Loop Communication and Tracking Tool. American Journal of Roentgenology. 2023;220(3):429–440. doi: 10.2214/AJR.22.28378. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Wandtke B, Gallagher S. Reducing Delay in Diagnosis: Multistage Recommendation Tracking. American Journal of Roentgenology. American Roentgen Ray Society; 2017;209(5):970–975. doi: 10.2214/AJR.17.18332. [DOI] [PubMed] [Google Scholar]
  • 11.Mabotuwana T, Hall CS, Hombal V, Dalal S, Gunn ML. Impact of Follow-Up Imaging Recommendation Specificity on Adherence. Advances in Informatics, Management and Technology in Healthcare. IOS Press; 2022. p. 87–90. doi: 10.3233/SHTI220667. [DOI] [PubMed] [Google Scholar]
  • 12.Guenette JP, Kapoor N, Lacson R, et al. Development and Assessment of an Information Technology Intervention to Improve the Clarity of Radiologist Follow-up Recommendations. JAMA Netw Open. 2023;6(3):e236178. doi: 10.1001/jamanetworkopen.2023.6178. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Guenette JP, Lynch E, Abbasi N, et al. Actionability of Recommendations for Additional Imaging in Head and Neck Radiology. J Am Coll Radiol. 2024;S1546–1440(24)00007–3. doi: 10.1016/j.jacr.2024.01.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.White T, Aronson MD, Sternberg SB, et al. Analysis of Radiology Report Recommendation Characteristics and Rate of Recommended Action Performance. JAMA Network Open. 2022;5(7):e2222549. doi: 10.1001/jamanetworkopen.2022.22549. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Hammer MM, Kapoor N, Desai SP, et al. Adoption of a Closed-Loop Communication Tool to Establish and Execute a Collaborative Follow-Up Plan for Incidental Pulmonary Nodules. American Journal of Roentgenology. 2019;212(5):1077–1081. doi: 10.2214/AJR.18.20692. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Jhala K, Lynch EA, Eappen S, et al. Financial Impact of a Radiology Safety Net Program for Resolution of Clinically Necessary Follow-up Imaging Recommendations. Journal of the American College of Radiology. 2023; doi: 10.1016/j.jacr.2023.12.016. [DOI] [PubMed] [Google Scholar]
  • 17.Brown TB, Mann B, Ryder N, et al. Language Models are Few-Shot Learners. arXiv; 2020. doi: 10.48550/arXiv.2005.14165. [DOI] [Google Scholar]
  • 18.Fink MA, Bischoff A, Fink CA, et al. Potential of ChatGPT and GPT-4 for Data Mining of Free-Text CT Reports on Lung Cancer. Radiology. Radiological Society of North America; 2023;308(3):e231362. doi: 10.1148/radiol.231362. [DOI] [PubMed] [Google Scholar]
  • 19.Nori H, King N, McKinney SM, Carignan D, Horvitz E. Capabilities of GPT-4 on Medical Challenge Problems. arXiv; 2023. doi: 10.48550/arXiv.2303.13375. [DOI] [Google Scholar]
  • 20.Abbasi N, Lacson R, Kapoor N, et al. Development and External Validation of an Artificial Intelligence Model for Identifying Radiology Reports Containing Recommendations for Additional Imaging. American Journal of Roentgenology. American Roentgen Ray Society; 2023;221(3):377–385. doi: 10.2214/AJR.23.29120. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Ji Z, Lee N, Frieske R, et al. Survey of Hallucination in Natural Language Generation. ACM Comput Surv. 2023;55(12):248:1–248:38. doi: 10.1145/3571730. [DOI] [Google Scholar]
  • 22.Pedregosa F, Varoquaux G, Gramfort A, et al. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research. 2011;12:2825–2830. [Google Scholar]
  • 23.Seabold S, Perktold J. Statsmodels: Econometric and Statistical Modeling with Python. Austin, Texas; 2010. p. 92–96. doi: 10.25080/Majora-92bf1922-011. [DOI] [Google Scholar]
  • 24.Mabotuwana T, Hall CS, Tieder J, Gunn ML. Improving Quality of Follow-Up Imaging Recommendations in Radiology. AMIA Annu Symp Proc. 2018;2017:1196–1204. [PMC free article] [PubMed] [Google Scholar]
  • 25.Lau W, Payne TH, Uzuner O, Yetisgen M. Extraction and Analysis of Clinically Important Follow-up Recommendations in a Large Radiology Dataset. AMIA Jt Summits Transl Sci Proc. 2020;2020:335–344. [PMC free article] [PubMed] [Google Scholar]
  • 26.DeSimone AK, Kapoor N, Lacson R, et al. Impact of an Automated Closed-Loop Communication and Tracking Tool on the Rate of Recommendations for Additional Imaging in Thoracic Radiology Reports. Journal of the American College of Radiology. 2023;S154614402300399X. doi: 10.1016/j.jacr.2023.05.004. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Desai S, Kapoor N, Hammer MM, et al. RADAR: A closed-loop quality improvement initiative leveraging a safety net model for incidental pulmonary nodule management. Jt Comm J Qual Patient Saf. 2021;47(5):275–281. doi: 10.1016/j.jcjq.2020.12.006. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.OpenAI Platform. https://platform.openai.com. Accessed March 14, 2024.
  • 29.Dutta S, Long WJ, Brown DFM, Reisner AT. Automated detection using natural language processing of radiologists recommendations for additional imaging of incidental findings. Ann Emerg Med. 2013;62(2):162–169. doi: 10.1016/j.annemergmed.2013.02.001. [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material

Table S1: Likert Scale assessment of Recommendation for Additional Imaging Rationale extracted by Large Language Models (LLMs)

Table S2: Evaluation Guidelines for Large Language Model (LLM) Outputs

Table S3: Posthoc analysis of differences in accuracy between recommendation for additional imaging (RAI) details

Table S4: Posthoc analysis of differences in accuracy between subspecialties

RESOURCES