Skip to main content
Journal of Imaging Informatics in Medicine logoLink to Journal of Imaging Informatics in Medicine
. 2025 Jun 24;38(Suppl 1):1–49. doi: 10.1007/s10278-025-01501-x

Selected Abstracts from the SIIM 2025 Annual Meeting of the Society for Imaging Informatics in Medicine (SIIM)

PMCID: PMC12443643  PMID: 40555941

Title: SIIM 2025 Annual Meeting

Date: May 21–23, 2025

Venue: Oregon Convention Center – Portland, OR

Sponsorship: Publication of this supplement was sponsored by the Society for Imaging Informatics in Medicine (SIIM). All content was reviewed and selected by the Annual Meeting Program Committee, which held full responsibility for the abstract selections.

Applied Informatics Abstracts

001 - A Radiology Ontology for AI Datasets, Models and Projects (ROADMAP)

Presenter: Abhinav Suri, University of California Los Angeles

Abhinav Suri1, Safwan Halabi2, Hari Trivedi3, Charles Kahn4

1University of California Los Angeles, Los Angeles, CA, USA

2Northwestern University, Chicago, IL, USA

3Emory University, Atlanta, GA, USA

4University of Pennsylvania, Philadelphia, PA, USA

Background/Problem Being Solved: With the explosion of research involving the use of AI, systematic identification of categorization of AI models and datasets to enable discoverability remains a challenge. While many efforts have been established to standardize terminology for methods and datasets used in literature, no structured nomenclature exists to catalog AI models and datasets in the radiology realm to programmatically capture their relevant clinical and technical aspects.

Intervention(s): To address the lack of a structured nomenclature to describe radiology AI models and datasets, we developed an ontology to describe key information necessary for both clinicians and engineers to discover, search, and rapidly evaluate these resources.

Barriers/Challenges: Given the breadth of research in this area and its constantly evolving nature, it was critical to capture information that was generalizable, comprehensive, and standardized to encode model and dataset information.

Outcome: We created the Radiology Ontology for AI Datasets, Models and Projects (termed ROADMAP) that captures 3,457 classes and 4,541 subclass relationships. In brief, the ontology describes a top-level entity (a “project”) with 0 or more associated models and 0 or more associated datasets. Models are described using a living index of tasks and methods (reconstructed from the PapersWithCode) along with existing radiological ontologies such as RadLex, the RSNA Radiology Playbook, and Radiology Common Data Elements. Datasets are described in a comparable manner and contain additional information such as imaging types/sequences, PHI considerations, and patient demographics.

Conclusion/Statement of Impact/Lessons Learned: ROADMAP can be used to comprehensively describe models and datasets used in AI applications in radiology in a structured, searchable format. Future directions for this work include testing the comprehensiveness of ROADMAP on an extant set of publications and creating an initial set of structured model and dataset cards using the ontology.

Keywords: Applications; Artificial Intelligence/Machine Learning; Standards & Interoperability

002 - A Robust and Reliable Processing Pipeline on the Cluster for Quantitative Pediatric Neuroimaging

Presenter: Hyuk Jin Yun, Children's Mercy Hospital

Hyuk Jin Yun1, Matthew Genchev1, Colin Dietz1, Sarah Foster1, Maura Sien1, Sherwin Chan1, Avner Meoded1

1Children’s Mercy Hospital, Kansas City, MO, USA

Background/Problem Being Solved: Pediatric neuroimaging is essential for identifying brain anomalies in clinical practice. Quantitative measures from magnetic resonance imaging (MRI) have shown potential in detecting subtle brain abnormalities. However, integrating these quantitative MRI measures into clinical workflows faces significant challenges: 1) non-biological variance due to diverse processing methods across different age groups, and 2) long processing times, delaying clinical decision-making and impacting patient care.

Intervention(s): To address these limitations, we developed a robust and fast neuroimaging processing pipeline for structural and diffusion MRI applicable across all pediatric age groups. FreeSurfer and ACAPULCO were applied to structural MRI for obtaining morphological measures in cortical, subcortical, and cerebellar parcellations. Diffusion tensor imaging (DTI) is processed by FSL and DSI Studio to derive regional diffusion measures, such as fractional anisotropy and fiber tracking.

Barriers/Challenges: The primary challenges include ensuring consistency in processing methods across different age groups and reducing the processing time to facilitate timely clinical assessments. Implementing the pipeline on a high-performance computing (HPC) cluster, comprised of multiple nodes with 56 CPU cores, 768GB of RAM, and Nvidia T4 GPUs, was crucial to overcoming these barriers.

Outcome: Using our pipeline, MRIs from in-house 165 normative pediatric subjects aged 2 to 18 years were successfully processed. On HPC cluster, average processing time per patient was 5 hours and 33 minutes. This demonstrates that our pipeline is a reliable approach for analyzing MRI and DTI data across a wide age range, offering fast processing times that assist clinical workflows.

Conclusion/Statement of Impact/Lessons Learned: Our neuroimaging processing pipeline provides a robust and valuable resource for performing quantitative MRI in clinical practice. The fast-processing time is particularly beneficial in clinical settings where timely diagnosis and treatment can significantly impact patient outcomes. Future work will focus on validating the pipeline with larger datasets and exploring its potential integration into clinical workflows.

Keywords: Applications; Clinical Workflow & Productivity; Emerging Technologies; Imaging Research

003 - Addressing Workflow Pain Points in Radiology with AutoHotkey: Leveraging ChatGPT for Macro and Hotkey Creation

Presenter: Douglas Spaeth-Cook, Emory University Hospital

Douglas Spaeth-Cook1, Dan Cohen-Addad1

1Emory University, Atlanta, GA, USA

Background/Problem Being Solved: Radiologists frequently encounter workflow inefficiencies that disrupt productivity and increase cognitive load. Simple tasks, such as switching between systems or accessing frequently used resources, are common sources of frustration. Automation tools like AutoHotkey (AHK) offer a practical means of mitigating these inefficiencies, especially when paired with AI coding assistants like ChatGPT.

Intervention(s): This project explored how ChatGPT can assist radiologists in creating customized AHK scripts to streamline common tasks. These scripts provided functionality such as:

1. Task Switching: Rapid toggling between PACS, dictation software, web browsers, and paging systems using customizable hotkeys.

2. Quick Access Tools: One-step launching of frequently used websites such as Radiopaedia, UpToDate, and StatDx for on-demand references.

3. Menu Creation: Custom radial or dropdown menus for launching multiple tools in one interface.

Barriers/Challenges: Technical Barriers: Syntax or Coding Limitations.

1. ChatGPT may sometimes generate AHK code with minor syntax errors or inefficiencies. Users with limited coding experience might struggle to troubleshoot or debug scripts.

Practical Barriers: IT Restrictions.

1. Hospitals and health systems often have strict IT policies that prohibit the use of third-party scripts, macros, or automation tools like AutoHotkey.

2. Running AHK scripts may trigger antivirus software or be flagged as a security risk, requiring administrative permissions.

Outcome: ChatGPT served as a coding partner to create AHK scripts tailored to individual radiologist workflows. By providing natural language instructions to ChatGPT, users could generate scripts without requiring advanced programming knowledge. These scripts were tested in real-world radiology environments.

The integration of AHK scripting supported by ChatGPT demonstrated clear benefits, including:

1. Faster task switching.

2. Reduced cognitive disruptions during image interpretation.

3. Improved access to clinical resources and tools.

This approach highlighted the potential of AI tools to simplify workflows, reduce cognitive burden, and improve efficiency in practice.

Conclusion/Statement of Impact/Lessons Learned: The success of ChatGPT-enabled AHK scripting reflects broader efforts to integrate ergonomic and customizable tools into radiology. Building on strategies outlined by McGrath et al. (2022) and Grigorian et al. (2023), this work highlights the value of technology in improving productivity and reducing cognitive load. By automating repetitive tasks and streamlining access to resources, radiologists can better navigate today’s increasingly complex clinical environments.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Educational Systems; Enterprise Imaging; Provider Experience

004 - AI Assisted Annotation and Federated Learning: How to Standardize Data Labeling to Boots Training Performance

Presenter: John Garrett, University of Wisconsin

John Garrett1, Iman Zare Estakhraji2, Farzana Ali3, Gian Marco Conte4, Prerna Dogra5, Shahriar Faghani4, Mona Flores5, Reza Forghani6, Brad Generaux5, Amilcare Gentili7, Kristopher Kersten5, Meghan Lubner1, Andrew Missert4, Spencer Workman8, Joseph Yacoub9, Khaled Younis10, Yuankai Huo8

1University of Wisconsin, Madison, WI, USA

2GE Healthcare, Madison, WI, USA

3Stony Brook University, Stony Brook, NY, USA

4Mayo Clinic, Rochester, MN, USA

5NVIDIA, Santa Clara, CA, USA

6University of Florida, Gainesville, FL, USA

7University of California San Diego, San Diego, CA, USA

8Vanderbilt University, Nashville, TN, USA

9MedStar Georgetown, Washington, DC, USA

10MedAIConsult, Cleveland, OH, USA

Background/Problem Being Solved: Federated learning presents a promising approach to leveraging extensive datasets in medical imaging without central data collection, thus maintaining data privacy and reducing ePHI breach risks. This method involves simultaneous model training across multiple sites with only the network weights being shared. The focus of this project is on renal cell carcinoma (RCC) segmentation, where variability in data labeling across sites poses significant challenges. RCC is a pertinent choice due to its high occurrence and potential benign nature, necessitating accurate segmentation for subsequent diagnostic analyses.

Intervention(s): To address labeling inconsistencies, AI-assisted annotation tools will be employed to standardize the data annotation process across different sites. This project utilizes a hybrid dataset comprising a public dataset (TCGA-KIRC) for training and a comprehensive dataset from UW Madison for testing. The intervention includes manual and AI-assisted annotations using open-source tools like ITK Snap and 3D Slicer, followed by federated learning using open source libraries to evaluate model performance and data requirements.

Barriers/Challenges: The primary challenge is the inherent variability in training masks generated by different annotation tools and protocols used at various sites, potentially degrading model performance and increasing data requirements.

Outcome: The project aims to demonstrate whether AI-assisted annotations can enhance the efficiency and consistency of data labeling in federated learning setups, thus potentially improving model performance and reducing the number of samples needed for effective training.

Conclusion/Statement of Impact/Lessons Learned: By integrating AI-assisted annotation within a federated learning framework, this initiative expects to set a benchmark for improved segmentation accuracy and operational efficiency in medical imaging. Successful outcomes will provide crucial insights into optimizing deep learning tasks across diverse clinical environments, offering a scalable model for future multi-site medical imaging projects. Results and methodologies will be shared publicly to aid further research and development in this domain.

Keywords: Artificial Intelligence/Machine Learning; Imaging Research

005 - An Interactive NLP Approach for Improving Completeness and Annotation Efficiency in Prostate Screening Reports

Presenter: Jeroen Geerdink, University of Twente

Hridya nair Suresh1, Jeroen Veltman1, Lars Bosboom1, Maikel Viskaal1, Jeroen Geerdink1, Shenghui Wang1

1University of Twente, Enschede, Netherlands

Background/Problem Being Solved: Radiology is an important component of healthcare, playing a vital role in disease diagnosis. Thus, the completeness of these reports is essential, as minor errors can significantly affect the diagnosis and further treatment. The mistakes or missing fields in the report can arise due to factors such as increased workload, time constraints and inexperienced radiologists. This research focuses on automating the process of checking reports and providing radiologists with suggestions for any missing information. An interactive interface is also developed where the model is deployed for the radiologist to use, and also to derive annotation from the user interactions to solve the problem of limited annotated datasets.

Intervention(s): The prostate screening radiology reports we used for this study are Dutch semi-structured text data, thus Natural language Processing (NLP) techniques were used to extract the important information from the reports. Dutch Language models BERTje and MedRoBERTa.nl were tested for this task, but they exhibited overfitting due to a limited dataset. A hybrid Conditional Random Field model was implemented in identifying fields. The model was able to identify the majority of the fields. The lower performance for certain fields is attributed to the underrepresentation of these fields in the reports. To address the challenges of limited data and underrepresentation, we developed an interface that integrates the model into the radiologists’ workflow, allowing for both the application of the model and the collection of annotations through user interactions.

Barriers/Challenges: We were unable to integrate the interactive interface into the hospital system fully. We use the non-interactive interface that shows the model results into the system, but to know the performance of the interactive one in real-time instead of at the end of reporting is not still done.

Outcome: The Hybrid Conditional Random Field model was the effective model that identified the fields except for the fields that were underrepresented. This indicates that while the CRF model is adaptable, it requires more annotated data to improve accuracy in identifying all fields consistently. For quantitative evaluation, F1 scores were used to assess the accuracy of the CRF model, which ranged from 0.94 to 0.45, with the lower scores attributed to underrepresented fields like ”aspect” and ”grootte.”

Conclusion/Statement of Impact/Lessons Learned: The model developed can aid the radiologists in checking the report's completeness and compliance. By integrating the model into an interface radiologists can use, and also by leveraging the interface for collecting annotated data, we can increase the efficiency of the annotation process.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Quality Improvement & Quality Assurance

006 - Approaches for LLM Based Automated Categorization of Hospital Imaging Safety Event Reports

Presenter: Chintan Shah, Cleveland Clinic

Chintan Shah1, Brandon English1, Shelly Horvath1, Rekha Mody1, Po-Hao Chen1

1Cleveland Clinic, Cleveland, OH, USA

Background/Problem Being Solved: Organizing safety event reports by patterns is critical for quality improvement. Currently, these are manually categorized into one of 40 predefined categories by trained personnel, a time intensive process. Large Language Models (LLMs) offer a promising avenue for reducing manual effort.

Intervention(s): A dictionary of definitions and examples for 40 pre-defined categories was created by subject matter experts (SMEs). Approaches for LLM-based automated categorization were tested on 320 reports and compared with SME-assigned categories. Llama 3.1-8b was used and hosted locally via Ollama. Four approaches were evaluated:

1. Zero-Shot Categorization: A prompt was provided with all 40 categories.

2. Instructional Modification: To address frequent misclassifications involving “Near Miss,” clarification guidelines were introduced.

3. Top-3 Restriction: Top-3 most frequently used categories were identified. LLM was limited to selecting one of the top three when applicable, or "Other” when not applicable.

4. Retrieval-Augmented Categorization: Category definitions were embedded and used to retrieve the top 10 candidate categories, which were then passed to the LLM for final categorization.

Barriers/Challenges: Challenges included the large number of categories, nuanced definitions, and utility to manual workflow.

Outcome: Approach #1 resulted in 49% (158/320) correct category assignments. “Near miss” reports were incorrectly assigned (0/32). Approach #2 improved the categorization rate overall to 53% (170/320), and in the “Near miss” category to 63% (20/32). Approach #3 improved performance in each of the top 3 categories (51/52 “contrast extravasation”, 25/32 “delay in exam”, and 30/32 “near miss”). However, other reports were commonly mis-categorized into these categories, resulting in an overall lower performance of 44% correct (140/320). In approach #4 the correct category was retrieved in 89% (284/320), however, the final categorization was correct in 64% (205/320).

Conclusion/Statement of Impact/Lessons Learned: Integrating LLMs into safety event categorization workflows has highlighted the challenges above. Subsequent efforts focus on summarizing the event for trained personnel to expedite categorization.

Keywords: Artificial Intelligence/Machine Learning; Quality Improvement & Quality Assurance

007 - AtTheViewBox: No-code Solution in Creating Collaborative Interactive Cased-Based Presentations

Presenter: Michael Fei, Creighton University

Michael Fei1, Dane Van Tassel1, Vineeth Gangaram2

1Creighton University, Omaha, NB, USA

2University of Pennsylvania, Philadelphia, PA, USA

Background/Problem Being Solved: Radiology lectures are largely taught through PowerPoint presentations with 2D screenshot images. This experience does not mimic the workflow of reading radiological studies at the workstation. AtTheViewBox is a tool that can be used to embed DICOM files in presentations with the tools of scrolling, windowing, panning, zooming, and “multiplayer interaction” where learners can view cases and project their view onto the presenter screen. Our previous work demonstrated strong learner demand for this format and that such presentations can be built with existing open-source software libraries. However, we noted that coding as a requirement was a significant barrier to entry for adoption of these tools.

Intervention(s): A UI no-code solution was created around our web-based application, AtTheViewBox. The AtTheViewBox application allows users to load various DICOM cases. Educators can embed these cases into their presentations and create collaborative session rooms. Learners can scan a QR-code to join the session rooms and broadcast their screen to others in the room. We also created a Chrome extension that allows educators to easily import cases into AtTheViewBox. View https://mfei1225.github.io/AtTheViewBox_Demo_Site/ for demonstration.

Barriers/Challenges: Users still have to find and import cases into AtTheViewBox. Often finding cases online and/or exporting cases from PACs presents one of the largest barriers. In the future, we plan on including a list of publicly available cases that users can query.

Outcome: Residents and educators from the University of Pennsylvania, Creighton University, Emory University, and Mallinckrodt Radiology Departments, were surveyed after showing them a demo presentation with small bowel obstruction cases. The educators showed strong interest in this format with an average rating of 9.5/10 agreeing that it would be a strong asset to medical education, however, they reported an average probability of 58% for being able to create such a presentation themselves. The same educators were surveyed after showing them a video of the no-code interface used to build the demo and the probability improved significantly to an average of 83%. Additionally, a majority noted that it was similar to or easier than their experiences making case presentations using PowerPoint.

Conclusion/Statement of Impact/Lessons Learned: Creating a no-code UI around AtTheViewBox narrows the gap of demand in cased-based learning and the ability for educators to create such presentations. Ultimately this can enrich the educational experience by replicating reading at the workstation and oral boards. Based on these results, workshops to teach attendings how to use this tool have been scheduled at University of Pennsylvania and Creighton University.

Keywords: Applications; Educational Systems; Emerging Technologies

008 - Automated Generation of Patient-Friendly Multimedia Reports for Whole Body MR Studies from Text-Only Radiology Reports

Presenter: James Ledoux, Cornell University

James Ledoux1, Kurt Teichman1, Dana Galizia1, Keith Hentel1, Jacob Kazam1, George Shih1

1Cornell University, New York, NY, USA

Background/Problem Being Solved: Radiology reports traditionally are text-only and use terminology that patients have difficulty understanding. Patients can benefit from a patient-friendly multimedia report containing key images of findings.

Intervention(s): We created a system to assist in creating patient-friendly multimedia reports for Whole Body MR. The radiologists insert a Key Image macro in their report in the form “Key Image: {Series: 2, Image: 6, Description: 3mm left thyroid nodule, Follow-up: No}” for each finding. We use mirth-connect as an HL7-interface engine to monitor our HL7 results feed. Mirth find these dictated exams in prelim status and launches a REST API call to a flask HTTP listener that calls the script to construct the patient-friendly report.

We create a docx report from a template with python-docx. We parse the radiologist report into report sections and identify key images. Key images and their annotations are queried and pulled from PACS using pynetdicom and store data locally using SQLite. DICOM metadata is read using pydicom. Images and their annotations are recreated using matplotlib using window level and rescaling from the DICOMheader. We replace each body-part section of the report text with no identified key findings with the text “no significant findings.” We then upload the docx to an internal webserver and send an email with a link to the report to our internal team within about 30 seconds of the report being saved.

Barriers/Challenges: Manual proofreading and editing are currently part of the workflow to ensure that the multimedia text and images are accurate. We are also working on automating the insertion of patient-friendly explanations of important and incidental findings using a RAG-based LLM workflow.

Outcome: These patient-friendly multimedia reports can be distributed to patients via Epic including the Epic patient portal and app.

Conclusion/Statement of Impact/Lessons Learned: We have an automated system for generating patient-friendly multimedia reports from standard text-only radiology reports.

Keywords: Clinical Workflow & Productivity; Communication Data Management; Patient/Family Experience

009 - Automated Notification System That Facilitates AI Triage on a Legacy PACS System

Presenter: Adam Flanders, Thomas Jefferson University

Adam Flanders1, Paras Lakhani1, Ryan Lee1, Avi Sharma1, Tod Simons1, Vijay Rao1

1Thomas Jefferson University, Philadelphia, PA, USA

Background/Problem Being Solved: AI triage applications are designed to alert radiologists by setting a flag on the PACS worklist in order to ensure that the flagged study is read ahead of other exams. Many legacy PACS and RIS systems do not accept or cannot process an AI result which would prioritize specific exams with potential acute findings. We implemented a custom notification system that creates an alert on a legacy PACS to provide a means for the busy radiologist to address acute findings in a timely fashion.

Intervention(s): The notification system consists of four components: (1) AI result processing engine, (2) results database, (3) radiologist/workstation database and (4) notification poller. Integral is a real-time database called “presence management” which keeps a dynamic record of all working radiologists and active workstations currently logged into the Philips Intellispace PACS system. All radiologists/workstations are categorized by role and location. Contemporaneous AI results for acute intracranial hemorrhage are sent by HL7 to a MIRTH receiver which parses the message and stores the accession number, the location where the exam originated and the AI result. An active poller interrogates the AI results database for new positive results and matches the exam to active radiologists/workstations in that location. Shift variation is accommodated. A custom modal window and a bell sound on the specific workstation(s) to deliver the alert to the most appropriate radiologist(s) to avoid alert fatigue for others. The radiologist can launch the exam from the modal window or dismiss the window.

Barriers/Challenges: Challenges were primarily developing rules that would simultaneously accommodate a large academic core and multiple community practices to minimize alert fatigue. The capability of sending alerts to the right individual at the right time is inexorably tied to having an up-to-date inventory of active radiologists matched to workstation locations; this can be difficult to maintain in large, heterogeneous multi-specialty practices.

Outcome: Since its inception nine months ago, there have been 347 ICH alerts sent and acknowledged by specific radiologists throughout this practice covering eighteen hospitals in two states. There has been rapid adoption and acceptance of this relatively minimalistic alert mechanism which is tightly integrated into the core PACS viewer that delivers the alert to the most appropriate person augmenting care delivery and minimizing interruptions.

Conclusion/Statement of Impact/Lessons Learned: An alert system based on roles and locations for AI triage applications on a legacy PACS can minimize delays in care by quickly notifying the most appropriate radiologist.

Keywords: Administration & Operations; Applications

010 - Balancing Performance and Cost: The Role of COTS GPUs in Medical Imaging

Presenter: Megan Puertas, Cleveland Clinic - Florida

Megan Puertas1, Marvin Tucker1

1Cleveland Clinic Florida, Weston, FL, USA

Background/Problem Being Solved: Advanced imaging technologies in healthcare imaging imposes the use of high-performance graphics processing units (GPUs) that meet rigorous standards for reliability, precision, and regulatory compliance. However, medical-grade GPUs are often expensive. This financial burden can strain the budgets of healthcare institutions, potentially limiting access to imaging solutions. As healthcare systems worldwide struggle with increasing costs, finding cost-effective alternatives have become imperative. Commercial off-the-shelf (COTS) GPUs, primarily developed for gaming and general computing, present a potential solution to this problem. These GPUs are mass-produced, widely available, and less expensive than their medical-grade counterparts. However, adopting COTS GPUs in healthcare imaging raises concerns about their ability to meet performance standards and regulatory requirements.

Intervention(s): The analysis focused on computational speed, accuracy, and compatibility with existing imaging applications. Testing was completed to simulate real-world scenarios, such as 3D rendering and real-time processing. Additionally, compatibility testing was prioritized to demonstrate how well COTS GPUs integrate with existing hardware and software in imaging healthcare settings. This included evaluating their interoperability with picture archiving and communication systems (PACS) and DICOM (Digital Imaging and Communications in Medicine) images.

Barriers/Challenges: A primary concern was the inability to meet the stringent regulatory requirements imposed on medical devices, which are critical for ensuring patient safety and data integrity. Another significant challenge was ensuring consistent performance and reliability. While COTS GPUs excel in gaming and general computing tasks, their performance can be inconsistent under the high computational loads typical of medical imaging applications. Furthermore, resistance from stakeholders also posed a barrier, as they were hesitant to deviate from established norms and adopt unproven solutions.

Outcome: The study revealed the importance of rigorous testing and strategic implementation. Healthcare organizations that adopted COTS GPUs reported positive outcomes, including enhanced operational efficiency and optimized budgets. However, the study also reinforced the need for ongoing monitoring and periodic re-evaluation to ensure that performance and reliability remain aligned with evolving medical standards.

Conclusion/Statement of Impact/Lessons Learned: The adoption of COTS GPUs in medical imaging can provide cost reduction and improved access to advanced technologies. This approach enables healthcare providers to allocate resources more efficiently, potentially broadening access to innovative imaging solutions. This study emphasizes the importance of balancing cost savings with performance and regulatory compliance. Rigorous testing, stakeholder engagement, and the development of standardized evaluation protocols are crucial for the successful integration of COTS GPUs into medical workflows.

Keywords: Administration & Operations; Standards & Interoperability

011 - Can GPT-4 Do Your Systematic Review? Large Language Models as a Research Assistant for Literature Reviews in Radiology

Presenter: Dana Alkhulaifat, Children's Hospital of Philadelphia

Dana Alkhulaifat1, Satvik Tripathi2, Suhani Dheer2, Allison Brea3, Yohan Kim2, Dania Daye4, Tessa Cook2

1Children's Hospital of Philadelphia, Philadelphia, PA, USA

2University of Pennsylvania, Philadelphia, PA, USA

3Tufts University, Medford, MA, USA

4Harvard University, Cambridge, MA, USA

Background/Problem Being Solved: The increasing adoption of large language models (LLMs) in radiology research underscores the need to evaluate their utility in research workflows, particularly in evidence synthesis. This study focuses on assessing GPT-4’s ability to perform data exploration and visualization tasks, critical components of systematic reviews, while emphasizing its potential to streamline research processes in radiology.

Intervention(s): A systematic search was conducted across five databases: PubMed, EMBASE, SCOPUS, Web of Science, and IEEE Xplore. Boolean operators and targeted keywords, including “Large Language Models,” “Radiology,” and specific LLMs like GPT-4 and LLaMA, were used. Only original research articles published from 2022 onwards were included, excluding reviews, commentaries, editorials, and preprints. Two independent reviewers conducted the title and abstract screening, with adjudication by a third reviewer as needed. Data were extracted on application domains, specific LLMs used, and publication year. GPT-4 was employed to assist with data synthesis, and visualization in the context of systematic reviews, showcasing its ability to enhance efficiency and accuracy in these tasks.

Barriers/Challenges: Ensuring the accuracy and reliability of LLM outputs and validating the suitability of generated visualizations for scientific reporting remain key challenges.

Outcome: GPT-4 was evaluated for its performance in processing extracted data and generating figures such as bar plots, trend analyses, pie charts, and word clouds. The resulting visualizations accurately represented key trends, including adoption rates of LLMs, domain-specific applications, and keyword frequencies.

Conclusion/Statement of Impact/Lessons Learned: GPT-4 enhances systematic reviews by streamlining data exploration and visualization. However, limitations include potential biases, occasional inaccuracies, and inability to perform critical appraisal. Human oversight is essential, making GPT-4 a useful but complementary tool.

Keywords: Artificial Intelligence/Machine Learning

012 - CDE Definition Stubs: Building a Scalable Foundation for Structured Radiology Reporting

Presenter: Michael Hood, Massachusetts General Hospital

Michael Hood1, Tarik Alkasab1, Heather Chase2, Roshan Fahimi1

1Massachusetts General Hospital, Boston, MA, USA

2Microsoft Nuance, Burlington, MA, USA

Background/Problem Being Solved: Radiology reports are rich with clinical information but predominantly exist as unstructured free text. Standardized structured reporting, which represents findings as FHIR structures labeled with Common Data Element (CDE) identifiers, promises to revolutionize imaging workflows. However, developing CDE definitions is complex and time-consuming, and the lack of published CDE definitions covering the breadth of findings described in clinical radiology has hindered adoption. Recent work has demonstrated that LLMs can help to accelerate CDE development. The next step is to apply this capability to create usable CDEs at scale.

Intervention(s): To address this challenge, we introduce the concept of the CDE definition stub, a bare-bones data model of a radiology finding that includes:

1. The name of the finding and an AI-generated description.

2. Basic attribute definitions for Presence (whether the finding is present) and Change from Prior (unchanged, new, or changed).

While these stubs are simple and would not characterize findings' details, they encode sufficient information to enable applications based on structured imaging findings data. Furthermore, they serve as first drafts that can evolve into more detailed data models as processes mature to add more attribute definitions.

We developed a toolset that 1) allows users to enter the name of a finding, which is used to generate a JSON representation of the stub; and 2) converts the stub JSON into a format compliant with the ACR/RSNA CDE JSON schema, enabling submission for formal review.

Barriers/Challenges: While CDE definition stubs simplify the creation of structured data, validation and refinement still require radiologist input.

Outcome: Proof-of-concept testing demonstrated the toolset’s ability to rapidly generate JSON outputs for selected findings.

Conclusion/Statement of Impact/Lessons Learned: The concept of CDE definition stubs, combined with a supporting toolset, addresses key challenges in scaling structured radiology reporting, paving the way for radiology reports to power downstream vendor applications, automated workflows, and precision medicine.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Communication Data Management; Emerging Technologies; Imaging Research; Standards & Interoperability

013 - Chameleon Dataset- Large Synthetic Radiology Report Dataset: A Pilot Study

Presenter: Satvik Tripathi, University of Pennsylvania

Satvik Tripathi1, Dana Alkhulaifat1, Rithvik Sukumaran1, Charles Chambers1, Darco Lalevic1, Emiliano Garza2, Meghana Muppuri2, Dania Daye3, Tessa Cook1

1University of Pennsylvania, Philadelphia, PA, USA

2Massachusetts General Hospital, Boston, MA, USA

3Harvard University, Cambridge, MA, USA

Background/Problem Being Solved: The availability of open-access radiology report datasets remains a critical limitation in academic radiology research. With the growing adoption of large language models (LLMs), compliance with protected health information (PHI) regulations poses significant challenges, restricting the scope of research leveraging models like GPT. This limitation necessitates the development of PHI-free, high-quality text datasets to advance natural language processing (NLP) and LLM research in radiology.

Intervention(s): We introduce Chameleon, a large-scale synthetic radiology report dataset generated entirely using the GPT-4 API. The dataset was constructed without any PHI, providing a resource for training and evaluating NLP methods and LLM-based applications in radiology.

Barriers/Challenges: Key challenges encountered during dataset development included ensuring consistency in linguistic style, maintaining medical and anatomical accuracy, and adhering to structured reporting formats characteristic of clinical radiology.

Outcome: 500 synthetic radiology reports were generated, encompassing diverse imaging modalities and pathologies, including head CT, thoracic CT (with and without contrast), and abdominal CT. Four expert medical reviewers systematically evaluated thoracic reports based on content accuracy, format integrity, and stylistic coherence. Four prompting strategies were tested, with the optimal approach selected based on a cumulative scoring metric derived from the reviewers' evaluations. The mean time for generating a single report was 0.23 seconds, highlighting the efficiency and scalability of the methodology.

Conclusion/Statement of Impact/Lessons Learned: The Chameleon dataset addresses a critical gap in radiology research by providing a scalable and PHI-compliant resource for LLM and NLP applications. This synthetic dataset ensures ethical data usage while enabling advancements in radiology-specific AI research. Future efforts will aim to expand the dataset’s scope, incorporating additional imaging modalities and refining generation techniques to enhance clinical applicability and translational potential.

Keywords: Artificial Intelligence/Machine Learning; Emerging Technologies; Imaging Research

014 - Cost Justification Model for AI Triage Tool Integration

Presenter: Irene Lee, Emory University

Irene Lee1

1Emory University, Atlanta, GA, USA

Background/Problem Being Solved: Artificial intelligence (AI) tools are increasingly adopted in radiology to enhance diagnostic accuracy, optimize workflows, and reduce turnaround times. However, the financial implications of AI implementation, including initial costs, operational savings, and return on investment (ROI), remain underexplored. Understanding these factors is essential for strategic planning and justifying the adoption of AI solutions in clinical settings.

Intervention(s): This study evaluates the cost impact of integrating AI-based tools into radiology workflows across three domains: automated triage for critical findings, quality assurance (QA) of imaging protocols, and structured reporting enhancements. Data is collected from a large academic hospital that is implementing AI tools for these purposes.

Barriers/Challenges: There are certain challenges to doing a cost analysis. These include: Data

1. Availability and Quality: Access to accurate and comprehensive data on AI implementation costs, workflow metrics, and financial outcomes may be limited. Inconsistent documentation of AI tool performance and operational metrics can hinder analysis.

2. Generalizability Across Institutions: The findings may be specific to our institution, limiting the generalizability to other settings with different workflows, patient populations, or financial structures.

3. Cost Attribution Complexity: Distinguishing the financial impact of AI tools from other concurrent workflow changes, such as staffing adjustments or new policies, can complicate cost-effectiveness analysis.

4. Lack of Long-Term Data: Assessing ROI and cost-effectiveness requires longitudinal data, which may not be readily available since there are a lot of newly developed AI models in the past few years.

Outcome: We plan to analyze the cost associated with AI implementation including licensing fees, infrastructure upgrades, and personnel training, while financial benefits are assessed through metrics such as reduced reporting errors, improved radiologist productivity, and faster critical case turnaround. We also propose to conduct a time-motion study to measure workflow efficiency, and changes in reimbursement rates due to improved reporting quality are also evaluated.

Conclusion/Statement of Impact/Lessons Learned: Integrating AI into radiology workflows offers significant potential to enhance efficiency and improve patient outcomes. This study highlights the critical need for robust cost-benefit analyses to guide institutional decision-making and ensure sustainable AI adoption.

Keywords: Administration & Operations; Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Enterprise Imaging; Quality Improvement & Quality Assurance

015 - Design and Implementation of an Integrated EMR/RIS/PACS Platform in Africa

Presenter: Maruf Adewole, University of Pennsylvania

Kelvin Njeru1, Maruf Adewole2, Roselyne Okello3, Henry Rabah3, Udunna Anazodo4

1Sonar Imaging Centre, Nairobi, Kenya

2University of Pennsylvania, Philadelphia, PA, USA

3Kenyatta University, Nairobi, Kenya

4McGill University, Montreal, Quebec, Canada

Background/Problem Being Solved: Private and faith-based healthcare facilities account for the majority of health services in Africa. Access to affordable and Fast Healthcare Interoperability Resources (FHIR) compliant software hinders efficient patient care leading to poor health outcomes. This has precipitated manual record keeping and siloing of patient records within facilities. Current commercial platforms for record keeping: Electronic Medical Records (EMR), Radiology Information System (RIS), and Picture Archiving and Communication System (PACS) are expensive and not well integrated across the disparate clinical pathways/facilities. This challenge is a major contributor to delivery of timely and accurate diagnoses in Africa.

Intervention(s): We have developed the Sonar Informatics Platform (SIP) to offer an open, web-based RIS and PACS platform with a custom EMR integrated for cost-effective acquisition, sharing, and archiving of medical data and images. SIP leverages the open-source OpenMRS and Orthanc PACS functionalities and is hosted on our own servers with multiple redundancies ensuring compliance with local data protection laws, scalability and reliability. Secure access was implemented using Firebase. Individual facilities self-host a lite version of the EMR, synchronized with the central EMR.

Barriers/Challenges: High server set up cost.

Outcome: SIP was deployed in a pilot imaging center in Nairobi, Kenya and at two sister peripheral imaging centers in rural Kenya (400 km from Nairobi). It demonstrates improved efficiency in data management and accessibility. User feedback indicates enhanced workflow integration and significant cost savings compared to existing commercial RIS/PACS system s. Detailed performance metrics, including image archival and retrieval times and cost savings are under evaluation.

Conclusion/Statement of Impact/Lessons Learned

SIP enables timely access to life saving diagnostics and care with a secure single patient record across facilities in a resource limited setting. Future steps include embedding AI tools (e.g., Speech to text and large language models).

SIP enhances the secure sharing of patient data among facilities while ensuring affordability. This potentially improves patient outcomes and reduces the cost of care.

Keywords: Clinical Workflow & Productivity; Enterprise Imaging; Storage

016 - Developing a Standardized Ontology for Imaging Findings Using Large Language Models with Crowd-Sourced Iteration and Validation

Presenter: Yilun Zhang, Massachusetts General Hospital

Yilun Zhang1, Tarik Alkasab1

1Massachusetts General Hospital, Boston, MA, USA

Background/Problem Being Solved: Radiology reports contain valuable imaging findings and clinical recommendations, but inconsistent reporting styles and lack of a standardized structure and ontology hinder their integration into clinical workflows. A reliable, comprehensive, and consistent ‘lingua franca’ for imaging findings with extractable structure and features allows for bidirectional automation of report generation and understanding.

Intervention(s): In collaboration with the Radiological Society of North America (RSNA) and the American College of Radiology (ACR), we developed and leveraged the Common Data Elements (CDE) framework to evaluate and expand standardized radiology ontologies. A large language model (LLM) pipeline was developed using de-identified local report texts from Massachusetts General Hospital and publicly available datasets such as MIMIC, enabling the creation and refinement of a robust, scalable framework.

Barriers/Challenges: Key challenges included harmonizing diverse reporting styles and achieving consensus for CDE definitions among radiologists in a scalable, federated manner without exposing patient information.

Outcome: We developed an open-source, LLM-powered tool that enables users to access validated CDEs, contribute new or iteratively revise existing CDEs with the aid of an LLM, and ingest their local reports for analysis without risk of data exposure. The tool also unifies across and integrates with RadLex, SNOMED CT, LOINC, UMLS, and other ontologies. This allows for distributed and secure updates to the CDE framework and LLM-powered utilities in a scalable and asynchronous manner.

Conclusion/Statement of Impact/Lessons Learned: This work demonstrates the potential of an open-source tool that combines LLMs and the CDE framework to create a standardized, scalable ontology for radiology findings with additional utilities such as LLM-generated CDEs and reports and automatic extraction of CDEs from a local corpus of reports. This novel framework and tool drives human-in-the-loop automation, fosters interoperability, and revolutionizes clinical decision-making by bridging the gap between fragmented data silos and transforming the traditionally free-text narrative report into a structured output of actionable clinical insights.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Emerging Technologies; Quality Improvement & Quality Assurance; Standards & Interoperability

017 - Developing Interactive Systems for Enhanced Communication and Feedback Between Medical Experts and Explainable AI Models

Presenter: Presenter: Jeroen Geerdink, University of Twente

Xun Zhao1, Oswald Kessels1, Esmée Dijkstra1, Lars Bosboom1, Jeroen Veltman1, Jeroen Geerdink1, Johannes Hegeman1, Maurice van Keulen1, Shenghui Wang1

1University of Twente, Enschede, Netherlands

Background/Problem Being Solved: Despite the promise of explainable AI models like PIP-Net in medical imaging, no user-facing interface has been developed to bridge these models with real-world clinical workflows. While PIP-Net offers intrinsic interpretability by classifying images based on key visual concepts (prototypical parts) and addressing undesirable behaviors like biases and shortcut learning, its practical effectiveness and alignment with medical workflows remain unexplored. To fully realize the potential of such models in clinical decision-making, an interactive interface is needed to facilitate user engagement, feedback, and integration into real-world medical practices.

Intervention(s): The application framework contains three main components: (1) integration of PIP-Net and YOLO models for hip fracture inference, enabling users to compare outputs from interpretable prototype networks and segmentation-based models; (2) feedback tools allowing users to validate or reject model explanations and provide corrective annotations to refine the model’s performances; and (3) an intuitive interface equipped with a message center for guidance and explanations, as well as a "find similar images" feature to contextualize model outputs by referencing relevant training data. These components aim to align AI functionality with clinical workflows, enhancing usability, trust, and model accuracy.

Barriers/Challenges: One of the key challenges is guiding users on how to make accurate annotations, as different users may have varying standards. Some annotations may be too large and imprecise for further model training, while others may be too small, leading to a lower perceived performance of the model compared to human evaluation.

Outcome: The outcome of the study showed that most users were satisfied with the interface, particularly with the message center for the guidance and the "find similar images" feature for enhancing context. Additionally, all users found the YOLO model to perform better in both classification accuracy and explainability compared to other models.

Conclusion/Statement of Impact/Lessons Learned: In conclusion, the integration of PIP-Net and YOLO models into a user-friendly interface successfully facilitated user interaction and feedback, enhancing the model's alignment with clinical needs. The positive response to the interface's features, such as the message center and "find similar images" tool, demonstrates its potential to support clinicians in real-world applications. The preference for the YOLO model highlights its effectiveness in providing accurate classifications and transparent explanations, underscoring the importance of model performance and interpretability in medical imaging tasks. Further refinement of the annotation process and interface will improve usability and model accuracy.

Keywords: Applications; Clinical Workflow & Productivity; Imaging Research

018 - Development, Evaluation, and Assessment of Large Language Models (DEAL) Checklist: Reporting Methods and Results for LLM-Based Radiology Research

Presenter: Satvik Tripathi, University of Pennsylvania

Satvik Tripathi1, Dana Alkhulaifat1, Florence Doo2, Pranav Rajpurkar3, Rafe Mcbeth1, Dania Daye3, Tessa Cook1

1University of Pennsylvania, Philadelphia, PA, USA

2University of Maryland, Baltimore, MD, USA

3Harvard University, Cambridge, MA, USA

Background/Problem Being Solved: Large language models (LLMs) are increasingly employed in radiology for automated reporting, workflow enhancement, and decision support tasks. Despite their transformative potential, radiology research involving LLMs often suffers from inconsistent methodology reporting, limiting reproducibility, generalizability, and clinical integration.

Intervention(s): We developed the Development, Evaluation, and Assessment of Large Language Models (DEAL) Checklist to address these challenges. This standardized reporting framework is designed specifically for LLM applications and builds on existing guidelines, such as the Checklist for Artificial Intelligence in Medical Imaging (CLAIM) and EQUATOR network. It provides a comprehensive structure for documenting methodologies, evaluation metrics, and ethical considerations in radiology-focused LLM research.

Barriers/Challenges: Radiology presents unique challenges in the development and application of LLMs. These include dataset biases arising from differences in radiology reports across various imaging modalities, subspecialties, and institutions, as well as the need for models to be compatible with combined textual and imaging information. Additionally, the inherent stochastic variability in LLM outputs complicates the reliability and clinical applicability of results. The absence of standardized protocols for fine-tuning LLMs and optimizing prompt engineering furthers the disparities in model performance, complicating their deployment across diverse clinical settings.

Outcome: The DEAL Checklist offers two reporting pathways: DEAL-A, for studies focusing on the development and fine-tuning of LLMs tailored to radiology tasks, and DEAL-B, for studies utilizing proprietary or pre-trained models. Both pathways emphasize the reporting of model specifications, dataset preparation, evaluation metrics, and ethical considerations, ensuring transparency and reproducibility.

Conclusion/Statement of Impact/Lessons Learned: The DEAL Checklist is a crucial step in advancing rigorous and reproducible LLM-based radiology research. Improving standardization facilitates the reliable integration of LLMs into clinical practice, ultimately enhancing diagnostic accuracy, workflow efficiency, and patient care outcomes.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Emerging Technologies; Imaging Research

019 - Dexa Automation into Powerscribe Reports: An Implementation Guide for Radiologists, PACS Administrators and Technicians

Presenter: Florence Doo, University of Maryland

Nathan Bumbarger1, Devina Chatterjee2, Brandon Cofield2, Olga Haan2, Stephanie Jo2, Florence Doo2

1United States Air Force, San Antonio, TX, USA

2University of Maryland, Baltimore, MD, USA

Background/Problem Being Solved: Osteoporosis-related fractures are a significant health concern, particularly hip fractures, which have high mortality rates in older adults. Dual-energy X-ray absorptiometry (DEXA) is the gold standard for diagnosing osteoporosis but requires manual transcription of data into radiology reports, increasing inefficiency and error risk. In the context of radiologist shortages, optimizing workflows is essential to reduce burnout and improve productivity.

Intervention(s): A vendor-neutral structured reporting (SR) system was implemented using Hyland’s PACSgear ModLink and Nuance PowerScribe 360. This system automated the transfer of DEXA data directly into radiology reports, eliminating manual transcription. Custom fields were created to map DEXA values to structured templates, allowing seamless integration into the radiology reporting process.

Barriers/Challenges: Setting up the SR system required approximately 10–15 hours of PACS administrator time over two months, including mapping fields and testing workflows. Radiologists needed to review and delete irrelevant fields in the templates. Minimal training was required for DEXA technologists to adapt to the new workflow, and technical challenges were addressed by support teams.

Outcome: Implementation of the SR system led to a significant reduction in report generation time, with radiologists experiencing a 2 to 5 fold improvement in efficiency. No significant errors were observed in mapped values. The process did not add measurable workload for technologists and was well integrated into existing workflows.

Conclusion/Statement of Impact/Lessons Learned: Automating DEXA data integration into radiology reports significantly improves efficiency, reduces transcription errors, and supports radiologists in managing increasing workloads. The setup process requires modest effort but offers meaningful return on investment for radiology practices. This approach has the potential to improve access to radiologists for patient care and provides a scalable solution for addressing inefficiencies in reporting workflows.

Keywords: Clinical Workflow & Productivity; Enterprise Imaging; Quality Improvement & Quality Assurance; Standards & Interoperability

020 - Diagnostic Radiology Workflow with “at-a-glance” Air Traffic Control-like Methods Ensuring Patient Safety

Presenter: Les Folio, James A. Haley Veterans Hospital

Les Folio1, Kristi DuBois2, Krista Edgar2, Ann Folk1

1James A. Haley Veterans Hospital, Tampa, FL, USA

2Moffitt Cancer Center, Tampa, FL, USA

Background/Problem Being Solved: Radiology departments have a variety of processes for managing patients undergoing general diagnostic exams, for example ICU CXR, OR miscounts, interdepartmental x-rays, and fluoroscopy. We present a physical process/ workflow that maximizes patient safety, similar to a companion high reliability organization: air traffic control and the methodology applied to controlling aircraft with handoffs ensured.

Intervention(s): We describe general diagnostic workflows that apply an “at-a-glance” process similar to air traffic control methods where physical flight strips have shown to be effective in combat trauma imaging triage. In addition to scheduled exams, each new exam study has an audible indicator and a physical location for displaying printed exam requests representing patients in various phases/locations. Our CT section also has a visual indication based on printed requests and color and location of folders, however, for this work, we focus on general diagnostic imaging.

Barriers/Challenges: Paper driven workflows in radiology are continually questioned, for example “why don't we go 100% digital?” or “why do we move these requests from here to there?” Also, some might think that paper driven workflows are not "informatics" since no computers are involved; on the contrary, workflow processes ARE INFORMATICS by definition; computers or not.

Outcome: Upon opening an additional hospital, we refined a paper driven workflow with consistent success in managing patients undergoing diagnostic radiology exams and procedures. We describe a positive chain of custody process of each patient with little to no opportunity to “drop the ball.”

Technologist inputs that have worked in nearby medical centers and those recently out of training, overwhelmingly favor paper driven workflows over those completely digital. They share stories of where the ball was dropped, a patient had been waiting where technologists essentially forgot about them, where paper would have prevented the oversight. We plan on "a show of hands" during Q&A for those claiming to be 100% digital (since we believe there will be fewer, easy to count!).

Conclusion/Statement of Impact/Lessons Learned: Patients have been known to be overlooked in some completely digital radiology workflows; for this reason, our center is not doing away with paper driven processes in the foreseeable future. Our paper requisitions maintain a physical chain of custody of each patient with an instant "visual contract” as to status at all times.

We share a clever display of paperwork driven by priority, location and status. If nothing else, our successful experience may help justify other imaging departments that continue to have a paper-driven workflow.

Keywords: Administration & Operations; Clinical Workflow & Productivity; Provider Experience; Quality Improvement & Quality Assurance

021 - Display or Dismay? Initial Experience in Deploying 52 Workstations with Consumer-grade Displays for Remote Diagnostic Workstations

Presenter: Katie Hulme, Cleveland Clinic

Katie Hulme1, Jen Arnold1, Ryan Thomas1, Douglas Nachand1, Namita Gandhi1, Po-Hao Chen1

1Cleveland Clinic, Cleveland, OH, USA

Background/Problem Being Solved: Institutional demand necessitated a large-scale cost-effective solution be provided within a consolidated time frame to facilitate hybrid work for our radiologists

Intervention(s): 52 remote workstations with commercial-grade displays (WCD) were deployed alongside 30 workstations with diagnostic-grade displays (WDD) over 12 months. The selected WCD display (Dell Technologies, G3223Q) met the specifications for diagnostic (non-mammography) displays outlined in the 2022 ACR-AAPM-SIIM Technical Standard. Displays were calibrated to the Grey Scale Standard Display Function (GSDF) at a white point of 375 cd/m^2 using a handheld photometer (X-Rite, i1Display Pro) and software (QUBYX, Perfectlum 4). Displays were evaluated for uniformity, visual integrity, and GSDF conformance prior to deployment. Quality control (QC) training was provided to radiologists. QC consisted of display calibration and visual assessment of a test pattern and was required to be performed monthly.

Barriers/Challenges: While WDDs have built-in photometers and can automatically run scheduled QC tasks, QC for WCDs require user-initiation; increased rates of non-compliance (overdue QC or unresolved failures) coupled with frequent QC failures for WCDs require workflow improvements and support effort from imaging informaticists and physicists.

Outcome: QC test histories were exported for analysis from all remote workstations. Stations had been deployed, on average, for 128 days (WCD) and 238 days (WDD). Failures in GSDF compliance occurred 20% of the time for WCD compared to 1.5% for WDD, with maximum absolute deviations from GSDF of 10.8% for WCD, on average, compared with 3.4% for WDD. WCD were able to maintain a white point of 375 cd/m^2 over the evaluated period, but with higher variability than WGD.

Conclusion/Statement of Impact/Lessons Learned: The deployed WCDs performed within specifications but with more variability and higher failure rates, suggesting post-deployment operational expenses may offset initial capital savings. Diagnostic-grade displays may prove more cost-effective over time through a comparative total cost analysis.

Keywords: Administration & Operations; Quality Improvement & Quality Assurance; Standards & Interoperability

022 - Early Experience with 3D LC-OCT (Line Field Confocal-Optical Coherence Tomography) Imaging Show Potential to Reduce Skin Biopsies

Presenter: Thomas Beachkofsky, James A. Haley Veterans Hospital

Thomas Beachkofsky1, Les Folio1

1James A. Haley Veterans Hospital, Tampa, FL, USA

Background/Problem Being Solved: Line Field Confocal-Optical Coherence Tomography (LC-OCT) is an FDA cleared emerging, non-invasive, non-ionizing imaging modality capable of cellular and molecular-level resolution. This technology has potential to reduce invasive biopsies of suspicious cancerous skin lesions by enabling precise, 3D real-time visualization.

Traditional dermatologic imaging primarily relies on visible light-based tools such as specialized visible light photography and dermoscopy. Challenges include bulkiness, operator variability, and technical limitations. LC-OCT addresses these barriers by providing real time three-dimensional, cross-sectional, and en-face imaging, enabling dermatologists/dermatopathologists to assess tissue microstructure detail only previously available in histologic sections. LC-OCT is currently being clinically implemented in several medical centers globally, including three centers in the United States.

Intervention(s): We present LC-OCT as a groundbreaking imaging modality with applications in diagnosing and managing common skin cancers, including basal cell carcinoma, melanoma, and squamous cell carcinoma. We share clinical use cases, example 3D image data, review planar and imaging terminology and video demonstrations of real-time volumetric imaging.

Barriers/Challenges: Currently, there is limited accessibility to LC-OCT, compounded by high equipment and training expenses, combined with low reimbursement rates.

Technical limitations include small field of view (0.05 cm x 0.12 cm and depth of 0.05 cm) and artifacts caused by structures e.g. hair. Also, highly pigmented skin absorbs the imaging source, and can lead to lower resolution of deeper structures.

Regulatory challenges include insurance reimbursement for LC-OCT (remains limited), similar to other imaging modalities in dermatology.

Outcome: LC-OCT introduces a new era of non-invasive dermatologic imaging, offering a potential alternative to traditional biopsies for diagnosing potentially cancerous lesions. This modality enables mapping of superficial, spreading tumors before surgery, potentially reducing the need for wide excisions. It also provides real-time visualization of tumor margins, streamlining surgical procedures by eliminating delays associated with subspecialty pathology confirmation.

Conclusion/Statement of Impact/Lessons Learned: This abstract features cross-sectional and 3D imaging examples, real-time 3D volumetric imaging videos, and device graphics to illustrate LC-OCT’s capabilities. These advanced imaging tools, previously unique to radiology, are now transforming dermatologic practice. sectional and 3D images in addition to a video of real time 3D volumetric imaging including cut plane reformats similar to advanced PACS tools no longer unique to radiology.

Keywords: Applications; Artificial Intelligence/Machine Learning; Emerging Technologies; Enterprise Imaging; Imaging Research; Provider Experience

023 - Efficient Imaging Data Management for Enhanced Research Collaboration

Presenter: John Garrett, University of Wisconsin

Orhan Unal1, John Garrett1, Richard Bruce1

1University of Wisconsin, Madison, WI, USA

Background/Problem Being Solved: Current workflows for managing research imaging data are error-prone, non-standard, and usually unsuitable for modern large-scale projects. These limitations impede data analysis, sharing, integrity, and reproducibility, crucial for advancing medical research.

Intervention(s): The project introduces a two-phase workflow designed to optimize the operational processes for research imaging data management:

1. Data Capture and Anonymization:

Imaging data in RAW and DICOM formats are directly captured from MRI scanners, guided by user-specified parameters. Anonymization tools ensure compliance with privacy regulations at the scanner level before secure transfer to a staging area. Opting in by specifying individual exams rather than bulk captures reduces bandwidth strain and aligns with system limitations, making the process scalable and efficient.

2. Ingestion to Data Warehouse:

Data in the staging area is efficiently organized and pre-processed before entering the data warehouse. A single-command workflow allows users to initiate ingestion with just the exam number on MRI scanners, simplifying adoption for MR technologists. The containerized pipeline categorizes and ingests data securely into the centralized warehouse, ensuring streamlined access to organized datasets while minimizing redundancy.

Barriers/Challenges: Managing vast imaging datasets, such as tens of gigabytes of data per study and hundreds of thousands of images, places significant strain on systems. These challenges are further compounded by network bandwidth limitations and outdated infrastructure, which hinder secure, efficient data transfer, and storage.

Outcome: Initial findings indicate expedited workflows and enhanced data accessibility for downstream processing and analysis. This automated workflow improves operational efficiency, reduces data handling errors, standardizes processing, and supports collaboration and reproducibility.

Conclusion/Statement of Impact/Lessons Learned: This approach addresses key challenges in imaging data management, enhancing security, accessibility, and operational efficiency. By enabling easy exam-specific data capture and efficient staging processes for ingestion to data warehouse, it makes large-scale imaging research feasible and adaptable to other modalities like PET/MRI and CT.

Keywords: Administration & Operations; Clinical Workflow & Productivity; Enterprise Imaging; Imaging Research; Security; Storage

024 - Empowering Novice Developers in Radiology Informatics: Chat-Oriented Programming

Presenter: Patricia Wu, Beth Israel Deaconess Medical Center

Patricia Wu1, Mohamad Hotait1, David Kwan2, Seth Berkowitz1

1Beth Israel Deaconess Medical Center, Boston, MA, USA

2Insygnia Consulting Inc, Vaughan, Ontario, Canada

Background/Problem Being Solved: Radiology informatics often requires software development skills for facilitating research and clinical workflows, posing challenges for those without technical expertise. This project explored how a researcher with no prior coding experience leveraged AI assistance to create a fully functional application.

Intervention(s): Using a Large Language Model chatbot as a coding assistant, a functional web application was developed as a reading worklist for a research study. The application listed hyperlinks to a custom OHIF image viewer for participants to complete a research task. Upon completion of the task, results were saved to a database. The chatbot was used to create interface code to query a pre-existing API to display unread studies to the user. The process involved crafting clear and specific prompts to guide AI output, using the chatbot for debugging, and integrating the web application in the existing infrastructure under the supervision of an experienced mentor.

Barriers/Challenges: The primary challenge was the lack of prior coding experience. Deploying the code created by AI required knowledge of existing infrastructure.

Outcome: The project culminated in the successful deployment of a fully functional, user-friendly web portal for a reader study. The initiative demonstrated that individuals without traditional coding expertise can effectively contribute to informatics projects with the assistance of AI tools and proper mentorship.

Conclusion/Statement of Impact/Lessons Learned: This experience highlights the transformative potential of AI-assisted development and the importance of mentorship in bridging skill gaps, empowering researchers from diverse backgrounds to engage in and contribute to the field of radiology informatics. Tutorials, videos, technical forums, and search engines have historically guided interested students to learn programming skills. Large language models are a massive paradigm shift for motivated individuals to quickly create functional software prototypes. Chat-oriented programming may also have value in increasing the productivity of experienced developers.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Educational Systems; Imaging Research; Organizational & Professional Development

025 - Enhancing Diversity in Imaging Data: A Framework for Inclusive AI Development

Presenter: Lawrence Guan, Yale School of Medicine

Lawrence Guan1, Sophie Chheang1, Irene Dixe De Oliveira Santo1

1Yale University, New Haven, CT, USA

Background/Problem Being Solved: Artificial intelligence (AI) has rapidly transformed radiology, with algorithms excelling in tasks from pathology detection to guiding therapeutic interventions. However, the effectiveness, fairness, and generalizability of these tools hinge on the diversity of training datasets. Despite advances, radiology datasets often lack sufficient demographic, geographic, and disease-specific diversity, reinforcing biases and limiting applicability, particularly in underrepresented populations. This study aims are to: (1) identify gaps in publicly available radiology datasets by analyzing their demographic composition, geographic distribution, and disease variability; (2) propose actionable strategies for recruiting diverse patient populations to improve dataset inclusivity.

Intervention(s): We analyzed 15 publicly available radiology datasets (e.g., NIH ChestX-ray14, UK Biobank) to evaluate demographic and geographic representation. Metadata, including patient age, gender, race/ethnicity, and geographic origin, were extracted and compared against global population distributions using the Dataset Representation Index (DRI). To address identified gaps, we developed recruitment strategies focused on underrepresented populations and evaluated the role of federated learning for secure international data sharing.

Barriers/Challenges: The analysis uncovered significant demographic imbalances: rural populations, women, and racial minorities were underrepresented. Data primarily originated from North America and Europe, with limited contributions from Africa and South Asia, restricting applicability to low- and middle-income countries (LMICs).

Strategies that yield potential for mitigate these imbalances include: (1) implementation of targeted recruitment strategies and global collaborations to increase representation across key metrics, improving dataset diversity, and (2) federated learning enabled secure and privacy-preserving data sharing between international institutions, enhancing inclusion without compromising legal or ethical standards.

Outcome: Our findings reveal the critical need to address disparities in radiology datasets to mitigate risks of biased AI systems. Collaborative efforts with global health institutions, community-based recruitment initiatives, and the adoption of federated learning frameworks are vital for achieving equitable representation. Standardized demographic reporting and bias monitoring are essential to ensure ongoing fairness and transparency in AI development.

Conclusion/Statement of Impact/Lessons Learned: This study highlights the importance of fostering diversity in imaging datasets to enhance the reliability and fairness of radiology AI tools. By adopting our proposed framework, stakeholders can build inclusive systems that better serve global populations and contribute to more equitable healthcare delivery.

Keywords: Artificial Intelligence/Machine Learning; Patient/Family Experience; Quality Improvement & Quality Assurance

026 - Enhancing Efficiency: Time Savings from Autopopulating Ultrasound Measurements

Presenter: Dylan Sadowsky, Tulane University

Dylan Sadowsky1, Gautam Dua1, Kelsey Berman1, Yousef Dawoud1, Mitchell Jackson1, Parisha Babuta2, Mohammad Mousa1, Matthew Parker1, Jane Ball1, Mandy Weidenhaft1, Jeremy Nguyen1, Matthew Abella1, Cynthia Hanemann1

1Tulane University, New Orleans, LA, USA

2University of California, Irvine, CA, USA

Background/Problem Being Solved: Ultrasound reporting workflows, especially for abdominal ultrasounds—the most common type in many practices—are often inefficient. There is little data quantifying the time savings from automating the transfer of measurements in complete abdominal ultrasounds. Radiologists spend significant time manually dictating or inputting these measurements, a clerical task that takes time away from analyzing images and adds to cognitive load and burnout.

Manual entry is also prone to errors, compromising report accuracy and requiring corrections. With increasing imaging volumes and a radiologist shortage, maximizing efficiency is essential. Automating the population of ultrasound measurements into reports allows radiologists to focus on high-value tasks, improve accuracy, and reduce burnout. This study aims to quantify the potential time savings for abdominal ultrasound procedures through automation.

Intervention(s): This study implemented software to automate the transfer of ultrasound measurements into the reporting system. Automation was tailored to our complete abdominal ultrasound protocol, incorporating 13 standardized fields.

Two radiologists manually dictated 30 measurements for complete abdominal ultrasound studies. They timed the process from the first to the last measurement and recorded any errors.

Barriers/Challenges: Configuring various ultrasound machines was a logistical challenge, as each required multiple adjustments to ensure proper transmission of structured data due to differing settings.

When technologists recorded multiple measurements for the same organ, we adjusted the protocol to avoid defaulting to the mean. Instead, technologists were required to select the most accurate value for the data sheet.

Outcome: In our study, manually dictating measurements for complete abdominal ultrasounds took an average of 1 minute and 38 seconds, with a standard deviation of 23.6 seconds. For institutions performing 50 studies daily, this totals approximately 81.7 minutes per day. Over a year this equals about 354 hours.

The study also revealed a 13.3% error rate in manual dictation, including:

Major digit omissions: e.g., "1.47" instead of "15.47" and "3.156" instead of "31.56."

Rounding or minor discrepancies: e.g., "208.54" instead of "208.51."

Substitution of qualitative terms: e.g., "5 point" instead of "5.8."

Significant errors, such as "1.47" instead of "15.47," could misrepresent findings and impact clinical interpretation, while even minor discrepancies require time-consuming corrections. Substituting qualitative terms further highlights the variability in manual transcription. Automating this process eliminates these risks, improving both efficiency and accuracy.

Conclusion/Statement of Impact/Lessons Learned: Automating the measurement process eliminates this error-prone step, improves reporting accuracy by 13.3%, reduces cognitive load, and decreases professional burnout. Recovering 354 hours annually and reducing errors can significantly enhance workflow efficiency.

Keywords: Administration & Operations; Applications; Clinical Workflow & Productivity; Organizational & Professional Development; Provider Experience; Quality Improvement & Quality Assurance; Systems Management

027 - Enhancing Image Library Operations Through an Interactive Dashboard Improving Productivity and Decision-Making

Presenter: Gloria Hwang, Stanford University

Amy Bui1, Gloria Hwang1, Roniela Turingan1, Shantika Devi1

1Stanford University, Stanford, CA, USA

Background/Problem Being Solved: The previous image exchange system frequently experienced downtimes and was cumbersome for employees to use. In response to this, a new image exchange system was implemented to support the daily operations of the radiology image library. While the new system improved certain processes, it lacked reporting capabilities needed to provide visibility into the department’s operational and productivity metrics.

Intervention(s): To address this limitation, a dashboard was developed to complement the system by enabling the tracking of key metrics. The team defined key metrics such as study volumes from all image management locations, studies sent to the Picture Archiving and Communication System, employee productivity metrics, and study rejection rates. Working collaboratively with the vendor and internal teams, the team mapped the necessary webhook events, application programming interface endpoints, and production gateways to enable real-time data retrieval for these metrics. The data was consolidated into a subject-specific data mart, forming the foundation for the dashboard.

Barriers/Challenges: A key challenge was ensuring the accuracy and reliability of the data during the development phase, requiring thorough validation using test data and uploading test images to ensure the data was consistent with the expected operational reality.

Outcome: The dashboard visualized all of the key metrics and provided a comprehensive view to enable data-driven decision-making for daily task assignments and staffing allocation. By analyzing historical trends and real-time study volumes, the tool facilitated proactive resource planning and operational efficiency. With robust data available, leaders were able to anticipate busy periods and optimize staffing thus improving operational management.

Conclusion/Statement of Impact/Lessons Learned: The implementation of the dashboard has transformed image library operations by enabling quicker decision-making and providing automated daily data. This supports leadership in operational planning and performance measurement, driving continuous improvements in efficiency and service delivery.

Keywords: Administration & Operations; Clinical Workflow & Productivity; Emerging Technologies; Patient/Family Experience; Provider Experience; Standards & Interoperability

028 - Haske: An Open PACS Platform Integrating AI for Affordable Medical Image Archiving and Diagnosis in Resource-Constrained Settings

Presenter: Maruf Adewole, University of Pennsylvania

Maruf Adewole1, Ayomide Oladele2, Aanu Gbadamosi2, James Ajigbotosho3, Charity Umoren2, Kelvin Njeru4, Oluyemisi Toyobo3, Abiodun Fatade3, Farouk Dako1, Udunna Anazodo5

1University of Pennsylvania, Philadelphia, PA, USA

2Medical Artificial Intelligence Laboratory, Lagos, Nigeria

3Crestview Radiology Limited, Lagos, Nigeria

4Sonar Imaging Centre, Nairobi, Kenya

5Montreal Neurological Institute, Montreal, Quebec, Canada

Background/Problem Being Solved: Medical imaging workflows across Africa are hindered by the lack of accessible and affordable Picture Archiving and Communication System (PACS) platforms that adhere to Fast Healthcare Interoperability Resources (FHIR) standards. Existing commercial solutions are cost-prohibitive for many centers in resource-constrained environments. This gap restricts the development and deployment of AI solutions, further compounding the challenges of delivering timely and accurate diagnoses in underserved areas.

Intervention(s): Haske (meaning ‘Light’) is an open, web-based PACS platform designed to enable cost-effective acquisition, sharing, analysis and archiving of medical images. Haske leverages the open-source Orthanc PACS functionalities and is hosted on Amazon World Services (AWS) for equitable access. Secure access was implemented using Google's Firebase, while the Mercure DICOM orchestrator (mecure-imaging.org) was integrated for seamless interoperability with AI analysis tools. Haske includes an embedded reporting module, making it a comprehensive ecosystem for medical imaging management.

Barriers/Challenges: Implementation faced significant challenges from poor existing medical imaging infrastructure to unreliable internet connectivity. Cultural reliance on traditional practices like physical image storage, film printing and darkroom techniques is also a challenge. This is being resolved through iterative improvements and user feedback.

Outcome: Haske has been deployed in a pilot imaging centre in Lagos, Nigeria, where preliminary evaluation demonstrates improved efficiency in image management and accessibility. Early user feedback indicates an enhanced radiodiagnosis workflow process and potential for cost savings compared to existing commercial PACS systems. Detailed performance metrics, including image archival and retrieval times and AI accuracy rates, are under evaluation.

Conclusion/Statement of Impact/Lessons Learned: Haske’s integration of FHIR-compliant PACS functionalities with AI-enabled diagnostics offers a transformative solution for medical imaging in resource-constrained environments. It facilitates seamless image management and supports faster, more accurate diagnoses thereby reducing barriers to access and affordability leading to better patient outcomes and addressing critical healthcare challenges in Africa and beyond.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Enterprise Imaging

029 - Improving Radiologist Access to Critical Resources via a PACS-Integrated Web Platform

Presenter: Morgan McBee, University of South Carolina

Morgan McBee1, Wendy Ketchum1, David Gaspard1

1University of South Carolina, Charleston, SC, USA

Background/Problem Being Solved: At our institution, radiologists must navigate multiple disconnected resources to access various types of information and utilize different tools. These include internal websites, multiple cloud storage platforms, HR systems, policy repositories, and ticketing systems. Accessing them requires leaving PACS and is therefore disruptive to workflow.

Intervention(s): Our PACS allows embedding of web content within the worklist and information window. A web platform was created and styled to blend with the PACS interface and was embedded directly within the PACS. It consolidates essential clinical, technical, and administrative information into one location, enabling radiologists to remain within PACS while accessing information. Resources include things such as protocols, guidelines, policy documents, IT support, and ticketing. The platform’s familiar interface reduced training needs and improved accessibility.

Barriers/Challenges: Challenges included organizing content from multiple repositories and ensuring accuracy. Like any information repository, ongoing maintenance will be the largest challenge going forward.

Outcome: The platform reduced time spent searching for information. Informal feedback cited intuitive navigation and seamless design as key benefits. Radiologists reported less frustration with finding essential and relevant resources.

Conclusion/Statement of Impact/Lessons Learned: Embedding an integrated web resource system directly into PACS improved efficiency and increased satisfaction. Consolidating disparate resources into a single, accessible interface addressed a key informatics challenge for our radiologists. This approach offers a reproducible model for other institutions. Future efforts will focus on content updates and potentially expanding support to additional roles.

Keywords: Administration & Operations; Applications; Clinical Workflow & Productivity; Communication Data Management; Provider Experience; Systems Management

030 - Integrating Medical Informatics Training in Radiology Residency: A Pathway to Enhanced Clinical Impact

Presenter: Irene Dixe de Oliveira Santo, Yale University

Irene Dixe de Oliveira Santo1, Lawrence Guan1, Sophie Chheang1

1Yale University, New Haven, CT, USA

Background/Problem Being Solved: As artificial intelligence (AI) and informatics tools advance in radiology, integrating training in these areas into residency programs has become a topic of increasing interest. This study evaluates a semi-structured medical informatics mini-fellowship undertaken by two senior radiology residents (IDdOS, LG) during their final year at a tertiary academic center. The fellowship is designed to provide four months of focused training throughout their final year, guided by a primary mentor (SC) with additional support from the institution’s Program for Innovation in Imaging Informatics (PI3).

Intervention(s): The mini-fellowship includes structured training with protected time and institutional funding to complete the SIIM bootcamp, prepare for the Certified Imaging Informatics Professional (CIIP) exam, and finish Epic Physician Builder modules focused on data analytics and workflow optimization. Additionally, the residents are tasked with presenting selected articles at a journal club and complete individualized projects based on their interests. IDdOS further expanded her expertise by auditing machine learning courses at our associated university.

Barriers/Challenges: By the midpoint of the fellowship, both residents have made substantial progress. LG completed, and IDdOS nearly finished, the Epic Physician Builder course. IDdOS also completed the SIIM bootcamp and passed the CIIP exam. LG’s project focuses on developing a dashboard to streamline radiology-pathology correlation, aiming to improve diagnostic accuracy and patient outcomes. IDdOS is working on a project to optimize navigation between maternal and neonatal records to improve access to prenatal imaging in congenital disorder cases. Both residents selected journal club topics, and article authors were invited to enhance discussions.

Outcome: The mini-fellowship has the potential to equip residents with essential informatics skills, fostering innovation in clinical decision-making, workflow optimization, and data analysis. Projects addressing maternal-neonatal records and radiology-pathology correlation illustrate the real-world impact of informatics training on patient care. By including certification programs like CIIP and Epic Physician Builder, the fellowship enhances residents' ability to implement advanced solutions in clinical practice.

Conclusion/Statement of Impact/Lessons Learned: This mini-fellowship illustrates and emphasizes the importance of integrating informatics training into residency programs, with the potential to preparing radiologists to lead in precision medicine, optimize clinical workflows, and deliver patient-centered care.

Keywords: Administration & Operations; Clinical Workflow & Productivity; Educational Systems; Organizational & Professional Development; Provider Experience

031 - Iterative Improvement of PACS-Integrated AI Algorithm for Automatic Brain Metastasis Detection and 3D Segmentation

Presenter: Nazanin Maleki, Children's Hospital of Philadelphia

David Weiss1, Nazanin Maleki2, Khaled Bousabarah3, Cornelius Deuschl4, Sven Schoenherr3, Sahil Chadha1, Julian Lautenschlager3, Wolfgang Holler3, Malte Westerhoff3, Spyridon Bakas5, Ajay Malhotra1, Nagaraj Moily3, Veronica Chiang1, Sanjay Aneja1, Fatima Memon6, Elizabeth Schrickel7, MingDe Lin3, Mariam Aboian2

1Yale University, New Haven, CT, USA

2Children's Hospital of Philadelphia, Philadelphia, PA, USA

3Visage Imaging, GmbH, Berlin, Germany

4University Hospital Essen, Essen, Germany

5Indiana University, Bloomington, IN, USA

6Carolina Radiology Associates, Myrtle Beach, SC, USA

7The Ohio State University, Columbus, OH, USA

Background/Problem Being Solved: Accurate detection and segmentation of brain metastases (BM) are pivotal for diagnosing, treating, and surveilling patients with BM. Nevertheless, manual lesion measurement is time-consuming. In this study, we iteratively improved the performance of a deep learning-based algorithm (Model 1–3, M1-3) for BM detection and 3D segmentation by leveraging a research instance of our PACS that streamlines continual learning and clinician-in-the-loop feedback.

Intervention(s): In this retrospective single-center study, 156 pre- and post-Gamma Knife radiosurgery (GKR) MRI studies of patients with BM were de-identified and sent from clinical production to a research instance of our PACS (AI Accelerator, AIA, Visage Imaging, Inc.). Reference standard 3D BM segmentations were performed by a board-certified neuroradiologist using AIA, providing same tools as in clinic. Initially, nnU-Net (M1) was trained on a 227 single institution dataset. A team of clinicians, researchers, and data scientists assessed the model’s performance visually and quantitatively using AIA. In particular, false positive and negative segmentations were analyzed, and MR image quality was assessed by a second board-certified neuroradiologist using AIA. Algorithm was retrained and tested based on the feedback. This methodology was iterated. Modified nnU-Net (M2) and STU-Net with brain masking (M3) trained on BraTS-METS2023 dataset were investigated. Precision and sensitivity were used as metrics for BM detection.

Barriers/Challenges: A team of clinicians, researchers, and data scientists had to assess the MR image quality, ground truth segmentation masks, and AI-predicted segmentation masks manually to determine the strengths and weaknesses of the algorithm, which was then used to inform the iterative improvement of the AI algorithm.

Outcome: A total of 607 cerebral metastases from 156 studies from 40 patients were investigated. In BM detection, precision and sensitivity were 0.859 and 0.633 for M1, 0.921 and 0.825 for M2, and 0.924 and 0.827 for M3, respectively. False positive AI-predicted BMs outside of the brain were eliminated in M3.

Conclusion/Statement of Impact/Lessons Learned: AIA enables the development and iterative improvement of segmentation models. Our network, translatable to clinical PACS, provides precise automated detection and segmentation of brain metastasis.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Imaging Research

032 - Large Language Model Integration with Radiology Report Query Tool Enables Complex Searches

Presenter: Kurt Teichman, Cornell University

Kurt Teichman1, James Ledoux1, Chenglin Zhu1, Keith Hentel1, Jacob Kazam1, Martin Prince1, Shih George1

1Cornell University, New York, NY, USA

Background/Problem Being Solved: Most radiology report search tools (eg, document storage databases like solr and elastic search combined with basic NLP) allow for useful queries of radiology reports (eg, Find all the positive cases of appendicitis on CT studies in the ED for the last 12 months). However, they are limited in their capacity to do data extraction and/or filtering on documents for more complex requirements (eg, extracting discrete measurements out of free text in a structured format like JSON and/or complex queries like if lesions measure >= 2.5 cm).

Intervention(s): Here we provide methodology for integrating large language models (LLMs) on top of radiology report an initial query (eg, renal cell carcinoma on CT studies), and allow a user-defined LLM prompt (‘Find all renal tumors >= 2.5 cm and provide in JSON format’) to be run on the initial query result. Our approach uses a template mechanism, composing a report alongside a user prompt, to be sent to the LLM to be processed. This is similar to typical retrieval augmented generation (RAG) templating. We also provide means of storing query history and persisting LLM results as they can often take a long time to finish.

Barriers/Challenges: Accessing a LLM (e.g. OpenAI ChatGPT) is often difficult from inside institutions following HIPAA compliance, and thus, running your own LLM (e.g Llama 3.1 8b) is ideal. Providing multiple users reliable access to LLM services is also challenging as result processing takes time (i.e. each report needs to be passed in with your query). Asynchronous methodology is required. We utilize 2 queues, redis based as well as a queue provided with vLLM.

Outcome: Users are able to obtain refined search results often saving them considerable amounts of time. We should note, however, that these results are “mostly” accurate. The usage of smaller language models (e.g Llama 3.1 8B) can sometimes hallucinate results when input is of longer context size

Conclusion/Statement of Impact/Lessons Learned: We leverage LLMs along with modern software engineering practices to allow users to obtain more granular results searching radiology report archives. Future work will incorporate Multimodal LLMs to provide additional options for queries.

Keywords: Artificial Intelligence/Machine Learning

033 - Leveraging AI to Optimize Radiology Workflow: A ChatGPT-Based Organizational Assistant

Presenter: Nicholas Mynarski, Northwell Health

Nicholas Mynarski1, Matthew Barish1, Eran Ben-Levi1, Ritesh Patel1

1Northwell Health, New Hyde Park, NY, USA

Background/Problem Being Solved: Radiology departments are complex institutions with numerous protocols, policies, and procedures that are essential for appropriate patient care and diagnostic accuracy. However, troubleshooting protocols and accessing and understanding documentation can often be time-consuming for radiologists. These interruptions can lead to errors, decrease productivity, and delay patient care.

Intervention(s): To optimize the radiologist’s workflow, a ChatGPT (GPT-4o) chatbot was trained on a dataset of up-to-date departmental policies, protocols, procedures, standardized clinical support tools (LI-RADS criteria, etc.), contact information, and important links. The chatbot was instructed to operate within the bounds of its trained dataset and alert users if the information provided was not within its dataset.

Barriers/Challenges: Several barriers and challenges were encountered during the development and implementation of this chatbot. These included the need to accurately represent the complex and nuanced nature of departmental policies and procedures in a machine-readable format, ensuring the chatbot’s ability to understand and respond to a vast array of queries, and maintain the chatbot’s accuracy and relevance over time.

Outcome: The chatbot was successfully implemented and demonstrated significant benefits to the radiology department. The AI has been used extensively by the department’s body and chest divisions and radiology residents, actively being utilized by 150+ users with thousands of chatbot inquiries. Over time, interdepartmental interest has grown with other divisions and technicians seeking to integrate the chatbot into their workflow.

Conclusion/Statement of Impact/Lessons Learned: The development and implementation of a ChatGPT chatbot for radiology departments represents a significant advancement in the field of applied informatics. By leveraging the power of artificial intelligence, we can improve the efficiency and accuracy of clinical workflows (standardization), leading to improved patient care. However, it is crucial to carefully consider the challenges associated with developing and maintaining such a chatbot, including data quality and accuracy.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Emerging Technologies; Provider Experience; Standards & Interoperability

034 - Leveraging Data Analytics to Drive Improvements in Lung Cancer Screening Rates

Presenter: Amy Bui, Stanford University

Amy Bui1, Ken Lim1, Thomas Cartwright1

1Stanford University, Stanford, CA, USA

Background/Problem Being Solved: As part of efforts to establish a comprehensive lung cancer screening (LCS) program, the oncology department aimed to improve early-stage lung cancer detection by increasing screening rates among eligible patients. The initiative focused on tracking screening rates and identifying gaps in patient outreach to address the low national average rate of just 5%. To help solve this problem, an analytics solution was built to track progress and intervention effectiveness.

Intervention(s): The team integrated population health, primary care, and radiology datasets to identify two key populations: (1) patients eligible for screening based on the U.S. Preventive Services Task Force recommendations, and (2) eligible empaneled patients—those who had seen their PCP within the past 18 months. To better target different provider groups, the data was grouped by the enterprise's entities, representing distinct patient populations.

Barriers/Challenges: The most significant challenge while building the dashboard was accurately determining patient eligibility due to inconsistencies in smoking history documentation. It was also challenging to determine which of the eligible patients were empaneled.

Outcome: The dashboard effectively visualized lung cancer screening rates among eligible patients across the enterprise, distinguishing baseline screenings from annual follow-ups and tracking orders for newly screened patients. Additionally, it enabled providers to monitor the completion rate of smoking history documentation in the electronic health record (EHR) and validated the usage of the billing code for shared decision-making visits.

Conclusion/Statement of Impact/Lessons Learned: Access to comprehensive LCS data enabled the team to make timely adjustments, increasing the number of patients screened. The dashboard highlighted the gap in PCP knowledge about patient eligibility, leading to including LCS as a health maintenance topic in the EHR. Since launching the program and implementing monthly metric tracking, the average weekly screening rate has increased by 138%, from 8 to 19 patients per week over the past two years.

Keywords: Administration & Operations; Provider Experience; Quality Improvement & Quality Assurance

035 - Monitoring AI Performance After Deployment: Demonstrated in Pulmonary Embolism Detection

Presenter: Alicia Maehara, University of California Los Angeles

Alicia Maehara1, Gordon Guyant2, Melody Bounmasanonh2, Alexandria Uy2, Edward Zaragoza2, William Hsu2

1Marlborough School, Los Angeles, CA, USA

2University of California Los Angeles, Los Angeles, CA, USA

Background/Problem Being Solved: Commercially available AI systems have shown promising results in timely detecting pulmonary embolism. However, continuous monitoring of AI systems is necessary for optimal quality of care. This project aims to develop a system for continuously monitoring an imaging AI for detecting PE (Pulmonary Embolism).

Intervention(s): The system consists of: 1. A data pipeline that retrieves and transforms data from the imaging AI and CT reports from the medical record; 2. An open-source Large Language Model (Llama 3.1 8b) that extracts and structures PE results based on CT reports using a custom prompt; 3. An algorithm that processes and compares results; and 4. A dashboard to facilitate interpretation of the data and enable case review.

Barriers/Challenges: CT Pulmonary Angiogram has the highest specificity and sensitivity for detecting PE compared to V/Q scan, D-dimer, and compression ultrasound. Therefore, we develop a system to extract and structure PE results (PE and incidental PE AI) from contrast-enhanced CT reports by customizing Llama 3.1 8b and comparing those results with the imaging AI’s predictions.

Outcome: During the evaluation period between 07/08/2024 and 11/16/2024, the imaging AI detected 368/23,319 (1.6%) PE cases from CT exams. Concordance with CT report result was 98.8% (23,038/23,319). Positive predictive value was higher in PE CTA exams (PE CTA 84.9% vs PE Incidental 50%).

Conclusion/Statement of Impact/Lessons Learned: We have successfully developed an automated system for continuously monitoring an imaging AI tool for detecting PE and incidental PE. The pipeline is modularized to facilitate ingestion of data from diverse sources. This enables future monitoring of disparate AI systems in different domains where downstream data may be derived from clinical systems, registries, and patient-reported outcomes.

Keywords: Artificial Intelligence/Machine Learning; Quality Improvement & Quality Assurance

036 - Multiprotocol Analysis of Quantum Versus Classic Encryption Methods for Medical Imaging

Presenter: Young-Tak Kim, Massachusetts General Hospital

Youngtak Kim1, Synho Do1, John Mayfield1

1Massachusetts General Hospital, Boston, MA, USA

Background/Problem Being Solved: Given the sensitive nature of information within DICOM images and metadata, encryption is typically the initial step. Up until the recent technological era, AES/RSA have been highly resilient to traditional attacks. However, as quantum computing technologies have grown exponentially in the last few years, there is a growing concern that quantum computation methods such as Grover’s Algorithm can reduce the time to brute-force calculation of RSA/AES encrypted data to the square root of the traditional computation time.

Intervention(s): We propose a series of quantum key distribution (QKD) methods for encryption of DICOM based upon the unique tenets of quantum mechanics including the Heisenberg Uncertainty Principle (BB84, SARG04, E91, Measurement Device Independent and Twin Field Paradox protocols) and Phase Coherence Principle (Differential Phase Shift and Reference Frame Independent protocols). Additionally, we approached the problem from an adversarial point of view with exploration of evolving quantum hacking techniques including injection locking, power analysis, temporal ghost imaging, and wavelength control.

Barriers/Challenges: There are several limitations of the study including the simulated environment where generalizations were made in quantum implementation using a quantum hardware simulator, as well as the simulated photon within an optic fiber which may inherently be more prone to noise and signal loss over extended distances versus in open field implementations. Additionally, adversarial attacks were in a simulated environment and may not incorporate additional barriers such as air gaps, active firewalls, and AI countermeasures.

Outcome: In the experiments with QKD, there was 100% eavesdropping detection as any measurement of the quantum systems results in collapsing of the state into an expected value or observable. While RSA had a lower MITM success percentage, none of these attacks were detected. The potential trade-off is the longer key generation time of the QKD protocols. From the adversarial standpoint, injection locking demonstrated the greatest success rate in intercepting quantum keys without disrupting the QKD process, while power analysis exploited power consumption patterns to identify secret keys.

Conclusion/Statement of Impact/Lessons Learned: This pilot study is meant to start the conversation of future-proofing medical imaging given its inherent vulnerabilities with robust eavesdropping detection which may be of great utility in scenarios such as the recent CrowdStrike event. Implementing both the offensive and defensive strategies helped to identify potential vulnerabilities and opportunities for quantum cryptography across the medical imaging environment. Specific to the operational engineering aspect, the potential utility of these protocols could be readily implemented given the existing architecture of optic cables.

Keywords: Administration & Operations; Applications; Emerging Technologies; Enterprise Imaging; Imaging Research; Quality Improvement & Quality Assurance; Standards & Interoperability; Security; Systems Management

037 - Predicting Cholecystectomy Complexity Using Large Language Models: Enhancing Preoperative Decision-Making with AI

Presenter: Anagha Tirumalai Ramaswamy, Bangalore Medical College and Research Institute

Anagha Tirumalai Ramaswamy1, Mir Noor Hassan1

1Bangalore Medical College and Research Institute, Bengaluru, India

Background/Problem Being Solved: Cholecystectomy is a surgical procedure to remove the gall blader. The complexity of this surgery varies based on patient comorbidities, anatomical variations and interoperative challenges. Predicting cholecystectomy complexity preoperatively is crucial for optimizing surgical planning and improving patient outcome. Large language models (LLMs), offer a promising approach by leveraging vast amounts of unstructured clinical data. This study hypothesizes that large language models (LLMs) can accurately predict cholecystectomy complexity from preoperative ultrasound reports thereby aiding in precise preoperative planning and improving patient outcomes.

Intervention(s): An experienced laparoscopic surgeon (25 years) performed the surgery and stratified the intraoperative difficulty with Nassar grading scale (1–4), which was used as ground truth. Preoperative ultrasound reports of the abdomen were analyzed with Gemmav2 (locally run LLM) and the Nassar grading was predicted. This was repeated 7 times for each patient and the scores were averaged across trails. Kendall correlation between the intraoperative score and the predicted preoperative LLM score was evaluated. We also stratified the cases as high or low complexity and the accuracy of this classification was also computed.

Barriers/Challenges: Limited sample size, variability of USG reporting based on the radiologist, bias in the results due to predominantly female patients, limitations in the accuracy of LLM due to inadequate experience.

Outcome: Among the responses, 50% (15 patients) were accurate, 43.34% were acceptable and 6.66% (2 patients) were discordant. A Kendall’s tau coefficient of 0.67 (p < 0.001) was obtained with this model. When we evaluated the risk stratification as easy (Nassar2) we obtained an accuracy of 79.31%.

Conclusion/Statement of Impact/Lessons Learned: LLMs can predict the intraoperative risk scores using preoperative ultrasound reports and aid in precise preoperative planning.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Educational Systems; Imaging Research; Patient/Family Experience

038 - Pythia: An Integrated Large Language Model Platform for Enhanced Radiology Reporting

Presenter: Shawn Lyo, University of Pennsylvania

Shawn Lyo1, Satvik Tripathi1, Tessa Cook1

1University of Pennsylvania, Philadelphia, PA, USA

Background/Problem Being Solved: Large language models (LLMs) show significant promise in enhancing various aspects of radiology reporting, from clinical history synthesis and proofreading to differential diagnosis assistance. At our institution, access to LLM capabilities during reporting was previously limited to chat interfaces, which are cumbersome and inefficient. The lack of a centralized platform that efficiently integrates relevant clinical information with LLM tools reduces the potential impact of LLM tools in clinical practice.

Intervention(s): We developed Pythia, a comprehensive platform that integrates multiple LLM-powered tools to enhance radiology reporting. Built using Python Flask, an institutionally approved Azure OpenAI GPT-4 instance for prompt completions, and ChromaDB vector database, the platform accepts structured inputs including exam type, current report drafts, relevant prior reports, and clinical notes. Key functions include automated clinical history generation, proofreading, prior report comparison, and differential diagnosis assistance. Current implementation utilizes hotkey macros for data input; however, the architecture is designed for future automation and integration with radiology applications.

Barriers/Challenges: Development required navigating several challenges including operating within the constraints of hospital firewall restrictions, which limited available technical tools and necessitated careful architecture design. The user interface needed optimization to present information without overwhelming radiologists' attention. Prompt engineering with ensemble techniques was necessary to achieve high accuracy while limiting computational costs and response times.

Outcome: Pythia has been successfully implemented as a unified platform for integrating LLM-powered tools into radiology workflows. It is undergoing continued addition and refinement of features.

Conclusion/Statement of Impact/Lessons Learned: Pythia represents a significant step toward unified LLM integration in radiology workflow. By consolidating multiple reporting enhancement tools into a single platform, it has the potential to improve report quality and consistency while reducing radiologists' cognitive load. The platform's current implementation using hotkey macros is compatible with any reporting software stack but also maintains a framework for future direct system integration. This flexibility in deployment, combined with the platform's extensible design, creates a foundation for continuous improvement and adaptation to evolving needs in radiology reporting.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Emerging Technologies; Provider Experience; Quality Improvement & Quality Assurance

039 - RADAR (Real-time AI Data Assessment and Reporting): A Solution for Automated Monitoring of Commercial AI Model Performance Built with “Off-the-Shelf” Components

Presenter: Adam Flanders, Thomas Jefferson University

Adam Flanders1, Paras Lakhani1, Prahlad Menon1, Robyn Ball2, Luciano Prevedello3, George Shih4, Avi Sharma1, Ryan Lee1

1Thomas Jefferson University, Philadelphia, PA, USA

2The Jackson Laboratory, Bar Harbor, ME, USA

3The Ohio State University, Columbus, OH, USA

4Cornell University, New York, NY, USA

Background/Problem Being Solved: Most commercial AI customers have no independent ability to measure model performance or drift and must rely upon the vendor for this critical task. Presented herein is a prototype of a semi-autonomous application that continuously measures model performance for a triage brain CT hemorrhage model built with “off-the-shelf” components.

Intervention(s): The tool consists of four components: (1) an AI result receiver, (2) a result database, (3) a report interpreter and (4) a statistical engine. All results processed by CT hemorrhage inference engine were sent simultaneously to a MIRTH HL7 receiver for processing. The database was updated with the final radiology report when it was made available. The report impression was parsed and processed by an ensemble of five LLMs (llama3.2:1b, llama3.2:3b, codellama:7b, llama3.1:8b, granite3-dense: 2b) running in the Ollama framework. A “consensus” was reached when three or more LLMs agreed. Using the consensus as reference, a confusion matrix was created to generate AI model performance metrics. Fleiss’ and Cohen’s kappas were calculated to check agreement between the LLMs. Throughout, the administrator interacted with a web dashboard that provided the updated performance of the model and provided a means to inspect discordant results.

Barriers/Challenges: Automating report review requires a consensus of an ensemble of LLMs and an iterative approach using a combination of human review and prompt engineering as a means to minimize human evaluation. The challenge was to find a balance where only periodic human review of the automated report validation was necessary.

Outcome: The database monitored results from over ~16,000 BRAIN CT exams derived from eighteen hospitals and 35 scanners collected from nine months of continuous use. Due to heterogeneous inter-model agreement an ensemble of LLMs was chosen as consensus to confer more consistent results. An iterative process removed of low performing LLMs to boost performance. Finally, the chosen consensus was scored against expert evaluation of a subset of reports. Fleiss kappa for these LLMs: 0.73 and Cohen’s kappa ranged from 0.14 to 0.78.

Conclusion/Statement of Impact/Lessons Learned: An ensemble of LLMs was employed as a first pass to verify radiology report imaging findings and can be used to automate independent quality control of a triage AI application in the clinical setting.

Keywords: Administration & Operations; Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Quality Improvement & Quality Assurance

040 - RadOntology: An Ontology-Driven Intelligent Platform for Radiology Education and Clinical Insights

Presenter: Young-Tak Kim, Massachusetts General Hospital

Saul Langarica1, Young-Tak Kim1, Adham Mahmoud1, Jeanne Ackman1, Shaunagh McDermott1, Manisha Bahl1, Michael Lev1, Michael Gee1, Synho Do1

1Massachusetts General Hospital, Boston, MA, USA

Background/Problem Being Solved: Radiological services are experiencing unprecedented demand, and radiologists are increasingly pressured to manage a higher volume of image interpretations while maintaining high-quality diagnostic accuracy. Although traditional AI tools offer potential relief, their adoption in clinical practice has been hindered by challenges such as low reliability and lack of transparency. To address these issues, we introduce RadOntology, an enriched knowledge graph derived from structured radiology reports, designed to assist radiologists with image interpretation and support training and education. By integrating intuitive and accurate data retrieval through natural language queries, reliable clinical insights, and enhanced report generation assistance, RadOntology aims to alleviate workload pressures, improve diagnostic outcomes, and bridge the gap between AI innovation and clinical utility.

Intervention(s): RadOntology employs advanced large language models and Vision Transformers to transform unstructured radiology reports and images into an ontology-based knowledge graph. Through sophisticated entity extraction and structured data representation, the system enables users to perform intuitive natural language queries, explore relevant cases, and retrieve expert-authored radiology reports for guided report writing.

Barriers/Challenges: Key challenges include ensuring interoperability with existing radiology information systems, maintaining high-quality ontologies, and achieving broad acceptance among radiologists.

Outcome: RadOntology delivers an integrated and open-source platform that facilitates intuitive natural language queries to extract clinically relevant entities from radiology reports and retrieve structured data, such as similar expert-authored cases and imaging. These tools could reduce diagnostic ambiguity and increase confidence in clinical decision-making. The system’s educational capabilities enable radiologists and trainees to explore curated case collections grouped by anatomical region or pathology type, promoting deeper insights through comparative analysis.

Conclusion/Statement of Impact/Lessons Learned: RadOntology generates ontology-based knowledge graphs, facilitates advanced image pattern retrieval, and provides reliable references, empowering users with greater diagnostic confidence and educational value.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Educational Systems

041 - RadPath: Improving Radiology-Pathology Correlation Follow-Up for Radiologists via an Automated System

Presenter: Othria Ahmed, Northwell Health

Othria Ahmed1, Matthew Barish1, David Hirschorn1, Graham Keir, Ritesh Patel1, Gabriel Felder1

1Northwell Health, New Hyde Park, NY, USA

2Cornell University, New York, NY, USA

Background/Problem Being Solved: Radiology-pathology correlation is a critical part of the diagnostic process and patient care, enabling radiologists to validate imaging interpretation with biopsy-confirmed pathology. However, this process places a significant burden on individual radiologists and is often fragmented and time-intensive. These challenges are amplified within a large integrated healthcare network consisting of radiologists reading studies across more than 20+ facilities with over 2.5 million imaging studies annually. Additional complexities due to disparate clinical information systems, including multiple radiology information systems (RIS) and laboratory information systems (LIS), further complicate this process.

Intervention(s): We developed a fully automated system that notifies radiologists (attendings and residents) of relevant pathology correlations to their recently dictated studies. This system aggregates new pathology reports across the network, identifies relevant prior imaging, and provides a weekly notification email to the radiologist. The cases can be accessed via a web interface that displays final radiology and pathology reports side-by-side. The interface enables radiologists to provide immediate feedback on concordance or discordance, flag cases as interesting, and directly launch PACS for further analysis. This feedback is utilized to refine the system's matching algorithm and to build a comprehensive database of path-proven and interesting cases, serving as a resource for research and education.

Barriers/Challenges: Creating a system across a large healthcare network presented significant challenges, including the integration of different LIS platforms while accommodating the workflow of radiologists operating across multiple facilities. Additionally, creating an algorithm to accurately identify “relevant” imaging cases required nuanced analysis of pathology and radiology reports. Lastly, scalability and managing the sheer volume of data required unique solutions.

Outcome: We created a fully automated system for rad-path correlation and labeled a repository of pathology-proven radiology cases. Until Mar 2025, 17,518 cases have been reviewed with 162 distinct users. Additionally, users have marked 2,112 interesting cases and 214 discordant cases.

Conclusion/Statement of Impact/Lessons Learned: The implementation of Rad-Path Results has successfully streamlined the rad-path correlation process across a large healthcare network, enhancing diagnostic accuracy and facilitating radiologist workflow. With over 6,400 cases reviewed and a growing repository of pathologically proven cases, this system has demonstrated its value as a tool for improving patient care, research, and education. A similar system can also be implemented to facilitate pertinent imaging follow-up for flagged studies requiring further investigation.

Keywords: Clinical Workflow & Productivity; Educational Systems; Quality Improvement & Quality Assurance

042 - So You Want to Embrace Open Science in Research? How to Do This Responsibly

Presenter: Raisa Amiruddin, Children's Hospital of Philadelphia

Raisa Amiruddin1, Nazanin Maleki1, Nikolay Yordanov2, Pascal Fehringer3, Athanasios Gkampenis4, Ahmed Moawad5, Mariam Aboian1

1Children's Hospital of Philadelphia, Philadelphia, PA, USA

2Sofia University St. Kliment Ohridski, Sofia, Bulgaria

3Friedrich Schiller University Jena, Jena, Germany

4University of Ioannina, Ioannina, Greece

5Mercy Fitzgerald Hospital, Darby, Pennsylvania

Background/Problem Being Solved: Open science aims to develop multidimensional, generalizable knowledge that is findable, accessible, interoperable & reusable within the scientific community.

The neuroimaging community would benefit from high-quality, well-characterized, open-source databases, & publicly available tools for generating & analyzing curated data. Openness & connectivity on the design, performance, capture & assessment of research can create capacity for improvement & advocate for responsible & sustainable research and innovation.

Intervention(s): In the past decade, there has been an exponential increase in published neuroscience literature accompanied by raw data, encouraging the creation of fee-free data-sharing initiatives to be used for research like cell morphometry analysis, magnetic resonance (MR) image analysis, & genomic, proteomic & transcriptomic analysis. Availability of open datasets benefits the medical community, statisticians, data scientists & early career researchers with access to data without requiring access to an imaging center. Several platforms host & support a variety of deposited data.

Barriers/Challenges: Creation of open data comes with challenges in privacy & standardization. Modern search engines are adept at combing through public files at scale to uncover a variety of data previously thought to be absent.

Outcome: The ACR recommends that image data needs to be processed in a way that does not allow reconstruction of potentially recognizable faces, while also balancing against preserving the utility of images.

Skull stripping removes everything outside the cranial cavity while defacing removes or replaces facial structures only. In case of head & neck cancers involving the face, if the estimated re-identification risk exceeds acceptable thresholds, images cannot be released publicly without restrictions. It is uncertain how AI-based techniques might be repurposed for modification of facial anatomy instead of complete removal. It is proposed to "mask" the face by interposing barriers in 3D space between the face & observer. The mask would have to contact the skin surface & have the same pixel-value distribution as the data. This discourages casual re-identification attempts & preserves greater utility than traditional de-facing approaches.

Conclusion/Statement of Impact/Lessons Learned: AI can scrutinize medical images in ways that humans cannot. Models have predicted patient race from skeletal anatomy. Researchers are unable to isolate image features responsible for recognition or mitigate this ability by various efforts. While further research into this domain is needed, human-led efforts to detect racial biases & train models to equalize racial outcomes should be considered for medical safety.

Open science holds immense potential for neuroimaging by fostering collaboration & transparency in the neuroscience community.

Keywords: Clinical Workflow & Productivity; Educational Systems; Organizational & Professional Development

043 - Streamlining Radiology Workflow with an LLM-Powered Prior Report Summarization Tool

Presenter: Julie An, New York University

Julie An1, Daniel Reninnghoff1, Malte Westerhoff2, Gina Ciavarra1, William Moore1, Michael Recht1, Ankur Doshi1

1New York University, New York, NY, USA

2Visage Imaging, GmbH, Berlin, Germany

Background/Problem Being Solved: Comprehensive review of prior radiology reports during image interpretation can be a time-consuming process, especially in patients with complex medical histories and numerous prior studies. Prior reports provide critical context for accurate diagnosis. This challenge underscores the need for innovative solutions that streamline workflows and reduce cognitive load while improving diagnostic accuracy.

Intervention(s): With our PACS vendor, we developed a large language model (LLM) tool that can be integrated within our PACS and uses an institutional-compliant GPT-4 model to generate concise, comprehensive summaries of prior radiology report impressions. These summaries highlight relevantfindings to assist in the interpretation of current exams. The prompt was iteratively refined to enhance the relevance, usability, and accuracy of the output

Barriers/Challenges: LLMs may generate inaccurate information, with potential additions or omissions, which can undermine trust and applicability. Investigating the frequency and impact of such occurrences is necessary to ensure high standards are met for use in clinical care.

Outcome: The prompt was iteratively refined to generate an output that is clear, concise and comprehensive. We have received IRB approval to perform a retrospective analysis that will evaluate the performance of the LLM, and we will present data on:

Overall accuracy.

Presence of significant omissions and additions.

Potential clinical impact of errors.

Perceived clinical utility Workflow efficiency impact.

Conclusion/Statement of Impact/Lessons Learned: We propose a novel tool that has the potential to improve radiologist efficiency and reporting accuracy. Ongoing refinement and user feedback will improve clinical utility, paving the way for integration into routine practice and for improving patient care.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity

044 - Systematic Detection and Correction of DICOM Label Discrepancies in Large-Scale Chest X-ray Datasets

Presenter: Frank Li, Emory University

Frank Li1, Theo Dapamede1, Mohammadreza Chavoshi1, Bardia Khosravi2, Janice Newsome1, Aawez Mansuri1, Rohan Satya Isaac1, Hari Trivedi, Judy Gichoya1

1Emory University, Atlanta, GA, USA

2Yale University, New Haven, CT, USA

Background/Problem Being Solved: Artificial intelligence models require large, high-quality annotated datasets for development. Manual labeling of radiology datasets is expensive and time-consuming; hence most labels are extracted from existing radiology reports, with the most common tool being CheXpert labeler. Beyond labeling, radiology images also require harmonization and de-identification of DICOM metadata and pixel data. Verifying the quality of the resultant curation process remains extremely challenging due to dataset scale and heterogeneity, and manual review is time-prohibitive. We describe our curation process for our CXR dataset (containing > 2 million images) and highlight common pitfalls and solutions.

Intervention(s): Many errors were encountered within DICOM files during data curation that would introduce noise and affect downstream analyses, including 1) mislabeled DICOM tags for ViewPosition (PA, AP, and lateral); 2) images with Contrast Limited Adaptive Histogram Equalization (CLAHE) without differentiation in DICOM tags; and 3) duplicate images with unique SOP Instance UID within the same study.

To correct the view position, we developed an in-house view position deep learning classifier trained on CheXpert that achieved AUC=1.00 on test data. To identify duplicates and CLAHE images, cosine similarity was computed between image pairs within each study. A cosine similarity of 1 indicated identical images. For images with high similarity (but < 1), image noise was quantified using a high pass filter, with the noise signal's histogram fitted to a Poisson distribution to obtain the distribution parameter (λ). For image pairs exceeding the chosen similarity threshold, images with greater noise (or λ) were classified as CLAHE, based on the assumption that CLAHE processing increases image noise.

Barriers/Challenges: The large sample sizes of medical imaging datasets make pairwise comparison of images computationally intensive and time-consuming. Parallel computation techniques can significantly accelerate this process by distributing the workload across multiple processors.

Outcome: Our interventions identified 46,512 (1.87%) lateral images incorrectly labeled as AP view in the DICOM tags, 397,384 (15.98%) CLAHE images, and 1,782 (0.07%) duplicate images among the total 2,486,502 images.

Conclusion/Statement of Impact/Lessons Learned: DICOM tags are inaccurate as labels for dataset curation. Computational approaches leveraging deep learning and image processing techniques like noise quantification can improve curation at scale with minimal human input. Such tools can be incorporated in orchestration engines for routing medical images to AI platforms to improve match rate.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity

045 - Tool for Enhanced 3D Planning and Evaluation of AAA Procedures in Vascular Surgery

Presenter: Nazanin Maleki, Children's Hospital of Philadelphia

David Weiss1, Nazanin Maleki2, Thomas Hager3, Mariam Aboian2, MingDe Lin4, Khaled Bousabarah5, Wolfgang Holler5, Kathryn Simmons3, Johannes Haubold1, Sarah Loh3, Uwe Fischer3, Julius Chapiro3, Cornelius Deuschl1, Sanjay Aneja3, Edouard Aboian3

1University Hospital Essen, Essen, Germany

2Children's Hospital of Philadelphia, Philadelphia, PA, USA

3Yale University, New Haven, CT, USA

4Visage Imaging, Inc., San Diego, CA, USA

5Visage Imaging, GmbH, Berlin, Germany

Background/Problem Being Solved: Volumetric assessment of abdominal aortic aneurysms (AAA) offers precise pre- and post-endovascular aortic repair (EVAR) evaluation but is laborious. This study aimed to train and validate deep learning-based network facilitating automated segmentation and volume determination of pre- and post-EVAR infrarenal AAAs displayed on computed tomography angiographies (CTA) and to evaluate its role in accelerating clinical workflow.

Intervention(s): A HIPAA-compliant study was performed investigating de-identified pre- and post-interventional CTAs of patients who underwent EVAR for management of infrarenal AAA at our institution. In research instance of our PACS (AI Accelerator, AIA, Visage Imaging, Inc.), ground truth volumetric segmentations of total aneurysm and lumen were performed from lowest renal artery to aortic bifurcation. nnU-Net model was trained and validated on this dataset. External validation was performed using multi-institutional datasets. Efficiency gains provided by model were tested against two attending vascular surgeons and one vascular surgery resident who performed semi-automatic AAA segmentation on both internal and external validation datasets using AIA. Baseline patient demographics were recorded.

Barriers/Challenges: Time consuming Manual adjustments.

Ensuring the external datasets are sufficiently large, diverse.

Ensuring that the trained nnU-Net model performs reliably across multi-institutional datasets with different imaging protocols and scanner settings.

Outcome: A total of 110 patients with 84 (76.4%) males were included in internal dataset. Training and internal validation datasets comprised 176 and 44 pre- and post-EVAR CTAs; 60 validation studies from external institutions were included. For total aneurysm, mean Dice similarity coefficient was 0.972±0.013 and 0.960±0.035 in internal and external validation. AI-generated thrombus volumes showed a very strong correlation with ground truth in internal (r=0.996) and external validation (r=0.940). Mean algorithm-facilitated time savings of 117.1 seconds (56.0%) were demonstrated for total aneurysm.

Conclusion/Statement of Impact/Lessons Learned: Our PACS-based institution-agnostic network enables automated volumetric AAA analysis. Integration of an aorta segmentation algorithm into the model, including a manually adjustable bounding box for precise field-of-view selection, is being evaluated; the tool could be incorporated into routine clinical practice.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Educational Systems; Emerging Technologies; Quality Improvement & Quality Assurance

046 - Untapped Value at the Viewbox: Generating AI-Driven Insights from Radiology Readouts

Presenter: Karan Jani, Washington University in Saint Louis

Karan Jani1, Shinjini Kundu1, Martin Reis1, Sanjeev Bhalla1, Vamsi Narra1, David Ballard1, Anup Shetty1

1Washington University in Saint Louis, Saint Louis, MO, USA

Background/Problem Being Solved: Radiology readouts are a fixture of academic radiology practice and radiology education due to their information-rich nature; they enable attendings to convey didactic and experiential knowledge to trainees, covering case approach and impression synthesis, imaging protocoling and appropriateness criteria, and follow-up recommendations for various findings. However, analysis and utilization of radiology readout content remains underexplored in literature despite the wealth of information contained in each interaction.

Intervention(s): To capture information conveyed in radiology readouts, we developed an artificial intelligence (AI) system that extracts transcripts, case breakdowns, and teaching points from audio-recordings of readouts in real-time. Our fully automated pipeline uses an on-premises automatic speech recognition (ASR) model and cloud-based HIPAA-compliant large language models (LLM). A front-end application records audio and displays structured readout summaries to guide resident report generation and supplement learning.

Barriers/Challenges: Transcription accuracy of ASR models relies on audio recording quality, which varies with ambient conditions (e.g. ringing phones, other conversations). Accents and biomedical terminology may reduce ASR accuracy. Additionally, variability in attending teaching style may affect the quality of extracted insights.

Outcome: Our AI system extracts transcripts, case breakdowns, and teaching points from audio-recordings of readouts in real-time. Providing these automated outputs to residents improves engagement during readout, reducing typing time and increasing shared case-viewing time. Case breakdowns guide resident report generation, potentially minimizing attending revisions and improving workflows. Aggregated structured readout summaries can serve as curricular supplements to residents, present alternative teaching styles to educators, and reveal common knowledge gaps amongst trainees, highlighting areas for curricular improvement.

Conclusion/Statement of Impact/Lessons Learned: Our AI system leverages the underexplored content of radiology readouts to generate structured, case-based insights that enhance trainee engagement and learning, while providing educators opportunities to refine teaching style and curriculum. Importantly, this work explores a new class of AI tools for academic radiology practices, impacting both clinical workflows and educational missions.

Keywords: Administration & Operations; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Educational Systems; Emerging Technologies; Quality Improvement & Quality Assurance

047 - Using AI Agents to Generate Recommendations Based on ACR Whitepapers

Presenter: Paulo Kuriki, University of Texas Southwestern

Paulo Kuriki1, Yin Xi1, Ali Tejani1, Fernando Kay1, Krishna Sai Balineni1, Yee Ng1

1University of Texas Southwestern, Dallas, TX, USA

Background/Problem Being Solved: The American College of Radiology (ACR) whitepapers provide evidence-based guidelines to aid radiologists in recommending actions for incidental findings, such as adrenal, liver, and pulmonary nodules. However, consulting these documents during routine practice is time-consuming and labor-intensive. While models like GPT-4 may theoretically incorporate this knowledge, their use introduces concerns related to privacy, hallucinations, and errors in extracting accurate information. This study investigates the application of AI agents to automate the extraction of adrenal nodule features and generate recommendations based on ACR criteria.

Intervention(s): An agent-driven recommendation system was developed using Ollama and Llama 3.3 70B. A total of 769 Abdomen MRI reports were extracted. Using a separate dataset, prompt engineering was performed to optimize the extraction of structured JSON responses containing features to generate ACR-based recommendations. A Python function was developed to receive a JSON input and return an ACR-based recommendation. This function was integrated into Ollama as a tool, enabling it to generate recommendations upon request. The LLM was instructed to identify and extract adrenal nodule features and invoke the tool function to create structured recommendations. Results were manually reviewed by a 6-year experienced board-certified radiologist.

Barriers/Challenges: This pilot study was limited to reports from only two institutions. The use of a 70B model may pose computational challenges, while smaller models might achieve comparable results. Advanced prompt engineering techniques could further improve performance. Future efforts will focus on testing and fine-tuning smaller models to reduce computational demands.

Outcome: From 769 reports, 98 recommendations for adrenal nodules were generated. Of these, 79 (80%) were classified as correct. Errors in the remaining 19 cases were attributed to failures in correctly extracting features, especially fat content, stability, or cancer history, when not explicitly described. Additional failures arose from issues in the tool calling process or incorrect function classification. Users anticipate substantial time savings by avoiding consultation of whitepapers, with radiologists unfamiliar with abdominal imaging benefiting most from ACR-aligned recommendations. These results demonstrate the potential for integrating Agentic AI into existing report generation workflows.

Conclusion/Statement of Impact/Lessons Learned: AI agents can significantly improve the generation of systematic, consistent, and evidence-based recommendations in radiology reporting. By leveraging external APIs, databases, and decision-making tools, these agents augment the capabilities of LLMs, improving accuracy and reproducibility. This study highlights the significant potential of agentic AI models to support more efficient and precise report generation, offering promising opportunities for integration into clinical workflows.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Emerging Technologies; Patient/Family Experience

048 - Utility of CDE Definitions and FHIR Structure for Representing Chest CT Findings

Presenter: Roshan Fahimi, Massachusetts General Hospital

Roshan Fahimi1, Tarik Alkasab1, Michael Hood1, Heather Chase2

1Massachusetts General Hospital, Boston, MA, USA

2Microsoft, Boston, MA, USA

Background/Problem Being Solved: Accurate reporting of chest CT scans is critical for optimal patient care. The American College of Radiology/Radiologic Society of North America Common Data Elements (CDEs) project provides an organized ontology for radiology reports, facilitating clinical workflows by structuring diagnoses, anatomical locations, and other relevant details. Fast Healthcare Interoperability Resources (FHIR) offers a standard for transmitting and integrating this information with other patient data. Our study explores the potential of combining CDE and FHIR standards to enhance the exchange of structured radiology data.

Intervention(s): We collected a sample of outpatient chest CT reports using a specialized radiology report search engine (mpower Nuance). Findings from each report were manually identified and recorded with attributes such as measurements, anatomical location, disease severity, size, and volume. Each finding was associated with a CDE definition from the RadElement repository and a staging repository of proposed definitions. We documented whether findings could be encoded with a CDE definition and identified attributes not covered by existing CDEs. Negative findings, such as "clear lungs," were also noted.

Barriers/Challenges: The primary challenge was the incomplete coverage of existing CDE definitions, requiring the use of proposed definitions to encode a broader range of findings. Anatomy-based negative statements posed another challenge, as they are not consistently represented in current CDEs.

Outcome: In 82 chest CT reports, we identified 1,190 findings (average 14±3 per report). Using published and proposed CDE definitions, 83.2% of findings were encoded as FHIR Observations. This included common radiology findings such as pulmonary nodules and pleural effusions. Anatomy-based negative statements occurred at a rate of approximately six per report.

Conclusion/Statement of Impact/Lessons Learned: CDE-labeled FHIR Observations can encode most findings in chest CT reports, improving compatibility and data exchange across healthcare systems. As the CDE library expands, more comprehensive representation of radiology results will be feasible. The study also highlights the need for standard methods to represent anatomy-based negative statements.

Keywords: Applications; Clinical Workflow & Productivity; Enterprise Imaging; Standards & Interoperability; Systems Management

049 - Where Imaging and Clinical Informatics Meet: Software Development Life Cycle Methodologies & Terminology Review: Waterfall, Spiral, and Agile Approaches Applied to Use Cases

Presenter: Les Folio, James A. Haley Veterans Hospital

Gregg Cohen1, Nathan Bumbarger2, Kenda Fessock3, Les Folio4

1National Institute of Health, Bethesda, MD, USA

2United States Air Force, San Antonio, TX, USA

3Nemours Children's Health, Orlando, FL, USA

4James A. Haley Veterans Hospital, Tampa, FL, USA

Background/Problem Being Solved: The Certified Imaging Informatics Professional (CIIP) certification provides a foundational understanding of clinical informatics (CI), but practical applications often involve complex software development life cycle (SDLC) methodologies. Imaging informatics professionals may encounter terms like Waterfall, Extreme Programming (XP), scrum, spiral, or agile methods, which are not deeply covered in the CIIP curriculum, leading to challenges in selecting the appropriate approach.

Intervention(s): This article introduces SDLC basics to burgeoning imaging informatics professionals, focusing on methodologies such as Waterfall, scrum, spiral, and agile. Key terminologies like Product Owner, Scrum Master, Sprint, and Daily Scrum are explained. Use cases from various institutions highlight the importance of engaging radiology and clinical stakeholders early in the process to ensure successful product development or procurement.

Barriers/Challenges: A critical barrier is poorly defined user requirements at project onset, leading to costly "scope creep" and resource wastage. Success depends on clearly defined project scopes and active involvement of all stakeholders, including PACS and EI administrators, to prevent costly delays and misaligned project goals. Additionally, a lack of familiarity with SDLC methodologies and terminology can negatively impact project outcomes.

Outcome: The article offers a comparative overview of software development methodologies, supplemented by real-world use cases. For example, one center’s agile training enabled rapid development with participants assuming scrum master roles, while another leveraged grassroots development to integrate protocoling into a RIS vendor system. These experiences illustrate practical applications of SDLC approaches in imaging informatics.

Conclusion/Statement of Impact/Lessons Learned: Providing SDLC terminology and reviewing methodologies equips imaging informatics professionals with critical baseline knowledge. Sharing experiences from diverse institutions highlights practical challenges and strategies, preparing CIIPs to navigate the complexities of clinical informatics from their first day on the job.

Keywords: Applications; Artificial Intelligence/Machine Learning; Emerging Technologies

050 - Zero-shot Learning with RAG to Characterize Brain Radiation Necrosis from Aggregate Clinical Text

Presenter: Ricky Savjani, University of California Los Angeles

Ricky Savjani1, William Delery1, Ricky Savjani1, Eulanca Liu1, Tania Kaprealian1, Won Kim1

1University of California Los Angeles, Los Angeles, CA, USA

Background/Problem Being Solved: Radiation Therapy (RT) is often used to treat brain metastases (BM); however, Radiation necrosis (RN) is a side effect of radiotherapy in which surrounding healthy brain tissue becomes inflamed in around 5–25% of metastatic intracranial lesions. RN is difficult to diagnose and manage because it is often indistinguishable from tumor progression on imaging and has variable symptomatic rates. Treating physicians must manage patients with RN with limited, disorganized data. Here, we sought to use LLMs to aggregate all relevant Electronic Health Record (EHR) data to characterize radiation necrosis in over 1000 patients with BM treated with radiotherapy.

Intervention(s): We conducted SQL queries to identify all BM patients who underwent RT at our institution from 3/1/2013 to 10/22/2023. All Imaging reports, pathology reports, clinical notes, medications, problem lists, and operative notes were extracted into data frames and collated for each patient. We prompted Meta’s Llama 3.3 70B parameter model to identify which patients developed RN, their clinical progression, treatment history, and response to RN interventions on a large scale. We used zero-shot learning with Retrieval-Augmented Generation (RAG).

Barriers/Challenges: Although Llama3.3 is state-of-the-art at the time of this abstract writing, the overall window length is still limited to 128K tokens (~96,000 words).

Outcome: The LLM correctly identified the patient’s first occurrence of RN on imaging, medication administration, BM treatments, and the overall clinical response to RN. We are now validating these results through expert chart review.

Conclusion/Statement of Impact/Lessons Learned: The aggregation and summarization of patient data using clinical informatics with LLM integration effectively elucidates treatment history and outcomes tailored to each patient uniquely. By identifying a large cohort of patients, we are now characterizing risk factors for radiation necrosis that have been difficult to ascertain with tedious manual review alone.

Keywords: Applications; Artificial Intelligence/Machine Learning

Scientific Research Abstracts

051 - A Combined Deployable End-to-end Automated AI Detection and Quantitative Visualization Pipeline for Hemothorax on Admission Trauma CT

Presenter: Mehrdad Salimitari, University of Maryland

Ankush Jindal1, Mehrdad Salimitari1, Uttam Bodanapally1, Lei Zhang1, Wayne LaBelle1, Guang Li1, Bhavya Reddy1, Ozerk Turan1, Gabrielle Dickerson1, Abigail Corkum1, Melike Harfouche1, David Dreizin1

1University of Maryland, Baltimore, MD, USA

Introduction: Internal hemorrhage in the non-compressible torso is the leading reversible traumatic cause of death. Hemothorax (HTX) benefits from prompt diagnosis and quantitative visualization (QV). Volumes correlate with hemorrhage related outcomes. SOC methods rely on subjective radiologist assessment and WBCT interpretation times increase with injury severity. We address the unmet need for a DICOM/PACS-interoperable automated detection and precision diagnostics tool.

Hypothesis: Our approach will have high accuracy metrics and DSC in highly imbalanced in-the-wild data.

Methods: 8,157 CT consecutive cases from our trauma center included 838 HTX-positive and 7,319 HTX-negative cases. The dataset comprised 2,468,327 slices (64,570 w/ HTX). All cases underwent rigorous voxelwise human-in-loop labeling with multi-reader arbitration and systematic bias mitigation protocols. Our two-stage approach consisted of: 1) slice-level detection (SException) and focal loss for imbalanced data, with Embedding-Vision Transformer (E-VIT) for patient-level aggregation and 2) MedNeXt for QV. The model was trained on 8 H100 GPUs. Temporal validation set included 3,194 studies.

Results: We achieved 0.98 AUROC, 86.3% sensitivity, 91.2% specificity, 97.8% NPV, 59.9% precision, 70.8% F1-score and 0.79 DSC, with high HTX saliency. Lower performance was observed for small, less clinically significant HTX. Mean inference times were 27.6 seconds for detection, with preprocessing and feature extraction for E-VIT as the main computational bottleneck, and 81.2 seconds for segmentation.

Conclusion: Our method demonstrates high accuracy and overlap metrics in a temporal validation set with in-the-wild distribution. The approach provides the first quantitative, explainable detection/QV tool for HTX, aligning with established quantitative laboratory and vital sign standards in surgical care. The system's enterprise-ready orchestrator enables seamless clinical integration through low-latency pop-ups and IM alerts, DICOM SEG visualization, and automated volumetric reporting. Future work will focus on multicenter validation, bias analysis, and correlation with clinical outcomes.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Imaging Research; Quality Improvement & Quality Assurance

052 - AI-powered Extraction of Relevant Information from Prior Radiology Reports

Presenter: Yiting Xie, Merge by Merative

Yiting Xie1, Linda Bagley1, Marwan Sati1

1Merge by Merative, Chicago, IL, USA

Introduction: With increasing image volume and staffing shortages, radiologists have less time to find and adequately read patient history. This represents a significant risk to patients’ health and increases radiologist’s liability. Automatically finding relevant history is extremely challenging because it requires deep medical knowledge to identify terms that are not semantically similar. For example, there are over 100 conditions that could be related to a chief complaint of “abdominal pain”. Lung nodule mentions may be relevant for knee surgery prep due to their potential impact on surgery risks.

Hypothesis: Relevant information can be automatically extracted from radiology reports using machine learning (ML) and large language models (LLMs).

Methods: We algorithmically determined relevant priors based on the patient’s “Reason for Exam” in the imaging order. The framework uses a weighted combination of ChatGPT 4o and o1 LLMs and our proprietary ML model trained on extensive medical data from sources including PubMed and Wikipedia. Intelligent prompt engineering was employed to enhance LLM performance. The framework extracts relevant sentences and associated probabilities, with the highest probability sentence indicating the overall relevancy of each report.

We evaluated the framework on 1366 radiology reports from 300 patients using binary ground truth provided by 3 radiologists (single reader per report). Report level probability was compared to the ground truth. Performance of the framework was compared against a baseline model using ChatGPT 4o alone with a simple prompt.

Results: Out of the box ChatGPT 4o model achieved accuracy of only 74%. The combined framework achieved an AUC of 0.87 and an optimal accuracy was 85%.

Conclusion: This complex problem cannot be solved with the latest out-of-the-box LLMs however combining state of the art LLMs with medical-specific ML is promising for this critical yet very challenging problem.

Keywords: Artificial Intelligence/Machine Learning

053 - Applying Deep Learning Methods for Estimating the Volume of Pathological Regions in AD Brain Tissue

Presenter: Yujia Wei, Mayo Clinic

Yujia Wei1, Jason Cai1, Bradley Erickson1

1Mayo Clinic, Rochester, MN, USA

Introduction: Brain substructure volumes are valuable biomarkers for tracking Alzheimer's Disease (AD) progression and aiding diagnosis. Traditional manual MRI annotation is slow and inconsistent. AI-based semantic segmentation offers improved accuracy, speed, and reliability for structural analysis and volumetric assessment. However, current models are limited by the number of structures they can identify, and lesion volume data for AD patients is scarce. In this study, we attempt to develop a deep learning model to segment all gray matter structures on T1-weighted MRIs of AD patients, calculate the volumes of all brain regions, and compare the data between AD patients and age-matched healthy controls.

Hypothesis: We hypothesize that Umamba can be utilized for AI model development to segment all brain regions and identify potential features in T1-weighted MRIs.

Methods: A pre-trained model developed by the research group was applied for this study. The proposed AI model, Umamba-Bot, has been shown to achieve superior performance compared to other popular models. From the ADNI database, 134 age-matched elderly Cognitively Normal(CN) and 112 AD patients were selected. Using the trained model, the volumes of all 122 brain regions were calculated and analyzed statistically.

Results: Brain sub-structure volume calculations were performed, and the resultant volumes of normal controls and AD patients were compared. The volumes and volume changes of sixteen brain substructures all consistent with those reported in existing literature, reinforcing the reliability of the segmentation outputs.

Conclusion: We successfully segmented gray matter regions and accurately calculated the volumes of all brain substructures. Comparisons with age-matched normal controls identified abnormal brain regions in AD patients, consistent with findings in the existing literature. This demonstrates the potential of the proposed approach for clinical applications, including feature extraction, morphological analysis, and as a foundation for downstream diagnostic tools.

Keywords: Applications; Artificial Intelligence/Machine Learning; Imaging Research

054 - Assessing the Performance of ChatGPT-4 Omni in Classifying Fracture and Non-Fracture X-ray Images

Presenter: Nitin Chetla, University of Virginia

Swapna Vaja1, Nitin Chetla2, Alper Turgut3, Tejas Sekhar1, Tamer Hage4, Joseph Chang5, Mihir Tandon6, Jorge Chahla1, Arun Krishnaraj2

1Rush University, Chicago, IL, USA

2University of Virginia, Charlottesville, VA, USA

3University of Pittsburgh, Pittsburgh, PA, USA

4Virginia Tech University, Blacksburg, VA, USA

5University of Passau, Passau, Germany

6Albany Medical College, Albany, NY, USA

Introduction: Radiographic assessment is the standard of care for diagnosing suspected fractures. While LLMs like ChatGPT-4 Omni (ChatGPT-4o) show promise in augmenting radiographic workflows, supporting evidence is limited. This study aims to evaluate ChatGPT-4o’s performance in identifying various fracture types from the FracAtlas database.

Hypothesis: ChatGPT-4 Omni (ChatGPT-4o) will demonstrate high sensitivity but low specificity in classifying X-ray images as fracture or non-fracture, and its diagnostic accuracy will significantly vary depending on the prompt structure and wording used during the classification process.

Methods: The experiment evaluated ChatGPT-4 Omni (ChatGPT-4o) by classifying 1,000 X-ray images from the FracAtlas database (500 fracture, 500 non-fracture) of various body parts (hand, leg, hip, shoulder). Using a Python-based recursive loop, four prompts were tested: the first two (Test 1) asked if the image showed a fracture, with answer options reversed to test order effects. The other two prompts (Test 2) used different wording to test variations in response. Accuracy, precision, sensitivity, specificity, and F1 scores were calculated. Indecisive, incomplete, or refused responses were excluded from the analysis.

Results: The study found that ChatGPT-4 Omni (ChatGPT-4o) tends to over-identify fractures, with low to moderate specificity but high sensitivity. Changing the order of answer choices in prompts led to significant differences in accuracy and specificity. Prompt 1 achieved 0.672 accuracy and 0.547 specificity, while Prompt 2's accuracy dropped to 0.591, with specificity falling to 0.234. These results suggest ChatGPT-4o's outputs are dependent on prompt structure. Test 2 demonstrates similar differences across metrics of sensitivity, specificity, and F1 score.

Conclusion: Our study demonstrated statistically significant differences in intra- and inter-test accuracy, precision, sensitivity, specificity, and F1 score, suggesting that prompt outputs are dependent on user prompt input. Given ChatGPT-4o’s generally moderate to high sensitivity but low specificity, the tool poses a significant risk of false positives, with more serviceable use cases being highly dependent on prompt input.

Keywords: Artificial Intelligence/Machine Learning

055 - Augmented and Virtual Reality Integration with Digital Twin Technology

Presenter: Les Folio, James A. Haley Veterans Hospital

Sara Aghamiri1, João Santinha2, Rada Amin1, Pedro Gouveia2, Les Folio3

1University of Nebraska-Lincoln, Lincoln, NE, USA

2University of Lisbon, Lisbon, Portugal

3James A. Haley Veterans Hospital, Tampa, FL, USA

Introduction: A digital twin (DT) is a computational model visually representing a specific physical object, system, or process. While the DT provides virtual representations of real-world systems, Augmented, Mixed, and Virtual Reality (AR/MR/VR) provides an immersive experience.

In recent years, DT-AR/MR/VR integrations in healthcare have gained significant traction due to their potential to revolutionize training attending and resident physicians.

Hypothesis: DT-AR/VR integration creates a robust medical training platform for healthcare professionals, offering a more realistic and engaging learning environment, planning and image-guided-interventions.

Methods: This study reviews trends and applications of integrating cutting-edge technology of digital twins with augmented and virtual reality to discuss the opportunities, implementation, and challenges in medicine.

AI advancements in medical image processing and quantitative analysis are associated with augmented and virtual reality (AR/VR) head-mounted displays that allow enriched visualization and manipulation of these digital twins in various clinical scenarios.

Results: Medical imaging modalities, especially cross-sectional modalities (e.g., MRI, CT, PET/CT), provide high-resolution anatomical data essential for building virtual replicas that are essentially DTs.

DT-AR/VR integration also enhances the training of attending and resident physicians by enabling virtual environments to practice various medical procedures, improving their skills and decision-making. As an example, Surgical Planning and Intervention Guidance. DTs created from medical imaging data enable surgeons to virtually plan and simulate complex procedures, explore various approaches, identify potential risks, and optimize surgical strategies for improved outcomes. Additionally, these DTs can be overlayed on patients, enabling their use as a guide for interventions.

Challenges inherent to DT-AR/VR integration that must be addressed include:

Data Acquisition and Integration: Integrating real-time data from various sources, including diverse medical imaging modalities.

Procurement of sophisticated technology is still in its infancy, and convincing leadership to acquire is difficult.

Data Privacy and Security: Protecting sensitive patient information derived from medical imaging and complying with regulations is paramount.

Conclusion: DT-AR/VR integrations grounded in rich medical imaging data hold immense promise for advancing healthcare. The future trajectory of this technology relies on further integration of AI for advanced image analysis and predictive modeling, ultimately enabling real-time DT refinement for continuous monitoring, personalized interventions, and more accurate predictions of disease progression and treatment response. Addressing the challenges through collaborative efforts and robust frameworks will be crucial to fully realizing the transformative potential of DT-AR/VR in healthcare.

Keywords: Applications; Artificial Intelligence/Machine Learning; Emerging Technologies

056 - Automated Vestibular Schwannoma Detection Using YOLO-Based Models

Presenter: Sahika Betul Yayli, Mayo Clinic

Sahika Betul Yayli1, Parv Mehta2, Daniel Blezek1, Matthew Carlson1, Neetu Soni1, Milan Sonka3, Bradley Erickson1, Girish Bathla1

1Mayo Clinic, Rochester, MN, USA

2University of Texas, San Antonio, TX, USA

3University of Iowa, Iowa City, IA, USA

Introduction: Accurate localization of vestibular schwannomas prior to segmentation is critical for improving automated analysis. YOLO-based object detection models may offer robust, rapid tumor localization, streamlining subsequent segmentation tasks.

Hypothesis: We hypothesize that YOLO-based object detection models can reliably identify regions of interest encompassing vestibular schwannomas, ensuring complete tumor inclusion within a single 3D-bounding box.

Methods: T1-weighted contrast-enhanced MRI slices were preprocessed into 2D slices with normalized intensities, and bounding box annotations were generated from ground truth masks. YOLOv8 and YOLOv10 models, including lightweight (s) and medium (m) variants, were trained and evaluated on internal and external datasets for precision, recall, F1-score, and complete tumor containment. Complete tumor containment was assessed by calculating the percentage of tumor volume encapsulated within a 7 cm3 bounding box, derived by taking the 3D center of the detected bounding boxes on 2D slices.

Results: Both YOLOv8 and YOLOv10 achieved high detection accuracy, with F1-scores exceeding 0.93 and 100% complete tumor containment. YOLOv10m demonstrated slightly superior performance, with an Intersection-over-Union above 0.85 on both internal and external datasets. This reliable ROI determination enables consistent cropping for subsequent segmentation models, ensuring stable and improved segmentation results.

Conclusion: YOLO-based models are effective tools for vestibular schwannoma detection, guaranteeing that the entire tumor is contained within the predicted region. This automated localization step improves data preprocessing quality and forms a robust foundation for downstream segmentation tasks.

Keywords: Artificial Intelligence/Machine Learning

057 - Automating Linear Measurements of the Distal Ascending Aorta in CT Angiograms Using AI-based Heatmap Regression to Identify Oblique Planes

Presenter: Andrew Missert, Mayo Clinic

Andrew Missert1, Adam Dachowicz1, Jason Klug1, William Ryan1, Gian Marco Conte1, Wolfgang Holler3, MingDe Lin2, Alex Bratt1

1Mayo Clinic, Rochester, MN, USA

2Yale University, New Haven, CT, USA

3Visage Imaging, GmbH, Berlin, Germany

Introduction: Linear aorta measurements from CT angiograms are crucial indicators of cardiovascular disease. The current clinical standard involves manual annotations of doubly oblique multiplanar reformatted images, which can be challenging, time-consuming, and subject to substantial inter- and intra-reader variability. We propose automating this process through AI-based methods, applying them to the distal ascending aorta just proximal to the brachiocephalic trunk.

Hypothesis: An AI model optimized for heatmap regression can accurately identify the doubly oblique measurement plane. Combined with aortic semantic segmentation, this approach will enable automated linear measurements with accuracy in the range of inter-reader variability.

Methods: In this prospective study, 579 linear annotations of the distal ascending aorta were performed by 47 board-certified radiologists using a semantic annotation engine integrated into the clinical PACS workflow. We applied a 3D aorta segmentation model (nnU-Net) to each CTA volume and computed the centerline. A heatmap annotation was generated by convolving a Gaussian kernel (sigma=1.0) with the 2D surface defined by the intersection of the 3D aorta mask and the plane perpendicular to the centerline closest to the annotation. A regressive heatmap model was then trained using 148,224 3D image patches from the training set. Testing was performed on 29 reserved volumes. Automated measurements were obtained at the centerline location maximizing the heatmap and compared to manual annotations using Bland–Altman analysis.

Results: The automated measurements closely matched the manual results, with a mean difference of 1.35 mm ± 1.22 mm, which falls within reported inter-reader variability (4.7 mm).

Conclusion: Automated linear measurement of the distal ascending aorta just proximal to the brachiocephalic trunk is feasible using an AI-driven regressive heatmap approach, achieving accuracy on par with manual expert measurements.

Keywords: Artificial Intelligence/Machine Learning; Imaging Research

058 - Automating Pennsylvania Act 112 Follow-Up Recommendation Detection in Radiology Reports Using Large Language Models

Presenter: Satvik Tripathi, University of Pennsylvania

Shawn Lyo1, Satvik Tripathi1, Tessa Cook1

1University of Pennsylvania, Philadelphia, PA, USA

Introduction: Pennsylvania Act 112 requires patient notification when imaging reveals a finding that requires follow-up imaging within 90 days. While structured reporting macros exist to flag these cases, their inconsistent usage by radiologists can lead to missed notifications. This creates a potential compliance gap. Large language models (LLMs) have demonstrated strong capabilities in understanding medical text and context, suggesting they could reliably identify reports that meet Act 112 notification criteria, regardless of whether the macro was used. However, their effectiveness in this specific regulatory compliance use case has not been systematically evaluated.

Hypothesis: A large language model can accurately identify radiology reports that meet Act 112 notification criteria, independent of macro usage.

Methods: We collected 1,000 abdominal imaging reports from our radiology information system: 500 reports with documented follow-up recommendations within 90 days and 500 with either no follow-up or recommendations beyond 90 days, based on structured macro documentation. We developed a system using Azure OpenAI GPT-4, ensemble prompting and universal self-consistency techniques, to analyze report text, which was stripped of the structured macros, and classify whether each case met Act 112 notification criteria. The model's classifications were compared against the ground truth established by the structured macro documentation.

Results: The LLM achieved an F1-score of 0.72, driven by a relatively higher precision (83%) compared to recall (64%). Notably, the model's follow-up recommendation rate did not vary significantly based on the actual follow-up intervals specified in the macro.

Conclusion: Our findings demonstrate that GPT-4-based large language models can effectively identify radiology reports requiring Act 112 follow-up notification, achieving high precision. However, discrepancies in recall suggest opportunities for refinement. Automating this process with LLMs could enhance compliance with local policies. Future efforts will focus on optimizing sensitivity and validating the approach across other subspecialties and report types.

Keywords: Administration & Operations; Applications; Artificial Intelligence/Machine Learning; Emerging Technologies; Quality Improvement & Quality Assurance

059 - Capability of Multi-Modal Large Language Models for Matching Findings in Longitudinal CT Studies

Presenter: Tejas Sudharshan Mathai, National Institutes of Health

Tejas Sudharshan Mathai1, Boah Kim1, Praveen Thoppey Srinivasan Balamuralikrishna1, Ronald Summers1

1National Institute of Health, Bethesda, MD, USA

Introduction: Radiologists routinely compare findings between the prior and follow-up CT exams, and then assess interval changes. However, this task is currently manually performed, and it can become cumbersome when comparing multiple time points.

Hypothesis: To evaluate the capability of Multi-Modal Large Language Models (MLLM) for matching findings between two longitudinal exams (prior vs. follow-up) using report sentences and CT images.

Methods: In this retrospective study, the public CT-RATE dataset containing longitudinal non-contrast chest CT studies was used. CT volumes and reports from the prior and follow-up visits of 67 patients were included. Findings in the reports (e.g., nodules, pleural/pericardial effusion) were automatically extracted, and the slice in the CT volume containing the respective finding was manually identified. Given a finding and CT image from the follow-up study, the MLLM identified the matched finding in the prior study. Two MLLMs (GPT-4o and Gemini-1.5-Pro) were evaluated, and the use of report text alone (i.e., Gemini-R) was compared against the combined use of both images and text (i.e., Gemini-C). Agreement with a rater was measured using Cohen’s κ.

Results: Longitudinal CT studies and reports from 67 patients (M/F ratio: 44/23, ages: 24 - 89 years, 134 CT volumes, 134 reports) were used. Gemini-R obtained the best results with 98.7% precision, 98.4% specificity, 79.6% sensitivity, with substantial agreement (κ = 0.75) with the rater. GPT-4o-C achieved 92.7% precision, 90.8% specificity, 82.6% sensitivity with substantial agreement (κ = 0.72). No significant differences were observed (p> .05) between GPT-4o-C vs. GPT-4o-R, GPT-4o-C vs. Gemini-C, GPT-4o-R vs. Gemini-R, respectively. However, there was a significant difference between Gemini-R and Gemini-C (p= .03).

Conclusion: In this pilot study, the Gemini-R MLLM (using report text alone) matched findings across longitudinal CT studies and showed potential for interval change assessment.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Emerging Technologies; Enterprise Imaging; Imaging Research

060 - Chest Radiography and Mammography-Derived Imaging Biomarkers Enhance Cardiovascular Risk Prediction Beyond ASCVD Risk Score

Presenter: Theo Dapamede, Emory University

Theo Dapamede1, Frank Li1, Bardia Khosravi2, Aisha Urooj3, Mohammadreza Chavoshi1, Chad Robichaux1, Aawez Mansuri1, Beatrice Brown-Mulry1, Rohan Isaac1, Chadi Ayoub3, Charles O'Neill1, Imon Banerjee3, Judy Gichoya1, Hari Trivedi1

1Emory University, Atlanta, GA, USA

2Yale University, New Haven, CT, USA

3Mayo Clinic, Rochester, MN, USA

Introduction: Cardiovascular disease (CVD) remains underdiagnosed in women in the United States and current risk prediction models underperform in female populations. Emerging deep learning techniques can quantify both breast arterial calcification (BAC) in mammograms and imperceptible features in chest radiographs (CXR) which presents a new possibility for enhanced opportunistic screening for CVD risk prediction. In this study, we evaluate CXR embeddings and BAC quantification as imaging biomarkers to enhance CVD risk prediction

Hypothesis: Automated imaging biomarkers derived independently from routine chest radiographs and mammograms may enhance traditional clinical risk scores for cardiovascular risk stratification.

Methods: We identified a cohort of women who underwent both screening mammography and CXR (frontal and lateral views) (N&#3f9,552) at our institution. The 10-year ASCVD risk score was calculated for patients using electronic health record data. We applied an in-house BAC quantification model to mammograms to extract BAC scores and utilized RAD-DINO to extract embeddings from CXRs. Two XGBoost models were trained on the resultant CXR embeddings to predict 10-year CVD event risk. We then developed a multivariable Cox proportional hazard model incorporating ASCVD score, BAC score, and both CXR-derived risk scores (frontal and lateral) as independent covariates to predict 10-year CVD risk and compare their hazard ratios (HR). Subgroup analysis was performed for patients aged 40–60 years for whom early detection may be most beneficial.

Results: Both BAC and CXR embeddings demonstrated higher HR (BAC: 1.14 [95%CI 1.01–1.28]; CXR Frontal: 4.91 [2.82–8.55]; CXR Lateral: 2.37 [1.27–4.39]) compared to ASCVD score (1.02 [1.01–1.03]), all p < 0.05. Kaplan-Meier analysis revealed significantly higher event rates in women with moderate-to-severe BAC versus zero-to-mild BAC (p < 0.005). Similarly, CXR model-positive patients showed increased CVD event rates (p < 0.005). While frontal CXR demonstrated higher HR and time-dependent AUCs compared to lateral views (0.68 [0.64–0.71] vs 0.65 [0.62–0.69], respectively), these differences were not significant, and both views independently predicted CVD events, suggesting that either view could be utilized for CVD risk assessment. Our findings were consistent across patients aged 40–60 years, suggesting strong potential for early risk stratification in younger women.

Conclusion: Automated analysis of mammograms and chest radiographs enhances cardiovascular risk prediction in women beyond ASCVD scores, enabling early intervention through existing screening programs.

Keywords: Applications; Artificial Intelligence/Machine Learning; Imaging Research

061 - Cohort Selection of MIDRC Exams Using the LOINC RSNA Radiology Playbook

Presenter: Paul Kinahan, University of Washington

Paul Kinahan1

1University of Washington, Seattle, WA, USA

Introduction: To develop artificial intelligence (AI) methods with improved repeatability and reproducibility, a recognized need is for large, publicly available, and curated imaging data sets. Notable public repositories of DICOM images exist, including the recent NIBIB-supported Medical Image and Data Resource Center (MIDRC). MIDRC currently has over 500,000 studies (exams) in process, of which over 180,000 have been processed and publicly available. In addition 20% of studies are sequestered for testing and validation studies. A challenge for users is to select appropriate cohorts using the highly variable Study Descriptions in the DICOM metadata supplied by the providing imaging centers.

Hypothesis: We hypothesized that using a restricted subset of the Logical Observation Identifiers Names and Codes (LOINC) RSNA radiology playbook could be used for efficient cohort selection from the MIDRC collection. This would make use of the algorithmically-generated and unique LOINC Long Common Name as an adjunct to the DICOM Study Description.

Methods: We used a restricted set, called the 'MIDRC-LOINC Mapping Table', from the ~10,000 Long Common Names from the LOINC RSNA radiology playbook. This set was selected in a hierarchical manner to balance the number of codes used versus the level of detail that is anticipated for cohort selection. A sample of 146,600 DR, CR, DX and CT exams containing 1,400 unique DICOM Study Descriptions in a highly-skewed long-tailed distribution was mapped with the MIDRC-LOINC Mapping Table.

Results: We were able to match over 97% of the Study Descriptions to 65 unique LOINC Long Common Names.

Conclusion: Using DICOM metadata, most incoming imaging data to MIDRC can be mapped to a restricted set of LOINC Long Common Names that are suitable for cohort selection from DICOM exams pooled from multiple imaging centers. The MIDRC-LOINC Mapping Table and a companion LOINC code attribute table are being regularly updated and are publicly available on GitHub.

Keywords: Artificial Intelligence/Machine Learning; Imaging Research; Standards & Interoperability

062 - Comparing General Vision Transformer Embeddings versus Task-Specific Deep Learning Models for Knee Osteoarthritis Grading

Presenter: Mohammadreza Chavoshi, Emory University

Mohammadreza Chavoshi1, Frank Li1, Theo Dapamede1, Bardia Khosravi1, Aawez Mansuri1, Rohan Satya Isaac1, Janice Newsome1, Hari Trivedi1, Judy Gichoya1

1Emory University, Atlanta, GA, USA

Introduction: Vision transformers are powerful tools for extracting image embeddings. However, the specificity and relevance of these embeddings for specialized diagnostic tasks remain unclear. The Kellgren-Lawrence Grade (KLG) scoring system for knee osteoarthritis (OA) assessment uses specific radiographic features -osteophytes, joint space narrowing, and bone deformity – to grade severity of knee OA, making it a suitable context to evaluate the performance of embeddings for specific diagnostic image analysis.

Hypothesis: To evaluate whether general-purpose medical image transformer embeddings (BioMedCLIP) can capture task-specific radiographic features as effectively as specialized deep learning models for KLG scoring.

Methods: We analyzed bilateral PA fixed-flexion knee radiographs from the NIH Osteoarthritis Initiative dataset. Bilateral images were cropped to unilateral. From 4,796 patients followed up at 12, 24, 36, 48, 72, and 96 months, we included 4,507 patients (38,199 images). We compared three approaches: (1) ConvNeXt and (2) ResNet18, both trained specifically for KLG classification, versus (3) a custom neural-network classifier trained on BioMedCLIP vision transformer embeddings. This design allowed us to contrast the performance of models learning task-specific features directly from images against a model using pre-extracted general medical image embeddings.

Results: Models trained directly on radiographs significantly outperformed the transformer embedding-based approach. ConvNeXt achieved the highest performance (quadratic weighted kappa: 0.8089, adjacent accuracy: 0.9413), followed by ResNet18 (kappa: 0.7474, adjacent accuracy: 0.9110). The BioMedCLIP embedding-based model showed notably lower performance (kappa: 0.5960, adjacent accuracy: 0.7890). All models performed better on extreme grades (0 and 4), with ConvNeXt achieving F1-scores of 0.709 and 0.835 respectively. Notably, the embedding-based approach completely failed to identify grade 1 cases (F1-score: 0.000) and showed substantial degradation in detecting intermediate grades (F1-scores: 0.407 and 0.389 for grades 2 and 3), suggesting limited capture of subtle radiographic features crucial for KLG scoring.

Conclusion: While general medical vision transformers like BioMedCLIP can recognize broad anatomical structures, their embeddings fail to capture the nuanced radiographic features essential for specialized tasks like KLG grading, particularly in distinguishing intermediate disease stages. The superior performance of task-specific deep learning models, especially in detecting subtle grade differences, emphasizes two key points: first, the current limitations of general-purpose medical image embeddings for fine-grained feature detection, and second, the need for developing specialized transformer architectures with focused pre-training on specific anatomical regions and imaging modalities. These results suggest that the successful application of vision transformers in medical imaging may require a shift from general to organ-specific pre-training strategies to ensure clinically relevant feature extraction.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Imaging Research

063 - DeepSPINE: A Comprehensive Deep Learning Model for Multi-Task Lumbar Spine MRI Analysis

Presenter: Kay Wu, University of Toronto

Kay Wu1, Satvik Tripathi2, Christopher Bridge3, Stuart Pomerantz4

1University of Toronto, Toronto, Ontario, Canada

2University of Pennsylvania, Philadelphia, PA

3Massachusetts General Hospital, Charlestown, MA, USA

4Transparent Imaging, Boston, MA, USA

Introduction: A lumbar spine MRI is vital for diagnosing persistent low back pain etiologies. Spinal MR interpretation is time-consuming and subject to inter-reader variability.

Hypothesis: We extend upon a deep learning model tailored for comprehensive, automated analysis of lumbar spine MRI to detect and grade nine degenerative spinal conditions.

Methods: DeepSPINE was trained and evaluated on a dataset of 54739 T2-weighted lumbar MRI studies (31439 female, 23300 male, mean age 58.3 years) to predict the presence and severity of spinal pathologies: left (LFS) and right foraminal (RFS) and spinal canal stenosis (SCS), disc bulging (DB), disc osteophyte complex (DOC), left (LFA) and right facet arthropathy (RFA), ligamentum flavum thickening (LFT), and epidural lipomatosis (EL). Using natural language processing, intervertebral level-by-level ground-truth labels of pathological processes from associated radiology reports were extracted. From the studies, the vertebral bodies were segmented. Using these, the intervertebral discs were localized and image volumes in the axial and sagittal planes of each disc were extracted. Each was fed into a convolutional neural network based on ResNeXt with softmax activation and categorical cross-entropy loss to perform the classification tasks.

Results: DeepSPINE demonstrated within-one class accuracies of 96.1%, 96.1%, and 97.0% and quadratic Cohen’s kappa of 0.745, 0.750, and 0.781 in classifying the severity of LFS, RFS, and SCS, respectively. For binary DB, DOC, LFA, RFA, LFT, and EL classification, AUC scores were 0.861, 0.838, 0.628, 0.632, 0.669, and 0.638.

Conclusion: We successfully trained an efficient deep learning model to automatically predict and grade various spinal pathologic processes. DeepSPINE achieved strong performance across classification tasks at each spinal level. To our knowledge, this is the first model trained on such a large and robust dataset to generate more comprehensive, descriptive level-by-level predictions of lumbar spine disease. DeepSPINE's comprehensive analysis of lumbar spine MRI shows potential to improve patient care by enhancing diagnostic accuracy for spinal diseases, providing standardized interpretations, streamlining workflow, and facilitating tailored treatment planning for spinal diseases, ultimately alleviating the burden on radiologists and enhancing efficiency and timely access to care.

Keywords: Artificial Intelligence/Machine Learning

064 - Detection of Aberrant Anterior Tibial Artery on Knee MRI Using Deep Learning

Presenter: Eduardo Farina, Federal University of São Paulo,

Cassiano Barros1, Eduardo Farina1, Daisy Kase1, Paulo de Tarso Kawakami Perez1, Leonardo Kazunori Tsuji1, Lucas Medeiros1, Adham do Amaral e Castro1, Felipe Kitamura1, Andre Aihara1

1Federal University of São Paulo, São Paulo, Brazil

Introduction: The aberrant anterior tibial artery (AATA) is a rare anatomical variant with an increased risk of injury during orthopedic procedures such as high tibial osteotomy, revision total knee arthroplasty, lateral meniscal repair, posterior cruciate ligament reconstruction, and tibial tubercle osteotomy screw fixation. Accurate preoperative detection of the AATA on knee magnetic resonance imaging (MRI) can guide surgical planning and reduce complications

Hypothesis: A deep learning algorithm trained on axial T2-weighted knee MRI can reliably detect the AATA with high sensitivity and specificity.

Methods: A retrospective dataset from a unified multi-center institution was acquired after IRB approval. The dataset comprised 70,260 axial T2-weighted images from 2,315 MRI exams (1,441 without AATA and 874 with AATA). The dataset was split into training (42,488 images; 866 non-AATA/562 AATA), validation (13,865 images; 286 non-AATA/182 AATA), and testing (13,907 images; 289 non-AATA/166 AATA) folds.

The dataset was annotated by five musculoskeletal radiologists with 1 to 7 years of MSK experience, under the supervision of a musculoskeletal radiologist with 24 years of expertise. Preprocessing involved pixel normalization between the 0.25th and 99.75th percentiles. A deep learning model (architecture: ResNet-based custom CNN) was trained to identify the presence of AATA, with metrics computed at the patient level. The model was developed using Python version 3.8.10 and PyTorch version 2.0.1.

Results: A slice-level analysis achieved an F1-score of 0.838, while patient-level classification applied a probability threshold of 0.172, yielding F1-scores of 0.966 on the validation set and 0.979 on the test set. AUC for the test set was 0.99.

Conclusion: This deep learning algorithm shows promise for automating the detection of the aberrant anterior tibial artery on knee MRI, potentially improving preoperative risk assessment and surgical outcomes. Further validation in larger, diverse datasets is warranted.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Imaging Research; Quality Improvement & Quality Assurance

065 - Developing Image-based Search Capabilities Within Hospital Databases to Diagnose Rare Brain Tumors

Presenter: Nazanin Maleki, Children's Hospital of Philadelphia

Sahil Chadha1, Nazanin Maleki2, Marc von Reppert3, Klara Willms3, Arman Avesta4, Tal Zeevi1, MingDe Lin5, Karl-Titus Hoffmann3, Sanjay Aneja1, Mariam Aboian2

1Yale University, New Haven, CT, USA

2Children's Hospital of Philadelphia, Philadelphia, PA, USA

3University of Leipzig, Leipzig, Germany

4Massachusetts General Hospital, Boston, MA, USA

5Visage Imaging, San Diego, CA, USA

Introduction: Current neuroradiology reference materials don’t fully capture the wide variety of brain tumor appearances, and text-based image search tools struggle due to inconsistent radiology report formats. previous studies attempting to create image-based search tools for brain tumor MRIs faced practical challenges. Techniques like feature extraction, image fusion, and transfer learning were explored, but these required expert neuroradiologists to manually delineate tumors, making them too complicated for routine clinical use. Additionally, previous algorithms didn’t involve clinicians in validating the usefulness of the search results, and they lack accurate matching of specific brain tumor subtypes. There has been a gap in addressing brain tumor subtype classification with greater precision in retrieval.

Hypothesis: An automated image-based search pipeline integrating convolutional neural networks and dimensionality reduction can efficiently retrieve clinically relevant brain tumor cases with high precision, improving diagnostic support and educational resources.

Methods: 295 patients (mean age and SD, 51 ± 20 years) with primary brain tumors who underwent surgical and/or radiotherapeutic treatment between 2000 and 2021 were included in this retrospective study. Semi-automated convolutional neural network-based tumor segmentation was performed, and radiomic features were extracted. The dataset was split into reference and query subsets, and four dimensionality reduction techniques including PCA, t-SNE, UMAP, and PHATE were applied to cluster reference cases. Radiomic features extracted from each query case were projected onto the clustered reference cases, and nearest neighbors were retrieved. Retrieval performance was evaluated using mean average precision at k, and the best-performing dimensionality reduction technique was identified. four expert neuroradiologists independently rated visual similarity using a five-point Likert scale.

Results: t-SNE with six components was the highest-performing dimensionality reduction technique, with mean average precision at 5 ranging from 78% to 100% by tumor type. PCA and UMAP achieved comparable results to t-SNE for Glioblastoma,Astrocytoma, and Meningioma, but results differed for Pilocytic Astrocytoma. PCA achieved a maximum mAP@5 of 73% for PA, and UMAP achieved 68%, compared to 78% for t-SNE. Across tumor types, the poorest performance was observed for PHATE. The top five retrieved reference cases showed high visual similarity Likert scores with corresponding query cases (76% ‘similar’ or ‘very similar’).

Conclusion: There is a critical need for a method to retrieve historical medical images based on visual similarity to an image query. We introduce an image-based search algorithm that automatically retrieves similar reference cases based on extracted query image features without any text input and this algorithm is the first to separate glioma subtypes.

Keywords: Artificial Intelligence/Machine Learning; Educational Systems; Imaging Research

066 - Dynamic Modeling for Breast Cancer Risk Prediction and Demographic Stratification

Presenter: Shamiha Binta Manir, Marquette University

Shamiha Binta Manir1, Priya Deshpande1

1Marquette University, Milwaukee, WI, USA

Introduction: Breast cancer represents a substantial health threat to women worldwide, necessitating innovative predictive tools to enable early detection and personalized care. This research leverages a time-adjusted risk prediction model combined with demographic clustering to support targeted public health interventions. Utilizing the Breast Cancer Surveillance Consortium (BCSC) dataset, comprising over 6 million mammograms and 1 million detailed demographic records, we explore dynamic risk modeling to enhance prediction accuracy and reveal critical demographic trends.

Hypothesis: We hypothesize that incorporating temporal demographic variables into machine learning frameworks will improve predictive accuracy and identify high-risk population clusters. By tracking changes in age, BMI, and menopausal status over time, we aim to construct dynamic risk trajectories and design customized interventions for diverse demographic groups.

Methods: A time-adjusted risk scoring system will be implemented using recurrent neural networks (RNNs) and transformer-based architectures to capture temporal variations in risk factors. Annual demographic changes, including age and hormonal transitions, dynamically update risk profiles. Clustering methods, such as K-means and hierarchical clustering, will be applied to identify unique population subgroups with distinct risk profiles. To address data imbalance, resampling techniques like Tomek Links are employed to optimize classifier training. Robustness and generalizability are ensured through cross-validation across demographic segments.

Results: Preliminary findings indicate that the Random Forest classifier achieves 87.47% accuracy using Tomek Link resampling. Clustering analysis identifies distinct high-risk cohorts, including post-menopausal women with elevated BMI, enabling precision-focused public health strategies. Additionally, the ROC curve demonstrates the model's predictive capability, achieving an AUC of 0.81, signifying strong classification performance.

Conclusion: This study highlights the potential of combining time-adjusted modeling with demographic clustering for breast cancer risk prediction. By dynamically refining risk profiles and identifying high-risk groups, the proposed approach advances precision medicine and offers scalable solutions for early detection and strategic public health planning.

Keywords: Applications; Artificial Intelligence/Machine Learning; Emerging Technologies; Imaging Research; Quality Improvement & Quality Assurance

067 - EdgeDeID: Advancing De-identification with Small LMs (SLMs) and Synthetic Data on Edge Devices

Presenter: Bardia Khosravi, Yale University

Bardia Khosravi1, Theo Dapamede2, Frank Li2, Pouria Rouzrokh1, Aawez Mansuri2, Elham Mahmoudi3, Rohan Satya Isaac2, Amirali Khosravi3, Mohammadreza Chavoshi2, Hari Trivedi2, Janice Newsome2, Bradley Erickson3, Judy Gichoya2

1Yale University, New Haven, CT, USA

2Emory University, Atlanta, GA, USA

3Mayo Clinic, Rochester, MN, USA

Introduction: De-identification of medical text reports is crucial for maintaining patient privacy while enabling data sharing for research and analysis. Existing tools, primarily based on BERT, face two key challenges: (1) they process the entire text sequentially, leading to suboptimal performance when dealing with long reports containing minimal protected health information (PHI), and (2) they require significant computational resources, making them impractical for edge devices. This study presents EdgeDeID, a novel approach leveraging small language models (SLMs) and synthetic data for efficient de-identification on edge devices.

Hypothesis: We hypothesize that by utilizing a small language model fine-tuned on synthetically generated data, EdgeDeID can accurately extract PHI entities, significantly reducing processing time compared to BERT-based methods, especially on edge devices with limited resources.

Methods: Synthetic data was generated using the Hermes model, a fine-tuned LLaMA 405B, and augmented with techniques such as error introduction and name variations to create a diverse training set. The original training set contained 18,000 samples, which was further expanded to 23,000 through augmentations. The Qwen 2.5 Coder decoder-only transformer (0.5B parameters) was fine-tuned on this dataset using supervised learning. The model supports a context length of 8k tokens, surpassing other models limited to 512 tokens. Temperature zero was used for all inference cases. EdgeDeID's performance was evaluated on 100 manually annotated reports.

Results: EdgeDeID processed 100 reports in under 20 seconds on a GPU with 8GB RAM and 4.2 seconds per report on devices without a GPU. The model achieved an overall recall (sensitivity) of 98.47% across all entities, with the lowest recall of 95.0% for date-type entities, which can be further optimized using rule-based methods to ensure complete anonymity. Model’s precision was 99.0%. Additionally, the model achieved >95% sensitivity across the 21PHI entities.

Conclusion: EdgeDeID presents a feasible solution for de-identifying medical reports on edge devices, addressing the challenges of data privacy when using LLMs that require external data transmission. By leveraging SLMs and synthetic data, EdgeDeID achieves rapid and accurate de-identification, facilitating secure data sharing for research and analysis.

Keywords: Applications; Artificial Intelligence/Machine Learning; Emerging Technologies; Security

068 - Emory Breast Imaging Dataset (EMBED) v2 – a Racially Diverse, Multi-modal Dataset of 1.2M Breast Imaging Exams and Associated Histopathology

Presenter: Rohan Satya Isaac, Emory University

Rohan Satya Isaac1, Beatrice Brown-Mulry1, Aawez Mansuri1, Theo Dapamede1, Chad Robichaux1, Frank Li1, Mohammedreza Chavoshi1, Judy Gichoya1, Hari Trivedi1

1Emory University, Atlanta, GA, USA

Introduction: We developed the EMory BrEast Imaging Datset (EMBED) in 2023, which is currently used at over 300 institutions worldwide and has been used to train or validate multiple FDA-cleared models. This dataset contained 350,000 2D and synthetic-2D mammograms along with patient demographics, risk factors, and pathologic outcomes for 116,000 patients. Since then, we have expanded EMBED with four additional years of data and added digital breast tomosynthesis (DBT), US, and MRI modalities, to create EMBEDv2.

Hypothesis: Substantial expansion of EMBED will enhance its utility for training and validating advanced breast imaging models, improving generalizability and performance across diverse patient populations.

Methods: We queried Magview software (Fulton, MD) for net new breast imaging exams since December, 2020. Relevant patient demographics, pathology reports, receptor, and recurrence information were extracted from the EHR. Discrepant data, such as changes in patient ID, breast density, and pathologic outcomes were manually reviewed and resolved. Outcomes are harmonized with the Georgia Department of Public Health Cancer Registry. Enrichment steps for both clinical and metadata were performed similar to EMBEDv1. Digital histopathology was aggregated for previously digitized patients, and will be collected prospectively beginning October, 2024.

Results: EMBEDv2 now encompasses 2013-2024 and has expanded from 116,177 to 260,815 patients and from 383,421 to 1,090,637 exams. The dataset contains 103,054 (40.3%) African American, 73,250 (28.6%) White, and 9,456 (4.2%) Hispanic patients. There are 767,500 (70.4%) screening mammograms, 204,246 (18.7%) diagnostic mammograms, 96,940 (8.9%) US, and 21,984 (2.0%) MRI exams. Ground truth pathologic outcomes are available for all biopsied patients with 4,959 (1.9%) invasive and 1,650 (0.6%) non-invasive cancers. 5 years of follow-up is available for 85,665 (32.8%) patients.

Conclusion: EMBED V2 contains 1.2M exams from 2020-2024, including two additional clinical sites, and expands to include digital breast tomosynthesis (DBT), US, and MRI exams. A subset of the dataset will again be released for researchers and be made available for commercial use.

Keywords: Artificial Intelligence/Machine Learning; Enterprise Imaging; Imaging Research

069 - Enhancing Access to Prenatal Studies Through EMR Customization: A Quality Improvement Initiative in Infants Aged 0–3 months

Presenter: Irene Dixe de Oliveira Santo, Yale University

Irene Dixe de Oliveira Santo1, Anne Gormley1, James Ha1, Lawrence Guan1, Cicero Silva1, Sophie Chheang1

1Yale University, New Haven, CT, USA

Introduction: Access to prenatal studies provides invaluable insights for timely and accurate diagnosis and management in young children, especially infants with congenital anomalies that are undergoing their first post-natal imaging studies. However, navigating electronic medical records (EMRs) to retrieve prenatal studies can be time-consuming and inefficient. Epic, as a widely used EMR platform, offers opportunities for customization to enhance workflow efficiency and improve patient care.

Hypothesis: The goal of our QI project was to customize Epic in children aged 0–1 year, incorporating a direct link to the mother’s chart to facilitate rapid access to prenatal studies.

Methods: In collaboration with our institution’s Epic team, a new print group was designed to appear automatically on the opening page of the child’s chart for patients aged 0–1 year. This section included a hyperlink to the mother’s chart, enabling direct and easy access to prenatal imaging. Time to access prenatal studies and the number of clicks required for 20 patients, were measured before and after implementation. Measurements were performed independently by three radiology residents.

Results: The implementation of the embedded print group and link to the mother’s chart significantly reduced the average time to access prenatal studies by an average of 27 seconds and decreased the number of clicks required by an average of 6.2 clicks. Radiologists’ feedback indicated improved satisfaction with workflow and enhanced ability to make informed clinical decisions in a timely manner.

Conclusion: Integrating a direct link to prenatal studies into the child’s chart on Epic proved to be an effective solution to a longstanding workflow inefficiency. This intervention highlighted the utility of EMR customization in optimizing clinical workflows and improving access to critical prenatal data. The benefits were particularly pronounced in cases of congenital abnormalities, where prompt access to prenatal studies facilitated better-informed diagnostic and therapeutic decisions. The creation and implementation of a print group section in Epic for children aged 0–1 year, with a direct link to the mother’s chart, significantly improved efficiency in accessing prenatal studies. This quality improvement initiative underscores the importance of collaborative efforts between clinical teams and EMR developers in enhancing patient care through targeted technological interventions.

Keywords: Clinical Workflow & Productivity; Provider Experience; Quality Improvement & Quality Assurance

070 - Enhancing Breast Cancer Diagnostics: Swin Transformer for Histopathology Image Classification

Presenter: Nadir Khan Yusufzai, Stony Brook University

Nadir Khan Yusufzai1, Jerome Liang1, Marc Pomeroy1

1Stony Brook University, Stony Brook, NY, USA

Introduction: Breast cancer remains a leading cause of cancer-related mortality worldwide, with early and accurate diagnosis critical for improving patient outcomes. Histopathological image analysis is a cornerstone of breast cancer diagnosis but is labor-intensive and subjective, relying heavily on pathologists’ expertise. Advances in artificial intelligence (AI), particularly deep learning, offer potential solutions to automate and enhance diagnostic workflows. The Swin Transformer, a state-of-the-art vision model, leverages hierarchical feature extraction and self-attention mechanisms to achieve superior performance in image classification tasks. This study aimed to evaluate the effectiveness of a pretrained Swin Transformer model in classifying benign and malignant breast cancer histopathology images using the BreakHis dataset.

Hypothesis: We hypothesized that the Swin Transformer, fine-tuned on the BreakHis dataset, would achieve high classification accuracy, sensitivity, and specificity in distinguishing between benign and malignant breast tumor images across multiple magnifications.

Methods: Data Preparation and Preprocessing:

The BreaKHis dataset, a publicly available collection of breast histopathology images captured at 40x, 100x, 200x, and 400x magnifications, was reorganized into two simplified directories: benign and malignant. A CSV file was used to track image metadata, including class labels and magnification levels. Images were renamed to align with patient IDs. Data were split into training (80%), validation (10%), and test (10%) sets.

Data Augmentation and Model Training:

To simulate real-world diagnostic variability, data augmentation was applied, including resizing, cropping, flipping, rotation, and color jittering. The Swin Transformer model, pre-trained on the ImageNet dataset, was retrained to classify histopathology images. The model was trained using the Adam optimizer with a learning rate of 1e-4.

Model Validation:

The model was validated against test datasets, and key performance metrics such as accuracy, sensitivity, specificity, and AUC were used to assess its potential integration into clinical workflows.

Results: The Swin Transformer model achieved an accuracy of 97.72%, sensitivity of 99.27%, and specificity of 92.15%, with an AUC of 0.9571. These results suggest that the model could provide significant clinical value by aiding pathologists in making more accurate and timely breast cancer diagnoses.

Conclusion: We demonstrate the potential of advanced deep learning models, such as the Swin Transformer, in augmenting breast cancer diagnostics. The integration of such technologies in healthcare settings could improve diagnostic accuracy and speed, reduce the dependency on specialized pathologists, and allow pathologists to make quicker treatment decisions, directly impacting patient outcomes. Future studies will focus on using the Swin Transformer to classify MRI scans of breast cancer into benign and malignant categories.

Keywords: Applications; Artificial Intelligence/Machine Learning; Imaging Research

071 - Enhancing Imaging Appropriateness in Acute Care: Aligning Large Language Models with the American College of Radiology Appropriateness Criteria

Presenter: Hersh Sagreiya, University of Pennsylvania

Michael Yao1, Charles Kahn1, Walter Witschey1, James Gee1, Osbert Bastani1, Hersh Sagreiya1, Allison Chae1

1University of Pennsylvania, Philadelphia, PA, USA

Introduction: Diagnostic imaging plays an essential role in managing acute patient care, yet many ordered studies are misaligned with established medical guidelines, such as the American College of Radiology (ACR) Appropriateness Criteria (AC). As a result, many ordered imaging studies lead to unnecessary costs, patient risk, and a burden to the overloaded healthcare system. This study explores the potential of large language models (LLMs) as clinical decision support tools to help clinicians order more appropriate imaging studies in acute healthcare settings.

Hypothesis: Prior work has primarily focused on leveraging language models to directly assign imaging studies to input patient case descriptions. In contrast, our method instead asks an LLM to predict the most appropriate ACR AC guideline title, or “Topic,” to describe an input patient scenario. Separately, we then parse through the textual ACR guidelines to determine the most appropriate, evidence-based imaging study(s) based on the predicted ACR AC Topic. We hypothesize that this inference strategy will enable LLMs to more accurately recommend appropriate imaging studies for acute patient presentations.

Methods: To experimentally validate our novel LLM inference strategy, we empirically evaluate six state-of-the-art LLMs on their ability to predict the correct ACR AC Topic label for an input patient case. We then further improve the performance of LLMs using zero-shot prompting techniques such as retrieval-augmented generation and chain-of-thought prompting.

Results: Our results demonstrate that our novel inference strategy can improve the accuracy of LLMs in predicting the most appropriate diagnostic imaging study by up to 50%. In a retrospective study with real patient case descriptions, autonomous LLM agents ordered more accurate imaging studies while simultaneously reducing the rate of unnecessary imaging orders compared with physicians.

Conclusion: Overall, our results underscore the potential of AI-driven support to improve clinical workflows. Future work will explore how similar LLM inference strategies can be applied to other areas of evidence-based medicine, where adherence to guidelines is critical for quality care.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity

072 - Enhancing Radiology Resident Education Through Locally Deployed AI/ML-Driven Report Feedback: A Pilot Study to Evaluate Feasibility and Impact

Presenter: John Garrett, University of Wisconsin

Orhan Unal1, Taylor Sellers1, Erik Winterholler1, Tan Nicholas1, Isha Pathak1, Allison Grayev1, John Garrett1

1University of Wisconsin, Madison, WI, USA

Introduction: Attending radiologist feedback on residents' preliminary reports is vital for radiology education. However, increasing imaging volumes have reduced opportunities for direct feedback, leaving attending revisions as an underutilized learning resource. Extracting structured, actionable feedback from these revisions is challenging due to variability and lack of guidance.

Hypothesis: A locally deployed AI/ML solution utilizing large language models (LLMs) can efficiently analyze and compare preliminary and final radiology reports, providing actionable feedback while safeguarding PHI through local deployment. This solution could serve as a practical educational tool for radiology training.

Methods: Impressions from overnight preliminary reports and corresponding final reports for chest X-ray and head CT studies were extracted via SQL queries and formatted into anonymized CSV tables. Leading open-source LLMs, including LLaMA3.1, Gemma2.5, and Mixtral, were tested with various zero-shot prompts designed to analyze differences between preliminary and final reports. Key aspects included clarity, relevance to training, clinical accuracy, and specificity. Initial subjective assessments evaluated the outputs’ alignment with educational goals, refining prompt design to optimize model performance in addressing radiology report nuances.

Results: Our initial findings demonstrate that LLMs effectively identify differences in clarity, training relevance, clinical accuracy, and specificity between preliminary and final reports. Performance varied with prompt structure, emphasizing the importance of prompt engineering in achieving meaningful results. Certain prompts generated outputs more aligned with clinical and educational objectives, confirming the feasibility of using LLMs to produce structured feedback. These results support the potential of locally deployed AI/ML tools to enhance traditional feedback mechanisms while maintaining privacy and data integrity. Quantitative analysis is ongoing to further validate these findings and optimize model outputs.

Conclusion: Locally deployed LLMs demonstrate robust performance in providing structured feedback to radiology residents while ensuring PHI protection. These findings highlight the potential for AI-driven tools to supplement traditional feedback, advancing radiology education and maintaining high standards of privacy and quality.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Educational Systems; Emerging Technologies; Provider Experience; Quality Improvement & Quality Assurance

073 - Evaluating Consistency in CT Derived Body Composition Analysis: A Comparative Study of Two Deep Learning Algorithms at L3 Vertebral Level

Presenter: Adam Dachowicz, Mayo Clinic

Adam Dachowicz1, Jason Klug1, Timothy Kline1, Panos Korfiatis1, Daniel Blezek1, Bill Ryan1, Eric Williamson1, Steve Langer1, Jeremy Collins1, Francis Baffour1, Andy Missert1, Gian Marco Conte1

1Mayo Clinic, Rochester, MN, USA

Introduction: The increasing availability of deep learning algorithms for body composition analysis from medical imaging presents opportunities to improve clinical workflows. However, potential variability in algorithm output raises concerns about the consistency of measurements. This study compares the performance of two state-of-the-art algorithms for body composition analysis, evaluating their agreement on key segmentation metrics.

Hypothesis: We hypothesize that despite high segmentation accuracy, differences in training will lead to measurable differences in output between the algorithms.

Methods: The study included 1,132 abdomen-pelvis CT scans from 615 patients. The analysis focused on the L3 vertebral level, a standard reference for body composition studies, using an internally developed model and Comp2Comp, an open-source tool. Both algorithms were applied to segment skeletal muscle (SKM), visceral adipose tissue (VAT), subcutaneous adipose tissue (SAT), and inter-muscular adipose tissue (IMAT). Segmentation was performed on identical CT slices taken at the L3 vertebra level. Segmentation accuracy was evaluated using the Dice Similarity Coefficient (DSC). Additionally, Bland-Altman analysis was used to assess agreement in total segmented volume. We also compare agreement across demographic and scanner variables of interest, including patient sex, age, and slice thickness.

Results: The algorithms demonstrate variable segmentation agreement across tissues, with highest agreement for SKM and lowest for IMAT, yielding DSC of 0.957±0.038 (mean +/- standard deviation) (SKM), 0.900+/−0.12 (VAT), 0.911+/−0.071 (SAT), and 0.843+/-.0.128 (IMAT). We see absolute segmented area differences of 4.32+/−6.55 cm^2 (SKM), 9.01+/−11.52 cm^2 (VAT), 18.22+/−26.34 cm^2 (SAT), and 1.18+/−1.16 cm^2 (IMAT). We observed evidence of a significant difference in the distribution in DSC for patient sex across all segmentation tissues (2-Sample KS p-value < 0.05), for age across all tissues (multi-sample A-D p-value < 0.05 and KW p value < 0.05), and for slice thickness across all tissues (2-Sample KS p-value < 0.05), suggesting segmentation performance is sensitive to these parameters.

Conclusion: The algorithms yield high but variable segmentation agreement across tissues. We observe evidence of varying levels of agreement across several demographic variables of interest, holding the L3 slice and originating CT scan constant. These result highlights the utility of benchmarking algorithms with similar outputs against each other to capture where disagreements occur. Such differences are important to be aware of in clinical practice.

Keywords: Artificial Intelligence/Machine Learning; Emerging Technologies; Imaging Research

074 - Evaluating Large Language Models for Multi-Institutional Radiology Report Annotation: A Prompt-Engineering Approach

Presenter: Mana Moassefi, Mayo Clinic

Mana Moassefi1, Les Folio2, Ghulam Rasool3, Sina Houshmand4, Peter Chang5, Katherine Andriole6, Maryellen Giger7, Jessyca Wagner8, Judy Gichoya9, Bradley Erickson1

1Mayo Clinic, Rochester, MN, USA

2James A. Haley Veterans Hospital, Tampa, FL, USA

3Moffitt Cancer Center, Tampa, FL, USA

4University of California San Francisco, San Francisco, CA, USA

5University of California Irvine, Irvine, CA, USA

6Harvard University, Cambridge, MA, USA

7University of Chicago, Chicago, IL, USA

8Midwestern State University, Wichita Falls, TX, USA

9Emory University, Atlanta, GA, USA

Introduction: The rapid evolution of large language models (LLMs) offers promising opportunities for radiology report annotation, aiding in determining the presence of specific findings. This study evaluates the effectiveness of a human-optimized prompt in labeling radiology reports across multiple institutions using LLMs.

Hypothesis: A human-optimized prompt can enable LLMs to accurately and consistently annotate radiology reports across multiple institutions, regardless of variations in report structures and institutional practices.

Methods: A multi-institutional dataset was curated, comprising 500 radiology reports per site from Mayo Clinic, University of California, San Francisco(UCSF), Massachusetts General Hospital(MGH), Unniversity of California-Irvine (UCI) and Emory. The reports’ findings included five categories: liver metastases (CT abdomen), subarachnoid hemorrhage (CT brain), pneumonia (chest X-ray), cervical spine fracture (CT), and glioma progression (MRI brain). A standardized Python script was distributed to participating sites, allowing the use of different locally executed LLMs and a human-optimized prompt. The script executed the LLM's analysis for each report, using a predefined answer set (e.g., ['Yes', 'No'] or ['Progression', 'Stable', 'Improved']), and compared predictions to ground truth labels provided by local investigators. Models’ performance using accuracy were calculated and results were aggregated centrally.

Results: The human-optimized prompt demonstrated high consistency across sites and pathologies, with overall performance surpassing initial expectations. Preliminary analysis indicates significant agreement between the LLM's outputs and investigator-provided ground truths across multiple institutions. At Mayo Clinic, eight LLMs were systematically compared, with Llama 3.1 70b achieving the highest performance in accurately identifying the specified findings. Comparable performance with Llama 3.1 70b was observed at two additional centers, demonstrating the model's robust adaptability to variations in report structures and institutional practices. We also note that for a small percentage of cases, the LLMs did not respond with the required short answer (e.g. ‘Yes’ or ‘No’) but provided a long explanation that usually was correct. However, these were counted as incorrect because the LLM did not follow the prompt.

Conclusion: Our findings illustrate the potential of optimized prompt engineering in leveraging LLMs for cross-institutional radiology report annotation. By eliminating the need for federated learning, this approach simplifies implementation while maintaining high accuracy and adaptability. Future work will explore model robustness to diverse report structures and further refine prompts to improve generalizability.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Emerging Technologies; Standards & Interoperability

075 - Evaluating Spatial Computing Platforms for Diagnostic Imaging: A Feasibility Study with the Apple Vision Pro

Presenter: Josh Volin, Emory University

Josh Volin1, Vidya Viswanathan1, Nabile Safdar1, Vikram Narayan1, Elias Kikano1

1Emory University, Atlanta, GA, USA

Introduction: The integration of augmented reality (AR) and virtual reality (VR) in diagnostic radiology has the potential to transform workflows by enhancing visualization and user experience. Spatial computing devices with advanced screen resolution, color contrast, and eye/finger-tracking features may improve educational and clinical applications.

Hypothesis: The primary aim was to survey participating radiologists about the usage of the AR/VR device.

Methods: This IRB-approved prospective study aims to include 25 radiology physicians and trainees at a large academic institution. Participants completed a preliminary survey assessing familiarity with AR/VR devices, followed by a guided experience using a AR/VR headset (Apple Vision Pro). This included a user interface demonstration, review of diagnostic medical images (Visage VP), and simulated electronic medical records (Epic). Additional metrics included time to competency (from donning the headset to PACS interaction) and total usage time recorded during the demo. Finally, a post-use survey was conducted which included binomial, 5-level scale, and open-ended prompts focused on ease of use, ergonomics, and workflow while using the headset. Summary statistics were then calculated.

Results: Preliminary data from 12 participants showed an average time to competency of 4.25 ± 1.07 minutes and total usage time of 13.13 ± 4.88 minutes. Pre-use surveys revealed 58% were unfamiliar with AR technology, though 67% saw its educational potential, and 83% emphasized ease of use for integration. Post-use surveys indicated 75% rated usability highly, 87.5% found the device comfortable, and 62.5% cited efficiency and quality of care as key benefits. Drawbacks included ergonomics (71%) and haptics (28%). Most participants (75%) indicated interest in clinical use, and 87.5% would recommend the headset to colleagues. Favored features included 3D visualization, anatomy education, and pathology correlation.

Conclusion: The Apple Vision Pro headset demonstrates significant potential for educational and clinical applications, offering advanced resolution and intuitive interaction. Ergonomic refinements remain necessary for broader adoption.

Keywords: Clinical Workflow & Productivity; Educational Systems; Emerging Technologies; Provider Experience; Quality Improvement & Quality Assurance

076 - From Data RAGs to Riches: Harmonizing MRI Sequence Classification with LLMs

Presenter: Briana Malik, St. Jude Children's Research Hospital

Briana Malik1, Dharmam Savani1, Asim Bag1, Silu Zhang1, Paul Yi1

1St. Jude Children's Research Hospital, Memphis, TN, USA

Introduction: MRI sequence classification in pediatric neuroimaging is hindered by inconsistent DICOM header metadata, complicating multi-institution clinical trials. To address this, we developed a Retrieval-Augmented Generation (RAG) framework using LLMs and evaluated if such a framework could harmonize heterogeneous metadata across diverse imaging protocols and scanners – offering a data-agnostic, scalable solution.

Hypothesis: RAG-based-LLM will improve pediatric MRI sequence classification.

Methods: Our end-to-end RAG pipeline facilitates pediatric MRI sequence classification by retrieving relevant context from a vector database before using GPT-4o for decision-making. As DICOM metadata is often inconsistent/incomplete, providing additional context using RAG theoretically helps resolve ambiguities.

The pipeline included four stages:

1. Extracting DICOM headers

2. Tokenizing metadata into embeddings

3. Storing embeddings in a FAISS vector database

4. Querying with and without RAG to compare classification performance

The dataset comprised 357 pediatric brain MRI series spanning T1, T2 (pre-/post-contrast), FLAIR, ADC, and “other” sequences, collected across heterogeneous scanners/protocols. Experiments were conducted using all 27 DICOM tags and the SeriesDescription tag alone to evaluate robustness. Accuracy of our RAG-enhanced pipeline was compared to baseline models using 1) all DICOM tags without RAG and 2) Series Description tags alone (with and without RAG).

Results: With all 27 tags, our RAG-based model achieved 83.7% accuracy, outperforming the non-RAG baseline (72.7%). In contrast, using only Series Description yielded lower accuracy (58.8% with RAG vs. 57.7% without). Most misclassifications occurred in the “other” category, often overlapping with ADC characteristics, reflecting the complexity of subtle sequence variations. Nonetheless, RAG reduced errors across all categories, including 17% improvement in T1 post-contrast classification.

Conclusion: Our RAG-based-LLM framework provides a generalizable, scalable solution to MRI sequence classification challenges. This method excels in resolving ambiguous cases, demonstrating potential to harmonize metadata interpretation and advance imaging informatics. Future work includes refining semantic retrieval for complex imaging tasks.

Keywords: Applications; Artificial Intelligence/Machine Learning; Emerging Technologies; Imaging Research

077 - Generation of Patient-friendly Radiology Video Reports by an Integrated AI System

Presenter: Luyang Luo, Harvard University

Luyang Luo1, Jenanan Vairavamurthy2, Mike Moritz3, Xiaoman Zhang1, Hong-Yu Zhou1, Sung Eun Kim1, Julian Acosta1, Subathra Adithan4, Stuart Schrof5, Ramon Ter-Oganesyan5, Mingxiang Wu6, Brady Chrisler3, Kent Kleinschmidt3,Sathvik Suryadevara3, Pranav Rajpurkar1

1Harvard University, Cambridge, MA, USA

2Icahn School of Medicine, New York, NY, USA

3Saint Louis University, Saint Louis, MO, USA

4Jawaharlal Institute of Postgraduate Medical Education and Research, Pondicherry, India

5Los Angeles General Medical Center, Los Angeles, CA, USA

6Shenzhen People’s Hospital, Guangdong Province, China

Introduction: Patient-centered radiology requires accessible communication of medical findings. While current efforts using large language models focus on simplifying reading levels, they often lack integration with imaging data that could enhance patient comprehension. We developed and evaluated an integrated AI system that generates patient-friendly video reports combining simplified explanations with highlights of findings on radiology images.

Hypothesis: We hypothesized that radiologists would rate the generated video reports as accurate and suitable for integration into their clinical workflow.

Methods: The system integrated GPT-4o for translating medical terminology into plain language, a grounding model for automated lesion highlighting and 3D anatomy rendering, and an avatar generation system for a virtual presenter interface. We evaluated the system using ten video reports generated for diverse cases, each associated with a detailed survey comprising ten questions using a 5-point Likert scale. Five radiologists conducted the surveys by reviewing the video reports together with the original reports.

Results: Radiologists strongly endorsed the system's core functionalities and the video reports' utility. The integration of displaying images was highly rated (100% positive: 68% strongly agree, 32% agree), along with the identification of the important CT findings (100% positive) and the comparison to normal CT scans (100% positive). The explanation's clarity (98% positive) and the avatar’s natural conversation style (98% positive) garnered universal positive feedback, with most feedback confirming the reports achieved an 8th-grade reading level comprehension (98% positive) and that the video sufficiently reviewed the findings (96% positive). Overall, radiologist users showed comfort in using these videos to help patients understand their reports (80% positive). Concerns remained about sharing videos before patient visits (60% positive, 20% neutral, and 20% negative). The 3D rendering feature showed mixed utility, with 50% finding it helpful, 44% remaining neutral, and 6% disagreeing with the usefulness.

Conclusion

This study demonstrates the feasibility of an AI-driven pipeline for generating patient-friendly video reports, representing a significant step toward enhanced patient-centered radiology communication.

Keywords: Applications; Artificial Intelligence/Machine Learning; Patient/Family Experience

078 - Harnessing AI for Medical Informatics: A Comparative Study of GPT and Traditional Methods in Expanding Medical Acronyms and Shorthands

Presenter: John Moon, Emory University

John Moon1, Emily Patel1, Hanzhou Li1, Zachary Bercu1, Janice Newsome1, Hari Trivedi1, Judy Gichoya1

1Emory University, Atlanta, GA, USA

Introduction: Medical acronyms and shorthand are widely used in healthcare to shorten word count and alleviate reading burden. However, varying interpretations across clinical contexts pose challenges, potentially disrupting workflows and paradoxically increasing the time required to understand the intended communication. Traditional methods of acronym and shorthand expansion, such as looking up medical dictionaries or conducting internet searches, often fail to account for contextual nuances, leading to further inefficiencies and misinterpretation.

Hypothesis: Large language models, such as ChatGPT 4.0, will outperform traditional lookup methods of medical acronyms and shorthand expansion (Taber’s Dictionary, OpenMD, and Google Search) in terms of accuracy.

Methods: Fifty de-identified History of Present Illness (HPI) statements between August and October 2024 containing >2 acronyms from a cross-sectional procedures workup database were selected. Acronym and shorthand expansions were generated using ChatGPT-4.0 by feeding in the HPI statement and comparing to three traditional lookup methods: Taber’s Dictionary, OpenMD, and Google Search. Accuracy was calculated as the proportion of correct expansions relative to the total acronyms analyzed. The phrase was marked “N/A” if it did not exist in the lookup method. Two physicians independently evaluated expansions for accuracy. A paired two-sample t-test was used to compare the performance of ChatGPT with each traditional method with statistical significance set at p < 0.05.

Results: A total of 498 acronyms/shorthands were identified from the 50 HPI statements by ChatGPT-4.0. After excluding 21 entries, which were either brand names or mislabeled, 489 (160 unique) acronyms/shorthands remained. On average, each HPI statement contained 9.5 ± 4.0 acronyms or shorthands, with a range of 4 to 26 per statement. ChatGPT-4.0 was correct on 160/160 (100.0%) compared to 47/160 (29.4%, p < 0.05) for Taber’s Dictionary, 87/160 (54.4%, p < 0.05) for OpenMD, and 127/160 (79.4%, p < 0.05) for Google Search. One such example of a less commonly used acronym in which ChatGPT outperformed traditional methods given the context of the HPI statement is the expansion of “SS.” ChatGPT correctly identified it as surgical scar, Google expanded it as sliding scale, Taber’s expanded it as a half, and OpenMD offered multiple incorrect expansions, including one-half, sliding scale, and Sjogren’s syndrome.

Conclusion: Context-aware capabilities of ChatGPT-4.0 in expanding medical acronyms and shorthands highlights its superiority in addressing inefficiencies inherent in traditional lookup methods. By incorporating AI-powered tools, healthcare systems may streamline communication, minimize errors, and boost clinical efficiency. Further studies should explore real-world implementation and evaluate their influence on clinical outcomes.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Patient/Family Experience; Quality Improvement & Quality Assurance

079 - High Performance Prompting for LLM Extraction of Findings from Radiology Reports

Presenter: Mohammed Kanani, University of Washington

Mohammed Kanani1, Arezu Monawer1, Lauryn Brown1, William King1, Zach Miller1, Nitin Venugopal1, Patrick Heagerty1, Jeffrey Jarvik1, Trevor Cohen1, Nathan Cross1

1University of Washington, Seattle, WA, USA

Introduction: Extracting information from radiology reports can provide critical data to empower clinical, QA/QI, research, and educational radiology workflows. For spinal compression fractures, this data can identify at risk populations and facilitate evidence-based follow up and treatment. Manual extraction from free-text radiology reports is laborious, and prone to errors. Large language models (LLMs) have shown promise. Fine tuning strategies are time and resource intensive, but a variety of prompting strategies have achieved similar results with less resources, data and annotation. Our study pioneers the use of Meta’s Llama 3.1 for automated extraction from free-text radiology reports to detect spinal compression fractures, outputting structured data without model training.

Hypothesis: An open-source LLM can extract imaging findings on specific pathologies from free-text radiology reports using prompt-based strategies, eliminating the need for specific model training.

Methods: We tested performance on a time-based sample of CT exams covering the spine from 2/20/2024 to 2/22/2024 acquired across our healthcare enterprise (637 anonymized reports, age 18–102, 47% Female). Ground truth annotations were manually generated by a group of attending/fellow/resident physicians, and medical students. These annotations were compared against the performance of three models (Llama 3.1 70B, Llama 3.1 8B, and Vicuna 13B) with nine different prompting configurations for a total of 27 model/prompt experiments.

Results: Among the 637 reports, 49 were classified as true by annotators (prevalence: 7.69%). The highest F1 score (0.91) was achieved by the 70B Llama 3.1 model when provided with a radiologist-written background, with similar results when the background was written by a separate LLM (0.86). The addition of few-shot examples had variable impact on these prompts (0.89, 0.84 respectively). Comparable ROC-AUC and PR-AUC performance was observed.

Conclusion: An open-source LLM excelled at extracting diagnostic information from free-text radiology reports using prompt-based techniques without model training.

Keywords: Artificial Intelligence/Machine Learning; Emerging Technologies; Imaging Research; Quality Improvement & Quality Assurance

080 - Identifying Gaps in Care Using AI for Vertebral Body Compression Fracture Detection on Chest X-Rays

Presenter: Mohamed Ibrahim, University of Alabama at Birmingham

Mohamed Ibrahim1, Omar Safarini1, John Eddins1, Lawrence Ngo2, Richard Brown2, Russell Stewart2, Srini Tridandapani1, Steven Rothenberg1

1University of Alabama at Birmingham, Birmingham, AL, USA

2Covera Health, New York, NY, USA

Introduction: Vertebral compression fractures (VCF) are often underreported as this is often not the purpose of the imaging examination. Identifying VCF is crucial for risk-stratifying patients for pharmacologic therapy in osteoporosis management. This retrospective study estimates the gap in care by estimating the incidence of missed VCF using AI as a subsequent reader.

Hypothesis: VCF is underreported on posterior-anterior (PA) and lateral chest radiographs performed in routine clinical practice.

Methods: An IRB-approved retrospective analysis of consecutive PA and lateral CXR exams (N = 16,066) was conducted using a commercially available FDA-cleared AI model for vertebral fracture detection. AI-flagged exams were classified as positive or negative for VCF. Radiology reports were analyzed using natural language processing (NLP) to determine whether VCF was reported. A board-certified radiologist reviewed the CXR images to establish the reference standard in a subset of discordant results (N&#3f337). Exams were considered positive with mild, moderate, or severe VCF using the Genant classification. Counts and enhanced detection rates (absolute, EDRa, and relative, EDRr) were calculated.

Results: The AI model flagged 13.2% (2,134/16,066) exams for VCF, of which 79.4% (1,695/2,134) were discordant with the radiology reports. 1.48% (237/16,066) exams were positive for VCF but missed by AI. In the subset of exams reviewed by a board-certified radiologist, the incidence of VCF was 9.6% (199/2,072). VCF was only reported in 30.1% (60/199) of exams and VCF was not reported in 69.8% (139/199) of confirmed VCF. The EDRa and EDRr of VCF was 6.7% (139/2,072) and 231% (139/60) respectively.

Conclusion: VCF is underreported on CXRs, and AI has the potential to close gaps in care for population health interventions. Further investigation is needed to stratify these results by clinically significant VCF with moderate to severe compression fractures which require further workup and management.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Imaging Research; Quality Improvement & Quality Assurance

081 - IDH-mutant Glioma Subtypes Classification Using Unsupervised Dimensionality Reduction of MRI Biomarkers

Presenter: Nazanin Maleki, Children's Hospital of Philadelphia

Klara Willms1, Nazanin Maleki2, Tal Zeevi3, Sahil Chadha4, Marc von Reppert1, Leon Jekel5, Jan Lost6, Sarah Merkaj7, Niklas Tilmanns6, Mariam Aboian2

1University of Leipzig, Leipzig, Germany

2Children's Hospital of Philadelphia, Philadelphia, PA, USA

3Yale University, New Haven, CT, USA

4University Hospital Essen, Essen, Germany

5University of Duisburg-Essen, Duisburg, Germany

6Heinrich-Heine-University, Düsseldorf, Germany

7University of Ulm, Ulm, Germany

Introduction: IDH-mutant glioma share many similarities on MRI and histopathology but have distinct differences in their survival based on their molecular subtypes. Radiomic biomarkers provide insights into glioma imaging phenotypes that may not appear in clinical evaluation . This study aimed to evaluate whether MRI-based biomarkers of IDH-mutant gliomas can differentiate 2021 CNS5 WHO classification molecular subtypes and identify distinct patterns of overall survival.

Hypothesis: To explore the use of radiomic biomarkers to classify IDH-mutant gliomas subtype

Methods: We included 179 adult patients with IDH-mutant, 1p/19q co-deleted oligodendrogliomas (WHO CNS5 grade 2–3) and IDH-mutant, non co-deleted astrocytoma (WHO CNS5 grades 2–4) with pre-treatment MRI on FLAIR , T1 post-contrast and T2 . Segmentations were performed using a UNETR algorithm trained on the BRaTS 2021 dataset and manual modification. M=94 texture biomarkers were extracted from the segmentation region of each scan using the PyRadiomics 3.1.0 pipeline. Dimensionality reduction techniques (PCA, t-SNE, UMAP, and PHATE) were employed to the radiomic biomarkers of (i) FLAIR; (ii) FLAIR and T1 post-contrast imaging (PGGE/PGSE); (iii) incorporating one-hot encoded lesion location information; and (iv) FLAIR and T2 imaging. K-Means clustering was used to identify relevant clinical clusters in the reduced space.

Results: (i) For FLAIR-based biomarkers, all dimensionality reduction techniques produced identical distributions between clusters B and C. The highest representations included 27.1% oligodendroglioma, grade 2, 24.8% astrocytoma, grade 2, and 19.4% oligodendroglioma in one cluster, and 31.1% astrocytoma, grade 2, 24.4% oligodendroglioma, grade 3, and 17.8% oligodendroglioma, grade 2 in the other. (ii) Adding radiomic features from T1 post-contrast, PCA differed: Cluster A had 46.7% astrocytoma, grade 2, versus 37.8% in Cluster B and 23.2% in Cluster C. TSNE, PHATE, and UMAP showed 46.2% astrocytoma, grade 2 in Cluster B, and 34.6% in Cluster C. (iv) Adding T2/FLAIR mismatch sign revealed no significant distribution differences: 20.9% positive mismatch in Cluster A and 12.8% in Cluster B. Kaplan-Meier survival analysis of UMAP clusters (log-rank test) showed no significant survival differences: p=0.781 (FLAIR), p=0.054 (FLAIR+T1Gd), p=0.455 (FLAIR+T2).

Conclusion: We show that there is a spectrum of imaging features of IDH-mutant gliomas that have significant overlap on FLAIR and T1 post-contrast imaging, but cluster into two groups only on FLAIR and three clusters when adding T1WI post-gadolinium radiomic features.

Keywords: Artificial Intelligence/Machine Learning; Imaging Research

082 - Impact of Algorithm-Driven Slice Selection on Body Composition Analysis: A Comparison of Two Deep Learning Models Using CT Images

Presenter: Jason Klug, Mayo Clinic

Jason Klug1, Adam Dachowicz1, Timothy Kline1, Panos Korfiatis1, Daniel Blezek1, Bill Ryan1, Eric Williamson1, Steve Langer1, Jeremy Collins1, Francis Baffour1, Andy Missert1, Gian Marco Conte1

1Mayo Clinic, Rochester, MN, USA

Introduction: Advancements in deep learning have enabled automated algorithms for body composition analysis on CT, including autonomous slice selection for segmentation. While this enhances automation, it raises concerns about consistency and comparability of outputs. This study evaluates the agreement of two body composition algorithms at the L3 vertebra level.

Hypothesis: We hypothesized that independent slice selection could introduce variability in body composition metrics, impacting consistency.

Methods: CT scans from 1076 patients across 2009 abdomen/pelvis series were analyzed using an internally developed model and Comp2Comp, an open-source tool. Both algorithms autonomously selected slices for L3 segmentation, including skeletal muscle (SKM), visceral fat (VAT), subcutaneous fat (SAT), and intermuscular fat (IMAT). The selected slice index and segmentation outputs were compared using Bland-Altman analysis and interclass correlation coefficients (ICCs). Differences were evaluated across demographic and scanner variables.

Results: The mean slice index difference between algorithms was 1.37 ± 6.27, with identical slices selected in 17.4% of cases. The absolute slice index difference was mildly correlated with the difference in SKM area (spearman rank 0.22; r2=0.03). The segmentation agreement was high, with mean differences (LOA) in SKM = −0.23cm2 (18.17, −18.63), SAT = −19.95cm2 (30.27, −70.16), VAT = 10.40cm2 (36.51, −15.72), and IMAT = −16.44cm2 (1.78, −34.65). ICCs for area agreement were excellent for SKM, SAT, and VAT but lower for IMAT (ICC: 0.983, 0.982, 0.995, 0.595, respectively). Absolute differences in SKM area between the algorithms were larger in males compared to females (F: 4.94 =/- 7.22 cm2; M 6.16 +/- 8.26 cm2; Mann Whitney U-test; p = 5.4e-0.5). Absolute differences in SKM area between AI model segmentation (Kruskal-Wallis test) for BMI (p = 1.4e-0.3), age (p = 4.1e-0.5), patient race group (p = 1.1e-0.2), and site (p = 5.4e-0.5) were significantly different. Absolute differences in SKM area were not significantly different for scanner manufacturer (p = 5.8e-0.2) and slice thickness (p = 6.1e-0.1).

Conclusion: Allowing algorithms to select slices autonomously did not significantly affect SKM, VAT, and SAT areas, while IMAT showed only moderate agreement. Absolute differences in SKM area between the models show differential effects on demographic features, including age, sex, BMI and site. Scanner variables did not show significant effects. Overall, automated model selection of the L3 slice may depend on the quality of vertebrae segmentation, but the L3 slice selected may play a minor role in body composition segmentation variability. The clinical impact of these differences will be evaluated in the future.

Keywords: Artificial Intelligence/Machine Learning; Emerging Technologies; Imaging Research

083 - Improving Radiology Report Conciseness and Structure via Locally Run Large Language Models

Presenter: Ghulam Rasool, Moffitt Cancer Center

Ghulam Rasool1, Iryna Hartsock1, Cyrillo Araujo1, Les Folio2

1Moffitt Cancer Center, Tampa, FL, USA

2James A. Haley Veterans Hospital, Tampa, FL, USA

Introduction: Radiology reports often suffer from verbosity and lack of standardized structure, hindering efficient interpretation. This study explores locally run, open-source large language models (LLMs) to improve report conciseness and structure while ensuring data privacy and compliance with regulatory frameworks.

Hypothesis: Open-source LLMs, when deployed locally, can streamline radiology reports by reducing redundancy and organizing findings into a structured format, enhancing their readability and clinical utility.

Methods: We analyzed 814 de-identified radiology reports from seven board-certified body radiologists at Moffitt Cancer Center. Locally implemented LLMs, including Mixtral, Mistral, and Llama, were evaluated using the Ollama framework. Five prompting strategies were tested to restructure and condense reports, and the Signal-to-Noise Ratio (SnR) metric was developed to quantify meaningful content relative to redundant information. Key metrics included formatting accuracy, adherence to structure, and improvement in SnR.

Results: The "Structure + Conciseness (Findings, Impressions)" and "Conciseness >> Structure" prompting approaches performed best, achieving the highest SnR values and significantly reducing redundancy while maintaining or enhancing clarity. These methods also demonstrated fewer formatting errors compared to other strategies. Mixtral outperformed Mistral and Llama in adhering to structural instructions and producing concise outputs.

Conclusion: Locally run, open-source LLMs like Mixtral can securely and effectively enhance the clarity, conciseness, and structure of radiology reports. These findings demonstrate the potential for LLMs to improve radiology workflows while addressing critical data privacy concerns.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Emerging Technologies; Security

084 - Large Language Model Sensitivity to Data Perturbations in Radiology Report Classification

Presenter: Vera Sorin, Mayo Clinic

Vera Sorin1, Jeremy Collins1, Panagiotis Korfiatis1

1Mayo Clinic, Rochester, MN, USA

Introduction: Large language models (LLMs) are increasingly evaluated and applied to radiology reporting tasks. This study aimed to assess the impact of different types and levels of input text perturbations on LLM performance in classifying radiology reports.

Hypothesis: LLM performance may vary under different noise levels introduced into the radiology reports.

Methods: This was a retrospective IRB-approved study. We evaluated two Google LLMs, Gemini-1.5-Flash-001 and Gemini-1.5-Flash-002, on a balanced dataset of 2,200 CT pulmonary angiography reports (1,100 positive and 1,100 negative for pulmonary embolism). Three forms of noise were introduced: (1) Character noise: random removal of 20, 30, 60, or 120 characters, (2) Symbol noise: random insertion of 3, 9, 12, 24, or 64 symbols, and (3) Word shuffle: random rearrangement of 10, 30, or 50 words. Performance metrics for both models under each level of noise were calculated.

Results: Without noise, Gemini-1.5-Flash-001 achieved accuracy 0.967, recall 0.935, and F1-score 0.966. Gemini-1.5-Flash-002 performed at accuracy 0.984, recall 0.971, and F1-score 0.983. At the highest noise levels, Gemini-1.5-Flash-001’s accuracy declined to 0.937 with character noise, 0.958 with symbol noise, and 0.932 with word shuffle. Gemini-1.5-Flash-002 had higher accuracy under the same conditions: 0.975 for character noise, 0.981 for symbol noise, and 0.975 for word shuffle. Overall, Gemini-1.5-Flash-002 demonstrated more resilience to noise, with smaller drops in accuracy across all perturbation types and levels.

Conclusion: Our results show that LLMs may be sensitive to data perturbations, including typos, formatting errors, and word shuffle. This vulnerability raises concerns about these models’ performance when handling imperfect clinical data, as well as a potential sensitivity to cyber-attacks. Understanding the robustness and carefully validating LLMs is necessary prior to integrating into clinical practice.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Emerging Technologies

085 - Large Language Models for Patient-Centered Care: A Proof of Concept Study for a Drain Management and Clinical Follow-Up Database in Interventional Radiology

Presenter: Krishnaveni Parvataneni, Massachusetts General Hospital

Krishnaveni Parvataneni1, Felix Dorfner1, Meredith Preziosi1, Patrick Sutphin1, Christopher Bridge1, Dania Daye2

1Massachusetts General Hospital, Boston, MA, USA

2Harvard University, Cambridge, MA, USA

Introduction: Effective management of abscess drains is essential for patient care and resource utilization in Interventional Radiology.

Hypothesis: Our study aims to create a drain management database using Large Language Models to extract key information from medical notes, reducing reliance on manual chart reviews for abscess drain management in Interventional Radiology.

Methods: In this IRB-approved, HIPAA-compliant study, we identified patients who underwent drain procedures in Interventional Radiology at a quaternary referral hospital. All procedural reports were collected from the EMR. We used the vLLM library to prompt one of three pre-trained publicly-available Large Language Models (LLMs), Qwen1.5-72B-Chat-AWQ, LLAMA-2-13b-chat-hf, and Mistral-7B-Instruct-v0.2, to extract 15 key data fields from the procedural notes using zero-shot prompting (prompting without examples). These fields were Number Drains, Drain Types, Output Volume, Material Quality, Drain Indication, Attending Operator, Drain Location, Is New Drain, Is Abscess Drain, Exchange/Reposition/Upsize, Injection, Removal, Purulent, Fistula, and Persistent Collection. Performance of this new model was compared to traditional NLP algorithms, including spaCy and Regular Expression (RegEx). Model performance was assessed using F1-Score for columns with binary classification, and Accuracy for columns with information extraction tasks.

Results: We identified a total of 1,990 reports that met our inclusion criteria from the period between October 2, 2022, and December 8, 2023. Each model's performance, including processing time and accuracy, was recorded: Qwen took 30 minutes, LLAMA 15 minutes, and Mistral 10 minutes to process 1,990 cases using 4 Nvidia A100 GPUs. Qwen had the highest accuracy (96.89%), outperforming LLAMA (85.10%) and Mistral (94.52%), and was chosen for its superior performance.

Overall, the utilization of LLMs significantly improved the reliability of data extraction from unstructured medical notes compared to traditional NLP methods such as RegEx and spaCy. The Qwen model achieved a notable accuracy of 96.89%, with an F1-Score of 0.9722, compared to 29.33% accuracy and an F1-Score of 0.7246 for the spaCy/RegEx model.

Conclusion: Our study demonstrates that using LLMs for automated data extraction in a clinical setting, particularly in IR abscess drain management, can enhance the efficiency and accuracy of information retrieval from unstructured medical notes. This approach has the potential to improve patient care and optimize resource utilization in Interventional Radiology.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Emerging Technologies; Quality Improvement & Quality Assurance

086 - Leveraging Foundation Model Embeddings in Adapter Training for Radiography Classification

Presenter: Xue Li, University of Wisconsin-Madison

Xue Li1, Jameson Merkow2, Noel Codella2, Alberto Pang2, Naiteek Sangani2, Alexander Ersoy2, Christopher Burt2, John Garrett1, Richard Bruce1, Joshua Warner1, Tyler Bradshaw1, Ivan Tarapov2, Matthew Lungren2,Alan McMillan1

1University of Wisconsin, Madison, WI, USA

2Microsoft, Redmond, WA, USA

Introduction: Foundation models, pre-trained on vast datasets, have revolutionized machine learning by providing transferable embeddings for diverse downstream applications, like image classification, search, and report generation. In this study, we utilize embeddings from both general-purpose and medical domain-specific foundation models to train adapters for the multi-class classification of radiography images, with a focus on tube placement to enhance diagnostic accuracy and efficiency.

Hypothesis: Foundation model embeddings can be effectively utilized to train lightweight adapter models for multi-class classification of radiography images.

Methods: 8,842 radiographs were labeled across seven categories. They were rescaled into a [0, 255] intensity range, and then fed into six foundation models for embedding extraction: DenseNet121, BiomedCLIP, Med-Flamingo, MedImageInsight, Rad-DINO, and CXR-Foundation model. Adapters were trained on the training dataset using traditional machine learning models, including K-Nearest Neighbors (KNN), logistic regression (LR), Support Vector Machines (SVM), random forest (RF), and Multi-Layer Perceptron (MLP). Model parameters were optimized based on performance on the validation dataset, and the trained adapters were subsequently evaluated on the test dataset to assess classification performance.

Results: Mean area under the curve (mAUC) metrics and computational efficiency were computed for each foundation model embedding paired with various adapter models. MedImageInsight achieved the highest mAUC of 93.85% with SVM while Med-Flamingo embeddings performed the worst, peaking at 77.49% with RF. Rad-DINO and CXR-Foundation model embeddings delivered strong results, achieving mAUC values of 91.12% and 89.02%, respectively, both with SVM. DenseNet121 and BiomedCLIP showed moderate performance, with mAUCs of 81.84% and 83.04%. Training and inference times for all adapters were within seconds, except for SVM and RF, which required under one minute for training.

Conclusion: Foundation model embeddings, like MedImageInsight, can effectively train lightweight adapter models for multi-class radiography classification, achieving high accuracy and computational efficiency supporting practical deployment.

Keywords: Artificial Intelligence/Machine Learning; Imaging Research

087 - Leveraging Pre-trained Medical Image Embeddings for Knee Pain Score Prediction: A Comparative Analysis of Vision Transformer and CNN Approaches

Presenter: Mohammadreza Chavoshi, Emory University

Mohammadreza Chavoshi1, Frank Li1, Theo Dapamede1, Bardia Khosravi2, Brandon Price1, Janice Newsome1, Aawez Mansuri1, Rohan Satya Isaac1, Hari Trivedi1, Judy Gichoya1

1Emory University, Atlanta, GA, USA

2Mayo Clinic, Rochester, MN, USA

Introduction: Knee radiographs are the most commonly used imaging modality to assess osteoarthritis. Despite their widespread use, correlating radiographic findings with patient-reported pain remains challenging due to the complex and subjective nature of pain experience. Studies show variable associations between radiographic osteoarthritis severity and pain intensity.

Hypothesis: We hypothesized that general pre-trained medical image embedding extractors can capture subtle radiographic features associated with pain as effectively as task-specific convolutional networks, potentially offering insights into subtle imaging characteristics that correlate with pain reporting.

Methods: Using the eMoRy Knee Radiograph (MRKR) dataset of 83,011 patients (503,261 knee radiographs), we extracted a subset of 17,157 patients (33,138 unilateral knee images) after excluding cases with inflammatory arthritis, recent trauma, infection, or prior arthroplasty. Pain scores (0–10 scale) were collected during clinical care within 7 days of imaging. We compared four deep learning approaches: ConvNeXt (a CNN architecture), and two vision transformer-based embedding extractors - RAD-DINO (trained via DINOv2 self-supervised learning) and BiomedCLIP (evaluated in both image-only and image-text modes). For BiomedCLIP's multimodal analysis, we generated standardized descriptions based on automatically extracted Kellgren-Lawrence grades.

Results: All models demonstrated comparable performance in predicting pain scores (RMSE 2.47–2.60, MAE 2.02–2.16). Detailed subgroup analyses revealed consistent prediction patterns across age groups (#60; 45, 45–70, >70), sex, and race. The pain score distribution was relatively uniform across demographic subgroups, with median scores ranging from 4 to 6. All models showed similar error patterns: slight underprediction for higher age groups and minor variations in prediction bias across sex and racial subgroups. Notably, BiomedCLIP with image and text inputs showed the most balanced error distribution across subgroups, suggesting that multimodal analysis may help mitigate demographic biases in pain prediction.

Conclusion: General-purpose medical image embedding models can match task-specific CNNs in predicting knee pain from radiographs, with consistent performance across demographic subgroups. The balanced error distribution in multimodal analysis suggests potential advantages in combining imaging features with structured clinical information. However, moderate prediction errors across all approaches underscore both the inherent complexity of pain assessment from imaging alone and the need for comprehensive clinical evaluation beyond radiographic findings.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Imaging Research

088 - Leveraging Volume Embeddings for Accurate and Efficient Glioma Classification

Presenter: Xue Li, University of Wisconsin-Madison

Xue Li1, Alan McMillan1

1University of Wisconsin, Madison, WI, USA

Introduction: Foundation models have shown remarkable success in generating embeddings for downstream tasks, like image retrieval and classification. While 2D embeddings are widely utilized, constructing effective representations for 3D datasets often involves aggregating embeddings into volume features. This study explores multiple approaches to construct volume embeddings from brain MRI scans and evaluate their performance in classification of low-grade glioma (LGG) versus high-grade glioma (HGG).

Hypothesis: Volume embeddings are effective and computationally efficient for identifying LGG and HGG from brain MRI scans.

Methods: This study utilized the BraTS 2020 dataset, including 369 subjects with four MRI series (T1, T2, T1CE, and FLAIR). Slice-level embeddings were first extracted from each series using the MedImageInsight foundation model and then aggregated into volume embeddings using various pooling methods, including mean, max, and median. All possible combinations of the series were explored by concatenating their respective volume embeddings. A Multi-Layer Perceptron (MLP) classifier, with a hidden layer size of 200, was trained on 276 subjects, while the remaining 93 subjects were reserved for testing.

Results: Performance was assessed using accuracy, precision, recall, and F1-score. The top section presents the performance using volume embeddings from different MRI series and pooling methods. The best-performing combination of one to four modalities is reported, with the highest accuracy achieved by max pooling on T1CE+FLAIR (96.8% accuracy and F1-score). The bottom section shows the performance of state-of-the-art (SOTA) models trained directly on images, where ResNet-50 attained the best results (98.9% accuracy and F1-score).

Conclusion: This study demonstrates that volume embeddings can accurately distinguish LGG from HGG, with all metrics exceeding 91%. The best embedding-based model lags behind SOTA only by 2.1%, which is inspiring and highlights the exciting promise for the future development of embedding-based methods as an efficient alternative for classification tasks.

Keywords: Artificial Intelligence/Machine Learning; Imaging Research

089 - LLM Showdown: Benchmarking Performance, Cost, and Speed of Granular Ground Truth Extraction in Radiology Reports for Post-Deployment Monitoring of AI Models

Presenter: Aawez Mansuri, Emory University

Aawez Mansuri1, Theo Dapamede1, Hanssen Li1, Wasif Bala1, John Moon1, Bardia Khosravi2, Chad Robichaux1, Frank Li1, Mohammedreza Chavoshi1, Beatrice Brown-Mulry1, Rohan Isaac1, Dan Cohen1, Ninad Salastekar1, Judy Gichoya1, Hari Trivedi1

1Emory University, Atlanta, GA, USA

2Yale University, New Haven, CT, USA

Introduction: Large language models (LLMs) show promise for extracting granular clinical detail from radiology reports, yet their performance, cost, and inference speed relative to smaller models remain unclear. We evaluated large and small open-source and proprietary LLMs—including GPT-4o, GPT-4o-mini, Meta Llama 3.1 8B, Llama 3.1 70B, Llama 3.3 70B, Microsoft Phi 3.5-mini, and Phi 3.5-moe—on extracting granular labels of clinically relevant subtypes from intracranial hemorrhage (ICH) and pulmonary embolism (PE) reports. This includes information on acuity, location/depth, size, and presence of complications. By comparing model outputs to human-annotated ground truths and assessing cost-to-performance trade-offs, we provide insights to guide clinical NLP deployment.

Hypothesis: We hypothesized that proprietary models (e.g., GPT-4o, GPT-4o-mini) would surpass open-source counterparts and that smaller variants (e.g., GPT-4o-mini) would achieve near-equivalent performance to larger models at lower cost.

Methods: We selected 600 radiology reports (300 ICH and 300 PE) each annotated by three radiology residents for ground truth. Models received identical prompts, and outputs were compared to annotations. We measured inference times and evaluated cost-to-performance for proprietary models using token-based pricing.

Results: GPT-4o excelled in extracting detailed ICH labels with variable performance for PE. GPT-4o-mini delivered comparable accuracy at lower cost and faster inference. Among open-source models, Llama 3.3 70B emerged as top performer, exceeding smaller open-source variants (Llama 3.1 8B, Phi 3.5-mini), but at a high computational cost. Overall, GPT-4o-mini offered a strong balance of accuracy, speed, and cost, while most smaller models did not match their larger counterparts’ performance.

Conclusion: This benchmarking shows that proprietary models like GPT-4o lead in label extraction, but smaller, more cost-effective options like GPT-4o-mini achieve nearly the same accuracy at lower cost. Although Llama 3.3 70B performs well among open-source models, its high computational demands limits practical use. Ultimately, selecting an LLM for post-deployment radiology AI monitoring should consider accuracy, cost, and speed, leveraging these insights for more balanced, resource-conscious decisions.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Quality Improvement & Quality Assurance

090 - LLM-Generated Clinical Histories: Evaluating Readability and Efficiency for Radiologic Interpretation

Presenter: Shawn Lyo, University of Pennsylvania

Shawn Lyo1, Satvik Tripathi1, Ali Tejani2, David Vu3, Siddhant Dogra4, Hamza Alizai5, Tessa Cook1

1University of Pennsylvania, Philadelphia, PA, USA

2University of Texas Southwestern Medical Center, Dallas, TX, USA

3Scripps Clinic Medical Group Inc., La Jolla, CA, USA

4New York University, New York, NY, USA

5Children's Hospital of Philadelphia, Philadelphia, PA, USA

Introduction: Radiologists face significant time pressure when interpreting imaging studies, especially during on-call shifts. Reviewing lengthy clinical documentation is time-consuming, with critical information often buried within extensive clinical notes. Large language models (LLMs) have demonstrated capabilities in synthesizing and summarizing medical text, suggesting potential utility in automatically generating focused, relevant clinical histories. However, the ability of LLMs to produce clinically useful and accurate history summaries for radiology interpretation has not been systematically evaluated.

Hypothesis: LLM-generated clinical histories will be more readable and efficient to review than manually extracted clinical documentation while maintaining clinical accuracy.

Methods: A two-stage summarization process was developed using an institutionally approved Azure OpenAI GPT-4o model utilizing ensemble prompting and universal self-consistency techniques. For each of 30 imaging studies, concise clinical histories were generated by synthesizing information from imaging orders, recent clinical notes, and prior imaging reports. Stage 1 generated a detailed timeline-based summary, while Stage 2 produced a single-paragraph condensed summary. Text length, readability scores (Flesch-Kincaid, SMOG, Coleman-Liau), and clinical completeness were analyzed.

Results: The generated clinical histories significantly reduced text while maintaining clinical utility. Source texts averaged 2,021 words. Stage 1 and 2 summaries reduced text length by 82.6% and 93.8% respectively. Increasing levels on readability indices suggest that requisite technical complexity for interpretation was preserved. Preliminary evaluations suggest that essential medical information was preserved.

Conclusion: LLM-based clinical history generation approach effectively reduces text length by over 93% while retaining essential clinical information. This has the potential to enhance radiologist workflow efficiency by minimizing the time spent reviewing histories and ensuring that critical details are accessible rather than buried in extensive clinical documentation. Further studies will evaluate its direct impact on time savings and diagnostic accuracy.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Emerging Technologies; Provider Experience; Quality Improvement & Quality Assurance

091 - Long-Context and Short-Context Supplemented GPT-4 Surpass Base GPT-4 in Adhering to the American College of Radiology Appropriateness Criteria for Ordering Neuroradiology Imaging

Presenter: Yasmine Eichbaum, Thomas Jefferson University

Yasmine Eichbaum1, Julietta Gervase1, Harish Appiakanna2, Rishi Gadepally3, Adam Flanders1

1Thomas Jefferson University, Philadelphia, PA, USA

2Mary Washington Hospital, Fredericksburg, VA, USA

3Hospital Corporation of America Healthcare, Inc., Thousand Oaks, CA, USA

Introduction: The American College of Radiology (ACR) Appropriate Use Criteria (AUC) are extensive, making them difficult to use efficiently. Foundational models have shown promise in providing appropriate recommendations, but the role of context length and order of information remains unclear.

Hypothesis: Compared to the base ChatGPT (GPT 4, 12/2023) which is not supplemented with AUC, our context supplemented versions of ChatGPT will provide more accurate imaging recommendations for neuroradiology clinical vignettes. Models with greater context will outperform ones with shorter context.

Methods: Four experimental models were created through the ChatGPT custom assistant web interface. Two were supplemented with the full AUC corpus (FC), while two were supplemented with tables only (TO). Two versions of each (FC and TO) were made by varying the order of AUC documents provided (FC1, FC2, TO1, TO2). These four, context-supplemented models plus the ChatGPT4 base model each processed fifty-one neurological clinical vignettes. Outputs were scored in accordance with the grading schema, with the final score being an average of the three runs per scenario. Kruskal-Wallis test was used to evaluate performance between models.

Results: All context-supplemented models performed significantly better against the base model in terms of percent correct: FC1 (73%), FC2 (68.8%), TO1 (70.6%), or TO2 (73.3%) versus base (42.6%) (p < 0.001 for each). There were no significant differences between the FC and TO models (p = 1.000) or between the two version orders (p = 1.000).

Conclusion: Customized context models supplemented with AUC guidelines significantly outperformed the base model in providing appropriate imaging recommendations. There was no significant difference between the FC and TO models, nor between the models with varying orders of provided context.

Keywords: Applications; Artificial Intelligence/Machine Learning; Emerging Technologies; Imaging Research

092 - Multidomain Sequence Classification of Brain, Knee, and Prostate MRIs Through 3D-CNN and Gradient Boosted Trees

Presenter: Yu-Cherng Chang, University of Miami

Yu-Cherng Chang 1

1University of Miami Coral Gables, FL, USA

Introduction: Numerous unique MR sequences are utilized in imaging each body part. Given human error and time constraints, DICOMs often contain incorrect image information, slowing radiologists’ workflows but also limiting the accessibility of MR data for machine learning. Automated classification could lessen the manual burden of labeling sequences and improve accuracy. While specialized algorithms for classification of MR images for individual body parts have been reported, an algorithm for multiple body domains is demonstrated in this abstract.

Hypothesis: Classification of MRI sequences from multiple domains can be achieved above the 85% accuracy estimated of DICOM headers in practice (Guld et al, 2002).

Methods: A subset of brain, knee, and prostate MR images were obtained from the publicly available fastMRI dataset (Zbontar et al, 2018). Sequences included 1000 brain axial T1, T1 post-contrast, and T2 images; 1000 knee axial T2 fat suppressed (FS), coronal PD, coronal PD FS, sagittal PD, and sagittal T2 FS images; and 300 prostate axial T2 and ADC images. Images were partitioned 80/20 between training/validation and testing.

Metadata-agnostic sequence classification was performed using a 3D Convolutional Neural Network (CNN)-XgBoost approach. Briefly, images were first run through a 3D CNN based on a ResNet-18 architecture with default ImageNet weights, implemented from an open source library (Soloyev et al, 2021). Output before the final network layer was extracted as a feature vector and subsequently input into gradient boosted decision trees implemented with the XgBoost algorithm, allowing classification into each sequence/body part. Training/validation of the decision trees was optimized with 5-fold cross validation with a grid search to determine the hyperparameters producing the highest validation accuracy.

Results: An accuracy of 91.3% was achieved on the test dataset. The largest error rate occurred in misclassifying brain axial T1 noncontrast versus contrast images.

Conclusion: Accurate automated classification of multidomain MR sequences was demonstrated through a 3D CNN-XgBoost approach.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Imaging Research

093 - Multi-stage Multimodal Deep Learning for Harmonization of Radiology Study Descriptions

Presenter: Zihan Li, University of Washington

Zihan Li 1, Paul Kinahan1

1University of Washington, Seattle, WA, USA

Introduction: Radiology study descriptions within DICOM image headers exhibit high variability, complicating efficient data harmonization and patient cohort selection for pooled collections. This challenge is accentuated by the long-tail distribution of descriptions, where a few categories dominate, leaving others underrepresented. Existing methods, including basic NLP tools, struggle with scalability and accuracy in such contexts.

Hypothesis: A multi-stage multimodal deep learning model can effectively harmonize radiology study descriptions by addressing the challenges of long-tail data distribution and enhancing the reliability of automated categorization.

Methods: We propose MFFNet (Multi-stage Feature Fusion Network) utilizing BERT (Bidirectional Encoder Representations from Transformers) for meta data processing and TotalSegmentator for image analysis. In Stage 1, patient exam-level features are used for coarse-grained classification. Stage 2 builds on these predictions, employing specialized models to refine the classification of LOINC Code. Stage 3 introduces image scan-level features including the image information to re-evaluate low-confidence cases flagged in earlier stages. We also adopt Focal Loss and Weighted Data Sampling to mitigate long-tail distribution issues. A confidence prediction mechanism further improves the reliability of classifications by escalating low-confidence cases for reassessment or expert review.

Results: MFFNet significantly outperformed BERT, achieving near-perfect accuracy across validation categories. The model demonstrated a high ability to handle diverse data, reducing misclassifications from 379 with BERT to one. In Stage 1, MFFNet achieved 100% accuracy in coarse-grained classification with no low-confidence cases. Stage 2 refined these predictions with 99.9% accuracy, identifying 28 low-confidence cases. In the final stage, MFFNet resolved all low-confidence cases, eliminating the need for expert intervention in most scenarios.

Conclusion: MFFNet is a 'Helper AI' tool for radiology study description harmonization that can address the long-tail challenge, thus boosting workflow efficiency for curation and selection of image cohorts for research using large collections of radiological images.

Keywords: Applications; Artificial Intelligence/Machine Learning; Emerging Technologies

094 - Nationwide Analysis of CTPA Yield Rates Using the EPIC Cosmos Database: Insights into Overutilization Patterns

Presenter: Vidya Sankar Viswanathan, Emory University

Vidya Sankar Viswanathan1, Joshua Volin1, Colin Segovis1, Nabile Safdar1, Elias Kikano1

1Emory University, Atlanta, GA, USA

Introduction: Computed Tomography Pulmonary Angiogram (CTPA) is the gold standard for diagnosing pulmonary embolism (PE). Despite its critical role, concerns regarding its overutilization persist. Previous evaluations of CTPA yield rates, often limited to single or a few institutions, have primarily focused on institutional-level data. By examining nationwide CTPA yield rates across a large multi-institutional dataset, this study aims to evaluate variability across different demographic, geographic, and institutional factors.

Hypothesis: By examining nationwide CTPA yield rate trends, the study seeks to provide a more comprehensive understanding of CTPA use patterns, including potential overutilization in specific patient subgroups and regions. These findings could inform targeted interventions, improve adherence to clinical decision-support tools, and ultimately enhance resource utilization and patient outcomes.

Methods: A retrospective cohort study was conducted using EPIC Cosmos, encompassing over 6.4 million CTPA scans performed across 1,614 hospitals in the United States between November 2021 and 2024. Patients included in the study underwent CTPA with an encounter diagnosis of acute pulmonary embolism, and yield rates were calculated as the ratio of positive PE diagnoses to total CTPA scans during an encounter. Subgroup analyses were conducted based on sex, race, age, geographic region, urban versus rural settings, and institutional factors.

Results: The study included a total of 6,465,645 CTPA scans. While the national positive CTPA yield rate was 7.88%, our home institutional yield rate was marginally higher at 8.61%. Yield rates varied significantly by sex, with males exhibiting a higher rate (8.38%) than females (7.47%). Geographic disparities were notable, with the Northeast reporting the highest yield (8.65%) and the South the lowest (7.55%). Urban and rural settings demonstrated nearly identical rates (7.98% and 7.98%, respectively). Minority groups, including Asians and Hispanics, had the lowest yield (6.48%). Yield increased with age, peaking at 9.28% for patients older than 85 years.

Conclusion: This is the first study leveraging the EPIC Cosmos database to evaluate CTPA yield nationwide. The findings highlight the potential overutilization of CTPA, particularly among women, minority groups, and younger patients. These disparities suggest a need for improved adherence to evidence-based guidelines to optimize resource utilization and reduce unnecessary radiation exposure. Addressing these issues through targeted interventions is crucial for improving diagnostic efficiency and patient outcomes.

Keywords: Administration & Operations; Clinical Workflow & Productivity; Quality Improvement & Quality Assurance

095 - On the Feasibility of Chest X-ray Reconstruction from Foundation Model Vector Embeddings

Presenter: Frank Li, Emory University

Frank Li1, Theo Dapamede1, Mohammadreza Chavoshi1, Bardia Khosravi2, Janice Newsome1, Aawez Mansuri1, Rohan Satya Isaac1, Hari Trivedi1, Judy Gichoya1

1Emory University, Atlanta, GA, USA

2Yale University, New Haven, CT, USA

Introduction: Foundation models are large AI systems pre-trained on vast amounts of data that can be efficiently adapted for diverse tasks through zero/few-shot learning fine-tuning. Their advantages include robust transfer learning capabilities, superior generalizability compared to specialized models, and reduced requirements for task-specific data and resources, offering both enhanced capability and cost-effectiveness. However, the potential for foundation models to encode protected health information (PHI) raises privacy concerns. As an initial investigation, we examined the feasibility of reconstructing original chest X-rays (CXRs) from vector embeddings extracted by a CXR-specific fine-tuned foundation model.

Hypothesis: Vector embeddings, being compact yet information-dense representations, should enable the reconstruction of original CXRs with high fidelity.

Methods: Vector embeddings (n=2,486,502) were extracted from a private CXR dataset using RAD-DINO, a foundation model trained exclusively on medical imaging data. A Wasserstein Generative Adversarial Network (WGAN) was developed for image reconstruction using the vector embeddings, where Wasserstein distance was employed for distribution matching, perceptual loss was utilized for semantic feature preservation, and a hybrid L1+L2 loss function was implemented to maintain both structural integrity and fine details. An additional 14,140 images from 2,000 MIMIC CXR patients were randomly selected for external validation.

Results: The reconstruction quality, measured by Frechet Inception Distance (FID), achieved a high distance of 102.42 between the original and reconstructed MIMIC CXRs, implying challenges to reconstruct images to their original form. Nonetheless, visual inspection of reconstructed images demonstrated preserved major anatomical features and view positions, while exhibiting slightly reduced contrast and enhanced smoothness. While the reconstruction process partially preserved some burned-in markers and annotations, these elements became indistinct and unidentifiable in the reconstructed images.

Conclusion: Our study demonstrates the feasibility of reconstructing images from vector embeddings, even for images unseen by the WGAN but used in RAD-DINO's training, revealing a potential of exposing training data used to train foundation models. We demonstrate that the burned-in text and markers are partially reconstructed but rendered unrecognizable, which is important to maintain fidelity of data anonymization. The ability to view some form of images can aid in model auditing especially when foundation models are trained on multimodal datasets as well as ablation studies.

Keywords: Artificial Intelligence/Machine Learning; Imaging Research

096 - Performance and Adjustment of an AI-Assisted Fracture Detection Tool in a Real-World Post-Clinical Deployment Setting

Presenter: Sameed Khan, Cleveland Clinic

Sarang Ingole1, Charit Tippareddy2, Sameed Khan3, Amy Zhou4, Kacey Pagano5, Mia Zivkovic4, Orlando Martinez4, Navid Faraji2

1CARPL.ai, New Delhi, Delhi, India

2University Hospitals, Cleveland, OH, USA

3Cleveland Clinic, Cleveland, OH, USA

4Case Western Reserve University, Cleveland, OH, USA

5Northeast Ohio Medical University, Rootstown, OH, USA

Introduction: AI fracture detection tools high performance in research settings is established but their effectiveness after deployment in clinical practice is understudied. Post-deployment analysis offers the opportunity to study the tool before and after on-the-fly modifications to operational parameters based on observed performance “in the wild.”

Hypothesis: Fracture detection tool sensitivity and specificity will be similar to levels assessed in preclinical studies (86.5% and 82.6% respectively). Corrective modifications will improve performance.

Methods: This study analyzed 2378 appendicular trauma studies referred to a tertiary ED following deployment of the fracture detection tool. After corrective modification to increase threshold points A (low-suspicion detection) and B (high-suspicion detection), a further 2023 patients were analyzed. All radiograph reports were reviewed by two board-certified radiologists. Performance was evaluated before and after modifications to the AI tool. Significance testing was performed via chi-square test for proportions.

Results: Initially, the AI tool achieved a sensitivity of 89.5%, specificity of 76.0%, accuracy of 79.6%, and a negative predictive value (NPV) of 95.0%. The precision was 0.58, and the F1 score was 0.71. Notably, 93.2% of false positives came from the low-suspicion group. After refining the tool, sensitivity increased significantly to 94% (p=0.008), specificity to 87% (p=1.9 x 10–15), and accuracy to 88.7% (p=3.4 x 10–16), with NPV of 97.7% (p=0.004). Precision improved significantly to 0.71 (p=8.72 x 10-8), and the F1 score to 0.81. Common false negatives involved fractures near complex joints or in cases of severe osteoporosis, while false positives were often associated with misidentified sesamoid bones, artifacts, and external hardware.

Conclusion: Initially, the AI tool demonstrated high sensitivity but a high rate of false positives. Refining the tool to adjust thresholds led to improved sensitivity, specificity, and overall performance, highlighting the importance of ongoing AI tool refinement for clinical deployment.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Imaging Research; Quality Improvement & Quality Assurance

097 - Practical Use of LLMs to Adjudicate Radiology Reports to Assess Performance of a Clinical AI Triage Tool

Presenter: Adam Flanders, Thomas Jefferson University

Adam Flanders1, Paras Lakhani1, Prahlad Menon1, Robyn Ball2, George Shih3, Luciano Prevedello4

1Thomas Jefferson University, Philadelphia, PA, USA

2The Jackson Laboratory, Bar Harbor, ME, USA

3Cornell University, New York, NY, USA

4The Ohio State University, Columbus, OH, USA

Introduction: The majority of commercial computer vision vendors provide limited data to customers to evaluate local performance of deployed solutions. Legislation mandates that customers monitor performance of these systems for drift/bias yet the effort needed to monitor these systems at scale is not trivial and can require substantial resources. Large language models (LLMs) applied to the diagnostic report may be useful in automating this process.

Hypothesis: To determine whether an ensemble of LLMs can be used effectively to monitor a commercial triage AI system.

Methods: The AI inference results for a commercial CT brain hemorrhage detector and the diagnostic reports were retrospectively collected on 16,172 ED and outpatient exams derived from 18 hospitals and 35 CT scanners in two states. The diagnostic reports were parsed to extract only the impression section and were presented to an ensemble of five LLMs (llama3.2:1b, llama3.2:3b, codellama:7b, llama3.1:8b, granite3-dense:2b) employing a single-shot prompt to confirm if presence of hemorrhage was documented in the report. The results returned presence of hemorrhage and if hemorrhage was present, the hemorrhage subtype. Each LLM was compared to the consensus (majority vote) across LLMs and to a subset of manually reviewed exams. Finally, the LLM consensus was compared to the manually reviewed exam set to assess overall performance.

Results: Agreement among the five LLMs varied considerably, with kappas ranging from 0.1 to 0.79. The smallest LLM (nlpheme1) performed substantially worse compared to both the LLM consensus and the manually reviewed exam set. Similarly, there was a wide range of Cohen's kappa (0.16–0.9) comparing each LLM to the consensus with the smaller sized model as the outlier. Comparison of the consensus to a random subset of 390 manually reviewed reports showed agreement of 0.68 which was augmented to 0.75 after removal of the two low performing LLMs. There was no substantial difference in F1 score for the ICH AI model using the two consensus schemes or the single best performing LLM.

Conclusion: A consensus ensemble of LLMs reviewing reports may have promise in verifying AI model performance in an automated fashion. Model selection, custom prompt engineering and manual verification are critical in ensuring useful results.

Keywords: Administration & Operations; Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Enterprise Imaging; Quality Improvement & Quality Assurance

098 - Prioritiz-IR - Design and Evaluation of Evaluation of an LLM-PhoneBot for Triaging Interventional Radiology Consults

Presenter: Hanzhou Li, Emory University

Hanzhou Li1, John Moon1, Wasif Bala1, Christopher Williamson1, Emily Patel1, Sampath Kumar2, Rayan Khan3, Syed Ali Mehdi3, Rithvik Pandibabu3, Varshith Chekka3, Aneesh Sabarad3, Hari Trivedi1, Zachary Bercu1, Janice Newsome1, Judy Gichoya1

1Emory University, Atlanta, GA, USA

2Temple University, Philadelphia, PA, USA

3Georgia Institute of Technology, Atlanta, GA, USA

Introduction: Fielding interventional radiology (IR) consult phone calls is essential; however, nonurgent calls can be extremely disruptive, with radiologists receiving an estimated 72 to 104 calls during a typical 12-hour shift. Moreover, interruptions paired with frustration may hinder effective communication and amount to additional calls to clarify or acquire more details. We developed a proof-of-concept, large language model (LLM)-based phonebot triage system (Prioritiz-IR), designed to triage consult phone calls and capture essential information. Through dynamic, human-like interactions, Prioritiz-IR filters non-urgent calls, prompts callers for critical details, and follows up on missing information. The collected data is summarized and displayed on a web-based user interface for provider review.

Hypothesis: We hypothesize that an LLM-based phonebot system can effectively triage interventional radiology (IR) consult phone calls by accurately capturing key information such as callbacks and patient identification.

Methods: Prioritiz-IR integrates an artificial intelligence (AI) phone agent (Bland AI, San Francisco, CA) and GPT-4 (OpenAI, San Francisco, CA) with a front-end Next.JS (Vercel, San Franscisco, CA) web application. A text-to-speech application (Luvvoice, New York, NY) was employed to simulate 15 non-urgent IR consultations with Prioritiz-IR. Prioritiz-IR was specified to capture the call priority, caller name, hospital name, department, call back number, patient name, age, medical record number (MRN), and consult summary, which comprises of the request reason and a one-liner summary of patient history. The accuracy of the captured information was verified through manual review by three investigators.

Results: Of the captured information, 100% accuracy was achieved for priority, caller name, hospital name, department, and patient age. One MRN had a minor formatting discrepancy, omitting a leading zero despite it being explicitly stated. All call back numbers were captured accurately, although 6 out of 15 were formatted as plain numbers as opposed to phone numbers. In 5 out of 15 instances, the caller’s title (e.g., nurse or doctor) was included alongside their name correctly although not asked to, and one patient’s name was misspelled as “Stephen” instead of “Steven.” For consult summaries, 7 out of 15 were missing key pieces of relevant medical history despite correct capture of reason for consult.

Conclusion: Prioritiz-IR accurately captures key elements of an IR consult, with minor formatting inconsistencies that do not hinder callbacks or patient identification. However, further work is needed to summarize consults effectively, ensuring the inclusion of relevant medical history without excessive details.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Provider Experience

099 - RadALIGN: Mitigating Hallucinations and Improving Reliability in AI-Generated Radiology Reports

Presenter: Krish Malik, Parkland High School

Krish Malik 1, Vedant Malik1

1Parkland High School, Allentown, PA, USA

Introduction: AI is transforming the field of radiology by enhancing multi-page report summarization complimenting radiologists. However, significant challenges with AI generated reports remain like speculative language, inaccuracies, deviations from factual content and lack of deep domain knowledge. LLMs can generate speculative conclusions, under adversarial conditions leading to flawed clinical decisions. RadALIGN framework aims to address these challenges by aligning LLMs, ensuring nonspeculative language, mitigating unwarranted assumptions to produce accurate, reliable and trustworthy radiology reports, thereby reducing hallucinations.

Hypothesis: If model alignment is designed to address issues like speculative language, factual inaccuracies, and compliance gaps with radiology-specific standards, then it will significantly enhance reliability, accuracy, and trustworthiness of AI-generated radiology reports.

Methods: A three-step framework using Indiana University Chest X-ray Radiology reports as input data with Llama3.2 3B Large Language Model (LLM) was implemented:

1.Baseline Generation: Establish baseline AI generated summary under standard conditions without model alignment: Reports generated using pre-trained LLM served as reference to identify factual gaps, clarity, speculative language.

2.Red Teaming: Identify vulnerabilities in AI generated report summaries.

3.Model Alignment: Align the AI based LLM model to generate safer, compliant, and factually accurate summaries.

Results: Model Alignment reduced Perplexity (prediction uncertainty) across all summarized reports, improving fluency and consistency. Perplexity dropped from range of 19.2655–71.0075 in the Baseline Model (unaligned) to 5.1829–6.2896 range with Model Alignment.

Token Diversity (vocab variability) decreased from range of 0.7778–1.0 to 0.5595–0.7099 range, restricting imagination, reducing hallucinations, and ensuring factual outputs. Red Teaming increased both Token Diversity (0.26 to 0.36–0.41) and Perplexity (1.1 to 2.1–2.6) exposing vulnerabilities in Baseline model generated radiology reports, which were mitigated through Aligning the Model.

Conclusion: Model alignment significantly improved output robustness and factual consistency by reducing perplexity, leading to more confident and factual radiology report summary generation. By addressing vulnerabilities exposed through Red Teaming, Model Alignment reduced variability and hallucinations, ensuring outputs aligned with radiology standards for reliability and safety.

Keywords: Applications; Artificial Intelligence/Machine Learning; Emerging Technologies

100 - Re-evaluating Right Heart Strain at CT Pulmonary Angiography with Artificial Intelligence-Driven Cardiac Chamber Segmentation

Presenter: Bardia Nadim, Brigham and Women's Hospital

Bardia Nadim1, Ian Pan1

1Brigham and Women's Hospital, Boston, MA, USA

Introduction: Pulmonary embolism (PE) is a life-threatening condition and the third most common cardiovascular disease. Prompt diagnosis is critical. Computed tomography pulmonary angiography (CTPA) is the mainstay of diagnosis, allowing for PE detection and assessment of right heart strain (RHS) by measuring the right ventricular (RV) to left ventricular (LV) diameter ratio. An RV:LV < 1 is considered normal; however, manual 2D assessment is prone to interobserver variability. Artificial intelligence (AI) offers promising automated cardiac segmentation methods. This study evaluates the concordance of AI volumetric segmentation with manual RV:LV ratings and explores AI's value in assessing RHS on CTPA.

Hypothesis: AI volumetric cardiac segmentation will demonstrate high concordance with manual assessments of the RV:LV on CTPA and offer a standardized evaluation of RHS.

Methods: 7,111 CTPAs from the RSNA PE Detection Challenge underwent automated cardiac segmentation by the publicly available TotalSegmentator algorithm. Manual review of a randomly selected subset demonstrated 98% of segmentations were satisfactory. RV:LV was calculated using: 1) the greatest short-axis chamber diameter (2D) and 2) segmentation volumes (3D). 2D RV:LVs were binarized, and concordance with the provided manual ratings of RV:LV ≥1 was calculated using Cohen’s kappa for PE-positive examinations. Two-sample t-test compared AI-driven ratios between manual groups. Pearson correlation coefficient compared the 2D and 3D AI ratios. One-way ANOVA compared AI ratios between negative PE, acute peripheral PE, and acute central PE examinations.

Results: 2,314 examinations were positive for PE. The 2D AI ratio means for the manual RV:LV < 1 or ≥1 groups were 0.92 and 1.14, respectively (p < 0.001). Cohen’s kappa between 2D AI ratio ≥1 and manual ratio ≥1 was 0.492. The 2D AI ratio means for negative and positive PE groups were 0.95 and 1.01, respectively (p < 0.001), whereas 3D ratio means were 1.36 and 1.55, respectively (p < 0.001). The Pearson correlation coefficient between the 2D and 3D ratios was 0.8. One-way ANOVA of negative PE (n = 5,380, mean = 0.954), acute peripheral PE (n=1,408, mean=0.98), and acute central PE (n=323, mean=1.12) examinations was significant (p < 0.001).

Conclusion: AI-driven cardiac segmentation is accurate and reproducible, though only fair-to-moderate concordance was seen between AI and manual RV:LV. There was high correlation between 2D and 3D AI RV:LV, indicating that greatest short-axis can serve as a proxy for volume; however, 2D methods overall underestimated RV:LV. Patients with acute central PE had the highest RV:LV, indicating that this may serve as a measure of RHS and PE severity.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Quality Improvement & Quality Assurance

101 - Report Interpretation Times for Radiology Residents and Attendings Before and After Deployment of a Triaging Model for Pulmonary Embolism and Intracranial Hemorrhage Detection in a Large Academic Center

Presenter: Aawez Mansuri, Emory University

Aawez Mansuri1, Theo Dapamede1, Hanssen Li1, Wasif Bala1, John Moon1, Bardia Khosravi2, Chad Robichaux1, Frank Li1, Mohammedreza Chavoshi1, Beatrice Brown-Mulry1, Rohan Issac1, Dan Cohen1, Ninad Salastekar1, Judy Gichoya1, Hari Trivedi1

1Emory University, Atlanta, GA, USA

2Yale University, New Haven, CT, USA

Introduction: Artificial intelligence (AI) models for detection of critical findings, such as pulmonary embolism (PE) and intracranial hemorrhage (ICH), hold potential to improve workflow efficiency. This is traditionally measured in turnaround time from exam completion, but little data exists on actual interpretation time by radiologists. In this project we evaluate the impact of an FDA-approved AI triage model for PE and ICH deployed in a large academic center for one year on report interpretation times for radiology residents and attendings, comparing pre- and post-deployment periods for both PE and ICH cases.

Hypothesis: We hypothesize that an FDA-approved triage model for ICH and PE reduces exam interpretation times for both attending radiologists and residents.

Methods: We performed a retrospective analysis comparing exam interpretation time (EIT) before and after implementation of an FDA-cleared AI tool for ICH detection on non-contrast head CT (NCCT) and PE detection on CT angiography of the chest (CTPA). The pre-AI period was April 17, 2022–April 16, 2023, and the post-AI period was April 17, 2023–April 16, 2024. The number of cases analyzed for attending radiologists and residents in each period is summarized. EIT time was defined as the timestamp from when the report was initially open to the time the report was signed (for an attending) or prelimed (for a resident). EIT was compared pre- vs. post-AI for attending radiologists and residents for each exam type. Statistical significance was assessed using the Kruskal-Wallis test.

Results: For NCCT, attending radiologists median EIT decreased from 193.81s (IQR: 208.36) pre-deployment to 182.63s (IQR: 212.57) post-deployment, while resident median EIT decreased from 282.06s (IQR: 200.06) to 271.32s (IQR: 204.44). For PE cases, attendings median EIT decreased from 351.73s (IQR: 207.09) to 341.61s (IQR: 195.23), and residents’ from 390.58s (IQR: 214.72) to 374.42s (IQR: 217.66) following implementation. All decreases were statistically significant (p < 0.05).

Conclusion: AI-assisted interpretation yielded statistically significant but modest reductions in exam interpretation time for both attending radiologists and residents across for NCCT and CTPA. Our results are consistent with other published studies showing minimal time savings to the radiologist. Reduction in cognitive burden may be a greater benefit; however, it is difficult to measure. Our ongoing work will evaluate model performance and EIT in exams that are positive and negative for ICH and PE to further delineate benefit.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Educational Systems

102 - Retrospective Comparison of Fracture Detection Performance Before and After Implementation of an AI Fracture Detection Tool

Presenter: Charit Tippareddy, University Hospitals

Charit Tippareddy1, Sarang Ingole2, Sameed Khan3, Kacey Pagano4, Amy Zhou5, Mia Zivkovic5, Orlando Martinez5, Navid Faraji1

1University Hospitals, Cleveland, OH, USA

2CARPL.ai, New Delhi, Delhi, India

3Cleveland Clinic, Cleveland, OH, USA

4Northeast Ohio Medical University, Rootstown, OH, USA

5Case Western Reserve University, Cleveland, OH, USA

Introduction: AI tools have been widely implemented in radiology departments promising improved efficiency, accuracy, and workload management of growing imaging volumes. This study evaluated a fracture detection tool’s effect on emergency worklist prioritization, resident sensitivity, specificity, and concordance of fracture detection.

Hypothesis: AI tool implementation will decrease time-to-first-read and increase fracture detection sensitivity, specificity and resident concordance with attending read.

Methods: The study analyzed 2159 patients with extremity radiographs, 1516 of which contained both resident and attending-authored reports. Time from exam completion to interpretation was collected. Studies with time-to-final-report greater than 120 minutes from exam completion were excluded as statistical outliers. Final attending report was used as ground truth. Resident concordance, specificity, sensitivity, and time-to-first-read were assessed before and after implementation of the AI fracture detector. Statistical significance was determined using a generalized linear mixed effects model that adjusted for scan anatomy (shoulder, humerus, elbow, wrist, hand, femur, ankle, or foot) and resident experience as fixed effects and the resident identity as a random effect. Bonferroni multiple comparisons adjustment was performed separately for each model across all coefficients.

Results: The average time from scan completion to initial interpretation was 38.0 minutes, non-significantly changed from 38.3 minutes before tool implementation (p = 0.10). Resident concordance did not differ, increasing from 94.1 percent before implementation to 95.2 percent (p = 0.36). However, resident fracture detection sensitivity, before and after adjusting for resident experience, increased from 83.7 percent to 93.1 percent (p = 4.76 x 10-4). Resident fracture specificity decreased from 98.5 percent to 96.0 percent (p = 0.06).

Conclusion: Software implementation did not affect time to first or final interpretation, likely due to de-prioritization of radiographs compared to other modalities and case types. However, the tool augmented resident sensitivity, indicating that it aids residents in identifying subtler fracture findings.

Keywords: Applications; Artificial Intelligence/Machine Learning; Imaging Research

103 - Robust Uncertainty-Informed Glaucoma Classification Under Data Shift

Presenter: Homa Rashidisabet, University of Illinois

Homa Rashidisabet1, R. V. Paul Chan1, Thasarat Vajaranant1, Darvin Yi1

1University of Illinois, Chicago, IL, USA

Introduction: Glaucoma is one of the leading causes of irreversible blindness globally. Deep learning (DL) has emerged as a promising approach for the automated diagnosis of glaucoma. However, challenges persist in translating these advancements to clinical settings. Conventional DL classification methods often exhibit overconfidence and lack robustness when faced with a shift in training data distribution, posing challenges in out-of-distribution (OOD) scenarios. These issues raise concerns about the suitability of current glaucoma DL diagnostic algorithms for real-world clinical deployment, potentially impacting patient safety.

Hypothesis: Our proposed approach, centered around uncertainty quantification, aims to effectively identify OOD samples, thereby enhancing the reliability of glaucoma predictions.

Methods: We proposed a novel DL framework, the Dirichlet model, for joint binary glaucoma classification and OOD detection. Our method incorporates uncertainty quantification, allowing the model to express uncertainty in its predictions, thereby providing a more reliable glaucoma assessment and addressing the overconfidence commonly seen in standard DL models. We compare the OOD detection performance of the Dirichlet model to the standard softmax-based DL approach. Trained on 712 fundus images from the Illinois Eye and Ear Infirmary, we evaluate both glaucoma classification and OOD detection on the RIMONE-DL and O-RIGA fundus datasets, as well as non-medical CIFAR-10 and Fashion-MNIST datasets.

Results: The Dirichlet model consistently outperforms the softmax model in OOD detection by 9.5% to 27.5% across datasets. Dirichlet achieved 82.1% and 78.7% AUC for detecting RIMONE-DL and O-RIGA glaucoma fundus datasets as OOD, with a strong performance of 100.0% AUC in detecting CIFAR-10 and Fashion-MNIST non-fundus datasets. Dirichlet maintains comparable glaucoma classification (AUC: 78.6% [78.2%, 79.1%], 54.3% [53.2%, 55.0%]) compared to softmax (AUC: 78.5% [78.2, 79.0], 60.2% [59.2%, 61.2%]) on RIMONE-DL and O-RIGA datasets.

Conclusion: The study demonstrates the effectiveness of our proposed uncertainty-aware Dirichlet model in OOD detection and glaucoma classification tasks across diverse domains, extending its utility beyond the initial training dataset. Furthermore, the incorporation of uncertainty scores in our model alerts users to instances where the model lacks sufficient information for a confident decision.

Keywords: Artificial Intelligence/Machine Learning

104 - Sharpness Matters. Rethinking the Impact of Image Resolution on Medical Image Classification

Presenter: Aditya Vikas Kulkarni, St. Jude Children's Research Hospital

Aditya Vikas Kulkarni1, Paul Yi1

1St. Jude Children's Research Hospital, Memphis, TN, USA

Introduction: Deep learning (DL) models commonly downsample medical images to reduce computational cost, but subtle pathologies may require higher-resolution inputs. Although previous work explored resolution effects in chest x-ray (CXR) DL classification, it did not evaluate explainability or out-of-distribution (OOD) generalizability with external datasets – key features of safe and trustworthy AI. We thus ask: Does increasing image resolution for training DL CXR classifiers improve not only in-distribution accuracy, but also explainability and OOD generalizability?

Hypothesis: Higher-resolution training improves OOD generalizability and explainability.

Methods: We trained DenseNet-121 classifiers on the SIIM-ACR Pneumothorax dataset (n=10,675 ) and set aside 10% for holdout testing. Six resolutions (64×64 to 1024×1024) were used for model training, each tuned via an independent hyperparameter search. The final models were validated through 5-fold cross-validation and tested on held-out and external OOD (n=726; 226 pneumothoraces) test-sets . For assessing explainability, Grad-CAM saliency maps, thresholded to binary masks were compared to radiologist-performed segmentations; mean IoU and percentage overlap (%-IN) were used to measure localization. AUROCs and saliency localization were compared between models using Delong’s and Wilcoxon Signed-Rank tests, respectively.

Results: As image resolution increased, both internal and external performance improved, ranging from AUROC 0.90 and 0.68 for 224x224 (internal and external, respectively) to 0.97 and 0.84 for 1024x1024 (p < .001, all). Generalizability was best for higher resolutions ( < 0.08 AUROC drop for 768x768 and 1024x1024 vs. ~0.2 drop for 64x64 and 128x128). Localization quality followed similar trends, with significantly higher mIoU and %-IN for higher image resolutions in internal and external test sets that were qualitatively more reliable.

Conclusion: Training DL classifiers with higher resolutions enhances OOD generalizability and explainability, alongside in-distribution accuracy. While radiology DL often uses lower resolutions (e.g., 224×224), our findings support adopting higher resolutions to maximize accuracy, trust, and safety, thus advancing more reliable clinical AI deployment.

Keywords: Artificial Intelligence/Machine Learning; Emerging Technologies; Imaging Research; Quality Improvement & Quality Assurance; Security; Storage

105 - Social Media for Accessible Education to Global Trainee Audiences: AI in Neuro-oncology

Presenter: Raisa Amiruddin, Children's Hospital of Philadelphia

Raisa Amiruddin1, Nazanin Maleki1, Nikolay Yordanov2, Pascal Fehringer3, Athanasios Gkampenis4, Anastasia Janas5, Ahmed Moawad6, Sanjay Aneja7, MingDe Lin7, Spyridon Bakas8, Mariam Aboian1

1Children's Hospital of Philadelphia, Philadelphia, PA, USA

2Sofia University St. Kliment Ohridski, Sofia, Bulgaria

3Friedrich Schiller University Jena, Jena, Germany

4University of Ioannina, Ioannina, Greece

5Charité – Universitätsmedizin Berlin, Berlin, Germany

6Mercy Fitzgerald Hospital, Darby, Pennsylvania

7Yale University, New Haven, CT, USA

8Indiana University, Bloomington, IN, USA

Introduction: Social media has helped build & share content in the medical community. Being part of users’ daily routines, it promotes engagement in learning & inspires interactions between learners & mentors.

Hypothesis: We sought to create a resourceful educational platform for the ASNR MICCAI BraTS Challenge, a landmark community benchmark event for brain tumor segmentation & analysis.

Methods: We launched a social media initiative in 2023 on X, LinkedIn, Instagram & WhatsApp to help learners build foundational knowledge of neuroimaging. We leveraged these platforms to host our lecture series, & upload them onto YouTube, summarize concepts as bite-sized information, relay the process of image annotation & algorithm development & guide learners towards open access resources for education. We utilized Hootsuite, Iconosquare & Chatilyzer to calculate metrics & analyze the impact of our online presence.

Results: We published 55 posts on X through 3 iterations (2023, 2024, 2025) of BraTS since February 2023 & garnered 112,460 impressions, with average engagement rate of 3.96% per post. The highest engagement rate for a post is 8.97%. There has been 264% growth of users viewing our content, 478% growth of users amplifying our posts, and audience growth rate of 8.45% since the day of broadcast of the 2025 Challenge. We published 8 LinkedIn & 8 Instagram posts for the 2025 Challenge, having average engagement rates of 11% & 2.99% per post respectively. We set up a WhatsApp community to support student coordinators through the 2025 image annotation pipeline with faculty annotators. We have exchanged 495 messages here, with an average of 12 messages/day. This annotation protocol is a unique opportunity for students to interact, learn & develop professional relationships with faculty while also simulating the experience of an MR image analysis internship.

Conclusion: Our initiative is using a dynamic learning framework to foster an expanding network of committed learners and mentors and increase awareness on the necessity of meticulously curated public databases for the development of precise and accurate algorithms for brain tumor analysis. This approach has become valued by our internationally distributed community of students and faculty for enhancing educational and professional growth in the world of neuroimaging and acknowledging their contributions to a revolutionary undertaking in the domains of open science and artificial intelligence.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Emerging Technologies; Enterprise Imaging; Quality Improvement & Quality Assurance

106 - Synergy of Harnessing a Pixel-AI Tool and NLP-based Patient Care Coordination Tool in the Diagnosis and Management of Brain Aneurysm

Presenter: Shlomit Goldberg-Stein, Northwell Health

Shlomit Goldberg-Stein1, Jonathan Scheiner1, Pina Sanelli1, Matthew Sagnelli1, Ritesh Patel1, David Hirschorn1, Matthew Barish1

1Northwell Health, New Hyde Park, NY, USA

Introduction: Integrating AI tools into clinical workflows requires change management strategies to ensure effective adoption and maximize benefits. This abstract describes the process employed during the implementation of a natural language processing (NLP) tool for patient care navigation of brain aneurysms, deployed synergistically with a pixel-based AI tool augmenting radiologist aneurysm detection on CTA.

Hypothesis: A structured, multi-phase approach was utilized, including the identification of a clinical gap in care, early multidisciplinary stakeholder engagement, selection of clinical champion, and implementation of pre- deployment trial with iterative feedback loops to assess operational impact and address concerns preemptively. Multiple consensus-building meetings, workflow mapping exercises, and training sessions were performed pre-deployment to tailor the navigation tool for specific clinical needs. Data integration was performed to maximize appropriate clinical capture.

Methods: Barriers included operational concerns such as high-volume workloads, alert fatigue, staffing shortages, and the need for customized data integration with the EHR for patient selection and to calculate downstream return on investment. Installation delays and the need for ongoing training also presented challenges. These barriers were addressed through proactive communication, appropriate resource allocation, and ongoing collaboration with stakeholders.

Results: Early data from NLP-based patient navigation implementations demonstrate successful integration into clinical workflows, appropriate patient capture, effective navigation to clinic, and high downstream ROI including additional new neurovascular procedures, surgeries, and follow-up imaging. In 6 weeks, 57 new aneurysms identified by AI alone (missed by radiologists) were navigated to care, resulting in 40 office visits, 40 follow up CT/MR, 12 diagnostic angiograms, and 2 neurosurgical procedures. The synergistic effects of combining patient navigation with a diagnostic AI pixel-based tool are demonstrated.

Conclusion: Successful AI implementation depends on optimizing change management processes. Our structured approach, emphasizing collaboration, customization, and EHR data integration, can be generalized to other AI implementations, facilitating further adoption and maximizing the benefits of AI in care delivery.

Keywords: Applications; Artificial Intelligence/Machine Learning; Imaging Research; Quality Improvement & Quality Assurance

107 - The Effects of Prompt Engineering and Consensus on Identification of Incidental Breast Findings by Large Language Models from Radiology Reports

Presenter: Benjamin Rush, University of Wisconsin

Benjamin Rush1, Thanh Nguyen1, Kayla Berigan1, Ryan Woods1, John Garrett1

1University of Wisconsin, Madison, WI, USA

Introduction: Radiology reports of imaging scans are unstandardized, leading to about 36.6% of incidental findings not receiving follow-up within 1 year. In addition, 7% of chest CT scans have incidental findings, of which 28% are malignant. Large language models (LLMs) can analyze reports and potentially identify cases for follow-up. Experiments testing LLMs’ capabilities of analyzing reports exist yet test few types of LLMs and prompts. The effects of incremental prompts, LLM size, and LLM training data remain largely unexplored. We compared the performance of identifying incidental breast findings from radiology reports by using multiple prompts for individual LLMs and by consensus on cases between LLMs.

Hypothesis: Optimizing prompting for incidental breast findings in radiology reports detection using consensus approaches can improve sensitivity and specificity.

Methods: We randomly selected 500 exams with “breast” in the radiology report from chest CTs obtained at our institution between 2015-2017 from female patients ages 40–72. We compared the performance of 126 combinations from 7 LLMs, 9 incremental prompts, and 2 LLM roles when identifying incidental breast findings in reports compared to a breast imaging fellow reader. Combinations and consensus between LLMs were evaluated by sensitivity, positive predictive value (PPV), specificity, negative predictive value (NPV).

Results: The reader identified 31 (6%) cases with incidental breast findings while individual LLM combinations ranged from identifying 98–478 cases with 0.67–1.00 sensitivity, 0.05–0.84 specificity, PPVs not exceeding 0.23, and NPVs above 0.95. Consensus on case identification reduced false positives: the consensus of 3 highly sensitive combinations identified 86 (17%) cases with 0.87 sensitivity, 0.31 PPV, 0.87 specificity, and 0.99 NPV.

Conclusion: Consensus from highly sensitive LLM, prompt, and role combination generally increased performance, though an optimization algorithm selecting high sensitivity combinations with dissimilar positive case labelling would likely further increase performance.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Imaging Research

108 - Using Few-shot Prompting of LLMs and icd-10 Codes for More Accurate and Readable Patient Friendly Explanations (pfx) for Incidental Findings on MRI for More Equitable Results Delivery to Patients

Presenter: Joseph Hentel, Briarcliff High School

Joseph Hentel1, Kurt Teichman2, Sophie Shih3, David Nauheim2, Jacob Kazam2, Martin Prince2, Dana Galizia2, James Ledoux2, Keith Hentel2, George Shih2

1Briarcliff High School; Briarcliff Manor, NY, USA

2Cornell University, New York, NY, USA

3Stuyvesant High School, New York, NY, USA

Introduction: Patient friendly explanations (PFx) of radiology findings that are accurate and easily understood may result in better understanding and improved health equity for patients. LLMs can greatly facilitate the generation of these PFx but given their stochastic nature and possibility for hallucination, they require expert review and proofreading for each PFx which is not easily done for every report.

Hypothesis: We used a few-shot prompting LLM workflow to generate PFx, based on examples of manually created PFx using the same style, tone, reading grade level, and also check accuracy by comparing ICD-10 codes of PFx with incidental finding ICD-10 codes. This LLM workflow minimizes the need for editing and proofreading.

Methods: We used the 25 most common incidental findings on MRI (brain, neck, chest, abdomen, pelvis) based on literature review (eg, cavernous hemangioma), to generate PFx using two prompting methods: (1) Zero shot prompting (2) Few shot prompting where we provide 5 examples of manually created and edited PFx and asked the LLM (GPT-4o-2024-05-13) to follow the length, style, and grade-level of those human created PFx. We also prompt the LLM to provide ICD-10 codes for the PFx corresponding to incidental finding as a way to check for accuracy of the PFx.

Results: The PFx examples used in few-shot prompting had 64.6 average number of words and 3.2 average number of sentences. For zero shot prompting vs. few shot prompting, the average number of words per PFx was 10.8 vs. 55.8 and the average number of sentences per PFx was 1.0 vs. 3.0.

Conclusion: Accurate patient friendly explanations (PFx) of incidental findings on MRI studies that can be easily and accurately generated by LLMs using few-shot prompting based on patient needs, may avoid patient confusion and anxiety, and result in better understanding and more equitable health delivery for patients with differing education and backgrounds including for non-English speakers. Future radiology reports will likely contain patient friendly explanations of relevant findings that are customized to each patient, for improved health equity.

Keywords: Artificial Intelligence/Machine Learning

109 - Using Open-Source Large Language Models to Extract Labels from Radiology Reports for Machine Learning Applications: A Proof-of-Concept Study

Presenter: Michael Fei, Medical, Creighton School of Medicine

Michael Fei1, Satvik Tripathi2, Krishnaveni Parvataneni3, Felix Dorfner3, Christopher Bridge3, Dania Daye4

1Creighton School of Medicine, Omaha, NE, USA

2University of Pennsylvania, Philadelphia, PA, USA

3Massachusetts General Hospital, Boston, MA, USA

4Harvard University, Cambridge, MA, USA

Introduction: When developing deep learning models, labeled data is expensive and/or time-consuming to obtain, presenting as one of the largest barriers. Open-source Large Language Models (LLMs) present as a tool to extract labels from radiology reports cheaply and efficiently. This study assesses the efficacy of using open-source LLMs to extract extravasation injury labels from radiology reports.

Hypothesis: LLMs can efficiently extract extravasation injury labels from radiology reports.

Methods: This was an IRB approved study with 6024 radiology reports of abdominal CTs from Massachusetts General Hospital and Brigham and Women’s Hospital with the indication of abdominal trauma. The Meta-Llama-3.1-70B, Qwen2.5-72B-Instruct, and Mistral-7B-Instruct-v0.3 models were run locally. The models were prompted using a zero-shot prompt including the report impressions and instructions to evaluate if active extravasation was present and output either: “No”, “Yes”, or “Undefined”. Ground truth was manually generated for 300 reports for statistical analysis. Accuracy, precision, recall, and F1-score with respect to the ground truth were calculated for each model and an ensemble of all three models. On the reports where the Llama model identified active extravasation, the LLM models were further prompted to classify where the extravasation was located between bowel, liver, kidney, abdominal wall, gluteal/thigh, spleen, retroperitoneal, or other.

Results: A total of 125 of the 300 reports contained active extravasation. The Llama model presented with an F1-score of .990, followed by the Ensemble (0.980), Qwen (0.973), and Mistral (.932) models. Looking at localizing the extravasation, the LLM performed strongly in identifying extravasation in the bowel with the highest F1-score of .966.

Conclusion: Small open-source LLM models prove to be an effective tool for labeling radiology reports with high accuracy. This proof of concept yields promising potential to label other pathologies and extract other free texts from radiology reports, greatly reducing cost, labeling time, and burden.

Keywords: Clinical Workflow & Productivity; Emerging Technologies; Imaging Research; Quality Improvement & Quality Assurance

110 - Using Point-of-Care Visible Light Imaging to Quantify Misidentification Errors in Portable Radiography

Presenter: Srini Tridandapani, University of Alabama at Birmingham

Srini Tridandapani1, Carson Wick1, Nabile Safdar2

1University of Alabama at Birmingham, Birmingham, AL, USA

2Emory University, Atlanta, GA, USA

Introduction: Misidentification errors continue to be an insidious problem in radiology. These errors are often undetected or addressed without adequate reporting, presenting a challenge to downstream users faced with incomplete information.

Hypothesis: Point-of-care (POC) patient visible light (VL) images can be used to detect and quantify misidentification errors in portable radiography.

Methods: Misidentification errors were detected by retrospectively querying PACS using logs from an existing deployment of an automated POC patient VL imaging system. Studies for which VL images were acquired but contained radiographs no longer in PACS were manually reviewed. For each of these studies, VL images were compared to VL images from other studies with the same patient identification number. If the patient in the VL images did not match, a misidentification error was noted).

Results: Over a one-year period, 19,997 portable radiography studies with POC patient VL images were acquired. Of these, 257 (1.3%) had at least one radiograph missing during follow-up PACS querying, and 97 (0.5%) had all radiographs missing. For the 160 studies with one, but not all, radiographs missing, an error with an individual radiograph was likely. For the 97 studies with all radiographs missing, an error with the study, e.g., a misidentification error, was likely. From these 97 studies, 18 misidentification errors were found after manual review (1 in 1111 portable radiography studies).

Conclusion: A scalable method was used to quantify misidentification errors in portable radiography by reviewing discrepancies in studies in PACS over time. This method provides a vastly reduced set of radiography studies for manual review to quantify these errors. Analyses using this method based on POC patient VL imaging will be expanded to investigate both trends and the causes of these errors with the ultimate goal of providing guidance regarding which areas in radiography to direct improvement efforts.

Keywords: Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Communication Data Management; Emerging Technologies; Imaging Research; Standards & Interoperability

111 - Utilization of Electronic Health Record Embedded Enterprise Imaging Exchange: A Single Institute Experience

Presenter: Josh Volin, Emory University

Josh Volin1, Vidya Viswanathan1, Peter Harri1, Colin Segovis1, Nabile Safdar1, Elias Kikano1

1Emory University, Atlanta, GA, USA

Introduction: Efficient exchange of medical imaging between healthcare institutions is critical for improving diagnostic accuracy and reducing redundant imaging. Investments in scalable, interoperable platforms aligned with upcoming regulatory requirements, such as the HTI-2 bill, are essential for reducing redundant imaging, improving care quality, and optimizing healthcare resources. This study evaluates the demand, volume, and geographic distribution of image exchanges facilitated by an electronic health record (EHR)-based image exchange platform at a single US academic healthcare system.

Hypothesis: The primary aim of this study is to evaluate the demand, volume and geographic distribution of EHR- based image exchange.

Methods: A retrospective analysis was conducted from 1/1/2023 to 4/2/2024 using the Epic Imaging Exchange Advanced platform (Verona, WI). This platform enables a bi-directional exchange of reference-quality images across institutions. Inbound data, defined as the images requested and received by our institution, included the number of images exchanged categorized by image type (DICOM, non-DICOM, PDF). Outbound data reflected the volume of thumbnails and full images shared with external institutions.

Results: In 2023, over 1.6 million patients from 370 institutions were queried for available outside images. Of these, 542,152 (33%) had images available, totaling 4.8 million inbound images. Non-DICOM images made up the majority (51%), followed by DICOM (34%) and PDF (15%). Full retrieval rates were 2.4% non-DICOM, 4.5% DICOM, and 3.8% PDF.

Outbound data (3/24/2023-4/1/2024) showed 1.96 million thumbnails and 1.53 million full images were shared with 1,478 healthcare organizations. Geographically, 33% of inbound images were from in-state institutions, while 40.8% of thumbnails and 66.3% of full images were exported out-of-state.

Conclusion: Image exchange is in high demand as patients increasingly seek care across the US. Additionally, there is high demand for scalable, interoperable exchange solutions aimed at reducing redundant imaging and improving diagnostic confidence which align with emerging regulatory standards.

Keywords: Applications; Artificial Intelligence/Machine Learning; Clinical Workflow & Productivity; Communication Data Management; Emerging Technologies; Enterprise Imaging; Organizational & Professional Development; Provider Experience; Quality Improvement & Quality Assurance; Standards & Interoperability; Systems Management

Supplement Details

Journal name: Journal of Imaging Informatics in Medicine

Supplement title: 2025 Annual Meeting of the Society for Imaging Informatics in Medicine (SIIM) - Selected Abstracts

Conference data: Oregon Convention Center | Portland, OR | May 21 – 23, 2025

Chair, Guest editor(s), or Organizing-committee name:

Alan McMillan, PhD

Professor, Radiology, University of Wisconsin School of Medicine and Public Health

2025 SIIM Scientific Research Abstract Program Co-chair

Chris Roth, MD, CIIP, FSIIM, MMCI

Associate Chief Medical Officer, Imaging Service Line, Duke University

Vice Chair of Informatics, Duke University

2025 SIIM Annual Meeting Planning Committee Chair

Steven Rothenberg, MD

Assistant Professor, Diagnostic Radiology, University of Alabama-Birmingham Medicine

Associate Scientist, Center for Clinical and Translational Science, University of Alabama-Birmingham Medicine

2025 SIIM Scientific Research Abstract Program Co-chair

Michael Toland

Senior Director, Information Technology University of Maryland Medical System

IT Site Executive, University of Maryland Medical System

2025 SIIM Applied Informatics Abstract Program Co-chair

Audrey Verde, MD, PhD

Assistant Professor. Neuroradiology, Duke University Health System

2025 SIIM Applied Informatics Abstract Program Co-chair

Paul Yi, MD

Associate Member, Department of Radiology (DoR), St. Jude Children's Research Hospital

Director, Intelligent Imaging Informatics (I3) & Image Quantification and AI (IQAI), St. Jude Children's Research Hospital

2025 SIIM Annual Meeting Planning Committee Vice Chair

Sponsor (or Society name): Society for Imaging Informatics in Medicine (SIIM)

Sponsorship statement: Publication of this supplement was sponsored by the Society for Imaging Informatics in Medicine (SIIM). All content was reviewed and selected by the Annual Meeting Program Committee, which held full responsibility for the abstract selections.

Abstracts (quantity): 111

Tables: N/A

Figures: N/A

Collated page proofs: Mackenzie Mercer mmercer@siim.org

Target publication date: Closest to May 21, 2025 (before/after ok)

Publication format: Online

Footnotes

Publisher’s Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.


Articles from Journal of Imaging Informatics in Medicine are provided here courtesy of Springer

RESOURCES