Evaluating a Large Language Model’s Ability to Answer Clinicians’ Requests for Evidence Summaries

Mallory N Blasingame; Taneya Y Koonce; Annette M Williams; Dario A Giuse; Jing Su; Poppy A Krump; Nunzia Bettinsoli Giuse

doi:10.1101/2024.05.01.24306691

Abstract

Objective

This study investigated the performance of a generative artificial intelligence (AI) tool using GPT-4 in answering clinical questions in comparison with medical librarians’ gold-standard evidence syntheses.

Methods

Questions were extracted from an in-house database of clinical evidence requests previously answered by medical librarians. Questions with multiple parts were subdivided into individual topics. A standardized prompt was developed using the COSTAR framework. Librarians submitted each question into aiChat, an internally-managed chat tool using GPT-4, and recorded the responses. The summaries generated by aiChat were evaluated on whether they contained the critical elements used in the established gold-standard summary of the librarian. A subset of questions was randomly selected for verification of references provided by aiChat.

Results

Of the 216 evaluated questions, aiChat’s response was assessed as “correct” for 180 (83.3%) questions, “partially correct” for 35 (16.2%) questions, and “incorrect” for 1 (0.5%) question. No significant differences were observed in question ratings by question category (p=0.39). For a subset of 30% (n=66) of questions, 162 references were provided in the aiChat summaries, and 60 (37%) were confirmed as nonfabricated.

Conclusions

Overall, the performance of a generative AI tool was promising. However, many included references could not be independently verified, and attempts were not made to assess whether any additional concepts introduced by aiChat were factually accurate. Thus, we envision this being the first of a series of investigations designed to further our understanding of how current and future versions of generative AI can be used and integrated into medical librarians’ workflow.

PERMALINK

This is a preprint.

Evaluating a Large Language Model’s Ability to Answer Clinicians’ Requests for Evidence Summaries

Mallory N Blasingame

Taneya Y Koonce

Annette M Williams

Dario A Giuse

Jing Su

Poppy A Krump

Nunzia Bettinsoli Giuse

Abstract

Objective

Methods

Results

Conclusions

Full Text Availability

ACTIONS

PERMALINK

RESOURCES

Cite

Add to Collections

PERMALINK

This is a preprint.

Evaluating a Large Language Model’s Ability to Answer Clinicians’ Requests for Evidence Summaries

Mallory N Blasingame

Taneya Y Koonce

Annette M Williams

Dario A Giuse

Jing Su

Poppy A Krump

Nunzia Bettinsoli Giuse

Abstract

Objective

Methods

Results

Conclusions

Full Text Availability

ACTIONS

PERMALINK

RESOURCES

Similar articles

Cited by other articles

Links to NCBI Databases