Abstract
Objective
To evaluate the quality of the answers and the references provided by ChatGPT for medical questions.
Patients and Methods
Three researchers asked ChatGPT 20 medical questions and prompted it to provide the corresponding references. The responses were evaluated for the quality of content by medical experts using a verbal numeric scale going from 0% to 100%. These experts were the corresponding authors of the 20 articles from where the medical questions were derived. We planned to evaluate 3 references per response for their pertinence, but this was amended on the basis of preliminary results showing that most references provided by ChatGPT were fabricated. This experimental observational study was conducted in February 2023.
Results
ChatGPT provided responses varying between 53 and 244 words long and reported 2 to 7 references per answer. Seventeen of the 20 invited raters provided feedback. The raters reported limited quality of the responses, with a median score of 60% (first and third quartiles: 50% and 85%, respectively). In addition, they identified major (n=5) and minor (n=7) factual errors among the 17 evaluated responses. Of the 59 references evaluated, 41 (69%) were fabricated, although they appeared real. Most fabricated citations used names of authors with previous relevant publications, a title that seemed pertinent and a credible journal format.
Conclusion
When asked multiple medical questions, ChatGPT provided answers of limited quality for scientific publication. More importantly, ChatGPT provided deceptively real references. Users of ChatGPT should pay particular attention to the references provided before integration into medical manuscripts.
Large language models (LLMs) constitute a branch of artificial intelligence (AI) at the intersection of linguistic and computer science.1 Trained on massive quantities of text-based data, LLMs have learned to interpret written input and produce language that is understandable to humans. LLMs incorporate several algorithms, such as generative pre-trained transformer (GPT). This type of neural network architecture is useful in chatbots, rendering them particularly effective at simulating human conversations. On November 30, 2022, the San Francisco-based company OpenAI released a freely available version of ChatGPT, a LLM based on their proprietary GPT (GPT-3).2 Since then, scientific articles have been written in part by ChatGPT and published.3, 4, 5 These publications were mainly to demonstrate the remarkable quality of the manuscripts written by ChatGPT. For example, a journal published an editorial for which the first 5 paragraphs were written by ChatGPT.4 Although the quality of the writing was excellent, there was no reference for the statements provided in these paragraphs.6 In another publication, the researcher asked ChatGPT to discuss the potential effect of taking rapamycin to increase longevity.5 The chatbot’s output constituted most of the article. Finally, an article describing the potential use of ChatGPT for medical writing was mainly written by the chatbot.3 In these articles, the parts produced by the chatbot did not include any references.
Although being low, the number of publications discussing ChatGPT has exploded in the past weeks. It went from only 24 publications identified in Pubmed with the term “ChatGPT” on February 7, 2023, to 92 publications on March 6, 2023 (date of first submission), to 237 on April 19, 2023 (date of revision of the manuscript). The first publications were mainly editorials or news from scientific journals. Most articles praised the quality of the writing1,7, 8, 9, 10, 11 or questioned the ethical aspects of using chatbot for scientific writing.6,12, 13, 14, 15 Other publications highlighted the potential of LLMs to improve health outcomes by “augmenting rather than replacing human expertise.”16
Although most articles suggest that the answers provided by ChatGPT are acceptable, there is no information regarding the sources of its knowledge and no reference is usually provided by ChatGPT, unless specifically requested. This limitation is acknowledged by OpenAI because the model was trained on multiple internet texts with a large number of sources.2 However, when conversing with ChatGPT, one can prompt it regarding the references backing its assertions.
To our knowledge, no study has evaluated the quality and appropriateness of the references provided by ChatGPT. This study aimed to evaluate the quality of the answers and corresponding references provided by ChatGPT, when responding to a wide range of medical questions. A secondary objective arose during the construction of the study because many citations provided by ChatGPT were not found. Hence, we aimed to assess the validity of the suggested references.
Patients and Methods
This was an experimental observational study conducted in February 2023 evaluating the quality of the responses provided by ChatGPT (version 3.5; OpenAI) to 20 medical questions from diverse fields.
The medical questions were identified by selecting 5 research articles published at the end of 2022 in 4 high-impact factor medical journals (British Medical Journal,17, 18, 19, 20, 21 Canadian Medical Association Journal,22, 23, 24, 25, 26 the Lancet,27, 28, 29, 30, 31 and New England Journal of Medicine17, 18, 19, 20, 21). These 20 articles spanned different topics and fields. The questions asked to ChatGPT were related to the primary objectives of the 20 studies (Supplemental Appendix, available online at https://www.mcpdigitalhealth.org/). In most cases, the question asked to ChatGPT was the stated primary objective of the study preceded by “what is.” In a few instances where the objective was deemed too narrow to ensure a minimal breadth of references, we formulated the question in a broader context. For example, the primary objective “to critically examine the leadership experiences of African Nova Scotian nurses in health care systems?,”22 was changed to “What are the leadership experiences of African nurses in the United States health care systems?”
In general, questions to ChatGPT started as “What are…” (see the aforementioned example), without any word limit or other constraint. After the answer by ChatGPT, a follow-up question asked, “Do you have references for this?” All references were counted, but only the first 3 were used for the analysis. To promote the external validity of the study and considering that ChatGPT is sensitive to previous chats, questions were asked by the 3 coauthors on different computers and ChatGPT accounts.
The primary outcomes were the appropriateness of the references and the quality of the responses. We initially aimed to evaluate the appropriateness of the references by multiple raters using the following verbal numeric scale: “On a scale of 0% to 100% where 0% signifies that the references had no relationship with the study question and 100% is for the 3 most pertinent references for this topic, how would you rate the references provided?” Our initial plan was to provide the articles to the raters. However, given that we failed to find the first 6 articles, we modified this outcome to evaluate whether the reference existed. To verify this, we searched Pubmed using the title and authors. If unsuccessful, we searched in the journal’s website. To better describe the references provided by ChatGPT, we evaluated whether the authors listed had previous publication in the field. We also looked at the title (Is it a title that seems pertinent for the study question?) and the journal (Does the citation look a plausible article for this journal?).
The quality of the response was measured using the following verbal numeric scale from 0 to 100%: “On a scale of 0 to 100% where 0 is equal to no answer, 100% to a perfect answer and 50% to the minimum acceptable answer how would you rate the answer provided to the question?” Moreover, raters were asked to report any factual error found in the ChatGPT answer and any other relevant qualitative feedback.
To ensure domain expertise for the selected articles, we invited the corresponding author of each article to act as rater and determine the quality of the response. We contacted each corresponding authors by email, inviting them to provide feedback on the answer related to their respective study objective. When the corresponding author did not reply, we contacted other listed authors. Each corresponding author was contacted at least 3 times before saying it was unsuccessful. These authors were deemed content experts in their field of publication.
The primary analysis of this study determined the validity of the references by calculating the proportion of references that really existed among all evaluated references. We also reported the proportion of references that listed an author with previous publications in the field of interest. The other primary analysis pertained to the quality of ChatGPT responses, reported as the median and interquartile range on the verbal numeric scale. We also reported the proportion of responses that contained factual errors according to the raters. Minor errors referred to erroneous details with limited effect to the quality of the overall response (eg, an overoptimistic affirmation that is not supported by any data), whereas major errors consisted of flagrant mistakes that invalidated the response (eg, wrong pathophysiologic explanation). However, we did not conduct the analysis using the verbal numeric scale for the appropriateness of the references owing to the small number of real references.
We had no prespecified idea of the median scores that would be obtained for the responses. It was estimated that the evaluation of at least 12 questions would provide varied study subjects and allow demonstrating the general quality of the responses or references. In addition, this would lead to at least 30 references to evaluate. On the basis of this and the premise that at least 12 (60%) of the invited authors would agree to help us, we invited 20 authors to rate 20 study questions to have at least 12 evaluations.
We did not seek an institutional review board approval given that all data were publicly available and no participants were involved. The study was completed without financial support. This article is an honest, accurate, and transparent account of the study being reported; no important aspects of the study have been omitted. Any discrepancies from the study as originally planned were explained (change in the outcome regarding the validity of the references).
Results
Each of the 20 study questions were asked to ChatGPT by a member of the study team (J.G., n=10; M.D.G., n=5; and E.O., n=5) and received a response that varied in length between 53 and 309 words. When prompted to provide its references, ChatGPT provided 2 to 7 references. A total of 59 references were included in the primary analysis.
When searching for the references suggested by ChatGPT, we noted that although they looked credible (Figure 1), most of them were fabricated by ChatGPT. Indeed, 56 of the 59 (95%) references contained authors with previous publications on a related topic found in Pubmed or were from recognized organizations (eg, Centers for Disease Control and Prevention; US Food and Drug Administration) (Table). Moreover, all titles seemed appropriate because it was related to the study question. However, 41 of the 59 (69%) references were fabricated. Among the 18 real references, 11 (61%) were titles of real published articles (including 3 with minor citation errors and 5 with major citation errors), 5 (28%) were existing websites, and 2 (11%) were books (Figure 2). The remaining references (n=41) did not exist. Of those, 29 (71%) of the fabricated articles were reportedly published in a known medical journal, website, or manuscript repository (eg, Centers for Disease Control and Prevention or MedRxiv) using an appropriate format of citation (ie, they reported a year, volume number, and page that were coherent with the journal). However, the reported volume and page range pertained to an unrelated article. All responses and references can be found in Supplemental Appendix, available online at https://www.mcpdigitalhealth.org/.
Figure 1.
Screenshot of an example of responses provided by ChatGPT.
Table.
Characteristic of the references provided by ChatGPT (N=59)
| n (%) | |
|---|---|
| Type of reference format | |
| Journal | 42 (71) |
| Website | 15 (25) |
| Book | 2 (3) |
| One or more author has published on the topic | 56 (95) |
| The title seems appropriate | 59 (100) |
| Journal/website exists and format adequate | 48 (81) |
| Title found on Pubmed | 11a (19) |
| Article found on the journal/website | 8 (13) |
Three articles with minor citation errors (volume and pages) and 5 with major citation errors (wrong year, volume, and pages; or wrong authors).
Figure 2.
Distribution of the references provided by ChatGPT (n=59).
When challenged about the accuracy of the references provided, ChatGPT provided variable answers. In one discussion, it confirmed that “references are available in Pubmed” and provided a weblink to Pubmed. This link was to other publications not related to our question (Figure 3). For another topic, ChatGPT responded, “I strive to provide the most accurate and up-to-date information available to me, but errors or inaccuracies can occur” (Figure 3).
Figure 3.
Example of responses from ChatGPT when questioned about the veracity of its references.
Of the 20 corresponding authors, 17 (85%) agreed to evaluate the responses. The median score they provided was 60% (first and third quartiles: 50% and 85%, respectively). Raters identified major factual error in 5 (29%) responses. For example, a rater identified that the mechanism of action of antipsychotic described by ChatGPT was incorrect. Another noted that ChatGPT overestimated the global burden of mortality associated with Shigella infections by a factor of 10. Finally, a rater identified that ChatGPT suggested the wrong corticosteroid administration route for patients with eosinophilic esophagitis: “The word ‘injections’ is inaccurate. The medicines are not injected. They are swallowed.” Minor errors were identified in 7 (41%) other responses.
Discussion
This study found that more than two-thirds of the references provided by ChatGPT to a diverse set of medical questions were fabricated, although most seemed deceptively real. Moreover, domain experts identified major factual errors in a quarter of the responses. These findings are alarming, given that trustworthiness is a pillar of scientific communication.
A previous study evaluated answers provided by ChatGPT using a scientific methodology in January 2023.32 This study reported that ChatGPT’s score (60%) was lower than that of Korean medical students (90%) in a single parasitology examination. Another study compared results of ChatGPT with those of 2 other LLMs on the United States Medical Licensing Examination Step 1 and Step 2 examinations using 2 question banks. With a mean score of 60%, they concluded that ChatGPT outperformed the other chatbots and “achieved the equivalent of a passing score for a third-year medical student.”33 A recent comment published in Nature reported that “ChatGPT fabricated a convincing response that contained several factual errors.”34 To our knowledge, this is the first study evaluating the quality of the references provided by ChatGPT. However, a recent article described a manuscript written by ChatGPT using mock data.35 When asked to conduct a literature search, ChatGPT suggested 9 references. The author of the article acknowledged, “Interestingly, at least some of these references that ChatGPT suggested do not exist in the form that is presented here.” We looked for these 9 articles, and none exist. A systematic review identified 60 early publications related to ChatGPT.11 This review identified 2 cases reports with inaccuracies in the references provided by ChatGPT. The first case report found that ChatGPT provided 3 nonexistent citations when asked to cite its information.36 In the second case report, 3 of the 7 references suggested by ChatGPT were duplicate, and most did not exist.37 Reference inaccuracies do not constitute a novelty in medical publishing but usually relate to minor mistakes.38, 39, 40, 41 For example, Browne et al38 noted that most of the errors in referencing among articles submitted to any of the 5 radiology journals were related to a failure to follow author submission guidelines. Another study reported that 15% of articles published in a single journal had an error, but it was most commonly a spelling or punctuation error.39 Finally, 2 studies reported that ∼15% of articles had errors, and ≤4.5% consisted of major errors. However, these errors did not relate to purely fabricated references.
The importance of proper referencing is undeniable. As suggested by Glick, “you are what you cite.”42 The quality and breadth of the references provided demonstrate that the researchers have performed a complete literature review and are knowledgeable about the topic. This process enables the integration of findings in the context of previous work, a fundamental aspect of medical research advancement.43 It limits the risk of biases. Failing to provide references is one thing but creating fake references would be considered fraudulent for researchers. When asked for references, ChatGPT made very appealing suggestions. It blended authors with a good research track to an interesting title in addition to a relevant journal like if the chatbot wanted to put the best of everything in a single reference. Some titles seemed to be the perfect article for our question. For example, for the question “What is the impact of haloperidol in intensive care unit patients with delirium?”, ChatGPT suggested the following inexistent title as reference: “Haloperidol in critically ill patients with delirium: a randomized, placebo-controlled trial.” In most cases, it suggested authors published multiple scientific articles on the subject of interest.
This study highlights an important shortcoming of LLMs, namely their risk of being “confidently wrong.” Although such models have now reached astounding performance in simulating human conversations, the next stages in their development must improve information validity and responsible deployment. Alhough a series of disclaimers on user registration highlight some of the limitations of ChatGPT,2 scientists considering the use of this tool to support manuscript preparation must be aware of its limitations.11 To promote responsible responses, chatbot developers should adapt the reward models that guide the algorithm’s output, so that they learn to optimize the validity of the information and references provided.44 OpenAI uses a supervised learning framework described as reinforcement learning from human feedback,2 which should be amenable to such modifications.
These findings can help orient the use of LLMs by scientists in ways that are safe and contribute to their well-being and that of their teams while targeting feasible applications.45 This includes the partial automation of tedious tasks, such as the initial steps in knowledge mapping and synthesis. This would free up time to critically appraise the information presented, thoroughly verify the corresponding sources, and orient subsequent searches. Future LLM iterations may also support manuscript formatting according to journal requirements and facilitate knowledge translation to diverse audiences (eg, translation, summarization, and adjustment of readability). On the long term, the use of LLM might improve knowledge translation by accelerating manuscript redaction and improve the quality of the writing. However, it could also threaten the scientific validity of manuscript if it incorporates inaccurate information and is misused.45 Researchers using ChatGPT may be misled by false information because clear, seemingly coherent, and stylistically appealing references can conceal poor content quality. Journals should consider clear guidelines regarding the allowed uses and reporting guidelines when tools such as ChatGPT are used. In the future, such tools may be redesigned to support human reviewers and publishers when appraising submitted articles and their corresponding bibliography. This use case highlights the enormous potential and alarming pitfalls of LLM integration in scientific writing. The inherent risk of overconfidence among LLMs also emphasizes the need for “humans in the loop” as a key element for their responsible implementation.
There are limitations to this study. First, we always asked the same question regarding references and did not specify to ChatGPT to limit references to published articles. However, when trying this a posteriori, we obtained similar answers. Second, the information provided was accurate as of February 2023 and may improve in the subsequent months or years. For example, there were 2 questions related to the COVID 19 pandemic. As ChatGPT was constructed on the basis of data published before September 2021, it may have been difficult to find sufficient knowledge for adequate answers. Moreover, ChatGPT has been launched publicly in November 2022, with the stated objective of iterative cycles of learning and improvement. It is probable that LLM will improve and learn to provide responses exempt from factual error. Finally, we used multiple raters to measure the quality of the responses. We preferred to have experts in each specific field than to have experts in scoring.
Conclusion
ChatGPT proposes undeniable progress. This study should alert the scientific community to be careful about the important risks of relying on its references because it is assembling very convincing yet often fabricated citations. These references seem constructed by merging existing references. To be useful for medical editing, chatbots should embrace the values of the scientific community such as integrity and completeness. Considering the speed of improvement of LLMs, we are hopeful that future versions of ChatGPT will suggest more accurate responses when asked to provide references. This will imply citing complete existing references instead of merging multiple publications.
Potential Competing Interests
The authors report no competing interests.
Acknowledgments
The authors acknowledge the contribution of Drs Changhai Ding, Areef Ishani, Yazdan Yazdanpanah, Nina Christine Andersen-Ranberg, Marc Rothenberg, Evan S Dellon, Ken Parhar, Jason Weatherald, Alberto Ruano Raviña, John Cleland, John H Krystal, Luciana Vercoza Viana, Kevin Ituka, Miriam Sander, Amanda Roberts, Fan Wang, and all others who agreed to rate the quality of the responses for this study.
Footnotes
Supplemental material can be found online at https://www.mcpdigitalhealth.org/. Supplemental material attached to journal articles has not been edited, and the authors take responsibility for the accuracy of all data.
Supplemental Online Material
References
- 1.Kitamura F.C. ChatGPT is shaping the future of medical writing but still requires human judgment. Radiology. 2023;307(2) doi: 10.1148/radiol.230171. [DOI] [PubMed] [Google Scholar]
- 2.ChatGPT Optimizing language models for dialogue. https://openai.com/blog/chatgpt/
- 3.Biswas S. ChatGPT and the future of medical writing. Radiology. 2023;307(2) doi: 10.1148/radiol.223312. [DOI] [PubMed] [Google Scholar]
- 4.O'Connor S. Open artificial intelligence platforms in nursing education: Tools for academic progress or abuse? Nurse Educ Pract. 2023;66 doi: 10.1016/j.nepr.2022.103537. [DOI] [PubMed] [Google Scholar]
- 5.Chat GPT Generative pre-trained transformer. Zhavoronkov A. Rapamycin in the context of Pascal's Wager: generative pre-trained transformer perspective. Oncoscience. 2022;9:82–84. doi: 10.18632/oncoscience.571. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Teixeira da Silva J.A. Is ChatGPT a valid author? Nurse Educ Pract. 2023;68 doi: 10.1016/j.nepr.2023.103600. [DOI] [PubMed] [Google Scholar]
- 7.Else H. Abstracts written by ChatGPT fool scientists. Nature. 2023;613(7944):423. doi: 10.1038/d41586-023-00056-7. [DOI] [PubMed] [Google Scholar]
- 8.Gao C.A., Howard F.M., Markov N.S., et al. Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers. Preprint. https://www.biorxiv.org/content/10.1101/2022.12.23.521610v1 Posted online December 27, 2022.
- 9.Cahan P., Treutlein B. A conversation with ChatGPT on the role of computational systems biology in stem cell research. Stem Cell Rep. 2023;18(1):1–2. doi: 10.1016/j.stemcr.2022.12.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Salvagno M., Taccone F.S., Gerli A.G. Can artificial intelligence help for scientific writing? Crit Care. 2023;27(1):75. doi: 10.1186/s13054-023-04380-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Sallam M. ChatGPT utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns. Healthcare (Basel) 2023;11(6) doi: 10.3390/healthcare11060887. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Thorp H.H. ChatGPT is fun, but not an author. Science. 2023;379(6630):313. doi: 10.1126/science.adg7879. [DOI] [PubMed] [Google Scholar]
- 13.Stokel-Walker C. ChatGPT listed as author on research papers: many scientists disapprove. Nature. 2023;613(7945):620–621. doi: 10.1038/d41586-023-00107-z. [DOI] [PubMed] [Google Scholar]
- 14.Tools such as ChatGPT threaten transparent science; here are our ground rules for their use. Nature. 2023;613(7945):612. doi: 10.1038/d41586-023-00191-1. [DOI] [PubMed] [Google Scholar]
- 15.Looi M.K. Sixty seconds on. ChatGPT. BMJ. 2023;380:205. doi: 10.1136/bmj.p205. [DOI] [PubMed] [Google Scholar]
- 16.Will ChatGPT transform healthcare? Nat Med. 2023;29(3):505–506. doi: 10.1038/s41591-023-02289-5. [DOI] [PubMed] [Google Scholar]
- 17.Weatherald J., Parhar K.K.S., Al Duhailib Z., et al. Efficacy of awake prone positioning in patients with covid-19 related hypoxemic respiratory failure: systematic review and meta-analysis of randomized trials. BMJ. 2022;379 doi: 10.1136/bmj-2022-071966. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Xie J., Wang M., Long Z., et al. Global burden of type 2 diabetes in adolescents and young adults, 1990-2019: systematic analysis of the Global Burden of Disease Study 2019. BMJ. 2022;379 doi: 10.1136/bmj-2022-072385. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Santer M., Muller I., Becque T., et al. Eczema Care Online behavioural interventions to support self-care for children and young people: two independent, pragmatic, randomised controlled trials. BMJ. 2022;379 doi: 10.1136/bmj-2022-072007. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.McIlroy D.R., Shotwell M.S., Lopez M.G., et al. Oxygen administration during surgery and postoperative organ injury: observational cohort study. BMJ. 2022;379 doi: 10.1136/bmj-2022-070941. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Candal-Pedreira C., Ross J.S., Ruano-Ravina A., Egilman D.S., Fernandez E., Perez-Rios M. Retracted papers originating from paper mills: cross sectional study. BMJ. 2022;379 doi: 10.1136/bmj-2022-071517. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Jefferies K., Martin-Misener R., Murphy G.T., Gahagan J., Bernard W.T. African Nova Scotian nurses' perceptions and experiences of leadership: a qualitative study informed by Black feminist theory. CMAJ. 2022;194(42):E1437–E1447. doi: 10.1503/cmaj.220019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Zhu Z., Huang J.Y., Ruan G., et al. Metformin use and associated risk of total joint replacement in patients with type 2 diabetes: a population-based matched cohort study. CMAJ. 2022;194(49):E1672–E1684. doi: 10.1503/cmaj.220952. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Rudoler D., Peterson S., Stock D., et al. Changes over time in patient visits and continuity of care among graduating cohorts of family physicians in 4 Canadian provinces. CMAJ. 2022;194(48):E1639–E1646. doi: 10.1503/cmaj.220439. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Skowronski D.M., Kaweski S.E., Irvine M.A., et al. Serial cross-sectional estimation of vaccine-and infection-induced SARS-CoV-2 seroprevalence in British Columbia, Canada. CMAJ. 2022;194(47):E1599–E1609. doi: 10.1503/cmaj.221335. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Naveed Z., Li J., Spencer M., et al. Observed versus expected rates of myocarditis after SARS-CoV-2 vaccination: a population-based cohort study. CMAJ. 2022;194(45):E1529–E1536. doi: 10.1503/cmaj.220676. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Kalra P.R., Cleland J.G.F., Petrie M.C., et al. Intravenous ferric derisomaltose in patients with heart failure and iron deficiency in the UK (IRONMAN): an investigator-initiated, prospective, randomised, open-label, blinded-endpoint trial. Lancet. 2022;400(10369):2199–2209. doi: 10.1016/S0140-6736(22)02083-9. [DOI] [PubMed] [Google Scholar]
- 28.Krystal J.H., Kane J.M., Correll C.U., et al. Emraclidine, a novel positive allosteric modulator of cholinergic M4 receptors, for the treatment of schizophrenia: a two-part, randomised, double-blind, placebo-controlled, phase 1b trial. Lancet. 2022;400(10369):2210–2220. doi: 10.1016/S0140-6736(22)01990-0. [DOI] [PubMed] [Google Scholar]
- 29.GBD 2019 Antimicrobial resistance collaborators. Global mortality associated with 33 bacterial pathogens in 2019: a systematic analysis for the global burden of disease study 2019. Lancet. 2022;400(10369):2221–2248. doi: 10.1016/S0140-6736(22)02185-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Sheikh J., Allotey J., Kew T., et al. Effects of race and ethnicity on perinatal outcomes in high-income and upper-middle-income countries: an individual participant data meta-analysis of 2 198 655 pregnancies. Lancet. 2022;400(10368):2049–2062. doi: 10.1016/S0140-6736(22)01191-6. [DOI] [PubMed] [Google Scholar]
- 31.Kramer C.K., Leitao C.B., Viana L.V. The impact of urbanisation on the cardiometabolic health of Indigenous Brazilian peoples: a systematic review and meta-analysis, and data from the Brazilian health registry. Lancet. 2022;400(10368):2074–2083. doi: 10.1016/S0140-6736(22)00625-0. [DOI] [PubMed] [Google Scholar]
- 32.Huh S. Are ChatGPT's knowledge and interpretation ability comparable to those of medical students in Korea for taking a parasitology examination? A descriptive study. J Educ Eval Health Prof. 2023;20:1. doi: 10.3352/jeehp.2023.20.1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Gilson A., Safranek C.W., Huang T., et al. How does ChatGPT perform on the United States medical licensing examination? The implications of large language models for medical education and knowledge assessment. JMIR Med Educ. 2023;9 doi: 10.2196/45312. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.van Dis E.A.M., Bollen J., Zuidema W., van Rooij R., Bockting C.L. ChatGPT: five priorities for research. Nature. 2023;614(7947):224–226. doi: 10.1038/d41586-023-00288-7. [DOI] [PubMed] [Google Scholar]
- 35.Macdonald C., Adeloye D., Sheikh A., Rudan I. Can ChatGPT draft a research article? An example of population-level vaccine effectiveness analysis. J Glob Health. 2023;13 doi: 10.7189/jogh.13.01003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Akhter H.M., Cooper J.S. Acute pulmonary edema after hyperbaric oxygen treatment: a case report written with ChatGPT assistance. Cureus. 2023;15(2) doi: 10.7759/cureus.34752. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Manohar N., Prasad S.S. Use of ChatGPT in academic publishing: a rare case of seronegative systemic lupus erythematosus in a patient with HIV infection. Cureus. 2023;15(2) doi: 10.7759/cureus.34616. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Browne R.F., Logan P.M., Lee M.J., Torreggiani W.C. The accuracy of references in manuscripts submitted for publication. Can Assoc Radiol J. 2004;55(3):170–173. [PubMed] [Google Scholar]
- 39.O'Connor A.E., Lukin W., Eriksson L., O'Connor C. Improvement in the accuracy of references in the journal Emergency Medicine Australasia. Emerg Med Australas. 2013;25(1):64–67. doi: 10.1111/1742-6723.12030. [DOI] [PubMed] [Google Scholar]
- 40.Montenegro T.S., Hines K., Partyka P.P., Harrop J. Reference accuracy in spine surgery. J Neurosurg Spine. 2020;34(1):1–5. doi: 10.3171/2020.6.SPINE20640. [DOI] [PubMed] [Google Scholar]
- 41.Al-Benna S., Rajgarhia P., Ahmed S., Sheikh Z. Accuracy of references in burns journals. Burns. 2009;35(5):677–680. doi: 10.1016/j.burns.2008.11.014. [DOI] [PubMed] [Google Scholar]
- 42.Glick M. You are what you cite: the role of references in scientific publishing. J Am Dent Assoc. 2007;138(1) doi: 10.14219/jada.archive.2007.0002. 12, 14. [DOI] [PubMed] [Google Scholar]
- 43.Gasparyan A.Y., Yessirkepov M., Voronov A.A., Gerasimov A.N., Kostyukova E.I., Kitas G.D. Preserving the integrity of citations and references by all stakeholders of science communication. J Korean Med Sci. 2015;30(11):1545–1552. doi: 10.3346/jkms.2015.30.11.1545. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Pause giant AI experiments: an open letter. Future of Life Institute. 2023. https://futureoflife.org/open-letter/pause-giant-ai-experiments/
- 45.Korngiebel D.M., Mooney S.D. Considering the possibilities and pitfalls of generative pre-trained transformer 3 (GPT-3) in healthcare delivery. NPJ Digit Med. 2021;4(1):93. doi: 10.1038/s41746-021-00464-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.



