Skip to main content
Springer logoLink to Springer
. 2025 Jun 12;194(4):1191–1198. doi: 10.1007/s11845-025-03971-y

Can artificial intelligence generate scientific discussion that passes peer review for publication in a high-impact orthopaedic journal?

Gerard A Sheridan 1,2,3, Lisa C Howard 1, Michael E Neufeld 1, Tom R Doyle 3,✉, Andrew J Hughes 2,3, Peter K Sculco 2, David E Beverland 4, Donald S Garbuz 1, Bassam A Masri 1
PMCID: PMC12413406  PMID: 40504456

Abstract

Background

There is huge interest in the use of artificial intelligence (AI) in the production and assessment of academic material; however, the role of AI remains unclear.

Aim

The purpose of this study was to perform a reviewer-blinded assessment of the quality of scientific discussion generated by an advanced AI language model (ChatGPT-4, Open AI) and determine whether this could be recommended for high-impact journal publication.

Methods

The introduction, methods and results sections of a recently published article from a high-impact journal were input into a current AI model. The AI application then produced a discussion and conclusion based on the provided text using a standardized prompt. Six experienced blinded reviewers scored all five sections of the hybrid article. A one-way analysis of variance (ANOVA) was used to assess significant differences between scores of each section. Reviewers recommended a decision regarding the suitability of the article for publication.

Results

AI composed a scientific discussion and conclusion. The median score was 80 (IQR 70–90) for introduction, 77.5 (IQR 70–90) for methods, 82.5 (IQR 50–90) for results, 60 (IQR 40–75) for discussion and 60 (IQR 40–80) for the conclusion. The median scores for the AI-generated sections were non-significantly lower than other sections (p = 0.37). The majority of reviewers (5/6, 83%) recommended “acceptance for publication after major revision”. One reviewer recommended “resubmission with no guarantee of acceptance”. There were no recommendations for rejection.

Conclusion

Current AI large language models are now capable of generating content that passes experienced peer review and is acceptable for publication in a high-impact orthopaedic journal, after revision. There are still many concerns regarding the integration of AI into the process of scientific writing, mainly the tendency of AI to rely on advanced pattern recognition and fabricated or inadequate references.

Level of evidence: Level IV

Keywords: AI, Artificial intelligence, Author, Chat GPT, Large language model, Peer review

Introduction

The role of artificial intelligence (AI) in medical research has been progressing at a significant rate in recent years. Large language models such as Chat GPT (Open AI, San Francisco, CA 94110, USA) have been shown to be capable of achieving the equivalent of a passing score for a third-year medical student on the United States Medical Licensing Examination (USMLE) [1]. As these large language models are evolving at an exponential rate, we are seeing improvements in their performance over very short periods of time, the latest of which can be seen with the release of ChatGPT-4 in March 2023. A study by Tagaki et al. demonstrated how GPT-4 outperformed GPT-3.5 in the Japanese Medical Licensing Examination (JMLE) [2]. GPT-4 was superior to GPT-3.5 in terms of accuracy and in answering difficult questions, and it also achieved the passing criteria for the JMLE, indicating that these large language models (LLMs) can match or outperform humans in knowledge-based medical fields already.

Within the orthopaedic community, the use of machine learning has already been described as a tool to assist in clinical decision making to improve patient care, for example, in predicting outcomes following Irrigation and Debridement (I&D) surgery for prosthetic joint infection (PJI) [3]. AI has also been shown to be capable of analysing radiographic images with significant precision in the context of hip and knee arthroplasty [4]. In addition to the role of AI in the clinical aspects of orthopaedic patient care, AI has also integrated with the orthopaedic research community. It is becoming a tool used appropriately and inappropriately in the production of orthopaedic research and in the generation of scientific content, similar to a human author.

In 1950, Alan Turing described the Turing Test (i.e. Imitation Game) where a machine passes the test provided its responses were indistinguishable from a human [5]. It has recently been shown that reviewers struggle to identify AI generated abstracts, highlighting the progress LLMs have made [6]. The role of LLMs in the production of scientific content is still unclear, and the scientific community is attempting to quickly implement safeguards in the interest of upholding scientific integrity and the quality of scientific content. The Editors-in-Chief for The Bone & Joint Journal, The Journal of Bone & Joint Surgery, Clinical Orthopaedics and Related Research and Journal of Orthopaedic Research published recommendations regarding the involvement of AI in the generation of scientific articles [7]. The recommendations include “AI applications cannot be listed as authors. Whether and how AI applications were used in the research or the reporting of its findings must be described in detail in the Methods section and should be mentioned again in the Acknowledgements section”.

The purpose of this study was to perform a reviewer-blinded assessment of the scientific quality of a discussion and conclusion section generated by an advanced AI LLM, and ultimately determined whether AI-generated content would be recommended for peer-reviewed publication. The hypothesis was that the discussion would pass peer review and would not be identified as computer-generated.

Methods

AI generation of discussion and conclusion

On May 10, 2023, a PubMed search was performed to identify recently published articles in the field of hip arthroplasty published in a Q1 high impact orthopaedic journal as identified in 2022 Journal Impact Factor, Journal Citation Reports (Clarivate, 2023). The article entitled “Survival of the Exeter V40 short revision (44/00/125) stem when used in primary total hip arthroplasty” published in Bone & Joint Journal on May 1, 2023, was arbitrarily selected [8]. The manuscript was reviewed in full (G.A.S & T.R.D). Following this, the introduction, methods and results sections from the article were inserted into the ChatGPT-4 application along with the references cited in the original discussion of the article [9–18] using the following specific instructions:

“Referencing all citations listed in the reference list, write a discussion section (1500 words total) and conclusion to follow the introduction, methods and results section listed here. This discussion and conclusion should be composed in a style that would be suitable for publication in the Bone & Joint Journal (BJJ). Reference list (Include all citations listed here in your discussion):”

The scientific discussion and conclusion generated by ChatGPT-4 were then copied directly, without any form of human editing, into a hybrid document to replace the original discussion and conclusion. This new hybrid document included the original introduction, methods and results section with the addition of the new AI-generated discussion and conclusion sections.

Reviewer selection process

Six fellowship-trained arthroplasty surgeons who work in large academic hospitals with experience in reviewing articles on the topic of total hip arthroplasty for high-impact orthopaedic journals were recruited to participate in the study (LCH, MEN, PKS, DEB, DSG, BAM). The reviewers were recruited from three different countries, in North America and Europe. The years of experience of Q1 peer review for each reviewer are summarized in Table 1.

Table 1.

Peer reviewer academic experience

Reviewer Journals Impact factor (June 2023) Experience (years)
1

Journal of Bone and Joint Surgery

Clinical Orthopaedics and Related Research

Journal of Arthroplasty

6.558

4.837

4.435

 > 20 years
2

Bone & Joint Journal

Clinical Orthopaedics and Related Research

Hip International

5.385

4.837

1.756

 > 25 years
3

Bone & Joint Journal

Journal of Arthroplasty

BMC Musculoskeletal Disorders

5.385

4.435

2.32

7
4 Journal of Bone and Joint Surgery 6.558 4
5

Bone & Joint Journal

Journal of Arthroplasty

Bone & Joint Open

5.385

4.435

2.8

4
6

American Journal of Sports Medicine

Hip International

Injury

6.057

1.756

2.687

3

All reviewers confirmed that they had not read the full article of the circulated manuscript and were unaware of its authorship. All reviewers had no affiliation to the article authors’ institution. All reviewers were blinded to the methodology of the current study and, as such, were completely unaware that part of the document they were reviewing was generated by AI.

Reviewer scoring

The hybrid document was then circulated to all reviewers with an accompanying reviewer score sheet (Appendix). The scoring system involved numerical grading of each individual section of the article (maximum score of 100 for each section). The reviewers were also asked to give one of the following overall decisions: (1) Accept, (2) Accept with minor revisions, (3) Accept after major revisions, (4) Do not accept yet (authors may resubmit with no guarantee of acceptance), (5) Reject. All score sheets were returned to the corresponding author and analysed.

Statistical analysis

Data distribution was assessed using the Shapiro–Wilk test, which confirmed that the data was not normally distributed. Therefore, scores were expressed using median values and inter-quartile ranges (IQR). Boxplot graphs were generated to demonstrate median values and IQRs. A one-way analysis of variance (ANOVA) was then performed to assess any significant difference between the five median scores. A p-value of < 0.05 was taken to be statistically significant. All statistical analyses were performed using Stata/IC 13.1 for Mac (64-bit Intel) (College Station, TX, 77,845). Institutional Review Board (IRB) approval was not required given the nature of the study.

Results

The individual reviewer scores for each section and overall decision are detailed in Table 2. The median score was 80.0 (IQR 70–90) for introduction, 77.5 (IQR 70–90) for methods, 82.5 (IQR 50–90) for results, 60.0 (IQR 40–75) for discussion and 60.0 (IQR 40–80) for the conclusion. There was no significant difference in the median scores for the discussion and conclusion compared to other sections, as shown in Fig. 1 (p = 0.37). No reviewer raised concerns in their free text comments regarding the potential use of AI in the manuscript.

Table 2.

Individual reviewer scores and overall decision

Reviewer Section Score
(100 max)
Overall decision
1

- Introduction

- Methods

- Results

- Discussion

- Conclusion

80

70

90

60

50

Accept after major revisions
2

- Introduction

- Methods

- Results

- Discussion

- Conclusion

40

40

40

40

40

Accept after major revisions
3

- Introduction

- Methods

- Results

- Discussion

- Conclusion

80

90

90

85

90

Accept after major revisions
4

- Introduction

- Methods

- Results

- Discussion

- Conclusion

90

70

80

75

80

Accept after major revisions
5

- Introduction

- Methods

- Results

- Discussion

- Conclusion

95

85

85

5

5

Do not accept yet (authors may resubmit with no guarantee of acceptance)
6

- Introduction

- Methods

- Results

- Discussion

- Conclusion

70

90

50

60

70

Accept after major revisions

Fig. 1.

Fig. 1

Boxplot of academic quality scoring (0–100)

There were no recommendations suggesting rejection of the manuscript, with all inviting edits and resubmission. The majority of reviewers (5/6, 83%) recommended “Accept after major revisions”. One reviewer recommended “Do not accept yet (authors may resubmit with no guarantee of acceptance)”. This reviewer noted in their free text comments “Further discussion of limitations, what the study actually showed, and comparison to pertinent literature is required. Since there is limited published evidence regarding the use of this stem in primary settings, the authors should state this as well. Needs to discuss patient selection problems with the stem”.

Discussion

The purpose of the current study was to assess the extent to which AI generated scientific writing would be able to withstand the scrutiny of experienced high-impact factor orthopaedic peer review. Based on the current results, AI was able to compose a nuanced scientific discussion and conclusion based on the introduction, methods and results sections. While there was a reduction in the median scores for discussion and conclusion, this difference was not statistically significant. Furthermore, the reviewers did not identify the use of AI or express criticism which warranted rejection of the manuscript. On the contrary, five of the six reviewers recommended publication after revisions, a verdict which will often proceed to manuscript publication. Hence, it was demonstrated that AI can generate scientific content that experienced reviewers identify as suitable for publication in Q1 orthopaedic journals.

Regarding the lower scores for the discussion and conclusion sections, it should be noted that at the time of the review of the hybrid manuscript, the introduction, methods, and results had already undergone peer review in the Bone & Joint Journal with subsequent revisions prior to publication in its final format. In contrast, the discussion and conclusion sections generated by AI were taken directly from the LLM without any human editing. Therefore, the human-generated sections had a distinct advantage having already undergone peer review and revision, while the AI sections had not. In a similar vein, the LLM was provided citations from the original article to assist the writing of the discussion, which likely had an impact on the quality of its output, as LLM task performance is known to be related to the prompts it is provided with [19].

If AI applications are to be involved in the process of scientific manuscript generation in the future, the time taken to edit an AI-generated discussion may be significantly lower than the time taken for a human to write the discussion section in its entirety. This time efficiency demonstrates one way in which AI may be able to improve the productivity of orthopaedic researchers in the future. However, concerning aspects of this method of content generation is the tendency of Chat-GPT to insert references throughout the body of text which are often fabricated, from low quality or non-peer-reviewed sources [19]. The generated discussion was instructed to include all references from the original article and did so with the exception of Hamilton et al. [18], producing a viable scientific discussion that was deemed to be coherent by expert review. However, this shows imperfect execution of the instructions and highlights a potential pitfall with LLM writing. The AI model used in this study did not have real-time internet access to the full articles for the references provided, nor does it have access to articles published after its data training set was compiled. So while the generated content may appear coherent and accurate, the model is generating a discussion without access to the latest evidence and inserting references where it predicts they should be inserted. This is a limitation of AI written discussions as they may not be able to properly assess the nuanced risks vs. benefits of a treatment or the context specific application of the findings. The limitations of the training dataset may prevent the model writing coherently about new advances which is has not encountered before. This was born out in the comments of reviewer 5 who highlighted that there is limited published evidence on this topic and they felt the discussion inadequately addressed this. This may suggest the LLM struggled with a novel topic. Newer versions of Chat-GPT will have real-time internet access and may be able to avoid some of these issues, but currently, this is a real and valid concern. However, despite this, the hybrid document passed the scrutiny of experienced peer reviewers asked to judge it to the standards of a Q1 publication.

Irrespective of any advantages of AI involvement in scientific writing, the question remains whether the involvement of AI is the appropriate direction to take and to what extent it should be involved in the future. Polisetty et al. addressed the pitfalls associated with the use of AI in hip and knee arthroplasty research, quoting the following issues: “vernacular conflation, repackaging limited registry data, prematurely releasing internally validated prediction models, appraising model architecture instead of inputted data, withholding code, and evaluating studies using antiquated regression-based guidelines” [20]. The appraisal of model architecture instead of inputted data is especially relevant when using AI to generate scientific discussion. Even though the discussion may appear to be coherent and appropriately referenced, there is a risk that the results are not comprehended by the LLM. The model is essentially using pattern recognition from its training data to create a discussion rather than deeply considering the implications of the results. It has previously been shown that experienced reviewers struggle to identify AI-generated scientific abstracts [6]. However, the present study demonstrates a further development in LLM capabilities as much greater complexity is contained in the discussion section compared to an abstract. The fact that this discussion may be the product of an entity without any “higher-order understanding” is an issue that needs further exploration, as it seems based on the current study that “advanced pattern recognition” may pass as a substitute for “intelligent composition” from the perspective of a human reviewer.

In addition to problems with data interpretation and comprehension, there may also be problems at a more foundational level with the data itself, regardless of whether the AI comprehension is accurate or not. Odouye et al. discuss the fact that large language models are limited by the data that are provided to them [21]. For example, if there is any biased or inaccurate data present in a dataset for analysis, this bias or inaccuracy may be reinforced by the algorithm since machine learning algorithms can only be as objective as the data they are trained on. Therefore, the scientific community must be ever-diligent regarding the quality of data being input for analysis. The capability and advancement of AI technology will never be able to compensate for poor data quality.

This current study appears to be the first to assess the quality of AI-generated discussion as determined by blinded experienced reviewers of high-impact orthopaedic journals. There was no significant difference in the appraised quality of the discussion and conclusion sections than the other human-generated sections. The article was flagged for revisions; however, this is standard practice for a vast majority of submissions to Q1 journals. The major concern is that AI-generated content may be masquerading as intelligently composed content when, in actuality, it is simply an excellent example of pattern recognition.

Limitations

It may be considered a limitation that introduction, methods and results had already undergone expert peer review and subsequent revision prior to the blinded review for this study, while the discussion section had not. Other potential limitations include the relatively small sample size of reviewers, the heterogeneity in the human peer review process and the potential biases of the reviewers.

Conclusion

Current AI large language models are now capable of generating content that passes experienced peer review and is acceptable for publication in a high-impact orthopaedic journal, after revision. There are still many concerns regarding the integration of AI into the process of scientific writing, mainly the tendency of AI to rely on advanced pattern recognition and fabricated or inadequate references.

Appendix: Reviewer score sheet

graphic file with name 11845_2025_3971_Figaa_HTML.jpg

graphic file with name 11845_2025_3971_Figab_HTML.jpg

Funding

Open Access funding provided by the IReL Consortium.

Declarations

Ethics approval

This study was deemed exempt from ethics.

Conflict of interest

The authors declare no competing interests.

Footnotes

Publisher's Note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Gilson A, Safranek CW, Huang T, Socrates V, Chi L, Taylor RA, Chartash D (2023) How does ChatGPT perform on the United States Medical Licensing Examination (USMLE)? The implications of large language models for medical education and knowledge assessment. JMIR Med Educ 9:e45312. 10.2196/45312 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Takagi S, Watari T, Erabi A, Sakaguchi K (2023) Performance of GPT-3.5 and GPT-4 on the Japanese Medical Licensing Examination: comparison study. JMIR Med Educ 9:e48002. 10.2196/48002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Shohat N, Goswami K, Tan TL et al (2020) 2020 Frank Stinchfield Award: identifying who will fail following irrigation and debridement for prosthetic joint infection. Bone Joint J 102-b:11–19. 10.1302/0301-620x.102b7.Bjj-2019-1628.R1 [DOI] [PubMed]
  • 4.Gurung B, Liu P, Harris PDR et al (2022) Artificial intelligence for image analysis in total hip and total knee arthroplasty: a scoping review. Bone Joint J 104-b:929–937. 10.1302/0301-620x.104b8.Bjj-2022-0120.R2 [DOI] [PubMed]
  • 5.Rider RE (1984) A mathematician: alan turing. Science 223:807. 10.1126/science.223.4638.807 [DOI] [PubMed] [Google Scholar]
  • 6.Stadler RD, Sudah SY, Moverman MA et al (2024) Identification of ChatGPT-generated abstracts within shoulder and elbow surgery poses a challenge for reviewers. Arthroscopy. 10.1016/j.arthro.2024.06.045 [DOI] [PubMed] [Google Scholar]
  • 7.Leopold SS, Haddad FS, Sandell LJ, Swiontkowski M (2023) Artificial intelligence applications and scholarly publication in orthopaedic surgery. Bone Joint J 105-b:585–586. 10.1302/0301-620x.105b.Bjj-2023-0272 [DOI] [PubMed]
  • 8.Evans JT, Salar O, Whitehouse SL et al (2023) Survival of the Exeter V40 short revision (44/00/125) stem when used in primary total hip arthroplasty. Bone Joint J 105-b:504–510. 10.1302/0301-620x.105b5.Bjj-2022-1124.R1 [DOI] [PubMed]
  • 9.No author. Orthopaedic Data Evaluation Panel. www.odep.org.uk/Home.aspx. 01/03/2023
  • 10.Woodbridge AB, Hubble MJ, Whitehouse SL et al (2019) The Exeter short revision stem for cement-in-cement femoral revision: a five to twelve year review. J Arthroplasty 34:S297-s301. 10.1016/j.arth.2019.03.035 [DOI] [PubMed] [Google Scholar]
  • 11.Ben-Shlomo Y, Blom A, Boulton C et al (2021) National joint registry annual reports. In: The National Joint Registry 18th Annual Report 2021. National Joint Registry© National Joint Registry 2021., London [PubMed]
  • 12.Wylde V, Blom AW (2011) The failure of survivorship. J Bone Joint Surg Br 93:569–570. 10.1302/0301-620x.93b5.26687 [DOI] [PubMed] [Google Scholar]
  • 13.Berg AJ, Hoyle A, Yates E et al (2020) Cement-in-cement revision with the Exeter short revision stem: a review of 50 consecutive hips. J Clin Orthop Trauma 11:47–55. 10.1016/j.jcot.2019.04.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Choy GG, Roe JA, Whitehouse SL et al (2013) Exeter short stems compared with standard length Exeter stems: experience from the Australian Orthopaedic Association National Joint Replacement Registry. J Arthroplasty 28:103-109.e101. 10.1016/j.arth.2012.06.016 [DOI] [PubMed] [Google Scholar]
  • 15.Duncan WW, Hubble MJ, Howell JR et al (2009) Revision of the cemented femoral stem using a cement-in-cement technique: a five- to 15-year review. J Bone Joint Surg Br 91:577–582. 10.1302/0301-620x.91b5.21621 [DOI] [PubMed] [Google Scholar]
  • 16.Wyatt MC, Poutawera V, Kieser DC et al (2020) How do cemented short Exeter stems perform compared with standard-length Exeter stems? The experience of the New Zealand National Joint Registry. Arthroplast Today 6:104–111. 10.1016/j.artd.2020.01.003 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Martin R, Clark N, James J, Baker P (2022) Clinical evaluation of the cemented Exeter Short 125 mm stem at a minimum of 3 years: a prospective cohort study. J Orthop 30:18–24. 10.1016/j.jor.2022.02.005 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Hamilton DF, Ohly NE, Gaston P (2018) Can arthroplasty stem influence outcome? (CASINO): a randomized controlled equivalence trial of 125 mm versus 150 mm Exeter V40 stems in total hip arthroplasty. Trials 19:226. 10.1186/s13063-018-2621-8 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Oeding JF, Lu AZ, Mazzucco M et al (2024) ChatGPT-4 performs clinical information retrieval tasks using consistently more trustworthy resources than does google search for queries concerning the Latarjet procedure. Arthroscopy. 10.1016/j.arthro.2024.05.025 [DOI] [PubMed] [Google Scholar]
  • 20.Polisetty TS, Jain S, Pang M et al (2022) Concerns surrounding application of artificial intelligence in hip and knee arthroplasty: a review of literature and recommendations for meaningful adoption. Bone Joint J 104-b:1292–1303. 10.1302/0301-620x.104b12.Bjj-2022-0922.R1 [DOI] [PubMed]
  • 21.Oduoye MO, Javed B, Gupta N, Valentina Sih CM (2023) Algorithmic bias and research integrity; the role of nonhuman authors in shaping scientific knowledge with respect to artificial intelligence: a perspective. Int J Surg 109:2987–2990. 10.1097/js9.0000000000000552 [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Irish Journal of Medical Science are provided here courtesy of Springer

RESOURCES