Skip to main content
Orthopaedic Journal of Sports Medicine logoLink to Orthopaedic Journal of Sports Medicine
. 2026 Sep 4;14(9):23259671261463933. doi: 10.1177/23259671261463933

Evolution of Artificial Intelligence Adoption in American Journal of Sports Medicine Abstracts: A Period and Nativity-Based Analysis

Zirvecan Güneş †,*, Esra Kutsal Mergen ‡
PMCID: PMC13554643  PMID: 42718969

Abstract

Background:

While large language models (LLMs) have made a tremendous impact on scientific workflow, there are also serious concerns about ethical integrity and nondiverse written language in medical and orthopaedic literature. Recent studies have yielded conflicting results regarding the involvement of artificial intelligence (AI) in writing processes.

Purpose:

To (1) determine the extent of AI involvement in American Journal of Sports Medicine (AJSM) abstracts across 4 different time periods; (2) to determine whether AI usage differs based on authorship nativity.

Study Design:

Cross-sectional study.

Methods:

We retrospectively reviewed abstracts from AJSM and organized these into 4 periods. We chose the year 2000 as a human baseline to establish a threshold for AI involvement. We also selected the years 2015 and 2020 to observe changes in AI probability scores. Considering the release of ChatGPT in November 2022, we included abstracts from 2023 to 2025 to determine the increase in AI usage. We classified articles into native English-speaking, non-native English-speaking, and mixed groups to determine authorship nativity based on institutional affiliations. We used Originality.AI to report AI probability scores. We evaluated differences across years and nativity groups using the Kruskal-Wallis and post hoc tests.

Results:

Across the 6 time periods, 1885 abstracts were analyzed. AI probability scores from the year 2000 (n = 124; mean, 1.68%) were significantly lower than all subsequent years (P < .001 for all pairwise comparisons). Scores remained stable between 2015 (n = 348; mean, 9.39%) and 2020 (n = 383; mean, 8.91%; P≥ .999). However, scores began a rapid and significant increase starting in 2023 (n = 345; mean 15.48%) and continued to rise through 2025 (n = 323; mean, 30.6%; P < .001) compared with all previous years. The percentage of abstracts exceeding the threshold increased consistently throughout the years, from 4% in the 2000 baseline to 62.5% in 2025. When nativity groups were compared within the Generative AI Era, no statistically significant differences in AI probability scores were observed in 2023 (P = .364), 2024 (P = .053), or 2025 (P = .454).

Conclusion:

Our study demonstrated that the prevalence of AI-like linguistic patterns in AJSM abstracts has increased since the release of ChatGPT. AI adoption across all nativity groups suggests that AI has become a universal standard for efficient article writing, regardless of the language barrier encountered by non-native speakers.

Keywords: artificial intelligence, authorship, bibliometrics, large language models, scientific writing


ChatGPT and other large language models (LLMs) revolutionized the way science is authored after November 2022.4,7,35 With LLMs, it is now possible for a machine to produce large quantities of text in a fluent manner similar to how humans write, think, and present their work. 14 The advantages of LLMs for academic productivity include the ability to quickly generate thought processes, edit language, format references, and provide concise summaries of a significant volume of literature.6,15,20,23 As a result, preparing manuscripts has become less time-consuming for researchers. LLMs allow researchers to devote their time and effort to interpreting their data rather than writing their manuscripts.9,12 However, the integration of LLMs into the medical literature introduces some important risks to scientific and ethical integrity.1,6,8 Often referred to as "hallucinations," it is common for LLMs to write a response with great confidence even though that response may be completely incorrect or drawn from nonexistent data.17,33 A recent study found that LLMs could create a manuscript acceptable for submission to an orthopaedic journal writing with fabricated results and references. 4 After blind review, the manuscript received a provisional acceptance decision by the journal, although reviewers pointed out some areas of inconsistency. 4

Currently, the medical community has difficulty distinguishing between true human authorship and artificial intelligence (AI)-generated writing. Hakam et al 14 reported that neither the authors nor any of the detection tools could consistently determine whether LLMs produced the text. 14 Reviewers also struggled to differentiate between human-written abstracts and those generated by AI. 29 This difficulty raises concerns regarding the potential for misinformation to enter the orthopaedic literature. Beyond fabrication, there is a risk of linguistic homogenization. Kobak et al 19 analyzed biomedical abstracts and found an increase in the frequency of specific words after the release of ChatGPT. 19 Words such as "delves," "showcasing," and "intricate" appeared with frequencies that exceeded historical baselines. 19 This trend inevitably leads to the erasure of the unique voices of scientific authors, both locally and globally.

Analyses aim to quantify the prevalence of LLMs in orthopaedic medical journals.5,24,25,27 Callanan et al 5 analyzed 3374 orthopaedic articles published in journals after the release of ChatGPT and found that 16.7% of articles exceeded the threshold for prominent AI involvement. 5 Miller et al 24 reviewed the Journal of Shoulder and Elbow Surgery and found that the presumed usage of AI in abstracts has increased significantly, from 21.1% in 2022 to 30.1% in 2024. They also found that there were differing rates of AI usage based on the section of the article. Detection rates in the abstracts increased, while they remained stable in the full-text manuscripts. This finding suggests that authors may use the tools to draft abstracts because they are shorter and more formulaic.

The American Journal of Sports Medicine (AJSM) is one of the top-ranked journals in the field of orthopaedics and sports medicine.26,30 According to research by Nassar et al, 26 AJSM had a much lower number of AI-related articles than other orthopaedic journals. They reported that only 3.34% of AJSM articles exceeded the predetermined AI threshold. However, their analysis included the full-text article; as a result, their estimate of AI usage may not be very accurate due to the dilution of AI signals throughout a full article, as previously stated in the study of Miller et al. 24 It would be best to focus an analysis on the abstract section of each article to achieve a more accurate estimate of AI usage.

The linguistic background of the authors is an important variable in our study. It is common for researchers from non-English-speaking countries to feel constrained by the language standards required by academic journals and publishers.2,16,21 According to Miller et al, 24 authors from these regions may be inclined to use AI tools to improve their writing clarity and grammar before submission. In addition, bias exists with detection tools that identify non-native English authors’ works as being AI-generated. 22 Liang et al 22 have shown that a significant number of detector algorithms classify these works as being AI-generated when non-native speakers actually write them because of the lower perplexity score relative to native English speakers’ writing.22,35 On the contrary, Hakam et al 14 found that standard criteria for identifying artificial text, such as a lack of nuance, did not correlate with the actual origin of the text.

This study aimed to track the evolution of writing patterns in AJSM abstracts and evaluate the progression of AI-based writing assistance using a period-based methodology across 4 distinct timeframes. We hypothesized that AI detection scores would show a statistically significant increase after the release of ChatGPT. We further hypothesized that this increase would differ based on the English nativity of the authorship.

Methods

Study Design

We retrospectively reviewed abstracts published in AJSM. We collected our data from the Web of Science Core Collection database. Our search was limited to documents published as either "Article" or "Review." We excluded all other document types, including "Editorial," "Letter to the Editor," "Correction," "Meeting Abstract," "Proceeding Paper," and "Biographical Item." These exclusions ensure that the analysis focuses on the standard approach rather than the different stylistic formats. We retrieved metadata for each record, including the publication year, article title, full abstract text, digital object identifier (DOI), and the institutional affiliations of all listed authors. This study did not include any protected health information or involve human subjects; instead, it used only publicly available data.

Temporal Period Selection

We identified the key time periods in the evolution of scientific writing. For our starting point, we selected the year 2000, before the widespread use of advanced automated writing tools. The second time period we selected was 2015, when several online grammar checkers, translation, and paraphrasing tools became available. The next time period we selected was 2020, as it was just before the introduction of ChatGPT. Furthermore, during this period, there were many changes in scientific publishing due to the coronavirus disease 2019 (COVID-19) pandemic. Finally, we selected 2023 to 2025 as the generative AI era. We subdivided the generative AI era into years (ie, 2023, 2024, and 2025) to see how rapidly large language models were adopted after the launch of ChatGPT in November 2022.

Authorship Nativity

We categorized the articles into native, non-native, and mixed groups. We used institutional affiliations as a proxy for linguistic background. The native English-speaking countries were defined according to the Kachru model of inner circle countries, which includes the United States, the United Kingdom, Canada, Australia, New Zealand, and Ireland. 18 An article by an author primarily affiliated with one of these native English-speaking countries was placed in the native group. Conversely, if an author was affiliated with an institution that was not located within any of these native English-speaking countries, and no co-authors were affiliated with a native English-speaking institution, it was placed into the non-native group. Lastly, an article written by an author whose institution was affiliated with a non-native country and had at least 1 co-author associated with a native English-speaking institution was placed into the mixed group. This distinction allows us to identify the possible humanizing effect of the native English-speaking co-author on the manuscript's refinement or polishing.

AI Detection Protocol

We selected Originality.AI (Version Turbo 3.0.2; Originality.ai, Inc) as an AI detection tool because it provides a measurable probability score ranging from 0 to 100, rather than simply a "yes," "no," or "mixed" answer. We extracted the abstract from each article. We removed copyright statements, keywords, conflict-of-interest statements, and other nontextual elements to prevent potential errors. We processed the cleaned text through the detector and recorded the AI probability score for each abstract.

Statistical Analysis

We conducted all statistical analyses in RStudio (R Foundation for Statistical Computing). We used the 2000 data set as our baseline, which was assumed to be entirely human-written, to establish an AI-positive threshold. We set this cutoff at the 97.5th percentile of the AI probability distribution for 2000 (≈3%). This specific cutoff was chosen as the optimal operating point because it keeps the false-positive rate strictly below the conventional 0.05 statistical tolerance margin. It yielded a high specificity for confirmed human-written text, effectively accounting for natural human stylistic variation without being overly conservative compared with higher thresholds. This percentile-based approach allowed us to define an empirical threshold derived from historical human writing patterns rather than relying on arbitrary cutoff values. Abstracts with scores ≥3% were classified as AI-positive. We examined the differences in AI probability scores across all years using the nonparametric Kruskal-Wallis test. When the Kruskal-Wallis test was significant, we applied Bonferroni-adjusted Dunn post hoc tests to compare nativity groups for 2023-2025 using the same approach. We tracked annual AI-positive abstracts with contingency tables and row percentages. We conducted 2-sided tests. We considered results significant if the P value was less than .05.

Results

We included 1885 abstracts in the study based on the inclusion criteria. The 124 abstracts from the year 2000 had a mean AI probability score of 1.68%. The 348 abstracts from 2015 had a mean score of 9.39%. The 383 abstracts from 2020 showed a similar mean score of 8.91%. The increase accelerated in 2023, reaching a mean of 15.48% across 345 abstracts. The 2024 cohort (n = 362) had a mean score of 18.24%. The score reached its highest value in 2025, when the 323 analyzed abstracts yielded a mean of 30.6%. The data show a continuous growth of AI-related characteristics in abstracts, which becomes more rapid over time (Table 1) (Figure 1).

Table 1.

Mean and Median Values of AI Probability Scores Throughout the Years a

Year N Mean ± SD b Median [IQR] c
2000 124 1.68 ± 11 00 [0-0]
2015 348 9.39 ± 21.3 1 [1-5]
2020 383 8.91 ± 20.6 1 [0-5]
2023 345 15.48 ± 28.5 2 [1-11]
2024 362 18.24 ± 30.9 2 [0-18]
2025 323 30.6 ± 38.6 6 [1-76]
a

AI, artificial intelligence; IQR, interquartile range.

b

Mean ± SD: arithmetic mean and standard deviation.

c

Median [IQR]: 25th-75th percentiles.

Figure 1.

Alt text: Distributions of AI probability scores by year, showing changes over 20+ years.

Distributions of artificial intelligence probability scores by year.

The Kruskal-Wallis test revealed statistically significant differences in AI probability scores across all years (P < .001) (Table 2). Consequently, we conducted pairwise comparisons using the Dunn post hoc test (Table 3). The analysis revealed that the 2000 cohort exhibited significantly lower AI probability scores compared with all subsequent years (P < .001 for all comparisons). We observed no significant difference between the 2015 and 2020 periods (P≥ .999). Although the 2023 cohort showed a numerical increase compared with 2015 and 2020, these differences were not statistically significant (2015 vs 2023: P = .232; 2020 vs 2023: P = .201). By 2024, the increase became statistically significant compared with both 2015 (P = .024) and 2020 (P = .019). We found no significant difference between the scores in 2023 and 2024 (P≥ .999). The 2025 cohort demonstrated a distinct surge in probability scores, resulting in a statistically significant difference compared with all other years (P < .001). These findings indicate a progressive acceleration in the influence of AI, with a distinct shift in the distribution by 2025.

Table 2.

Kruskal-Wallis Test for AI Probability Scores Across Years a

Outcome n H (df) P
AI probability 1885 242.17 (5) <.001
a

AI, artificial intelligence.

Table 3.

Pairwise Comparisons of AI Probability Scores Between Years a

Comparison, Years n1 n2 Z Statistic P Value b
2000 vs 2015 124 348 9.225 <.001
2000 vs 2020 124 383 9.341 <.001
2000 vs 2023 124 345 10.971 <.001
2000 vs 2024 124 362 11.549 <.001
2000 vs 2025 124 323 14.962 <.001
2015 vs 2020 348 383 0.005 ≥.999
2015 vs 2023 348 345 2.421 .232
2015 vs 2024 348 362 3.157 .024
2015 vs 2025 348 323 7.970 <.001
2020 vs 2023 383 345 2.474 .201
2020 vs 2024 383 362 3.228 .019
2020 vs 2025 383 323 8.147 <.001
2023 vs 2024 345 362 0.705 ≥.999
2023 vs 2025 345 323 5.578 <.001
2024 vs 2025 362 323 4.949 <.001
a

AI, artificial intelligence.

b

Dunn post hoc test with Bonferroni adjustment.

We evaluated differences among nativity groups using the Kruskal-Wallis test. We found no statistically significant differences in 2023 (P = .364), 2024 (P = .053), or 2025 (P = .454) (Table 4 and Figure 2). Although the 2024 cohort exhibited a marginal trend, the difference did not reach the 5% significance threshold. These data suggest that the variation in AI probability scores during the 2023-2025 period is driven primarily by the publication year rather than authorship nativity.

Table 4.

Kruskal-Wallis Tests for AI Probability by Nativity Within Each Year (2023-2025) a

Year Nativity n Mean ± SD Median_[IQR] H (df) P
2023 Native 193 16.5 ± 29 2 [1-15] 2.02 (2) .364
Non-native 118 15.4 ± 29.7 2 [0-8.5]
Mixed 34 9.9 ± 20.2 2 [0-6.5]
2024 Native 187 15 ± 27.7 2 [0-12.5] 5.87 (2) .053
Non-native 132 23.5 ± 34.1 3 [1-29.5]
Mixed 43 16.4 ± 32 1 [0-7.5]
2025 Native 185 28.1 ± 36.7 5 [1-53] 1.58 (2) .454
Non-native 98 32.7 ± 40.5 7.5 [1-80]
Mixed 40 37 ± 42.5 9.5 [2-87.5]
a

AI, artificial intelligence; IQR, interquartile range.

Figure 2.

Chart shows AI probability score distribution for nativity groups from 2023 to 2025. Values range mostly 40-75, highest in Mixed in 2025.

Artificial intelligence probability score distribution by nativity (2023-2025).

To assess the diagnostic characteristics of the AI detection tool, abstracts from 2000 were used as a human-written reference data set. Specificity and false-positive rates were calculated at different AI probability thresholds (1%, 3%, and 5%). These results are summarized in Table 5. Because the data set contained only human-written reference texts, sensitivity could not be estimated. Furthermore, to evaluate how different cutoff choices influence the overall trend, we analyzed the proportion of abstracts exceeding these alternative thresholds across all years (Table 6).

Table 5.

Specificity of the Artificial Intelligence Detection Tool Using the 2000 Abstracts as the Human- Written Gold Standard

Cutoff True Negatives False Positives Specificity (%) False-Positive Rate (%)
1 103 21 83.06 16.94
3 119 5 95.97 4.03
5 121 3 97.58 2.42

Table 6.

Proportion of Abstracts Exceeding Different Artificial Intelligence Score Cutoffs by Year

Year N ≥1 (%) ≥3 (%) ≥5 (%)
2000 124 16.9 4 2.4
2015 348 76.7 32.5 25.9
2020 383 74.7 34.5 27.4
2023 345 75.7 47 38
2024 362 74.6 49.2 41.7
2025 323 86.7 62.5 53.6

Focusing on our established primary threshold (3%), the number of abstracts identified as AI-positive continued to grow significantly over the years (Figure 3). The 2000 cohort contained 4% of abstracts that surpassed this baseline threshold. The proportion of abstracts exceeding the threshold reached 32.5% and 34.5% during 2015 and 2020, respectively. The percentage of AI-positive abstracts reached 47% in 2023 and 49.2% in 2024. The highest percentage occurred in 2025, when 62.5% of the abstracts exceeded the threshold. Rather than definitively proving authorship, our results strongly suggest a substantial increase in AI-assisted writing or AI-like linguistic patterns throughout the study period.

Figure 3.

Bar chart shows the percentage of AI-positive abstracts from 2000 to 2025.

Percentages of artificial intelligence-positive abstracts by year.

Discussion

The major findings of our study demonstrated a significant, rapid increase in the involvement of AI in AJSM abstracts. Analyzing 1885 abstracts, we found that the proportion of AI-positive manuscripts surged from a human baseline of 4% in 2000 to a peak of 62.5% in 2025. The mean AI probability scores reached their highest value in 2025 (30.6%), representing a statistically significant increase over all previous years (P < .001). Furthermore, the lack of significant differences across nativity groups in recent years suggests that AI adoption has become a standard in abstract writing, regardless of language barriers.

The academic field began questioning grammar and spell-checking tools when they first appeared10,31 in the early 2000s. Galletta et al 10 stated that these tools would eventually become standard for daily academic activities. The scientific workflow is now undergoing a major transformation through LLMs, which are profoundly changing academic and orthopaedic literature on a much greater scale than previous tools.3,19,28,36 Our data demonstrated this trend within the AJSM abstracts. The generative AI era exhibited distinct abstract composition patterns compared with the human baseline from the year 2000, according to our study. This shift matches with Kobak et al, 19 who reported that the impact of LLMs on scientific writing surpassed the effect of major global events such as the COVID-19 pandemic.

The 2025 cohort received higher AI probability scores and was completely separated from previous years. While the human baseline from the year 2000 showed a mean AI-probability score of 1.68% and only 4% of abstracts exceeded the predetermined threshold, 62.5% of 2025 abstracts were AI-positive, with a mean score of 30.6%. The pre-AI data set from Callanan et al 5 established a 32.875% threshold for AI involvement, resulting in 16.7% of manuscripts exceeding this level after ChatGPT became available. Our results show a much higher prevalence. This deviation from the 2000 baseline in the 2023-2025 cohort indicates a fundamental change in writing patterns distinct from the gradual evolution of language observed between 2000 and 2020. These findings suggest that the so-called hybrid author has become the standard writing practice rather than an exceptional case.

Our findings contrast with those of Nassar et al, 26 who showed that AJSM articles contained AI traces at a rate of 3.34%. This discrepancy is likely because of AI-signal dilution. Full-text manuscripts often undergo multiple layers of human intervention during the publication process, including peer-review revisions, editorial feedback, and professional copyediting. These editorial processes may substantially modify the manuscript and potentially dilute linguistic patterns that AI detection algorithms attempt to identify. We acknowledge that analyzing only abstracts may amplify AI detection signals. However, abstracts typically represent the manuscript and generally undergo fewer editorial modifications compared with full-text articles. Therefore, focusing on abstracts may provide a more consistent representation of the writing characteristics. Miller et al 24 validated this method by showing that abstract detection rates rose to 30.1% in 2024 while full-text detection rates remained steady. The abstract needs to present information in a brief and standardized format. The LLMs generate perfect results when performing these operations. Furthermore, abstracts are the first point of visibility for articles for readers. Authors likely use these tools to polish and improve this specific section to increase journal acceptance rates.

The evaluation of authorship nativity contradicts our initial hypothesis about which users tend to employ these tools. A common assumption, and our hypothesis, was that non-native English speakers use artificial intelligence more than their native counterparts do to overcome language barriers. However, we found no statistically significant differences in probability scores among native, non-native, and mixed authorship groups. It appears that native and non-native speakers use AI tools at similar rates to polish their text and summarize data. This challenges the claim that high scores result from bias against non-native speakers. 22 Instead, it indicates that AI-assisted writing is becoming common across all groups.

Editorial policies regarding the use of AI vary significantly14,26; AJSM allows its use for language correction but requires a declaration. However, Ganjavi et al 11 found that many journals lack clear guidelines. Yang et al 34 argue that the field needs shared standards for benchmarking and reproducibility to address redundancy and a lack of rigor. The increased prevalence of AI suggests that restrictive policies may be outdated. Banning these tools assumes that their use is a deviation from the norm. Our data suggest that their use is now the norm. Editorial boards should focus more on knowledge verification than detection.

This study has several limitations. First, detection tools do not provide absolute certainty. Originality.AI generates a probability score based on patterns of human writing. This inherent uncertainty is followed by the technical constraints of analyzing abstracts. Detection algorithms generally demonstrate higher reliability when analyzing longer texts. The limited word count of an abstract may reduce the precision of probability scores compared with full-text analysis. Moreover, scientific abstracts adhere to strict structural guidelines. LLMs are trained to mimic these stylistic norms. Consequently, high-quality human writing that prioritizes clarity and conciseness may trigger false positives due to low perplexity. This mechanism likely explains why a small proportion of the 2000 abstracts exceeded the predefined threshold in our data set. In some cases, highly standardized phrasing and predictable structures may even generate relatively high AI-probability scores despite the text being entirely human-written. Furthermore, the accuracy of these tools changes as the models evolve. A detector calibrated to earlier model versions may fail to detect AI-generated text or misclassify human text as AI-generated. The black box nature of these algorithms prevents a transparent analysis of which specific linguistic features triggered a high probability score. In the present study, the detector output was therefore interpreted as a probabilistic indicator of AI-like linguistic patterns rather than definitive proof of AI authorship. The availability of a continuous probability score also allowed us to analyze score distributions across time and conduct robustness analyses using alternative thresholds, which is particularly useful for large-scale bibliometric investigations of writing patterns.

Another important limitation concerns the use of institutional affiliation as a proxy for linguistic background. This approach does not capture the individual linguistic history or English proficiency of authors. The academic workforce is highly mobile; therefore, a researcher affiliated with an institution in the United States or the United Kingdom may be a non-native English speaker, whereas a researcher working in a non-native English region may possess near-native fluency because of previous education or training in English-speaking environments. In addition, authorship in scientific publications does not necessarily reflect writing responsibility. Although we assumed that the first author was primarily responsible for drafting the abstract and manuscript, collaborative studies often involve multiple contributors in the writing and editing process. Consequently, co-authors with different linguistic backgrounds may substantially influence the final text. These factors introduce unavoidable uncertainty when interpreting the relationship between linguistic background and AI-probability scores, and the nativity-based comparisons should therefore be interpreted with appropriate caution.

Third, our study is limited to AJSM. While AJSM is a high-impact publication, Nassar et al 26 observed significant variability in detection rates across sports medicine journals. Consequently, our findings may not generalize to other journals with different acceptance rates or geographic distributions of authors.

Fourth, we utilized the year 2000 to define human writing. However, academic writing evolves independently of AI. Changes in editorial standards, such as the shift toward structured abstracts and stricter word limits, occurred between 2000 and 2022. Some deviations in AI probability scores may reflect these trends in scientific writing rather than the involvement of LLMs alone.

Fifth, our methodology fails to consider techniques designed to avoid detection. Authors may combine AI-humanizing tools with traditional paraphrasing. These tools restructure the statistical patterns that detection algorithms rely on. Authors may keep editing texts using detectors until they are no longer flagged. 14 The text is still AI-generated even as detector scores change. More advanced strategies can also conceal AI usage. Users can direct LLMs to write like a human or add deliberate errors to avoid elevated AI detection scores. Therefore, our results may underestimate actual use.

Finally, another limitation of the present study is the use of a single AI detection tool. Although recent studies suggest that several detectors demonstrate comparable performance, differences in model architecture, training data, and scoring strategies may lead to variability across tools.13,32 However, we selected Originality.AI for 3 main reasons. First, independent evaluations demonstrate its high accuracy in detecting text generated by advanced large language models. It consistently minimizes false negatives compared with other tools. Second, Originality.AI provides a continuous probability score ranging from 0 to 100. This scoring system allowed us to perform nonparametric statistical analyses. It also enabled us to establish a specific threshold based on our year 2000 human baseline. Finally, the tool is widely recognized and continuously updated. This adaptability maintains its detection efficacy against newly released language models. Future studies may benefit from incorporating multiple AI detection systems and comparing their outputs across larger data sets further to validate the observed trends in AI-related writing patterns.

Conclusion

The prevalence of AI involvement in AJSM abstracts has increased since the release of ChatGPT. Similar AI adoption across all nativity groups suggests that AI has become a standard for efficient article writing, regardless of the language barrier encountered by non-native speakers. The evaluation of AI into the scientific workflow should require newly defined approaches and guidelines to verify and validate information and references rather than focusing only on detection. In the context of sports medicine, these safeguards are critical to ensuring that AI-generated hallucinations or fabricated data do not infiltrate the evidence-based literature that guides clinical decisions. Failure to validate AI-assisted content could result in the adoption of unreliable surgical techniques or rehabilitation protocols, directly compromising patient safety and the quality of clinical outcomes.

Footnotes

Final revision submitted May 16, 2026; accepted May 30, 2026.

The authors have declared that there are no conflicts of interest or sources of funding in the authorship and publication of this contribution.

Ethical approval was not sought for the present study. This research exclusively analyzed publicly available bibliographic data and did not involve human participants or animal subjects.

ORCID iD: Zirvecan GüneşInline graphic https://orcid.org/0000-0001-7833-0331

References

  • 1. Almufarreh A, Ahmad A, Arshad M, Onn CW, Elechi R. Ethical implications of ChatGPT and other large language models in academia. Front Artif Intell. 2025;8:1615761. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2. Arenas-Castro H, Berdejo-Espinola V, Chowdhury S, Rodríguez-Contreras A, James ARM, Raja NB, et al. Academic publishing requires linguistically inclusive policies. Proc Biol Sci. 2024;291(2018):20232840. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3. Bernstein E, Ramsamooj A, Millar KL, Lum ZC. Identification and categorization of the top 100 articles and the future of large language models: Thematic analysis using bibliometric analysis. Jmir ai. 2025;4:e68603. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4. Brameier DT, Alnasser AA, Carnino JM, Bhashyam AR, von Keudell AG, Weaver MJ. Artificial intelligence in orthopaedic surgery: Can a large language model “write” a believable orthopaedic journal article? J Bone Joint Surg Am. 2023;105(17):1388-1392. [DOI] [PubMed] [Google Scholar]
  • 5. Callanan T, Marquez J, Pisani C, Schmitt P, Pietro J, Chen M, et al. Evaluating artificial intelligence-based writing assistance among published orthopaedic studies: Detection and trends for future interpretation. J Bone Joint Surg Am. 2025;107(16):1887-1893. [DOI] [PubMed] [Google Scholar]
  • 6. Celik SU. Integrating artificial intelligence into scientific writing: a narrative review for clinical and surgical researchers. Am J Surg. 2025;250:116657. [DOI] [PubMed] [Google Scholar]
  • 7. Chan A, Rahimi-Ardabilli H, Rogers WA, Coiera E. The real-world impact of artificial intelligence ethics frameworks across a decade in healthcare: a scoping review. J Am Med Inform Assoc. 2025;32(11):1767-77. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8. Dergaa I, Chamari K, Zmijewski P, Ben Saad H. From human writing to artificial intelligence generated text: examining the prospects and potential threats of ChatGPT in academic writing. Biol Sport. 2023;40(2):615-22. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9. Fatani B. ChatGPT for future medical and dental research. Cureus. 2023;15(4):e37285. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10. Galletta DF, Durcikova A, Everard A, Jones BM. Does spell-checking software need a warning label? Commun ACM. 2005;48:82-6. [Google Scholar]
  • 11. Ganjavi C, Eppler MB, Pekcan A, Biedermann B, Abreu A, Collins GS, et al. Publishers' and journals' instructions to authors on use of generative artificial intelligence in academic and scientific publishing: bibliometric analysis. Bmj. 2024;384:e077192. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12. Granjeiro JM, Cury A, Cury JA, Bueno M, Sousa-Neto MD, Estrela C. The future of scientific writing: AI tools, benefits, and ethical implications. Braz Dent J. 2025;36:e256471. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13. Hadra M, Cambridge K, Mesbah M. Evaluating the accuracy and reliability of AI content detectors in academic contexts. Int J Educat Integ. 2026;22(1):4. [Google Scholar]
  • 14. Hakam HT, Prill R, Korte L, Lovreković B, Ostojić M, Ramadanov N, et al. Human-written vs AI-generated texts in orthopedic academic literature: Comparative qualitative analysis. JMIR Form Res. 2024;8:e52164. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15. Izquierdo-Condoy JS, Arias-Intriago M, Tello-De-la-Torre A, Busch F, Ortiz-Prado E. Generative artificial intelligence in medical education: Enhancing critical thinking or undermining cognitive autonomy? J Med Internet Res. 2025;27:e76340. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16. Jain VK, Iyengar KP, Vaishya R. Is the English language a barrier to the non-English-speaking authors in academic publishing? Postgrad Med J. 2020;98:234-235. [DOI] [PubMed] [Google Scholar]
  • 17. Ji Z, Lee N, Frieske R, Yu T, Su D, Xu Y, et al. Survey of hallucination in natural language generation. ACM Comput Surv. 2022;55:1-38. [Google Scholar]
  • 18. Kachru BB. World Englishes and applied linguistics. World Englishes. 1990;9(1):3-20. [Google Scholar]
  • 19. Kobak D, González-Márquez R, Horvát E, Lause J. Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Sci Adv. 2025;11(27):eadt3813. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20. Lee JM. ChatGPT: how to use it and the pitfalls/cautions in academia. Ann Pediatr Endocrinol Metab. 2025;30(5):229-241. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21. Li J, Zong H, Wu E, Wu R, Peng Z, Zhao J, et al. Exploring the potential of artificial intelligence to enhance the writing of English academic papers by non-native English-speaking medical students - the educational application of ChatGPT. BMC Med Educ. 2024;24. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22. Liang W, Yuksekgonul M, Mao Y, Wu E, Zou J. GPT detectors are biased against non-native English writers. Patterns (N Y). 2023;4(7):100779. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23. Lingard L. Writing with ChatGPT: An illustration of its capacity, limitations & implications for academic writers. Perspect Med Educ. 2023;12(1):261-70. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24. Miller AS, Tyagi A, Sudah SY, Rompala A, Nicholson AD, Srikumaran U, et al. Evaluation of the impact of large language learning models on publications in the Journal of Shoulder and Elbow Surgery. JSES Int. 2025;9(5):1803-1808. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25. Miller LE, Bhattacharyya D, Miller VM, Bhattacharyya M. Recent trend in artificial intelligence-assisted biomedical publishing: A quantitative bibliometric analysis. Cureus. 2023;15(5):e39224. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26. Nassar JE, Farias MJ, Singh M, Dinh PV, Sahhar M, Daher M, et al. Large language model-based writing in published sports medicine research: Uncovering a growing influence. Orthop J Sports Med. 2025;13(9):23259671251371234. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27. Porto JR, Morgan KA, Hecht CJI, Burkhart RJ, Liu RW. Quantifying the scope of artificial intelligence–assisted writing in orthopaedic medical literature: An analysis of prevalence and validation of AI-detection software. J Am Acad Orthop Surg. 2025;33(1):42-50. [DOI] [PubMed] [Google Scholar]
  • 28. Sheridan GA, Howard LC, Neufeld ME, Doyle TR, Hughes AJ, Sculco PK, et al. Can artificial intelligence generate scientific discussion that passes peer review for publication in a high-impact orthopaedic journal? Ir J Med Sci. 2025;194(4):1191-1198. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29. Stadler RD, Sudah SY, Moverman MA, Denard PJ, Duralde XA, Garrigues GE, et al. Identification of ChatGPT-generated abstracts within shoulder and elbow surgery poses a challenge for reviewers. Arthroscopy. 2025;41(4):916-24.e2. [DOI] [PubMed] [Google Scholar]
  • 30. Vaishya R, Shekhawat S, Vaish A, Migliorini F. Evolution of journal rankings in orthopedics and sports medicine (2000-2024): A SCImago-based bibliometric analysis. Orthopadie (Heidelb). 2025;54(10):795-803. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31. Vernon A. Computerized grammar checkers 2000: capabilities, limitations, and pedagogical possibilities. Comput Composit. 2000;17(3):329-49. [Google Scholar]
  • 32. Weber-Wulff D, Anohina-Naumeca A, Bjelobaba S, Foltýnek T, Guerrero-Dib J, Popoola O, et al. Testing of detection tools for AI-generated text. Int J Educat Integ. 2023;19(1):26. [Google Scholar]
  • 33. Wu S, Fei H, Pan L, Wang WY, Yan S, Chua T-S. Combating multimodal LLM hallucination via bottom-up holistic reasoning. ArXiv. 2024;abs/2412.11124. [Google Scholar]
  • 34. Yang AJ, Woo JJ, Ramkumar PN. Editorial commentary: Shifting from redundancy to rigor in orthopaedic large language model research. Arthroscopy. 2025;41(11):4946-4949. [DOI] [PubMed] [Google Scholar]
  • 35. Yoo JH. Defining the boundaries of AI use in scientific writing: A comparative review of editorial policies. J Korean Med Sci. 2025;40(23):e187. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36. Zhou L, Wu AC, Hegyi P, Wen C, Qin L. ChatGPT for scientific writing - The coexistence of opportunities and challenges. J Orthop Translat. 2024;44:A1-a3. [DOI] [PMC free article] [PubMed] [Google Scholar]

Articles from Orthopaedic Journal of Sports Medicine are provided here courtesy of SAGE Publications

RESOURCES