ABSTRACT
Background
As residency program applications become more homogenized due to pass-fail United States Medical Licensing Examination Step 1 scores, pre-clinical grading, and standardized letters of recommendation, personal statements remain a key subjective tool to assess applicant fit. With the rise of large language models like ChatGPT, applicants may increasingly use AI to draft these statements.
Objective
To determine whether ChatGPT can generate orthopaedic surgery personal statements that are indistinguishable from and comparable in quality to human-written statements.
Methods
In 2025 at a single urban academic university hospital, ChatGPT-4.0 was given a standardized prompt to generate 5 personal statements across distinct themes. Five de-identified statements from applicants who interviewed or matched at our institution from 2020 to 2021 were randomly selected. The 10 essays were randomized and independently reviewed by 8 blinded assessors, including 6 attending physicians and 2 resident physicians, all of whom serve on the resident admissions committee. Using a 9-question REDCap survey, evaluators rated each statement on readability, originality, authenticity, and overall quality on a 100-point scale; guessed whether each was AI- or human-generated (Yes/No); and indicated whether they would offer an interview (Yes/No). Chi-square tests were used for analysis.
Results
There were no significant differences between ChatGPT and human-written statements in readability (P=.63), originality (P=.48), authenticity (P=.79), or overall quality (P=.84). Assessors could not reliably identify AI- versus human-written statements (19 of 40 vs 22 of 39, P=.43) and interview offers were similarly comparable (29 of 40 vs 29 of 39, P=.85).
Conclusions
ChatGPT can produce orthopaedic surgery personal statements that are both indistinguishable from and comparable in quality to human-written ones.
Introduction
Residency program directors rank the personal statement as of “average importance” in determining interview invitations.1 Despite this, medical students often perceive the personal statement as a significant stressor, requiring multiple rounds of revision from mentors or even costly professional editing services.2,3 With the advent of large language models (LLMs) such as ChatGPT, applicants may increasingly turn to AI-generated assistance due to its accessibility and ease of use. However, studies across multiple medical specialties have shown that AI-generated personal statements, using ChatGPT-3 and -4, can be indistinguishable from and comparable in quality to human-written ones.4-6 Within orthopaedic surgery, a recent study found that ChatGPT-generated personal statements were correctly identified only 56% of the time, but were scored less favorably across several quality metrics.7 However, that study uploaded human-written statements as references and generated all essays within the same chatbot session, introducing AI self-referencing and bias, thereby not fully assessing ChatGPT’s independent capabilities. As AI continues to advance rapidly and ChatGPT’s ability to adopt a more humanistic tone improves, understanding the current capabilities of AI in personal statement generation is essential for residency selection committees to understand how to best interpret and utilize this potentially less personal aspect of the application.
This study aims to assess whether ChatGPT can generate orthopaedic surgery personal statements that are both indistinguishable from and comparable in quality to those from previous applicants who interviewed or matched at our institution.
KEY POINTS
What Is Known
As residency applications become more uniform and AI tools more accessible, personal statements remain one of the few subjective components still vulnerable to undisclosed AI assistance.
What Is New
At a single orthopaedic residency site, ChatGPT-generated personal statements were found to be indistinguishable to reviewers from human-written ones across readability, originality, authenticity, quality, and interview offer rates in a blinded assessment.
Bottom Line
Residency programs should be aware that AI-generated personal statements can closely mimic human submissions, prompting consideration of how to evaluate applicant authenticity and narrative elements moving forward.
Methods
This study was conducted in 2025 at a single urban academic university hospital with 40 residents. Participants included 8 assessors involved in orthopaedic residency admissions: 6 attending physicians (1 of whom is a program director) and 2 senior residents. Assessors were selected across different orthopaedic subspecialties.
ChatGPT-4.0 was prompted to generate 5 personal statements using a standardized prompt adapted from Karakash et al8:
“Please compose a personal statement for my application to orthopaedic surgery residency programs. I am currently in my last year of medical school in the United States. The statement should adopt an academic tone without being pretentious, maintaining clarity over sophistication, and avoiding predictable transitions and extraneous wording. Emphasize cohesiveness and a smooth narrative flow. The essay should be narrative-driven, showcasing a unique and original story. Avoid simply listing achievements from my resume. The statement should not exceed 600 words. Make the theme revolve around [theme].”
Each statement followed a distinct theme: athletic injury, family member orthopaedic injury, research, global health, and exposure to the field during medical school. These themes were selected based on the senior author’s experience as a residency admissions committee member, reflecting commonly encountered narrative frameworks in personal statements. Each essay was generated in a separate chat session to minimize AI self-referencing or bias on the same day using a paid personal account. The daily prompt limit was not reached. No applicant CV or background materials were included in the prompts. The essays generated are included in the online supplementary data. Additionally, 5 de-identified personal statements from applicants who interviewed and/or matched at our institution prior to the introduction of LLMs were randomly selected as comparators from the 2020-2021 residency application cycle.
The 10 essays were randomized and administered online to assessors through a 9-question REDCap survey (provided as online supplementary data). Eight blinded participants assessed each statement using a 100-point Likert scale across 4 metrics: readability, originality, authenticity, and quality. Assessors also classified each essay as ChatGPT- or human-written and indicated whether they would grant the applicant an interview based on the personal statement. No additional formal training was provided prior to evaluation, to reflect real-world residency application review practices.
At the end of the survey, assessors answered supplemental questions regarding AI in residency applications, as reflected in the online supplementary data. Assessors were given one month to complete the survey.
One assessor recognized a human-written statement and was therefore unable to provide an unbiased evaluation. As a result, their response for that specific essay was removed from the final analysis. This resulted in 40 completed surveys for the ChatGPT-written essays and 39 for the human-written ones.
Statistical Analysis
All statistical analyses were performed using IBM Statistical Package for Social Sciences (Version 30). Independent samples t tests were used to compare quality metrics between ChatGPT-generated and human-written personal statements. Pearson’s chi-square tests were used to analyze the proportion of correctly identified statement types and to compare the proportion of interviews granted between groups. A formal power analysis was not performed for this study. Differences were considered statistically significant when P<.05.
This study did not involve patients and was deemed exempt from institutional review board review.
Results
There were no significant differences between ChatGPT-generated and human-written essays across all evaluated metrics: readability (73.90±19.59 vs 75.90±17.44, P=.63), originality (71.95±19.77 vs 68.55±22.23, P=.48), authenticity (71.95±23.17 vs 70.55±23.37, P=.79), and quality (73.13±18.84 vs 73.95±18.17, P=.84; Table).
Table.
Assessor Responses to Personal Statements
| Human-Written | ChatGPT-Written | P value | |
|---|---|---|---|
| Readability | 73.90±19.59 | 75.90±17.44 | .63 |
| Originality | 71.95±19.77 | 68.55±22.23 | .48 |
| Authenticity | 71.95±23.17 | 70.55±23.37 | .79 |
| Quality | 73.13±18.84 | 73.95±18.17 | .84 |
| Essays Correctly Identified, n/N (%) | |||
| ChatGPT | 19/40 (47.5) | .43 | |
| Human-written | 22/39 (56.4) | ||
| Grant Interview, n/N (%) | |||
| ChatGPT | 29/40 (72.5) | .85 | |
| Human-written | 29/39 (74.4) | ||
| Supplemental Questions | Range | ||
| Question 1 | 24.75±22.34 | 0-50 | |
| Question 2 | 48.25±22.42 | 10-84 | |
| Question 3 | 30.25±9.24 | 16-50 | |
Note: P<.05 is considered statistically significant.
Additionally, the proportion of essays correctly identified as ChatGPT, 19 of 40 essays, versus human-generated, 22 of 39 essays, did not significantly differ between the 2 groups (47.5% vs 56.4%, P=.43; Table). The likelihood of an essay leading to an interview was also comparable between ChatGPT-, 29 of 40 essays, and human-generated, 29 of 39 essays, and human-written statements (72.5% vs 74.4%, P=.85; Table).
Assessor responses for supplemental questions are summarized in the Table.
Discussion
This study found that ChatGPT can generate orthopaedic surgery personal statements that are indistinguishable from and comparable in quality to human-written statements. Using open-ended prompts and separate chat sessions, ChatGPT-generated statements performed similarly to human-written ones across readability, originality, authenticity, and overall quality. Assessors were unable to reliably identify the origin of each statement, and both AI- and human-generated essays were equally likely to be deemed interview-worthy.
These findings align with prior studies across multiple specialties demonstrating that AI-generated personal statements are often difficult to distinguish from human-written essays and perform similarly across common quality metrics. In general surgery, plastic surgery, and anesthesiology, evaluators frequently misclassified AI-generated statements and reported comparable readability, authenticity, and overall quality with similar interview selection rates.4-6 Conversely, one internal medicine study found lower performance among ChatGPT-generated essays; however, methodological differences, including incorporating applicant CVs and generating all responses within a single chatbot session, may have constrained the model’s independence and produced more homogenized outputs.9
Within orthopaedic surgery, Lum et al found that ChatGPT-written statements deceived evaluators 56% of the time and performed similarly across many quality metrics, except for personal interests, reasons for choosing residency, career goals, compelling nature, and originality.7 Moreover, all assessors favored the real personal statements. However, similar to Nair et al,9 their study trained the chatbot by providing prior human-written examples and generated all essays within the same chat session, introducing AI bias and self-learning. This approach likely caused AI-generated essays to sound more similar to each other, potentially introducing reviewer unconscious bias and lower scores for AI-generated statements, particularly in humanistic metrics, such as compelling nature and originality. In contrast, our methodology allowed us to further explore ChatGPT’s independent capabilities and eliminate AI self-referencing. This may explain why we found no significant differences in quality metrics, particularly in originality, between AI- and human-generated essays.
This study has several limitations. As a large academic institution, our findings may not be generalizable to smaller, community-based institutions. Additionally, because the study focused on a small surgical subspecialty, the results may not be generalizable to larger, less competitive specialties. Subjectivity was inherent in the evaluation process, though this was mitigated by using experienced assessors, including attending physicians and senior residents. Additionally, this study utilized the paid version of ChatGPT-4o, while applicants using the free version may produce different results. This study represents a less typical application of AI, generating entire personal statements with a single prompt and no edits. This approach may differ from how most applicants use AI, relying on it primarily for grammatical edits or brainstorming, rather than full statement generation. Even though applicant essays in this study were sourced prior to the increased utilization of ChatGPT, these applicants potentially used other word-processing tools or bots when writing their personal statements. Furthermore, this study only utilized ChatGPT for generating personal statements, while other LLMs may produce statements of differing quality. Lastly, because a formal power calculation was not performed and the sample size was limited, we were unable to draw any conclusions regarding effect size.
As AI continues to evolve, there is a need to develop clear, standardized guidelines for AI use in residency program applications and to explore alternative approaches to evaluating applicant fit, such as targeted, program-specific questions that are less amenable to AI-generated responses.
Conclusions
This study demonstrated that ChatGPT can generate orthopaedic surgery personal statements that are both indistinguishable from and comparable in quality to those written by human applicants.
Supplementary Material
Acknowledgments
The authors would like to thank Dr Juliann Kwak, Dr Jennifer Bell, Dr Andrew Vega, Dr Joshua Gary, and Dr Ali Azad for their support on this article.
Author Notes
Funding: The authors report no external funding source for this study.
Conflict of interest: The authors declare they have no competing interests.
Editor’s Note
The online supplementary data contains the essays generated by ChatGPT and the survey used in the study.
References
- 1.Hasan LK, Cohen LL, Granger CJ, et al. Program directors’ perception of the role of personal statements in the orthopaedic surgery residency selection process. J Surg Orthop Adv. 2022;31(2):90–95. [PubMed] [Google Scholar]
- 2.White BA, Sadoski M, Thomas S, Shabahang M. Is the evaluation of the personal statement a reliable component of the general surgery residency application? J Surg Educ. 2012;69(3):340–343. doi: 10.1016/j.jsurg.2011.12.003. doi: [DOI] [PubMed] [Google Scholar]
- 3.Campbell BH, Havas N, Derse AR, Holloway RL. Creating a residency application personal statement writers workshop: fostering narrative, teamwork, and insight at a time of stress. Acad Med. 2016;91(3):371–375. doi: 10.1097/ACM.0000000000000863. doi: [DOI] [PubMed] [Google Scholar]
- 4.Patel V, Deleonibus A, Wells MW, Bernard SL, Schwarz GS. Distinguishing authentic voices in the age of ChatGPT: comparing AI-generated and applicant-written personal statements for plastic surgery residency application. Ann Plast Surg. 2023;91(3):324–325. doi: 10.1097/SAP.0000000000003653. doi: [DOI] [PubMed] [Google Scholar]
- 5.Whitrock JN, Pratt CG, Carter MM, et al. Does using artificial intelligence take the person out of personal statements? We can’t tell. Surgery. 2024;176(6):1610–1616. doi: 10.1016/j.surg.2024.08.018. doi: [DOI] [PubMed] [Google Scholar]
- 6.Johnstone RE, Neely G, Sizemore DC. Artificial intelligence software can generate residency application personal statements that program directors find acceptable and difficult to distinguish from applicant compositions. J Clin Anesth. 2023;89:111185. doi: 10.1016/j.jclinane.2023.111185. doi: [DOI] [PubMed] [Google Scholar]
- 7.Lum ZC, Guntupalli L, Saiz AM, et al. Can artificial intelligence fool residency selection committees? Analysis of personal statements by real applicants and generative AI, a randomized, single-blind multicenter study. JB JS Open Access. 2024;9(4):e24.00028. doi: 10.2106/JBJS.OA.24.00028. doi: [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Karakash WJ, Avetisian H, Ragheb JM, Wang JC, Hah RJ, Alluri RK. Artificial intelligence vs human authorship in spine surgery fellowship personal statements: can ChatGPT outperform applicants? Global Spine J. 2026;16(1):313–318. doi: 10.1177/21925682251344248. doi: [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Nair V, Nayak A, Ahuja N, et al. Comparing IM residency application personal statements generated by GPT-4 and authentic applicants. J Gen Intern Med. 2025;40(1):124–126. doi: 10.1007/s11606-024-08784-w. doi: [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
