Skip to main content
Journal of Graduate Medical Education logoLink to Journal of Graduate Medical Education
. 2026 Apr 15;18(2):172–175. doi: 10.4300/JGME-D-25-00676.1

CLEAR: Comparative Letter Examination and Analysis for Red Flags

Jaclyn Wiggins 1,✉, Melissa Jerdonek Sacco 2, Elizabeth Bradley 3, Jeremy Middleton 4, Jennifer Burnsed 5
PMCID: PMC13086159  PMID: 42005871

ABSTRACT

Background

Narrative letters of recommendation (LORs) remain a central element in fellowship selection. Programs screen these letters for language signifying a struggling learner or professionalism concerns, also known as “red flags,” when determining interview offers. However, thorough screening is time consuming for program directors.

Objective

To compare the speed and consistency of Microsoft Copilot to human reviewers in screening LORs for red flags.

Methods

A retrospective analysis was conducted using de-identified LORs submitted during the 2024-2025 neonatal-perinatal medicine fellowship application cycle at a single fellowship site. Two reviewers independently screened each letter for predefined red flags. Disagreements were resolved by consensus or third-party adjudication. A rule-based natural language processing (NLP) model, refined through prompt adjustments, screened the same letters. Time to completion and red flag detection were compared.

Results

A total of 195 LORs were reviewed. Following adjudication, red flags were confirmed in 21 letters. The NLP model flagged 16 letters and showed 76% (16 out of 21) agreement with the final adjudicated review. It processed all letters in 25 minutes, compared to the 554 minutes required by human reviewers. The model reliably identified terms “solid” and “good” with sentence-level context and showed consistency across the dataset, while humans varied more in detection, particularly with vague or indirect phrasing.

Conclusions

A rule-based NLP model offers an efficient and consistent method for initial LOR screening.

Introduction

Qualitative elements of a graduate medical education (GME) application, including letters of recommendation (LORs), are central to the selection process.1,2 With the increasing demands on an academic physician, little time is allocated for application review prior to recruitment of trainees. Many programs use these letters to screen for language signifying a struggling learner or professionalism concerns, also known as “red flags,” and to determine whether to invite the applicant for an interview.

Natural language processing (NLP) models have been cited in the literature as a successful screening tool when evaluating large quantities of text, with some studies quoting upwards of 90% accuracy.3-5 NLP models have been used in screening various aspects of resident applications or in predicting Match outcomes.6-9 Our proposed research question was, how does the performance of a commercially available NLP model compare with human reviewers when used for LOR screening for red flags?

Our hypothesis was that a commercially available NLP model is as accurate as and faster than a human reviewer in identifying predetermined red flags in a cohort of fellowship LORs.

Methods

Study Design and Sample

This was a retrospective comparative study evaluating the performance of an NLP model against human reviewers in identifying predefined concerning language (red flags) within 195 LORs submitted to our institution during the 2024-2025 neonatal-perinatal medicine fellowship application cycle. Fellowship applicants had completed a 3-year Accreditation Council for Graduate Medical Education–accredited pediatric residency prior to matriculation in fellowship. This study was conducted after the fellowship match was concluded. All letters were de-identified prior to analysis. Letters with templated or standardized formats were excluded.

Red Flag Term List Development

A list of red flag terms was developed through review of existing literature1,10 and expert consensus among faculty involved in undergraduate and graduate medical education. Terms included commonly cited phrases that may indicate concern regarding professionalism, performance, or interpersonal behavior. The final list was operationalized into Microsoft Copilot (online supplementary data Table 1).

Human Review Process

Two experienced fellowship selection committee members independently reviewed each letter to identify any of the predefined red flag terms, which is our current standard process. Reviewers each had over 5 years of experience in reviewing GME applicant LORs. However, reviewers did not receive any specialized training to compare current workflow state to a potential future workflow that would include NLP as a screening tool. Reviewers were blinded to each other’s results. Disagreements were resolved through discussion, consistent with standard selection committee procedures. If there was a disagreement between reviewers that could not be resolved with adjudication, a third reviewer (M.S.), who has expertise in reading GME letters, gave the final decision. The gold standard adjudicated letters consisted of the letters the human reviewers agreed upon, plus the third reviewer of any letters the NLP model identified that the human reviewers missed. Total review time was recorded given our standard of dual review.

NLP Model Application

We chose Microsoft Copilot, a commercially available NLP model, for ease of use and generalizability and used it to screen the same letters using a constrained, keyword-based approach. The model applied standardized text preprocessing and identified exact or variant forms of each red flag term. We utilized prompt engineering as an iterative process to develop our final prompt (online supplementary data Table 1). Total time to process the full dataset was recorded.

Outcome Measures and Analysis

Concordance between the NLP model and final human adjudication for red flag identification (binary, per letter), total review time, and interrated agreement between human reviewers before and after consensus was collected and analyzed using GraphPad Prism (version 10) and Cohen’s kappa. Sensitivity, specificity, Wilson’s confidence intervals (CIs), and McNemar’s test P value were calculated using R (Version 4.5.1). The total number of times a red flag was mentioned and counted by AI and human reviewers was recorded.

The University of Virginia Institutional Review Board reviewed this study, determined it to be exempt, and waived informed consent.

Results

A total of 195 letters were reviewed. After discussion and adjudication, red flags were confirmed in 21 letters identified by human review. The distribution of red flags identified by reviewer and AI are depicted in Figure 1. The red flags that the NLP model missed were both in letters where the specialty was wrong (pediatric intensive care unit instead of neonatal intensive care unit) and the level of learner was wrong (resident vs fellow). These discrepancies reflect errors in the letters themselves rather than differences in assessment of the applicants.

Figure 1.

Figure 1

Human vs AI Reviewer Red Flag Detection and Time to Review

Note: Reviewer 1 (R1) identified 16 letters with red flags; Reviewer 2 (R2) identified 18 letters. The first, second, and third iterations of the AI prompt identified 111, 68, and 16 letters with red flags, respectively (Figure 1A). R1 identified 35, and R2 identified 30 red flags across all letters. The AI model identified 27 red flags (Figure 1B). R1 took 199 minutes, and R2 took 355 minutes to complete the task (total 554 minutes). AI took 25 minutes to complete the task (Figure 1C). A human reviewer and the AI identified red flags in 15 letters. In 5 other letters, a human reviewer picked up the red flag, and in 1 letter the AI identified a red flag that the human reviewers both missed (Figure 1D).

Agreement between the 2 reviewers before discussion as described by Cohen’s kappa was (κ=0.81). The model’s findings matched the final human review in 16 of 21 cases (76%). One letter was identified by the model and not a human reviewer (Figure 1D). Comparing the model to each reviewer separately, it performed reliably (κ=0.85 [R1] and κ=0.87 [R2]). Sensitivity of the NLP compared to human review was 0.75 (95% CI, 0.53-0.89), and specificity was 1.00 (95% CI, 0.98-1.00). The correctly classified proportion was 0.97 (95% CI, 0.94-0.99), (P=.07, McNemar’s test P value), indicating no statistically significant difference between Copilot and human reviewers. The NLP tool processed all letters in 25 minutes. In contrast, Reviewer 1 took 199 minutes, and Reviewer 2 took 355 minutes to complete the task (total 554 minutes; Figure 1C).

The model successfully identified predefined red flag terms and phrases. Through refinement of the prompt, including a requirement to return the full sentence containing the term, the model was able to improve performance. It also demonstrated the capacity to recognize differences in sentiment when evaluating words that vary in interpretation, depending on tone and placement, such as “solid” or “good” (Figure 1B).

Red flag categories varied in how often they were identified. The model frequently captured the terms solid and good. Reviewers were more likely to identify language suggesting the letter was generic, not specific to the applicant, or not written for the correct specialty (Figure 2).

Figure 2.

Figure 2

Frequency of Red Flag Types by Reviewer

Discussion

This study demonstrates that a simple, keyword-based NLP tool can screen narrative letters of recommendation for language that may raise concern during the fellowship application process more efficiently than human reviewers and with moderate accuracy. With this approach, a program coordinator could run the initial screening step through Copilot, allowing program directors to focus on reviewing letters flagged for concerning language. This may reduce the time burden on faculty while preserving attention to important applicant characteristics. Prompt engineering was expedited by using previously published terms or phrases that indicated a potentially concerning applicant.1,10 This list was integrated into the NLP model prompt. The most time-intensive step in the process was de-identifying the letters, but this step could be avoided in the application review process if a program used a model that was within their institution’s firewall. Subsequently, once the terms list was inputted into the NLP model, it was able to categorize the terms, which helped in viewing areas in which the applicant would struggle. The categories included: lack of enthusiasm or vague praise (met expectations, solid); qualified praise (with additional supervision they could…); professionalism (lack of teamwork); warnings about clinical ability (will benefit from a structured environment); implied liability (would need continued mentorship).

One notable strength of the model is its consistency. Human reviewers, even when experienced, bring varied expectations and interpretations to a process shaped by limited variation in letter content. Many letters follow similar formats and use comparable phrases, which can make distinctions between applicants difficult to interpret. The model applied its rules without distraction or fatigue, offering a uniform review across all letters.

This study had several limitations, including the small sample size. We tested our prompt with a single application cycle of neonatal-perinatal medicine fellowship LORs. Further testing with larger GME specialties would be recommended as a next step. Another limitation with using current AI technologies such as Copilot is that they do not have perfect code and can hallucinate. That is why this technology must be used in conjunction with a human to verify results, and this LOR screening is just a tool in a larger algorithm to decide who to invite for an interview, not a standalone assessment. Finally, there could be potential biases in the original red flag list used in the prompt. Some of these terms have multiple connotations, and letter writers may have different uses of the same word.

These findings suggest that AI-supported screening may be a useful addition to fellowship selection workflows, particularly in programs with high applicant volumes. While human judgment remains essential for final decision-making, preliminary automation can streamline the process and reduce variability in early screening stages. Next steps are to create a pipeline for uploading letters to our firewall-protected NLP model in order to avoid time-consuming manual entry and integrate this process into program coordinator workflow in subsequent application cycles. For the purposes of this study, letters were de-identified and therefore unable to be correlated with fellow performance. Future research will include adapting the prompt to identify exceptional applicants and applicants likely to rank a program highly, or to predict those that will have academic success in the future.

Conclusions

The NLP model showed high specificity and moderate sensitivity in identifying predefined red flag terms and phrases, with agreement compared to human reviewers. Most discrepancies involved inaccuracies in the original letters rather than differences in classification. The model completed the task faster than human reviewers while producing comparable overall accuracy.

Supplementary Material

JGMED25006761.pdf (176.2KB, pdf)

Author Notes

Funding: The authors report no external funding source for this study.

Conflict of interest: The authors declare they have no competing interests.

Editor’s Note

The online supplementary data contains further data from the study.

References

  • 1.Saudek K, Treat R, Rogers A, et al. A novel faculty development tool for writing a letter of recommendation. PLoS One. 2020;15(12):e0244016. doi: 10.1371/journal.pone.0244016. doi: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Prager JD, Myer CM, 3rd, Pensak ML. Improving the letter of recommendation. Otolaryngol Head Neck Surg. 2010;143(3):327–330. doi: 10.1016/j.otohns.2010.03.017. doi: [DOI] [PubMed] [Google Scholar]
  • 3.Oami T, Okada Y, Nakada T-A. Performance of a large language model in screening citations. JAMA Netw Open. 2024;7(7):e2420496. doi: 10.1001/jamanetworkopen.2024.20496. doi: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Li M, Sun J, Tan X. Evaluating the effectiveness of large language models in abstract screening: a comparative analysis. Syst Rev. 2024;13(1):219. doi: 10.1186/s13643-024-02609-x. doi: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Dennstädt F, Zink J, Putora PM, Hastings J, Cihoric N. Title and abstract screening for literature reviews using large language models: an exploratory study in the biomedical domain. Syst Rev. 2024;13(1):158. doi: 10.1186/s13643-024-02575-4. doi: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Ortiz AV, Feldman MJ, Yengo-Kahn AM, et al. Words matter: using natural language processing to predict neurosurgical residency match outcomes. J Neurosurg. 2023;138(2):559–566. doi: 10.3171/2022.5.JNS22558. doi: [DOI] [PubMed] [Google Scholar]
  • 7.Drum B, Shi J, Peterson B, Lamb S, Hurdle JF, Gradick C. Using natural language processing and machine learning to identify internal medicine-pediatrics residency values in applications. Acad Med. 2023;98(11):1278–1282. doi: 10.1097/ACM.0000000000005352. doi: [DOI] [PubMed] [Google Scholar]
  • 8.Varman PM, Nicholas S, Conner A, Prabhu AS, French JC, Lipman JM. Feasibility of using AI to evaluate general surgery residency application personal statements. J Surg Educ. 2025;82(12):103655. doi: 10.1016/j.jsurg.2025.103655. doi: [DOI] [PubMed] [Google Scholar]
  • 9.Gudgel BM, Melson AT, Dvorak J, Ding K, Siatkowski RM. Correlation of ophthalmology residency application characteristics with subsequent performance in residency. J Acad Ophthalmol (2017) 2021;13(2):e151–e157. doi: 10.1055/s-0041-1733932. doi: [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Gross C, O’Halloran C, Winn AS, et al. Application factors associated with clinical performance during pediatric internship. Acad Pediatr. 2020;20(7):1007–1012. doi: 10.1016/j.acap.2020.03.010. doi: [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

JGMED25006761.pdf (176.2KB, pdf)

Articles from Journal of Graduate Medical Education are provided here courtesy of Accreditation Council for Graduate Medical Education

RESOURCES