Skip to main content
Plastic and Reconstructive Surgery Global Open logoLink to Plastic and Reconstructive Surgery Global Open
. 2026 Apr 22;14(4):e7618. doi: 10.1097/GOX.0000000000007618

A Blind Spot in the Algorithm: Assessing Bias in Artificial Intelligence–Generated Images of Plastic Surgery Patients

Parul Rai 1, Marie M Mina 1, Rolando J Casas Fuentes 1, Gabrielle C Rodriguez 1, Anitesh Bajaj 1, Emily George 1, Kathryn R Reisner 1, Arun K Gosain 1,
PMCID: PMC13102436  PMID: 42028106

Abstract

Background:

Artificial intelligence (AI) models can inherit biases from their training data, a phenomenon documented in other fields, including the corporate and legal realms. This study is the first to investigate such biases in AI-generated images of plastic surgery patients.

Methods:

Three AI image generators were used to generate 2600 images using the prompt “A photo of the face of a ___ patient,” encompassing various plastic surgery patient groups. Images were independently assessed by 3 blinded raters for demographic factors and compared with real-world patient demographics via Fisher exact tests (α = 0.05). Fleiss kappa was calculated to assess interrater reliability.

Results:

The AI-generated images displayed numerous biases. All platforms overrepresented non-White patients in cleft lip images and overrepresented female patients in these images (P < 0.001). Platforms exclusively depicted female patients in aesthetic surgery images involving facial cosmetic surgery and breast augmentation. Non-White patients and those older than 50 years were underrepresented in images related to aesthetic surgery (P < 0.0001). Non-White patients were also underrepresented in images of breast augmentation and breast reconstruction (P < 0.0001).

Conclusions:

Although AI offers valuable applications in education and surgical planning, its outputs necessitate critical evaluation by patients, physicians, and developers. Current depictions of plastic surgery patients by AI platforms can foster stereotypes on sex and ethnicity of patients seeking plastic surgery. This analysis highlights the need for user feedback–based models to prevent biased outputs, harness the power of AI responsibly, and ensure its ethical application in plastic surgery.


Takeaways

Question: What biases exist in artificial intelligence diffusion models when depicting plastic surgery patients?

Findings: This study used a comparative analysis of artificial intelligence–generated images against real-world patient demographics, revealing significant biases in depicting various plastic surgery patient groups and, notably, a consistent failure to accurately depict cleft lip. Such biases include overrepresentation of non-White subjects as cleft lip patients and young, White female subjects as aesthetic patients.

Meaning: As diffusion models become increasingly used in our field (eg, for marketing materials, images of surgical goals), we must be aware of their biases, to correct them, promote their responsible use, and prevent feelings of exclusion among patients.

INTRODUCTION

Artificial intelligence (AI) is a broad field of computer science dedicated to creating intelligent machines that can mimic human cognitive functions.14 Machine learning is a subfield of AI that focuses on developing algorithms that allow computers to learn from data without being explicitly programmed.58 Machine learning is a key enabling technology behind many modern AI applications, such as autonomous vehicles, image recognition, natural language processing, and recommendation systems.6,911 Generative AI specifically focuses on creating new content, such as images, text, music, or videos,12 by training from large datasets of existing content sourced from publicly available data from the internet,13 open-source datasets, consumer data, and user-generated content.14,15 The quality of these datasets directly impacts the performance and potential biases of AI models.16,17

Real-world consequences of AI bias underscore the urgent need for ethical development. For example, biased risk assessment algorithms in courtrooms disproportionately flag Black defendants as future criminals,18 whereas AI hiring tools have discriminated against women.19 Even facial recognition systems from leading companies exhibit higher error rates for women, young people, and people of color.20 The medical field is not immune to such biases: algorithms have underestimated the level of sickness in Black patients compared with White patients and in women compared with men.21,22

This study aimed to assess AI biases in representing various plastic surgery patient groups and to juxtapose how their biases manifest in aesthetic versus reconstructive realms. It investigates the degree to which various generative AI models accurately reflect patient demographics and their capability of depicting common plastic surgery conditions. By addressing these aims, this study provided valuable insights into the validity of AI in representing plastic surgery patients while exposing biases.

METHODS

Image Generation

The study used 3 prominent AI image generators: Midjourney (Midjourney, Inc., San Francisco, CA), DreamStudio (Stability AI, London, England), and Leonardo.Ai (Leonardo AI, Sydney, Australia). The former 2 platforms were recognized in the 2024 AI Index Report among the most notable generative AI models in 2023,23 and Leonardo.Ai was selected due to its distinct focus on AI art generation.24 One hundred images were generated on each platform using the prompt “A photo of the face of a ___ patient.” The blank in the prompt was filled with various patient categories and subcategories, listed in Table 1.

Table 1.

Patient Groups

Area of Plastic Surgery Patient Group
Facial plastic surgery Cleft lip
Cleft lip, before repair*
Cleft lip, after repair*
Facial plastic surgery
Breast surgery Breast reconstruction
Breast augmentation
Burn surgery Burn reconstruction
Burn scar
Pediatric burn scar

One hundred images per patient category were generated on each platform (Midjourney, DreamStudio, Leonardo.Ai) using the prompt “A photo of the face of a ___ patient,” yielding 2600 total images.

*For cleft lip, the modifiers “before repair” and “after repair” were added after initial prompts failed to generate representative images.

Among burn surgery prompts, the term “burn reconstruction” was initially used; however, the predominance of acute burn images necessitated the addition of “burn scars” and “pediatric burn scar” to better reflect the patient population encountered in plastic surgery. Prompt engineering was used to specifically test for prompt biases, as the AI models failed to depict cleft lip accurately when timing was not specified, occasionally producing images with a perioral lesion. This necessitated the addition of terms such as “before repair” and “after repair.” Data from cleft lip images categorized as “before repair,” “after repair,” and unspecified time of repair were aggregated into a single “pooled” cohort to allow for a broader evaluation of the AI models’ depictions of cleft lip. Likewise, data from the burn surgery subcategories were aggregated into a pooled cohort.

Image Selection and Exclusion Criteria

Images with a single, unambiguous subject in the foreground were included. Images containing a figure in the background were permissible if the primary subject in the foreground remained distinct. A subset of the AI-generated images used for analysis is displayed in Figure 1. To facilitate proper evaluation, the subject’s eyes and nose had to be clearly visible. Images featuring multiple figures were excluded from the analysis (Table 2, Fig. 2).

Fig. 1.

Fig. 1.

Sample of AI-generated patient images. Examples from each patient group across platforms. A, Midjourney. B, Stability AI (DreamStudio). C, Leonardo.AI.

Table 2.

Excluded Images by Platform and Reasons

Reason for Exclusion Midjourney, n (%) DreamStudio, n (%) Leonardo.Ai, n (%) Total by Reason
Insufficient identifiable facial features 45 (97.8) 0 (0.0) 1 (2.2) 46
Multiple subjects in foreground 1 (1.9) 2 (3.8) 50 (94.3) 53
Total exclusions by platform 46 (46.5) 2 (2.0) 51 (51.5) 99

Images lacking a single identifiable facial subject or containing multiple foreground subjects were excluded. Ninety-nine images were excluded based on predefined criteria.

Fig. 2.

Fig. 2.

Examples of excluded AI-generated patient images. Images excluded due to inadequate facial visualization or multiple foreground subjects. A, Obscured facial anatomy. B, Multiple individuals present.

Image Assessment

Three raters independently evaluated each image for the subject’s race (ie, White or non-White), sex (ie, male or female), and whether the subject was depicted in personal protective equipment (PPE), defined as garments or equipment worn to protect the wearer’s body from injury or infection, such as gloves, surgical masks, gowns, and eye protection. Raters assessed the subject’s apparent age (ie, older or younger than 50 y) in images depicting patients who underwent breast reconstruction, breast augmentation, and facial plastic surgery. Raters additionally evaluated the presence of facial lesions or perioral lesions in images of patients who underwent burn surgery and cleft lip repair, respectively. The perioral region was defined as the area within 1 cm of the vermilion border.

Statistical Analysis

Statistical analysis was performed using SPSS Statistics version 29.0.2.0 (IBM, Armonk, NY). Fisher exact tests were used to compare the demographic characteristics observed in the AI-generated images with real-world epidemiological data and to analyze whether a platform or prompt had a greater tendency to generate images with particular demographic or visual attributes. Fleiss kappa (κ) was calculated to evaluate interrater reliability across the following assessment categories: sex, race, age, presence of facial or perioral lesions, and presence of PPE. Statistical significance was defined as a P value less than 0.05.

Sourcing Epidemiological Data

Real-world patient demographics were sourced from various authoritative databases. Sex and racial demographic information for patients with cleft lip (n = 37,048) was retrieved from the TriNetX (sourced from electronic medical records from approximately 132 million patients from 68 healthcare organizations) using the International Classification of Diseases, 10th revision code Q36 and the National Birth Defects Prevention Network databases, respectively. Similarly, demographic data pertaining to breast reconstruction (n= 15,920) and burn surgery (n = 105,191) patients were extracted from the TriNetX database per their respective CPT codes. (See table, Supplemental Digital Content 1, which displays the TriNetX data, including CPT codes and procedure names for patients who underwent burn [n = 105,191] and breast reconstruction [n = 15,920], used to compare real-world demographics against AI-generated images. Because individual patients may possess multiple codes, these figures do not necessarily sum to the total category population, https://links.lww.com/PRSGO/E762.) Sex and age distribution data regarding the aesthetic patient groups—facial cosmetic surgery (n = 391,516) and breast augmentation (n = 249,560)—were obtained from The Aesthetic Society in collaboration with CosmetAssure.25 To determine the number of patients undergoing facial cosmetic surgery, procedures falling under the “Head & Face” category on The Aesthetic Society website were summed (Table 3). Because The Aesthetic Society does not publish data on racial demographics, these figures were estimated by extrapolating from percentages reported in the 2020 American Society of Plastic Surgeons Plastic Surgery Statistics Report, the most recent publicly available dataset describing this parameter.26

Table 3.

Sex and Age Distribution of Facial Cosmetic Surgery Patients

Procedure Female, n (%) Male, n (%) 17 y and Younger, n (%) 18–34 y, n (%) 35–50, y, n (%) 51–64, y, n (%) 65 y and Older, n (%) Total
Brow lift 26,809 (46.2) 2182 (3.8) 0 (0.0) 619 (1.1) 3522 (6.1) 14,399 (24.8) 10,451 (18.0) 57,982
Chin augmentation 6163 (35.5) 2521 (14.5) 202 (1.2) 5049 (29.1) 2020 (11.6) 1010 (5.8) 404 (2.3) 17,369
Ear surgery 26,633 (35.6) 10,808 (14.4) 8915 (11.9) 15,282 (20.4) 6368 (8.5) 3821 (5.1) 3056 (4.1) 74,883
Eyelid surgery 98,498 (43.2) 15,498 (6.8) 0 (0.0) 1705 (0.7) 17,868 (7.8) 54,678 (24.0) 39,745 (17.4) 227,992
Face lift 89,368 (46.5) 6693 (3.5) 0 (0.0) 294 (0.2) 7470 (3.9) 46,457 (24.2) 41,840 (21.8) 192,122
Fat transfer: face 18,460 (45.4) 1866 (4.6) 0 (0.0) 337 (0.8) 2897 (7.1) 9568 (23.5) 7524 (18.5) 40,652
Neck lift 37,031 (43.9) 5122 (6.1) 0 (0.0) 496 (0.6) 5264 (6.2) 20,676 (24.5) 15,717 (18.6) 84,306
Nose surgery 37,383 (42.6) 6481 (7.4) 1477 (1.7) 20,423 (23.3) 13,615 (15.5) 6486 (7.4) 1862 (2.1) 87,727
Total 340,345 (43.5) 51,171 (6.5) 10,594 (1.4) 44,205 (5.6) 59,024 (7.5) 157,095 (20.1) 120,599 (15.4) 783,033

Demographic data for facial cosmetic procedures categorized as “Head & Face” were obtained from The Aesthetic Society in collaboration with CosmetAssure. These data were used as epidemiological comparators for AI-generated images.

RESULTS

Demographics

Overall, 2600 images were generated across all patient groups and AI models (Fig. 1). Interrater reliability analysis revealed moderate to substantial agreement (0.5 < κ < 0.8) among the raters when assessing demographic factors in pediatric and adult populations and near-perfect agreement (κ > 0.8) when evaluating nondemographic visual elements (Table 4).27 A summary of the findings in this section may be found in Figure 3.

Table 4.

Interrater Reliability

Patient Group κ P
Sex, pediatric 0.669 <0.00001
Race, pediatrics 0.762 <0.00001
Sex, adults 0.697 <0.00001
Race, adults 0.505 <0.00001
Age 0.669 <0.00001
PPE 0.813 <0.00001
Lesion 0.874 <0.00001

Agreement among raters ranged from moderate to substantial for demographic variables and was near-perfect for nondemographic visual features. κ indicates Fleiss kappa; a P value less than 0.05 denotes statistical significance.

Fig. 3.

Fig. 3.

Demographic representation in AI images by patient group and platform. Comparison of AI-generated demographics with epidemiological reference data. A, Percentage of female patients by procedure. B, Percentage of non-White patients by procedure. C, Percentage of patients older than 50 years. “%50+” represents the percentage of patients aged 50 years or older. *P < 0.05, **P < 0.01, ***P < 0.001, ****P < 0.0001.

All platforms overrepresented non-White patients in cleft lip images (P < 0.0001 for Midjourney and DreamStudio) and female patients in these images (P < 0.001). All platforms exclusively depicted female patients in facial aesthetic surgery images, yet significantly underrepresented non-White patients and patients older than 50 years (P < 0.0001).

All platforms significantly underrepresented non-White patients undergoing breast augmentation (P < 0.0001). Although patients aged older than 50 years constitute only 14.0% of breast augmentation patients,25 Midjourney and DreamStudio accentuated this demographic skew by lacking representation of older individuals in their generated images entirely (P < 0.0001). All platforms demonstrated a significant underrepresentation of non-White patients in the context of breast reconstruction (P < 0.0001). Midjourney overrepresented women older than 50 years (P < 0.05), whereas DreamStudio and Leonardo.Ai underrepresented this demographic (P < 0.0001).

Upon pooled analysis of the adult burn surgery images, female patients were overrepresented (P < 0.0001), and racial representation varied by platform, with Midjourney and DreamStudio underrepresenting non-White patients (P < 0.0001) and Leonardo.Ai overrepresenting them (P < 0.0001) (Fig. 4). When isolating images of patients with pediatric burn scars, Midjourney and DreamStudio significantly underrepresented non-White individuals (P < 0.0001). This finding contrasts with the overrepresentation of non-White patients observed in cleft lip images, highlighting the inconsistency of biases across different surgical categories and AI platforms. No significant gender disparities were observed in the depiction of pediatric burn patients. Leonardo.Ai was excluded from the analysis of this patient group, as it flagged the prompt requesting “pediatric” patients as inappropriate content and consequently did not generate any images.

Fig. 4.

Fig. 4.

Heatmaps of demographic differences between AI output and real-world demographics. A, Percentage difference in women relative to real-world demographics. B, Percentage difference in non-White patients relative to real-world demographics. C, Percentage difference in the number of women aged 50 years or older relative to real-world demographics.

Nondemographic Visual Elements

PPE Use

Subjects were more likely to be shown wearing PPE after reconstructive procedures (ie, cleft lip surgery, breast reconstruction, burn reconstruction) compared with aesthetic procedures (ie, facial cosmetic surgery, breast augmentation) (P < 0.01). Although no significant difference in PPE use was observed between images of pediatric patients with burn scar and those of unspecified age, patients who underwent “burn reconstruction” were consistently more likely to be shown in PPE compared with patients with “burn scar” across all platforms (P < 0.001) (Table 5).

Table 5.

Comparisons in Depictions of Perioral/Facial Lesions

Patient Group A Patient Group B Midjourney DreamStudio Leonardo.Ai
% Lesion A % Lesion B P % Lesion A % Lesion B P % Lesion A % Lesion B P
Cleft lip, before repair Cleft lip, after repair 94 63 <0.0001 1 1 1 4 5 0.734
Cleft lip, pooled 78.5 1.0 4.5
Burn scar, unspecified Burn scar, pediatric 100 98 0.497 100 1 1 100
Burn scar, pooled Burn reconstruction 99.3 100 0.554 97.3 92 0.004 100.0 100 1

Paired groups (A versus B) assess differences by procedure type or terminology. Significant platform-dependent differences were observed in cleft lip and burn-related image depictions.

Lesions

Our analysis revealed 2 statistically significant findings regarding the depiction of lesions in AI-generated images. Midjourney was the only platform that demonstrated a rudimentary understanding of cleft lip, portraying patients with perioral lesions more frequently in “before repair” prompts (P < 0.0001). DreamStudio exhibited a significant difference in the depiction of facial lesions among patients who underwent burn surgery, with patients with “burn scar” more likely to exhibit such lesions compared with patients who underwent “burn reconstruction” (P = 0.004). No significant difference was observed in the depiction of facial lesions between patients with burn scars of unspecified age and pediatric patients with burn scars (Table 5).

DISCUSSION

Our investigation offers novel insights into the accuracy of demographic representation across various plastic surgery patient populations within 3 publicly available text-to-image generative AI models. As summarized in Figure 4, the models demonstrated significant overrepresentation of non-White patients in cleft lip images while underrepresenting them in other categories, such as aesthetic surgery and breast reconstruction. This discrepancy suggests a biased training dataset rather than an accurate reflection of epidemiological data. Unlike the other 2 platforms, Leonardo.Ai was not included in the 2024 AI Index Report report23; however, its distinct focus on AI art generation motivated its inclusion in our study.24 Our findings confirm a commonly discussed pitfall in AI algorithms—the sex and racial bias associated with certain prompts.1822 The biases we found demonstrate the specific manifestation of this problem within the context of plastic surgery, where outputs can perpetuate stereotypes regarding the sex and ethnicity of patients seeking care. Highlighting this bias is useful to allow for improvements to be made in the future, diminishing these biases and allowing for more equal representation.

Cleft Lip

The AI models showed a pronounced bias by overrepresenting non-White patients and women in cleft lip images, despite epidemiological data not supporting a higher incidence rate in these groups.28 This suggests a biased training dataset, possibly influenced by images from humanitarian missions. Although non-White children may be frequently featured in cleft lip–related charitable advertisements, this visibility often stems from efforts to provide surgical care in under-resourced regions worldwide and does not necessarily reflect higher incidence rates. Although one could argue that this bias might be attributed to the larger populations in countries such as India and China, resulting in a higher absolute number of cleft lip cases, this explanation is undermined by the observation that non-White subjects were underrepresented in AI-generated images depicting other patient groups.

Furthermore, none of the platforms could accurately depict a cleft lip, indicating a fundamental lack of understanding of the condition (Fig. 2). This observation raises concerns about the AI models’ fundamental comprehension of cleft lip. Notably, prompt engineering by specifying “before repair” versus “after repair” proved to be effective to some extent, as it significantly influenced the proportion of subjects depicted with perioral lesions in images generated by Midjourney.

Aesthetic Procedures

The platforms consistently portrayed a narrow and unrealistic demographic of aesthetic patients, namely those represented in breast augmentation and facial cosmetic surgery. Compared with their reconstructive counterparts, the aesthetic procedures tended to depict a disproportionately higher number of White female individuals (Figs. 3A, B).

They exclusively depicted female subjects in images of patients who underwent facial cosmetic surgery, despite a significant number of male patients in these groups.25 This discrepancy would likely be further augmented if nonsurgical procedures such as botulinum toxin injections were included in the epidemiological analysis. Moreover, the consistent depiction of women predominantly younger than 50 years in facial plastic surgery images across all 3 platforms contradicts real-world demographics, considering that the vast majority (70.9%) of this clientele is 50 years or older (Tables 3, 4). In addition to perpetuating biases and stereotypes, this misrepresentation of aesthetic patients can reinforce unrealistic beauty standards and hinder patients’ understanding of achievable surgical outcomes, potentially leading to dissatisfaction and unrealistic expectations.

PPE Use

The consistent trend of portraying patients after cleft lip surgery with PPE more frequently than patients before surgery suggests that the AI associates PPE with the presence of surgical intervention. Furthermore, the platforms’ greater propensity to depict patients undergoing reconstructive procedures in PPE compared with patients undergoing aesthetic procedures may reflect a subtle bias toward associating reconstructive procedures with a more medicalized context. This distinction, although subtle, may be a consequence of prompt bias triggered by the terminology used in the input and could inadvertently undermine the holistic care provided to all plastic surgery patients, as both reconstructive and aesthetic patients undergo surgical interventions with inherent risks and require the medical expertise of a plastic surgeon (Table 6).

Table 6.

Comparisons in Depictions of PPE Usage

Patient Group A Patient Group B Midjourney DreamStudio Leonardo.Ai
% PPE A % PPE B P % PPE A % PPE B P % PPE A % PPE B P
Cleft lip, pooled Facial cosmetic surgery 21.0 16.0 0.277 17.3 1.0 <0.001 23.3 4.0 <0.001
Cleft lip, before Cleft lip, after 11.0 38.0 <0.001 13.0 35.0 <0.001 25.0 38.0 0.048
Cleft lip, unspecified 14.0 4.0 7.0
Breast reconstruction Breast augmentation 87.0 4.0 < 0.001 2.0 0.0 0.156 10.0 1.0 0.005
Burn scar, unspecified Burn scar, pediatric 14.0 9.0 0.269 6.0 1.0 0.055 3.0
Burn scar, pooled Burn reconstruction 11.5 60.0 <0.001 3.5 22.0 <0.001 3.0 18.0 <0.001

Paired groups (A versus B) compare reconstructive and aesthetic procedures or terminology within the same anatomical region. PPE was more frequently depicted in reconstructive groups, including cleft lip and breast reconstruction.

Burn Surgery

Contrary to the hypothesis, the AI models showed a trend of overrepresenting White individuals in pediatric burn surgery images, despite socioeconomic factors that disproportionately affect non-White populations.29 Moreover, numerous socioeconomic factors, including living in low-income areas, single-parent households, and overcrowded conditions, contribute to an increased burn risk,3032 disproportionately affecting non-White populations.3335 This observation suggests that the AI models may lack a nuanced understanding of social determinants of health that contribute to disparities in certain medical conditions.

Ethical Considerations

Popular models such as DALL-E 3, accessed via ChatGPT (OpenAI, San Francisco, CA), and Imagen 3, accessed via Google Gemini (Google LLC, Mountain View, CA), were initially considered for this study but were ultimately excluded from the analysis. DALL-E 3, although incorporating feedback mechanisms to address biased images,36 does not generate images in response to prompts requesting patient images, citing ethical concerns related to patient privacy. This restriction, although well-intentioned, raises questions about the necessity of such limitations when other platforms readily generate similar images, highlighting the evolving landscape of ethical considerations in AI. In creating Imagen 2, developers prioritized diverse and inclusive image outputs; however, this well-intentioned effort inadvertently led to the production of inaccurate—even offensive—representations. Consequently, Google has suspended the generation of images depicting individuals, even using the more updated Imagen 3, underscoring the complexities and challenges inherent in developing ethical and responsible AI.37 Leonardo.Ai exhibited a different approach to ethical considerations by flagging prompts containing “pediatric patient” as inappropriate content and refusing to generate images of children. (See figure, Supplemental Digital Content 2, which displays various platform restrictions on patient and sensitive content, highlighting diverse ethical approaches and inconsistencies across the evolving AI landscape. These differences underscore the urgent need for clear guidelines to ensure the responsible and equitable development of generative models, https://links.lww.com/PRSGO/E763.)

The nascent nature of AI technology presents a unique challenge in terms of regulation. Although some degree of governmental oversight may mitigate potential harms, excessive intervention could stifle innovation and limit the potential benefits of AI. Currently, a patchwork of regulations exists, with varying degrees of stringency across different jurisdictions.38 Geolocation-targeted results may further complicate the regulatory landscape, as platforms adapt their outputs to comply with local laws and cultural norms.

Beyond formal regulation, it is crucial to recognize the influence of financial and political factors on AI development and deployment. The financial interests of funding organizations can shape the data used to train AI models, potentially introducing biases that reflect those interests.38,39 Similarly, political motivations, such as appeasing powerful groups or promoting specific ideological agendas, can influence AI policies and public pronouncements. Navigating these complex influences requires a balanced approach that fosters innovation while safeguarding ethical principles and promoting societal well-being.

Although the real-world comparative data in this study are based on American datasets, the AI platforms investigated are used globally, suggesting that the biases we identified are not confined to a single country. The underrepresentation of non-White patients in aesthetic surgery, for example, becomes even more jarring at the global scale. This study thus emphasizes the need for a global effort to create more diverse and representative datasets to ethically and accurately represent patients from all backgrounds and regions.

Dataset Composition and Future Directions

The biases observed in our study, such as the overrepresentation of young, White female patients in aesthetic surgery images and the inconsistent racial representation in other patient groups, are a direct result of the datasets used to train these generative AI models. These models are trained on vast amounts of data scraped from the internet,1315 which can perpetuate societal biases and stereotypes present in that data.16,17 For example, the disproportionate depiction of non-White children in cleft lip images may stem from a dataset oversaturated with images from humanitarian organizations or charitable campaigns. Similarly, the exclusive portrayal of women in aesthetic procedures likely reflects the disproportionate media representation of women undergoing such procedures.

To mitigate these biases, concrete solutions must be implemented by developers and users alike. Developers should prioritize the use of diverse and representative training datasets that accurately reflect global demographics and patient populations. Furthermore, they should increase transparency regarding the composition of their datasets, allowing for independent audits to identify and rectify biases. The development of user feedback–dependent algorithms is also crucial, enabling the models to learn from and correct biased outputs identified by users. On the user side, patients and physicians must be aware of these potential biases and critically evaluate AI outputs, recognizing that technological advancement does not inherently guarantee accuracy or freedom from bias.

Despite the biases observed in this study, AI has great potential for positive impact in health care, already being used in medicine to augment decision-making and enhance diagnostic capabilities. An application of AI algorithms, web scraping automates data collection from the web, drastically reducing the time it takes for research and expanding the evidence base in medicine and surgery.40 AI also has the potential to transform surgical care by augmenting critical decisions, such as the decision to operate and the informed consent process. For instance, the “My Surgery Risk” AI platform can use electronic health record data to predict postoperative complications and mortality with greater accuracy than physicians.41 Additionally, AI models are used for predictive analytics and diagnostics, capable of accurately predicting conditions such as sepsis and acute kidney injury.42 These beneficial applications highlight AI’s valuable role as a tool for augmenting human expertise, reinforcing the idea that although AI holds immense promise, its integration must be approached with a critical eye to address and mitigate biases such as those demonstrated in our study.

Limitations

This study is subject to some limitations that warrant consideration. Determining the sex of patients with pediatric cleft lip proved challenging due to prepubertal facial features, necessitating dependence on hair length and clothing style, which are not necessarily reliable indicators of sex. However, the high interrater reliability in this study strengthened the validity and consistency of the study’s findings. The analysis did not delve into a full racial breakdown, limiting it to White versus non-White individuals to mitigate the ambiguity associated with assessing race from images. Datasets used for comparison to real-world trends focus on the patient population in the United States and do not represent the global population, limiting the generalizability of these findings relative to other countries. Additionally, the American Society of Plastic Surgeons data were restricted to specific facial procedures, excluding nonsurgical procedures such as botulinum toxin and hyaluronic acid fillers, which are often performed at higher rates among older clientele. Furthermore, because Leonardo.Ai flagged prompts involving “pediatric” patients as inappropriate content, the assessment of the pediatric burn scar group was limited to Midjourney and DreamStudio. Consequently, Leonardo.Ai exclusively generated images of adults for cleft lip. These limitations highlight the need for future research with larger, more diverse datasets and more sophisticated AI models capable of accurately representing the full spectrum of plastic surgery patients.

CONCLUSIONS

This study provides compelling evidence of significant biases in AI-generated images of plastic surgery patients, raising concerns about their accuracy and potential to perpetuate harmful stereotypes. The overrepresentation of young, White female individuals in aesthetic surgery and the underrepresentation of non-White individuals in other surgical contexts, coupled with their inability to accurately depict cleft lip, highlight the limitations of current AI models. Additionally, this analysis highlighted implications of prompt bias, where the specific wording of the input can lead to inaccurate or stereotyped outputs. These biases can mislead patients and trainees, potentially influencing treatment decisions and perpetuating unrealistic expectations regarding surgical outcomes.

To use AI responsibly in medical practice, physicians should remain aware of how their patient population is represented by AI and recognize that technological advancement does not inherently guarantee accuracy or freedom from bias. This places an added responsibility upon healthcare providers to critically evaluate AI outputs, accurately represent these conditions for education and patient care, and advocate for the development of ethical and equitable AI tools. Although AI is revolutionizing information access, ongoing scrutiny and development are crucial to ensure equitable and ethical applications. Our findings highlighted the need for more diverse and representative training datasets, transparency regarding these datasets, and the development of user feedback–dependent algorithms. By addressing these challenges, we can harness the power of AI responsibly to enhance patient care and education in plastic surgery.

DISCLOSURE

The authors have no financial interest to declare in relation to the content of this article.

Supplementary Material

gox-14-e7618-s001.pdf (116.3KB, pdf)
gox-14-e7618-s002.pdf (1.6MB, pdf)

Footnotes

Published online 22 April 2026.

Disclosure statements are at the end of this article, following the correspondence information.

Related Digital Media are available in the full-text version of the article on www.PRSGlobalOpen.com.

REFERENCES

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

gox-14-e7618-s001.pdf (116.3KB, pdf)
gox-14-e7618-s002.pdf (1.6MB, pdf)

Articles from Plastic and Reconstructive Surgery Global Open are provided here courtesy of Wolters Kluwer Health

RESOURCES