Abstract
Background:
Artificial intelligence (AI) models can inherit biases from their training data, a phenomenon documented in other fields, including the corporate and legal realms. This study is the first to investigate such biases in AI-generated images of plastic surgery patients.
Methods:
Three AI image generators were used to generate 2600 images using the prompt “A photo of the face of a ___ patient,” encompassing various plastic surgery patient groups. Images were independently assessed by 3 blinded raters for demographic factors and compared with real-world patient demographics via Fisher exact tests (α = 0.05). Fleiss kappa was calculated to assess interrater reliability.
Results:
The AI-generated images displayed numerous biases. All platforms overrepresented non-White patients in cleft lip images and overrepresented female patients in these images (P < 0.001). Platforms exclusively depicted female patients in aesthetic surgery images involving facial cosmetic surgery and breast augmentation. Non-White patients and those older than 50 years were underrepresented in images related to aesthetic surgery (P < 0.0001). Non-White patients were also underrepresented in images of breast augmentation and breast reconstruction (P < 0.0001).
Conclusions:
Although AI offers valuable applications in education and surgical planning, its outputs necessitate critical evaluation by patients, physicians, and developers. Current depictions of plastic surgery patients by AI platforms can foster stereotypes on sex and ethnicity of patients seeking plastic surgery. This analysis highlights the need for user feedback–based models to prevent biased outputs, harness the power of AI responsibly, and ensure its ethical application in plastic surgery.
Takeaways
Question: What biases exist in artificial intelligence diffusion models when depicting plastic surgery patients?
Findings: This study used a comparative analysis of artificial intelligence–generated images against real-world patient demographics, revealing significant biases in depicting various plastic surgery patient groups and, notably, a consistent failure to accurately depict cleft lip. Such biases include overrepresentation of non-White subjects as cleft lip patients and young, White female subjects as aesthetic patients.
Meaning: As diffusion models become increasingly used in our field (eg, for marketing materials, images of surgical goals), we must be aware of their biases, to correct them, promote their responsible use, and prevent feelings of exclusion among patients.
INTRODUCTION
Artificial intelligence (AI) is a broad field of computer science dedicated to creating intelligent machines that can mimic human cognitive functions.1–4 Machine learning is a subfield of AI that focuses on developing algorithms that allow computers to learn from data without being explicitly programmed.5–8 Machine learning is a key enabling technology behind many modern AI applications, such as autonomous vehicles, image recognition, natural language processing, and recommendation systems.6,9–11 Generative AI specifically focuses on creating new content, such as images, text, music, or videos,12 by training from large datasets of existing content sourced from publicly available data from the internet,13 open-source datasets, consumer data, and user-generated content.14,15 The quality of these datasets directly impacts the performance and potential biases of AI models.16,17
Real-world consequences of AI bias underscore the urgent need for ethical development. For example, biased risk assessment algorithms in courtrooms disproportionately flag Black defendants as future criminals,18 whereas AI hiring tools have discriminated against women.19 Even facial recognition systems from leading companies exhibit higher error rates for women, young people, and people of color.20 The medical field is not immune to such biases: algorithms have underestimated the level of sickness in Black patients compared with White patients and in women compared with men.21,22
This study aimed to assess AI biases in representing various plastic surgery patient groups and to juxtapose how their biases manifest in aesthetic versus reconstructive realms. It investigates the degree to which various generative AI models accurately reflect patient demographics and their capability of depicting common plastic surgery conditions. By addressing these aims, this study provided valuable insights into the validity of AI in representing plastic surgery patients while exposing biases.
METHODS
Image Generation
The study used 3 prominent AI image generators: Midjourney (Midjourney, Inc., San Francisco, CA), DreamStudio (Stability AI, London, England), and Leonardo.Ai (Leonardo AI, Sydney, Australia). The former 2 platforms were recognized in the 2024 AI Index Report among the most notable generative AI models in 2023,23 and Leonardo.Ai was selected due to its distinct focus on AI art generation.24 One hundred images were generated on each platform using the prompt “A photo of the face of a ___ patient.” The blank in the prompt was filled with various patient categories and subcategories, listed in Table 1.
Table 1.
Patient Groups
| Area of Plastic Surgery | Patient Group |
|---|---|
| Facial plastic surgery | Cleft lip |
| Cleft lip, before repair* | |
| Cleft lip, after repair* | |
| Facial plastic surgery | |
| Breast surgery | Breast reconstruction |
| Breast augmentation | |
| Burn surgery | Burn reconstruction |
| Burn scar | |
| Pediatric burn scar |
One hundred images per patient category were generated on each platform (Midjourney, DreamStudio, Leonardo.Ai) using the prompt “A photo of the face of a ___ patient,” yielding 2600 total images.
*For cleft lip, the modifiers “before repair” and “after repair” were added after initial prompts failed to generate representative images.
Among burn surgery prompts, the term “burn reconstruction” was initially used; however, the predominance of acute burn images necessitated the addition of “burn scars” and “pediatric burn scar” to better reflect the patient population encountered in plastic surgery. Prompt engineering was used to specifically test for prompt biases, as the AI models failed to depict cleft lip accurately when timing was not specified, occasionally producing images with a perioral lesion. This necessitated the addition of terms such as “before repair” and “after repair.” Data from cleft lip images categorized as “before repair,” “after repair,” and unspecified time of repair were aggregated into a single “pooled” cohort to allow for a broader evaluation of the AI models’ depictions of cleft lip. Likewise, data from the burn surgery subcategories were aggregated into a pooled cohort.
Image Selection and Exclusion Criteria
Images with a single, unambiguous subject in the foreground were included. Images containing a figure in the background were permissible if the primary subject in the foreground remained distinct. A subset of the AI-generated images used for analysis is displayed in Figure 1. To facilitate proper evaluation, the subject’s eyes and nose had to be clearly visible. Images featuring multiple figures were excluded from the analysis (Table 2, Fig. 2).
Fig. 1.
Sample of AI-generated patient images. Examples from each patient group across platforms. A, Midjourney. B, Stability AI (DreamStudio). C, Leonardo.AI.
Table 2.
Excluded Images by Platform and Reasons
| Reason for Exclusion | Midjourney, n (%) | DreamStudio, n (%) | Leonardo.Ai, n (%) | Total by Reason |
|---|---|---|---|---|
| Insufficient identifiable facial features | 45 (97.8) | 0 (0.0) | 1 (2.2) | 46 |
| Multiple subjects in foreground | 1 (1.9) | 2 (3.8) | 50 (94.3) | 53 |
| Total exclusions by platform | 46 (46.5) | 2 (2.0) | 51 (51.5) | 99 |
Images lacking a single identifiable facial subject or containing multiple foreground subjects were excluded. Ninety-nine images were excluded based on predefined criteria.
Fig. 2.
Examples of excluded AI-generated patient images. Images excluded due to inadequate facial visualization or multiple foreground subjects. A, Obscured facial anatomy. B, Multiple individuals present.
Image Assessment
Three raters independently evaluated each image for the subject’s race (ie, White or non-White), sex (ie, male or female), and whether the subject was depicted in personal protective equipment (PPE), defined as garments or equipment worn to protect the wearer’s body from injury or infection, such as gloves, surgical masks, gowns, and eye protection. Raters assessed the subject’s apparent age (ie, older or younger than 50 y) in images depicting patients who underwent breast reconstruction, breast augmentation, and facial plastic surgery. Raters additionally evaluated the presence of facial lesions or perioral lesions in images of patients who underwent burn surgery and cleft lip repair, respectively. The perioral region was defined as the area within 1 cm of the vermilion border.
Statistical Analysis
Statistical analysis was performed using SPSS Statistics version 29.0.2.0 (IBM, Armonk, NY). Fisher exact tests were used to compare the demographic characteristics observed in the AI-generated images with real-world epidemiological data and to analyze whether a platform or prompt had a greater tendency to generate images with particular demographic or visual attributes. Fleiss kappa (κ) was calculated to evaluate interrater reliability across the following assessment categories: sex, race, age, presence of facial or perioral lesions, and presence of PPE. Statistical significance was defined as a P value less than 0.05.
Sourcing Epidemiological Data
Real-world patient demographics were sourced from various authoritative databases. Sex and racial demographic information for patients with cleft lip (n = 37,048) was retrieved from the TriNetX (sourced from electronic medical records from approximately 132 million patients from 68 healthcare organizations) using the International Classification of Diseases, 10th revision code Q36 and the National Birth Defects Prevention Network databases, respectively. Similarly, demographic data pertaining to breast reconstruction (n = 15,920) and burn surgery (n = 105,191) patients were extracted from the TriNetX database per their respective CPT codes. (See table, Supplemental Digital Content 1, which displays the TriNetX data, including CPT codes and procedure names for patients who underwent burn [n = 105,191] and breast reconstruction [n = 15,920], used to compare real-world demographics against AI-generated images. Because individual patients may possess multiple codes, these figures do not necessarily sum to the total category population, https://links.lww.com/PRSGO/E762.) Sex and age distribution data regarding the aesthetic patient groups—facial cosmetic surgery (n = 391,516) and breast augmentation (n = 249,560)—were obtained from The Aesthetic Society in collaboration with CosmetAssure.25 To determine the number of patients undergoing facial cosmetic surgery, procedures falling under the “Head & Face” category on The Aesthetic Society website were summed (Table 3). Because The Aesthetic Society does not publish data on racial demographics, these figures were estimated by extrapolating from percentages reported in the 2020 American Society of Plastic Surgeons Plastic Surgery Statistics Report, the most recent publicly available dataset describing this parameter.26
Table 3.
Sex and Age Distribution of Facial Cosmetic Surgery Patients
| Procedure | Female, n (%) | Male, n (%) | 17 y and Younger, n (%) | 18–34 y, n (%) | 35–50, y, n (%) | 51–64, y, n (%) | 65 y and Older, n (%) | Total |
|---|---|---|---|---|---|---|---|---|
| Brow lift | 26,809 (46.2) | 2182 (3.8) | 0 (0.0) | 619 (1.1) | 3522 (6.1) | 14,399 (24.8) | 10,451 (18.0) | 57,982 |
| Chin augmentation | 6163 (35.5) | 2521 (14.5) | 202 (1.2) | 5049 (29.1) | 2020 (11.6) | 1010 (5.8) | 404 (2.3) | 17,369 |
| Ear surgery | 26,633 (35.6) | 10,808 (14.4) | 8915 (11.9) | 15,282 (20.4) | 6368 (8.5) | 3821 (5.1) | 3056 (4.1) | 74,883 |
| Eyelid surgery | 98,498 (43.2) | 15,498 (6.8) | 0 (0.0) | 1705 (0.7) | 17,868 (7.8) | 54,678 (24.0) | 39,745 (17.4) | 227,992 |
| Face lift | 89,368 (46.5) | 6693 (3.5) | 0 (0.0) | 294 (0.2) | 7470 (3.9) | 46,457 (24.2) | 41,840 (21.8) | 192,122 |
| Fat transfer: face | 18,460 (45.4) | 1866 (4.6) | 0 (0.0) | 337 (0.8) | 2897 (7.1) | 9568 (23.5) | 7524 (18.5) | 40,652 |
| Neck lift | 37,031 (43.9) | 5122 (6.1) | 0 (0.0) | 496 (0.6) | 5264 (6.2) | 20,676 (24.5) | 15,717 (18.6) | 84,306 |
| Nose surgery | 37,383 (42.6) | 6481 (7.4) | 1477 (1.7) | 20,423 (23.3) | 13,615 (15.5) | 6486 (7.4) | 1862 (2.1) | 87,727 |
| Total | 340,345 (43.5) | 51,171 (6.5) | 10,594 (1.4) | 44,205 (5.6) | 59,024 (7.5) | 157,095 (20.1) | 120,599 (15.4) | 783,033 |
Demographic data for facial cosmetic procedures categorized as “Head & Face” were obtained from The Aesthetic Society in collaboration with CosmetAssure. These data were used as epidemiological comparators for AI-generated images.
RESULTS
Demographics
Overall, 2600 images were generated across all patient groups and AI models (Fig. 1). Interrater reliability analysis revealed moderate to substantial agreement (0.5 < κ < 0.8) among the raters when assessing demographic factors in pediatric and adult populations and near-perfect agreement (κ > 0.8) when evaluating nondemographic visual elements (Table 4).27 A summary of the findings in this section may be found in Figure 3.
Table 4.
Interrater Reliability
| Patient Group | κ | P |
|---|---|---|
| Sex, pediatric | 0.669 | <0.00001 |
| Race, pediatrics | 0.762 | <0.00001 |
| Sex, adults | 0.697 | <0.00001 |
| Race, adults | 0.505 | <0.00001 |
| Age | 0.669 | <0.00001 |
| PPE | 0.813 | <0.00001 |
| Lesion | 0.874 | <0.00001 |
Agreement among raters ranged from moderate to substantial for demographic variables and was near-perfect for nondemographic visual features. κ indicates Fleiss kappa; a P value less than 0.05 denotes statistical significance.
Fig. 3.
Demographic representation in AI images by patient group and platform. Comparison of AI-generated demographics with epidemiological reference data. A, Percentage of female patients by procedure. B, Percentage of non-White patients by procedure. C, Percentage of patients older than 50 years. “%50+” represents the percentage of patients aged 50 years or older. *P < 0.05, **P < 0.01, ***P < 0.001, ****P < 0.0001.
All platforms overrepresented non-White patients in cleft lip images (P < 0.0001 for Midjourney and DreamStudio) and female patients in these images (P < 0.001). All platforms exclusively depicted female patients in facial aesthetic surgery images, yet significantly underrepresented non-White patients and patients older than 50 years (P < 0.0001).
All platforms significantly underrepresented non-White patients undergoing breast augmentation (P < 0.0001). Although patients aged older than 50 years constitute only 14.0% of breast augmentation patients,25 Midjourney and DreamStudio accentuated this demographic skew by lacking representation of older individuals in their generated images entirely (P < 0.0001). All platforms demonstrated a significant underrepresentation of non-White patients in the context of breast reconstruction (P < 0.0001). Midjourney overrepresented women older than 50 years (P < 0.05), whereas DreamStudio and Leonardo.Ai underrepresented this demographic (P < 0.0001).
Upon pooled analysis of the adult burn surgery images, female patients were overrepresented (P < 0.0001), and racial representation varied by platform, with Midjourney and DreamStudio underrepresenting non-White patients (P < 0.0001) and Leonardo.Ai overrepresenting them (P < 0.0001) (Fig. 4). When isolating images of patients with pediatric burn scars, Midjourney and DreamStudio significantly underrepresented non-White individuals (P < 0.0001). This finding contrasts with the overrepresentation of non-White patients observed in cleft lip images, highlighting the inconsistency of biases across different surgical categories and AI platforms. No significant gender disparities were observed in the depiction of pediatric burn patients. Leonardo.Ai was excluded from the analysis of this patient group, as it flagged the prompt requesting “pediatric” patients as inappropriate content and consequently did not generate any images.
Fig. 4.
Heatmaps of demographic differences between AI output and real-world demographics. A, Percentage difference in women relative to real-world demographics. B, Percentage difference in non-White patients relative to real-world demographics. C, Percentage difference in the number of women aged 50 years or older relative to real-world demographics.
Nondemographic Visual Elements
PPE Use
Subjects were more likely to be shown wearing PPE after reconstructive procedures (ie, cleft lip surgery, breast reconstruction, burn reconstruction) compared with aesthetic procedures (ie, facial cosmetic surgery, breast augmentation) (P < 0.01). Although no significant difference in PPE use was observed between images of pediatric patients with burn scar and those of unspecified age, patients who underwent “burn reconstruction” were consistently more likely to be shown in PPE compared with patients with “burn scar” across all platforms (P < 0.001) (Table 5).
Table 5.
Comparisons in Depictions of Perioral/Facial Lesions
| Patient Group A | Patient Group B | Midjourney | DreamStudio | Leonardo.Ai | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| % Lesion A | % Lesion B | P | % Lesion A | % Lesion B | P | % Lesion A | % Lesion B | P | ||
| Cleft lip, before repair | Cleft lip, after repair | 94 | 63 | <0.0001 | 1 | 1 | 1 | 4 | 5 | 0.734 |
| Cleft lip, pooled | 78.5 | 1.0 | 4.5 | |||||||
| Burn scar, unspecified | Burn scar, pediatric | 100 | 98 | 0.497 | 100 | 1 | 1 | 100 | ||
| Burn scar, pooled | Burn reconstruction | 99.3 | 100 | 0.554 | 97.3 | 92 | 0.004 | 100.0 | 100 | 1 |
Paired groups (A versus B) assess differences by procedure type or terminology. Significant platform-dependent differences were observed in cleft lip and burn-related image depictions.
Lesions
Our analysis revealed 2 statistically significant findings regarding the depiction of lesions in AI-generated images. Midjourney was the only platform that demonstrated a rudimentary understanding of cleft lip, portraying patients with perioral lesions more frequently in “before repair” prompts (P < 0.0001). DreamStudio exhibited a significant difference in the depiction of facial lesions among patients who underwent burn surgery, with patients with “burn scar” more likely to exhibit such lesions compared with patients who underwent “burn reconstruction” (P = 0.004). No significant difference was observed in the depiction of facial lesions between patients with burn scars of unspecified age and pediatric patients with burn scars (Table 5).
DISCUSSION
Our investigation offers novel insights into the accuracy of demographic representation across various plastic surgery patient populations within 3 publicly available text-to-image generative AI models. As summarized in Figure 4, the models demonstrated significant overrepresentation of non-White patients in cleft lip images while underrepresenting them in other categories, such as aesthetic surgery and breast reconstruction. This discrepancy suggests a biased training dataset rather than an accurate reflection of epidemiological data. Unlike the other 2 platforms, Leonardo.Ai was not included in the 2024 AI Index Report report23; however, its distinct focus on AI art generation motivated its inclusion in our study.24 Our findings confirm a commonly discussed pitfall in AI algorithms—the sex and racial bias associated with certain prompts.18–22 The biases we found demonstrate the specific manifestation of this problem within the context of plastic surgery, where outputs can perpetuate stereotypes regarding the sex and ethnicity of patients seeking care. Highlighting this bias is useful to allow for improvements to be made in the future, diminishing these biases and allowing for more equal representation.
Cleft Lip
The AI models showed a pronounced bias by overrepresenting non-White patients and women in cleft lip images, despite epidemiological data not supporting a higher incidence rate in these groups.28 This suggests a biased training dataset, possibly influenced by images from humanitarian missions. Although non-White children may be frequently featured in cleft lip–related charitable advertisements, this visibility often stems from efforts to provide surgical care in under-resourced regions worldwide and does not necessarily reflect higher incidence rates. Although one could argue that this bias might be attributed to the larger populations in countries such as India and China, resulting in a higher absolute number of cleft lip cases, this explanation is undermined by the observation that non-White subjects were underrepresented in AI-generated images depicting other patient groups.
Furthermore, none of the platforms could accurately depict a cleft lip, indicating a fundamental lack of understanding of the condition (Fig. 2). This observation raises concerns about the AI models’ fundamental comprehension of cleft lip. Notably, prompt engineering by specifying “before repair” versus “after repair” proved to be effective to some extent, as it significantly influenced the proportion of subjects depicted with perioral lesions in images generated by Midjourney.
Aesthetic Procedures
The platforms consistently portrayed a narrow and unrealistic demographic of aesthetic patients, namely those represented in breast augmentation and facial cosmetic surgery. Compared with their reconstructive counterparts, the aesthetic procedures tended to depict a disproportionately higher number of White female individuals (Figs. 3A, B).
They exclusively depicted female subjects in images of patients who underwent facial cosmetic surgery, despite a significant number of male patients in these groups.25 This discrepancy would likely be further augmented if nonsurgical procedures such as botulinum toxin injections were included in the epidemiological analysis. Moreover, the consistent depiction of women predominantly younger than 50 years in facial plastic surgery images across all 3 platforms contradicts real-world demographics, considering that the vast majority (70.9%) of this clientele is 50 years or older (Tables 3, 4). In addition to perpetuating biases and stereotypes, this misrepresentation of aesthetic patients can reinforce unrealistic beauty standards and hinder patients’ understanding of achievable surgical outcomes, potentially leading to dissatisfaction and unrealistic expectations.
PPE Use
The consistent trend of portraying patients after cleft lip surgery with PPE more frequently than patients before surgery suggests that the AI associates PPE with the presence of surgical intervention. Furthermore, the platforms’ greater propensity to depict patients undergoing reconstructive procedures in PPE compared with patients undergoing aesthetic procedures may reflect a subtle bias toward associating reconstructive procedures with a more medicalized context. This distinction, although subtle, may be a consequence of prompt bias triggered by the terminology used in the input and could inadvertently undermine the holistic care provided to all plastic surgery patients, as both reconstructive and aesthetic patients undergo surgical interventions with inherent risks and require the medical expertise of a plastic surgeon (Table 6).
Table 6.
Comparisons in Depictions of PPE Usage
| Patient Group A | Patient Group B | Midjourney | DreamStudio | Leonardo.Ai | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| % PPE A | % PPE B | P | % PPE A | % PPE B | P | % PPE A | % PPE B | P | ||
| Cleft lip, pooled | Facial cosmetic surgery | 21.0 | 16.0 | 0.277 | 17.3 | 1.0 | <0.001 | 23.3 | 4.0 | <0.001 |
| Cleft lip, before | Cleft lip, after | 11.0 | 38.0 | <0.001 | 13.0 | 35.0 | <0.001 | 25.0 | 38.0 | 0.048 |
| Cleft lip, unspecified | 14.0 | 4.0 | 7.0 | |||||||
| Breast reconstruction | Breast augmentation | 87.0 | 4.0 | < 0.001 | 2.0 | 0.0 | 0.156 | 10.0 | 1.0 | 0.005 |
| Burn scar, unspecified | Burn scar, pediatric | 14.0 | 9.0 | 0.269 | 6.0 | 1.0 | 0.055 | 3.0 | ||
| Burn scar, pooled | Burn reconstruction | 11.5 | 60.0 | <0.001 | 3.5 | 22.0 | <0.001 | 3.0 | 18.0 | <0.001 |
Paired groups (A versus B) compare reconstructive and aesthetic procedures or terminology within the same anatomical region. PPE was more frequently depicted in reconstructive groups, including cleft lip and breast reconstruction.
Burn Surgery
Contrary to the hypothesis, the AI models showed a trend of overrepresenting White individuals in pediatric burn surgery images, despite socioeconomic factors that disproportionately affect non-White populations.29 Moreover, numerous socioeconomic factors, including living in low-income areas, single-parent households, and overcrowded conditions, contribute to an increased burn risk,30–32 disproportionately affecting non-White populations.33–35 This observation suggests that the AI models may lack a nuanced understanding of social determinants of health that contribute to disparities in certain medical conditions.
Ethical Considerations
Popular models such as DALL-E 3, accessed via ChatGPT (OpenAI, San Francisco, CA), and Imagen 3, accessed via Google Gemini (Google LLC, Mountain View, CA), were initially considered for this study but were ultimately excluded from the analysis. DALL-E 3, although incorporating feedback mechanisms to address biased images,36 does not generate images in response to prompts requesting patient images, citing ethical concerns related to patient privacy. This restriction, although well-intentioned, raises questions about the necessity of such limitations when other platforms readily generate similar images, highlighting the evolving landscape of ethical considerations in AI. In creating Imagen 2, developers prioritized diverse and inclusive image outputs; however, this well-intentioned effort inadvertently led to the production of inaccurate—even offensive—representations. Consequently, Google has suspended the generation of images depicting individuals, even using the more updated Imagen 3, underscoring the complexities and challenges inherent in developing ethical and responsible AI.37 Leonardo.Ai exhibited a different approach to ethical considerations by flagging prompts containing “pediatric patient” as inappropriate content and refusing to generate images of children. (See figure, Supplemental Digital Content 2, which displays various platform restrictions on patient and sensitive content, highlighting diverse ethical approaches and inconsistencies across the evolving AI landscape. These differences underscore the urgent need for clear guidelines to ensure the responsible and equitable development of generative models, https://links.lww.com/PRSGO/E763.)
The nascent nature of AI technology presents a unique challenge in terms of regulation. Although some degree of governmental oversight may mitigate potential harms, excessive intervention could stifle innovation and limit the potential benefits of AI. Currently, a patchwork of regulations exists, with varying degrees of stringency across different jurisdictions.38 Geolocation-targeted results may further complicate the regulatory landscape, as platforms adapt their outputs to comply with local laws and cultural norms.
Beyond formal regulation, it is crucial to recognize the influence of financial and political factors on AI development and deployment. The financial interests of funding organizations can shape the data used to train AI models, potentially introducing biases that reflect those interests.38,39 Similarly, political motivations, such as appeasing powerful groups or promoting specific ideological agendas, can influence AI policies and public pronouncements. Navigating these complex influences requires a balanced approach that fosters innovation while safeguarding ethical principles and promoting societal well-being.
Although the real-world comparative data in this study are based on American datasets, the AI platforms investigated are used globally, suggesting that the biases we identified are not confined to a single country. The underrepresentation of non-White patients in aesthetic surgery, for example, becomes even more jarring at the global scale. This study thus emphasizes the need for a global effort to create more diverse and representative datasets to ethically and accurately represent patients from all backgrounds and regions.
Dataset Composition and Future Directions
The biases observed in our study, such as the overrepresentation of young, White female patients in aesthetic surgery images and the inconsistent racial representation in other patient groups, are a direct result of the datasets used to train these generative AI models. These models are trained on vast amounts of data scraped from the internet,13–15 which can perpetuate societal biases and stereotypes present in that data.16,17 For example, the disproportionate depiction of non-White children in cleft lip images may stem from a dataset oversaturated with images from humanitarian organizations or charitable campaigns. Similarly, the exclusive portrayal of women in aesthetic procedures likely reflects the disproportionate media representation of women undergoing such procedures.
To mitigate these biases, concrete solutions must be implemented by developers and users alike. Developers should prioritize the use of diverse and representative training datasets that accurately reflect global demographics and patient populations. Furthermore, they should increase transparency regarding the composition of their datasets, allowing for independent audits to identify and rectify biases. The development of user feedback–dependent algorithms is also crucial, enabling the models to learn from and correct biased outputs identified by users. On the user side, patients and physicians must be aware of these potential biases and critically evaluate AI outputs, recognizing that technological advancement does not inherently guarantee accuracy or freedom from bias.
Despite the biases observed in this study, AI has great potential for positive impact in health care, already being used in medicine to augment decision-making and enhance diagnostic capabilities. An application of AI algorithms, web scraping automates data collection from the web, drastically reducing the time it takes for research and expanding the evidence base in medicine and surgery.40 AI also has the potential to transform surgical care by augmenting critical decisions, such as the decision to operate and the informed consent process. For instance, the “My Surgery Risk” AI platform can use electronic health record data to predict postoperative complications and mortality with greater accuracy than physicians.41 Additionally, AI models are used for predictive analytics and diagnostics, capable of accurately predicting conditions such as sepsis and acute kidney injury.42 These beneficial applications highlight AI’s valuable role as a tool for augmenting human expertise, reinforcing the idea that although AI holds immense promise, its integration must be approached with a critical eye to address and mitigate biases such as those demonstrated in our study.
Limitations
This study is subject to some limitations that warrant consideration. Determining the sex of patients with pediatric cleft lip proved challenging due to prepubertal facial features, necessitating dependence on hair length and clothing style, which are not necessarily reliable indicators of sex. However, the high interrater reliability in this study strengthened the validity and consistency of the study’s findings. The analysis did not delve into a full racial breakdown, limiting it to White versus non-White individuals to mitigate the ambiguity associated with assessing race from images. Datasets used for comparison to real-world trends focus on the patient population in the United States and do not represent the global population, limiting the generalizability of these findings relative to other countries. Additionally, the American Society of Plastic Surgeons data were restricted to specific facial procedures, excluding nonsurgical procedures such as botulinum toxin and hyaluronic acid fillers, which are often performed at higher rates among older clientele. Furthermore, because Leonardo.Ai flagged prompts involving “pediatric” patients as inappropriate content, the assessment of the pediatric burn scar group was limited to Midjourney and DreamStudio. Consequently, Leonardo.Ai exclusively generated images of adults for cleft lip. These limitations highlight the need for future research with larger, more diverse datasets and more sophisticated AI models capable of accurately representing the full spectrum of plastic surgery patients.
CONCLUSIONS
This study provides compelling evidence of significant biases in AI-generated images of plastic surgery patients, raising concerns about their accuracy and potential to perpetuate harmful stereotypes. The overrepresentation of young, White female individuals in aesthetic surgery and the underrepresentation of non-White individuals in other surgical contexts, coupled with their inability to accurately depict cleft lip, highlight the limitations of current AI models. Additionally, this analysis highlighted implications of prompt bias, where the specific wording of the input can lead to inaccurate or stereotyped outputs. These biases can mislead patients and trainees, potentially influencing treatment decisions and perpetuating unrealistic expectations regarding surgical outcomes.
To use AI responsibly in medical practice, physicians should remain aware of how their patient population is represented by AI and recognize that technological advancement does not inherently guarantee accuracy or freedom from bias. This places an added responsibility upon healthcare providers to critically evaluate AI outputs, accurately represent these conditions for education and patient care, and advocate for the development of ethical and equitable AI tools. Although AI is revolutionizing information access, ongoing scrutiny and development are crucial to ensure equitable and ethical applications. Our findings highlighted the need for more diverse and representative training datasets, transparency regarding these datasets, and the development of user feedback–dependent algorithms. By addressing these challenges, we can harness the power of AI responsibly to enhance patient care and education in plastic surgery.
DISCLOSURE
The authors have no financial interest to declare in relation to the content of this article.
Supplementary Material
Footnotes
Published online 22 April 2026.
Disclosure statements are at the end of this article, following the correspondence information.
Related Digital Media are available in the full-text version of the article on www.PRSGlobalOpen.com.
REFERENCES
- 1.Coleman JPH. AI and our understanding of intelligence. In: Arai K, Kapoor S, Bhatia R, eds. Intelligent Systems and Applications. Springer International Publishing; 2020:183–190. [Google Scholar]
- 2.Sennott SC, Akagi L, Lee M, et al. AAC and artificial intelligence (AI). Top Lang Disord. 2019;39:389–403. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Deng L. Artificial intelligence in the rising wave of deep learning: the historical path and future outlook. IEEE Signal Process Mag. 2018;35:180–177. [Google Scholar]
- 4.Salehi H, Burgueño R. Emerging artificial intelligence methods in structural engineering. Eng Struct. 2018;171:170–189. [Google Scholar]
- 5.Jordan MI, Mitchell TM. Machine learning: trends, perspectives, and prospects. Science. 2015;349:255–260. [DOI] [PubMed] [Google Scholar]
- 6.Janiesch C, Zschech P, Heinrich K. Machine learning and deep learning. Electron Mark. 2021;31:685–695. [Google Scholar]
- 7.Tarca AL, Carey VJ, Chen XW, et al. Machine learning and its applications to biology. Lewitter F, ed. PLoS Comput Biol. 2007;3:e116. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Schneider WF, Guo H. Machine learning. J Phys Chem B. 2018;122:1347–1347. [DOI] [PubMed] [Google Scholar]
- 9.Hosny A, Parmar C, Quackenbush J, et al. Artificial intelligence in radiology. Nat Rev Cancer. 2018;18:500–510. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Steedman M. Chapter 8—natural language processing. In: Boden MA, ed. Artificial Intelligence (Handbook of Perception and Cognition). Academic Press; 1996:229–266. [Google Scholar]
- 11.Subramanian HV, Canfield C, Shank DB, et al. Combining uncertainty information with AI recommendations supports calibration with domain knowledge. J Risk Res. 2023;26:1137–1152. [Google Scholar]
- 12.Muller M, Chilton LB, Kantosalo A, et al. GenAICHI: generative AI and HCI. Presented at: CHI Conference on Human Factors in Computing Systems Extended Abstracts. April 27, 2025. ACM; 2022:1–7. [Google Scholar]
- 13.Kar S, Roy C, Das M, et al. AI horizons: unveiling the future of generative intelligence. Int J Adv Res Sci Commun Technol. 2023;3:387–391. [Google Scholar]
- 14.Sampat S. Where do generative AI models source their data & information? Smith.ai. Available at https://smith.ai/blog/where-do-generative-ai-models-source-their-data-information. 2023. Accessed August 15, 2024. [Google Scholar]
- 15.National University Library. LibGuides. Research process: datasets. National University Library. Available at https://resources.nu.edu/researchprocess/datasets. Accessed August 15, 2024. [Google Scholar]
- 16.Mehrabi N, Morstatter F, Saxena N, et al. A survey on bias and fairness in machine learning. ACM Comput Surv. 2022;54:1–35. [Google Scholar]
- 17.Srinivasan R, Chander A. Biases in AI systems: a survey for practitioners. Queue. 2021;19:45–64. [Google Scholar]
- 18.Angwin J, Larson J, Mattu S, et al. Machine bias. ProPublica. Available at https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing. 2016. Accessed August 15, 2024. [Google Scholar]
- 19.Dastin J. Insight—Amazon scraps secret AI recruiting tool that showed bias against women. Reuters. Available at https://www.reuters.com/article/world/insight-amazon-scraps-secret-ai-recruiting-tool-that-showed-bias-against-women-idUSKCN1MK0AG/. 2018. Accessed August 15, 2024. [Google Scholar]
- 20.Buolamwini J, Gebru T. Gender shades: intersectional accuracy disparities in commercial gender classification. Presented at: Proceedings of the 1st Conference on Fairness, Accountability and Transparency, January 29–31, 2019, Atlanta, GA. PMLR; 2018:77–91. Available at https://proceedings.mlr.press/v81/buolamwini18a.html. Accessed August 15, 2024. [Google Scholar]
- 21.Obermeyer Z, Powers B, Vogeli C, et al. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366:447–453. [DOI] [PubMed] [Google Scholar]
- 22.Gender bias revealed in AI tools screening for liver disease. UCL News. Available at https://www.ucl.ac.uk/news/2022/jul/gender-bias-revealed-ai-tools-screening-liver-disease. 2022. Accessed August 15, 2024. [Google Scholar]
- 23.Available at https://aiindex.stanford.edu/wp-content/uploads/2024/05/HAI_AI-Index-Report-2024.pdf. Accessed November 19, 2024.
- 24.AI image generator—create art, images & video. Leonardo AI. Leonardo.Ai. Available at https://leonardo.ai/. Accessed November 19, 2024. [Google Scholar]
- 25.Available at https://cdn.theaestheticsociety.org/media/statistics/2022-TheAestheticSocietyStatistics.pdf. Accessed November 19, 2024.
- 26.Available at https://www.plasticsurgery.org/documents/News/Statistics/2020/cosmetic-procedures-ethnicity-2020.pdf. Accessed November 19, 2024.
- 27.Hartling L, Hamm M, Milne A, et al. Executive summary. In: Validity and Inter-Rater Reliability Testing of Quality Assessment Instruments. Agency for Healthcare Research and Quality (US); 2012. Available at https://www.ncbi.nlm.nih.gov/books/NBK92287/. Accessed November 8, 2024. [PubMed] [Google Scholar]
- 28.Taritsa IC, Ledwon JK, Bajaj A, et al. 12-year trends of orofacial clefts in the United States: highlighting racial/ethnic differences in prevalence of cleft lip and cleft palate. Cleft Palate Craniofacial J. 2024;62:820–828. [DOI] [PubMed] [Google Scholar]
- 29.Elrod J, Schiestl CM, Mohr C, et al. Incidence, severity and pattern of burns in children and adolescents: an epidemiological study among immigrant and Swiss patients in Switzerland. Burns. 2019;45:1231–1241. [DOI] [PubMed] [Google Scholar]
- 30.Dinesh A, Polanco T, Khan K, et al. Our inner-city children inflicted with burns: a retrospective analysis of pediatric burn admissions at Harlem Hospital, NY. J Burn Care Res. 2018;39:995–999. [DOI] [PubMed] [Google Scholar]
- 31.Hunter MA, Schlichting LE, Rogers ML, et al. Neighborhood risk: socioeconomic status and hospital admission for pediatric burn patients. Burns. 2021;47:1451–1455. [DOI] [PubMed] [Google Scholar]
- 32.Ozlu O, Basaran A. Sociodemographic factors and living conditions of pediatric burn patients. Paediatr Indones. 2022;62:149–155. [Google Scholar]
- 33.Amato PR, Keith B. Separation from a parent during childhood and adult socioeconomic attainment. Soc Forces. 1991;70:187. [Google Scholar]
- 34.Akee R, Jones MR, Porter SR. Race matters: income shares, income inequality, and income mobility for all U.S. races. Demography. 2017;56:999–1021. [DOI] [PubMed] [Google Scholar]
- 35.Rebbe R, Sattler KM, Mienko JA. The association of race, ethnicity, and poverty with child maltreatment reporting. Pediatrics. 2022;150:e2021053346. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.DALL·E 3 is now available in ChatGPT Plus and enterprise. Available at https://openai.com/index/dall-e-3-is-now-available-in-chatgpt-plus-and-enterprise/. Accessed November 19, 2024.
- 37.Raghavan P. Gemini image generation got it wrong. We’ll do better. Google. Available at https://blog.google/products/gemini/gemini-image-generation-issue/. 2024. Accessed July 31, 2024. [Google Scholar]
- 38.Zarsky T. The trouble with algorithmic decisions: an analytic road map to examine efficiency and fairness in automated and opaque decision making. Sci Technol Human Values. 2015;41:118–132. [Google Scholar]
- 39.Bender EM, Gebru T, McMillan-Major A, et al. On the dangers of stochastic parrots: can language models be too big? Presented at Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, March 3–10, 2021. ACM; 2021:610–623. [Google Scholar]
- 40.Goulas S, Karamitros G. How to harness the power of web scraping for medical and surgical research: an application in estimating international collaboration. World J Surg. 2024;48:1297–1300. [DOI] [PubMed] [Google Scholar]
- 41.Loftus TJ, Tighe PJ, Filiberto AC, et al. Artificial intelligence and surgical decision-making. JAMA Surg. 2020;155:148–158. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Loftus TJ, Shickel B, Ozrazgat-Baslanti T, et al. Artificial intelligence-enabled decision support in nephrology. Nat Rev Nephrol. 2022;18:452–465. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.




