Skip to main content
JAMA Network logoLink to JAMA Network
. 2023 Nov 9;141(12):1174–1175. doi: 10.1001/jamaophthalmol.2023.5162

Large Language Model Advanced Data Analysis Abuse to Create a Fake Data Set in Medical Research

Andrea Taloni 1, Vincenzo Scorcia 1, Giuseppe Giannaccare 1,2,
PMCID: PMC10636646  PMID: 37943569

Abstract

This quality improvement study evaluates the ability of GPT-4 Advanced Data Analysis to create a fake data set that can be used for the purpose of scientific research.


Since its public release by OpenAI in late 2022, ChatGPT (generative pretrained transformer) has gone through several updates. In March 2023, the large language model (LLM) GPT-4 came out, marking improvements in semantic understanding and response generation compared with GPT-3.5.1 More recently, the capabilities of GPT-4 were expanded with Advanced Data Analysis (ADA), a model that runs Python.2 This tool supports file uploads and downloads and can perform both statistical analysis and data visualization.3 Although these features may speed up scientific research, unethical uses of ADA are conceivable, including fabrication of data. We evaluated the ability of GPT-4 ADA (OpenAI; 09/11/2023 version) to create a fake data set that can be used for scientific research.

Methods

In this quality improvement study, a detailed prompt, providing criteria to compile each column of the database, was submitted to ADA (eFigure in Supplement 1). The LLM was asked to fabricate data for 300 eyes belonging to 250 patients with keratoconus who underwent deep anterior lamellar keratoplasty (DALK) or penetrating keratoplasty (PK). For categorical variables, target percentages were predetermined for the distribution of each category. For continuous variables, target mean and range were defined. Additionally, ADA was instructed to fabricate data that would result in a statistically significant difference between preoperative and postoperative values of best spectacle-corrected visual acuity (BSCVA) and topographic cylinder. ADA was programmed to yield significantly better visual and topographic results for DALK compared with PK. Upon completion, the fake database was analyzed with SPSS Statistics, version 29.0 (IBM Corp). A 2-tailed t test was performed to compare continuous variables. P < .05 was considered significant. This study followed the Standards for Quality Improvement Reporting Excellence (SQUIRE) reporting guideline. Because the study did not involve humans or animals, no ethical approval or informed consent was required.

Results

The LLM followed a step-by-step approach to build the data set. Sometimes answer generation stopped due to word count limit; in these cases, the prompt “Continue” was used to resume the process until a data set was successfully created (eFigure in Supplement 1). As requested, mean (SD) postoperative BSCVA and topographic cylinder for DALK were significantly better compared with PK (difference: BSCVA, −0.05 [0.15] logMAR [95% CI, −0.08 to −0.03 logMAR]; topographic cylinder, −0.91 [1.00] diopters [95% CI −1.08 to −0.76 diopters]; P < .001). Most statistical criteria provided in the prompt were met by the data set; however, data ranges of continuous variables were not always accurate (Table 1 and Table 2).

Table 1. Comparison Between Requested Values of the Submitted Prompt and Obtained Values of the Fake Data Set: Categorical Variables.

Categorical variable Requested value, % Obtained value, No. (%)
Sex
Male 55.0 160 (53.3)
Female 45.0 140 (46.7)
Surgery
DALK 50.0 150 (50.0)
PK 50.0 150 (50.0)
Bubble burst
DALK 1.0a 0a
PK 0 0
Microperforations
DALK 7.5 11 (7.3)
PK 0 0
Choroidal hemorrhage
DALK 0 0
PK 0.5a 0a
Stromal immune rejection
DALK 5 11 (7.3)
PK 5a 3 (2.0)a
Endothelial immune rejection
DALK 0 0
PK 10 15 (10.0)

Abbreviations: DALK, deep anterior lamellar keratoplasty; PK, penetrating keratoplasty.

a

Obtained numerical value that differs from the requested value.

Table 2. Comparison Between Requested Values of the Submitted Prompt and Obtained Values of the Fake Data Set: Continuous Variables.

Continuous variable Requested value, mean (range) P value Obtained value, mean (SD) [range] Difference between requested and obtained values, mean (SD) [range; 95% CI] P value
Age at surgery, y 35.0 (18 to 70)a NA 32.9 (13.1) [17 to 59]a NA NA
BSCVA, logMAR
Preoperative 1.00 (0.50 to 1.80)a <.05 1.00 (0.20) [0.50 to 1.60]a −0.79 (0.22) [−1.40 to −0.20; −0.81 to −0.76] <.001
Postoperative 0.20 (0.00 to 1.00)a 0.21 (0.11) [0.00 to 0.50]a
DALK postoperative NAb <.05 0.18 (0.11) [0.00 to 0.50] −0.05 (0.15) [−0.4 to 0.3; −0.08 to −0.03] <.001
PK postoperative NAb 0.24 (0.10) [0.00 to 0.50]
Topographic cylinder, D
Preoperative 4.50 (1.00 to 18.00)a <.05 4.45 (1.60) [1.00 to 8.85]a −1.36 (1.81) [−5.73 to 3.8; −1.57 to −1.16] <.001
Postoperative 3.00 (0.50 to 12.00)a 3.09 (0.83) [1.06 to 5.16]a
DALK postoperative NAb <.05 3.07 (2.22) [0.5 to 9.52] −0.91 (1.00) [−3.61 to 2.07; −1.08 to −0.76] <.001
PK postoperative NAb 3.83 (2.67) [0.50 to 11.19]

Abbreviations: BSCVA, best spectacle-corrected visual acuity; D, diopters; DALK, deep anterior lamellar keratoplasty; NA, not applicable; PK, penetrating keratoplasty.

a

Obtained numerical value that differs from the requested value.

b

Requested values for SDs and 95% CIs were not provided in the prompt submitted to Advanced Data Analysis. Similarly, no target data were provided for BSCVA and topographic cylinder of eyes that underwent DALK or PK.

Discussion

The LLM created a seemingly authentic database, showing better results for DALK than PK. Recently, we expressed concerns regarding the capability of an LLM to produce plagiarism-free scientific essays while evading artificial intelligence (AI) detection.4 ADA may pose a greater threat, being able to fabricate data sets specifically designed to quickly produce false scientific evidence, such as the better outcomes of DALK over PK that have not been proved, to our knowledge, by scientific evidence.5 Illegitimate data manipulation has been repeatedly reported in academia6; however, recognition of research misconduct is still an outstanding issue with no definitive solutions. Potential strategies to identify AI data fabrication may involve looking for peculiar statistical patterns in the data sets, similar to technology that detects AI-generated text by evaluating the likelihood of nonhuman patterns.4 On the other hand, sponsors may prevent data fabrication by equipping study centers with digital diagnostic devices that store measurements inside onboard memory. Thus, encrypted data backups could be uploaded directly to the sponsor’s online database platform, without allowing manual transcription of data by the investigators. Additionally, registration of clinical trials on ClinicalTrials.gov and adherence to CONSORT checklists, protocols, and statistical analysis plans could also safeguard data authenticity. Publishers and editors might want to consider the findings of this study in the peer-review process, to ensure that AI advancements will enhance, not undermine, the integrity and value of scientific research.

Supplement 1.

eFigure. (a) Prompt Submitted to Advanced Data Analysis. (b) Sample of 40 Data Entries From the Fake Data Set Fabricated by Advanced Data Analysis on Microsoft Office Excel 365

Supplement 2.

Data Sharing Statement

References

  • 1.Open AI. Accessed June 20, 2023. https://openai.com/
  • 2.Python.org. What is Python? Executive summary. Accessed August 23, 2023. https://www.python.org/doc/essays/blurb/
  • 3.Wang L, Ge X, Liu L, Hu G. Code interpreter for bioinformatics: are we there yet? Ann Biomed Eng. Published online July 23, 2023. doi: 10.1007/s10439-023-03324-9 [DOI] [PubMed] [Google Scholar]
  • 4.Taloni A, Scorcia V, Giannaccare G. Modern threats in academia: evaluating plagiarism and artificial intelligence detection scores of ChatGPT. Eye (Lond). Published online August 2, 2023. doi: 10.1038/s41433-023-02678-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Lucisano A, Scorcia V, Taloni A, Rossi C, Gioia R, Giannaccare G. Impact of topographic localization of corneal ectasia on the outcomes of deep anterior lamellar keratoplasty employing large (9 mm) versus conventional diameter (8 mm) grafts. Eye (Lond). 2023. Published online April 20, 2023. doi: 10.1038/s41433-023-02536-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Xie Y, Wang K, Kong Y. Prevalence of research misconduct and questionable research practices: a systematic review and meta-analysis. Sci Eng Ethics. 2021;27(4):41. doi: 10.1007/s11948-021-00314-9 [DOI] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplement 1.

eFigure. (a) Prompt Submitted to Advanced Data Analysis. (b) Sample of 40 Data Entries From the Fake Data Set Fabricated by Advanced Data Analysis on Microsoft Office Excel 365

Supplement 2.

Data Sharing Statement


Articles from JAMA Ophthalmology are provided here courtesy of American Medical Association

RESOURCES