Abstract
Aims
This proof-of-concept study aims to demonstrate the feasibility of using AI-Simulated Patients (ASPs) to test and illustrate the tailored selection of diabetes remission prediction models.
Methods
We created five initial ASPs representing diverse T2DM cases, each with key clinical parameters, then expanded to 100 ASPs as a convenience sample. Using a standardized process, GPT-4 applied six validated diabetes remission prediction models: DiaRem, A-DiaRem, ABCD, IMS, DiaBetter, and DRI. For each case, GPT-4 assessed input availability, estimated remission scores, discussed model strengths/limitations, and recommended the best-fit tool. Human experts supervised all outputs to ensure accuracy, reliability, and appropriate model selection for each ASP. Exact ranges and distributions were anchored in reported epidemiology of bariatric patient populations (e.g., age 18–70 years, BMI 30–65 kg/m², HbA1c 6–11%, C-peptide 0.2–5 ng/mL) with variability introduced for diversity. Sample size (n = 100) was chosen for breadth of illustration rather than statistical power.
Results
We applied a refined, criteria-based algorithm to 100 ASPs representing diverse T2DM cases considered for metabolic/bariatric surgery. Model selection incorporated key inputs, planned procedure, and model validation scope. The DRI was most common, followed by DiaBetter, ABCD Score, Advanced DiaRem, and IMS Score. Each recommendation reflected surgical type and data completeness. The resulting decision-making algorithm illustrates a reproducible methodology that may support future tailored application of remission prediction models, pending validation in real clinical datasets.
Conclusions
This study demonstrates that LLM-generated simulated patients can refine a reproducible, transparent, and ethically safe algorithm for selecting diabetes remission prediction models, with future validation required using real-world clinical data.
Supplementary Information
The online version contains supplementary material available at 10.1186/s13098-025-02012-z.
Keywords: Diabetes, Prediction model, Metabolic bariatric surgery, Artificial intelligence, AI-Simulated patient
Introduction
Metabolic and bariatric surgery (MBS) is widely recognized as a highly effective treatment for type 2 diabetes (T2DM) in subjects with obesity and/or uncontrolled diabetes. MBS exerts beneficial effects on glucose and lipid control; diabetic complications and diabetes remission compared to medical or lifestyle interventions alone [1, 2]. Many predictive models, such as DiaRem [3], advanced DiaRem [4], ABCD [5], IMS [6], DiaBetter [7], and newer tools like the Diabetes Remission Index (DRI) [8], have shown good accuracy for predicting short-term remission after MBS [9, 10]. Despite the variety of these models, clinicians face real-world challenges in selecting, applying, and interpreting them. No single model is universally superior, and each has limitations, especially when dealing with incomplete data or diverse patient populations. Missing or mismatched input data can reduce model accuracy. Some models require more detailed or specific variables than are routinely available [9].
The application of artificial intelligence (AI), including large language models (LLMs), particularly as a form of collaborative intelligence (CI), can now serve as a valuable assistant in medicine and surgery, including the field of MBS [11, 12]. AI-simulated patients (ASP), powered by LLMs and generative AI, are rapidly transforming medical education and clinical training [13]. ASPs can provide researchers with a safe means to test predictive models and trial interventions without relying on real patient data. It can speed up the development of new medications, devices, or treatment modalities, and help build trust in digital evidence for regulatory approval [14, 15]. AI algorithms can create synthetic datasets that mirror real patient populations, allowing researchers to train and validate models while protecting patient privacy [16]. Large ASP datasets can boost the accuracy and reliability of model performance estimates, including for external validation and testing generalizability. When AI models are validated on simulated or synthetic patient groups, they can show that they work well across different populations and clinical scenarios, supporting their integration into real-world healthcare [17, 18]. Despite the LLMs’ limitations, like medico-legal concerns and transparency [19], this strategy may be useful to create clinical scenarios to assess the current diabetes remission predictive models to find the best-matched model for every scenario.
This proof-of-concept study, using an exploratory methodology, aims to demonstrate the feasibility of using ASPs to test and illustrate the tailored selection of diabetes remission prediction models. Rather than establishing definitive comparative performance, our goal is to showcase a reproducible methodology for patient–model matching.
Methods
This is an exploratory, pre-clinical feasibility study using synthetic data only. No real patient data were used, and results are intended as an illustration of the framework rather than as a validation of model performance.
As an initial step, we designed five ASPs to reflect a realistic and diverse range of clinical scenarios commonly encountered in individuals with T2DM from a national database [20]. Each profile included key clinical parameters, such as age, body mass index (BMI), duration of T2DM, HbA1c, C-peptide levels, current antidiabetic medications, and the planned MBS procedure, organized in a table for clarity and transparency.
To ensure consistency, we used a standardized GPT-4 prompt (shown in Fig. 1) for all ASPs. GPT-4 was instructed to illustrate how the framework recommends a model for each patient profile, interpret expected outcomes, identify the strengths and limitations of each model, and recommend the most appropriate (best-fit) model for that scenario. For full reproducibility, the exact prompt is also provided in the supplementary materials. The prompting workflow employed GPT-4 with fixed parameters, including a deterministic temperature setting of 0.2 and a locked system prompt, to minimize variability and ensure reproducible outputs. Structured response constraints were applied through a standardized, table-based output template, and safeguards were implemented to prevent output drift across repeated runs. Prompt reproducibility was evaluated using a structured iterative workflow across three independent GPT-4 sessions, during which ASPs were processed with identical standardized prompts. All outputs were automatically archived without modification to ensure transparency and traceability. A verification script flagged any deviation in key tokens or response length greater than 2% from the baseline output, and human reviewers confirmed consistency in model selection logic and terminology. Across the three sessions, model recommendations were identical in 14 of 15 possible comparisons (93.3%), with a mean textual similarity score of 0.96 ± 0.02. Variations were limited to minor phrasing differences without affecting clinical content. Overall, the reproducibility of decision outcomes exceeded 90%, demonstrating high prompt stability and minimal drift across sessions.
Fig. 1.
GPT Prompt. (Temperature 0.2, deterministic constraints, and safeguards to prevent drift)
Finally, all GPT-4–generated outputs were independently reviewed by human experts to confirm alignment with model-specific input requirements and to minimize the risk of hallucinations. The review panel included two metabolic bariatric surgeons and one diabetologist, each with clinical experience in diabetes remission prediction and metabolic surgery. Experts independently evaluated the outputs for accuracy and coherence relative to established model criteria, and any discrepancies were resolved through group discussion until consensus was achieved. Inter-rater reliability was not formally quantified; thus, expert verification was qualitative and intended to confirm face validity rather than provide a statistical measure of agreement. This proof-of-concept was intentionally limited to GPT-4; future iterations will extend to GPT-5 and other LLMs (e.g., Gemini, DeepSeek) to assess reproducibility and performance differences.
To assess consistency in responses, we submitted the initial five ASPs, which were randomly selected from the Iran National Obesity Surgery Database (INOSD) [20] and their corresponding prompts to GPT-4 across three separate sessions. The outputs were largely consistent, with only minor variations in phrasing and model ranking.
In the next phase, GPT-4 was tasked with generating 95 additional ASPs to capture a broader spectrum of T2DM scenarios while maintaining medically plausible characteristics and a consistent format. These 100 ASPs, representing a wide variety of patients with T2DM scheduled for MBS, were then entered into ChatGPT using the same standardized prompt. Exact ranges and distributions were anchored in reported epidemiology of bariatric populations (e.g., age 18–70 years, BMI 30–65 kg/m², HbA1c 6–11%, C-peptide 0.2–5 ng/mL) with variability introduced for diversity that reflect common thresholds used to study type 2 diabetes. These thresholds are widely used to define inclusion criteria and reflect the diversity of real-world bariatric cohorts [21–23]. All ASPs were screened for biological and clinical plausibility using rule-based checks. Implausible cases (e.g., duration exceeding age, inconsistent C-peptide and insulin use, unrealistic HbA1c or BMI values) were automatically resampled. Logical dependencies, such as the inverse relation between diabetes duration and C-peptide, were preserved. After filtering, all 100 detailed ASPs met these criteria, with an average acceptance rate of ~ 85% and typically one resample per case, ensuring that all simulated patients were clinically coherent and realistic. Sample size (n = 100) was chosen for breadth of illustration rather than statistical power, as clarified with a formal power analysis statement.
We submitted the full set of 100 ASP profiles to GPT-4, and asked it to illustrate the selection process of six validated diabetes remission prediction models: DiaRem, Advanced DiaRem (A-DiaRem), ABCD Score, IMS Score, DiaBetter, and the Diabetes Remission Score (DRI Score). For each ASP, GPT-4 was directed to assess the availability of required inputs, provide qualitative or estimated remission scores, discuss the strengths and limitations of each model, and recommend the most appropriate prediction tool. To ensure accuracy and reliability, all AI-generated outputs were supervised and reviewed by human experts, who evaluated each response against the criteria of the respective models to minimize the risk of hallucinations. We did not quantify inter-rater agreement in this exploratory phase; future work should incorporate systematic reliability measures (e.g., kappa statistics) to strengthen reproducibility of human supervision.
Collaborative intelligence, enabled by human supervision, guided GPT-4 in tailoring model selection according to surgical type and data availability. This approach ensured that each remission prediction model was applied with clear use cases and justifications appropriate to the clinical context.
The Advanced DiaRem was applied to RYGB patients when C-peptide data were available, given its validation in this context and incorporation of both disease duration and beta-cell function. The ABCD Score was used for patients undergoing SG or RYGB, provided that C-peptide values were available, due to its broader validation across these procedures. The Diabetes Remission Score (DRI Score) was recommended for individuals scheduled for SADI-S or BPD-DS, particularly when C-peptide was present, as it is well-suited to more complex surgical interventions. The DiaBetter model was selected for OAGB patients, regardless of input completeness, as other models lacked validation in this group. Finally, the IMS Score served as a fallback option for patients with incomplete data, most commonly missing C-peptide, or when the planned procedure was outside the validated scope of the other models.
Based on the insights gathered from this process, we developed a simplified decision-making algorithm to assist clinicians in selecting the most suitable diabetes remission prediction model for specific patient scenarios.
Results
To demonstrate the framework’s application, we generated 100 ASPs and applied the algorithm to illustrate how it selects among six validated models. Model selection was guided by a refined, criteria-based algorithm that incorporated the availability of key clinical inputs (e.g., C-peptide, insulin use), the type of planned MBS procedure, and the validation context for each model. The final dataset included 100 ASPs (79% female), designed to reflect a broad spectrum of clinical scenarios commonly observed in individuals with T2DM being considered for standard MBS procedures. The cohort had a mean age of 43.3 ± 14.7 years and an average BMI of 48.2 ± 10.2 kg/m², indicating a population with severe obesity. The mean glycated hemoglobin (HbA1c) level was 8.6% ± 1.1, consistent with suboptimal glycemic control in MBS candidates. These profiles were constructed to mimic real-world clinical diversity while maintaining internal consistency across cases. These distributions and averages reflect the specific synthetic cohort created for this study, with no external benchmark, and should not be interpreted as evidence of real-world model performance.
Using the algorithm, the most appropriate remission prediction model was assigned to each ASP. In these simulated scenarios, the Diabetes Remission Score (DRI Score) was recommended in 50% of ASPs, especially for patients undergoing SADI-S, BPD-DS, or RYGB with available C-peptide data; however, this does not imply superior predictive ability.
The DiaBetter model was selected in 17% of ASPs, typically for those scheduled for OAGB, where other models lacked validation. The ABCD Score was used in 15% of ASPs, primarily for SG and RYGB candidates with complete input data. Advanced DiaRem was applied in 12% of ASPs, specifically for RYGB patients with full data, including C-peptide. In 6% of patients, where data were incomplete or the surgical procedure was not aligned with any model’s validation context, the IMS Score served as a conservative fallback. The reported frequencies reflect how often each model was recommended within this simulated dataset, rather than their actual performance.
A structured decision-making algorithm (Fig. 2) was developed to guide model selection based on clinical input availability and MBS type. This approach supports tailored model application in both clinical and simulation settings, promoting methodological rigor and contextual relevance.
Fig. 2.
Selection Algorithm Flowchart. Unlike conventional rule-based stratification, this ASP-driven workflow tests these decision nodes in a synthetic cohort, enabling reproducible simulation without patient data. While simplified, the flowchart is presented in a decision-support style with consistent color coding. A more advanced interface may further improve usability and alignment with clinical decision-support standards in future studies. Color Coding: Green: Validated prediction models. Orange: Fallback to IMS (Integrated Metabolic Score)
Representative examples of ASPs (selected from the full cohort of 100) are shown in Appendix. 2 to illustrate the model selection process. The complete GPT-4–generated responses and individualized model recommendations for all 100 ASPs are provided in Appendices 1 and 2, respectively. A summary of the best-fit model recommendation is shown in Appendix. 1. Exact ranges and distributions were anchored in reported epidemiology of bariatric populations with variability introduced for diversity.
Discussion
The primary contribution of this work is the demonstration of a reproducible, LLM-assisted workflow for patient-specific model selection using synthetic clinical scenarios, using exploratory methodology. While the algorithm’s outputs align with published model validation contexts, the emphasis here is on feasibility and potential scalability, not on head-to-head comparative performance.
Prediction models for diabetes remission after MBS are important tools that can support personalized care. They help clinicians estimate a patient’s likelihood of T2DM remission, aid in appropriate MBS procedure selection, and shape realistic expectations for both patients and clinician teams [9, 24]. These models can also play a role in tailoring pre- and post-operative management. However, each model has its limitations; for example, some models require data (e.g., C-peptide) that may not always be available [8, 25], and there is an increasing need for clear, patient-centered protocols to help clinicians choose the most appropriate model for individual scenarios.
The DiaRem and Advanced DiaRem (Ad-DiaRem) scores estimate the likelihood of diabetes remission after metabolic or bariatric surgery, particularly Roux-en-Y gastric bypass. Because DiaRem was developed using data from predominantly white U.S. populations, its accuracy declines in more diverse groups and often underestimates remission in patients labeled as “low probability.” [4, 26, 27]. Ad-DiaRem improves on this by adding factors such as diabetes duration and medication use, offering modestly better accuracy, especially for intermediate cases. However, both scores remain limited in predicting non-remission, long-term outcomes, and results after procedures other than RYGB, such as SG [4, 27–29].
The ABCD score is a simple and widely used tool for predicting diabetes remission after metabolic or bariatric surgery. Although easy to apply and validated in multiple settings, it is less accurate than newer models such as DiaBetter, Ad-DiaRem, or DiaRem2, particularly for patients undergoing sleeve gastrectomy or from diverse ethnic and regional backgrounds. The score performs better for short-term remission (1–2 years) but loses accuracy for long-term outcomes [30]. Its predictive value also varies across populations, and because it excludes key factors like HbA1c and medication use, its precision is limited for certain patients [31–33].
The IMS score is used to predict diabetes remission and guide the choice of MBS. It performs better at identifying patients unlikely to achieve remission than those who will, which means some potential responders may be missed [34]. Its accuracy varies by population and diabetes severity, performing less reliably in regions such as East Asia or among patients with moderate T2DM [32]. Although designed for both gastric bypass and SG, it does not always capture outcome differences between these procedures [32, 35]. The IMS score offers some insight into long-term outcomes (3–5 years), but like other models, its accuracy declines over time, and it does not account for post-surgical factors such as weight loss or relapse [5, 32, 35].
The DiaBetter score is a newer tool designed to predict T2DM remission after MBS. Research shows it performs better than older scoring systems like ABCD, DiaRem, and Ad-DiaRem, especially in the first few years after MBS [36]. It achieved some of the highest accuracy ratings among available models, with strong results in predicting both short-term and long-term outcomes. This makes it one of the more reliable tools currently available [37]. So, like all prediction models, its performance tends to fade over time. Its accuracy decreases beyond the two-year mark, which highlights the need for ongoing monitoring after MBS. Also, while it’s generally effective, the DiaBetter score, like others, relies heavily on pre-surgery data and doesn’t fully capture the complexity of each person’s journey or post-MBS changes like weight loss or relapse [34].
The Diabetes Remission Index (DRI) is another tool for predicting the likelihood of T2DM remission, after MBS [8]. While it shows some promise, it still needs more research and validation in broader patient groups before it can be widely trusted. One key limitation is that it doesn’t factor in BMI or whether the person is using diabetes medications, both important pieces of the puzzle. Also, it can’t be used at all unless C-peptide levels are available, which may not be part of standard pre-surgery testing in all settings.
The limitations of current diabetes remission prediction models mean that no single model works accurately or effectively for all patients. On the other hand, using multiple models at the same time can be time-consuming and sometimes confusing for both patients and healthcare providers.
While clinical judgment is essential for interpreting prediction models [10], it can be time-consuming and may vary between providers. To support more consistent and personalized decision-making, an algorithmic approach can help guide clinicians in selecting the most appropriate model based on specific patient scenarios. In this study, we used GPT-4 to evaluate AI-simulated patient profiles, demonstrating how collaborative intelligence can assist clinicians in choosing the best-fit model from six validated diabetes remission prediction tools. Based on the strengths and limitations of each model, ChatGPT-4 was guided to select prediction tools as follows: The Advanced DiaRem was used for RYGB patients with available C-peptide data, given its validation and inclusion of disease duration and beta-cell function. The ABCD Score was applied to SG or RYGB patients when C-peptide data were present, due to its broader validation. For SADI-S or BPD-DS candidates with C-peptide values, the DRI Score was preferred for its suitability to complex procedures. The DiaBetter model was used for OAGB patients regardless of data completeness, as no other models were validated for this group. Lastly, the IMS Score served as a fallback for cases with incomplete data—most often missing C-peptide—or when procedures fell outside other models’ validated scope.
In the simulated scenarios, GPT-4 most frequently recommended the DRI model within the synthetic cohort. This predominance does not indicate superior predictive performance but rather reflects the model’s broad applicability and suitability across diverse clinical contexts. DRI Score was favored when C-peptide data were available, particularly in patients where beta-cell function was critical to remission potential, although DRI does not incorporate BMI or medication use, limiting its scope in certain contexts. The DiaBetter Score was the second common best-fit model for ASPs and can be used for OAGB or when limited data is available. The ABCD score is another suitable prediction model when the patient is scheduled for SG or RYGB and when C-peptide is available. The Advanced DiaRem and IMS Score were consistently utilized by GPT-4 as alternative tools in specific scenarios where other models were less applicable. However, this reflects the algorithmic rules applied to the synthetic dataset, and should be interpreted as an illustration of feasibility rather than evidence of superiority.
Validating AI in medicine is essential to ensure that prediction models are accurate, reliable, and safe for clinical application. As AI tools become increasingly integrated into healthcare, establishing robust validation frameworks and clear quality standards is critical for building trust and supporting their real-world use. Key factors for successful clinical adoption and improved patient outcomes include external validation, transparency, and continuous quality assurance [38, 39]. On the other hand, integrating AI into healthcare raises important concerns about bias and ethics, which can affect patient safety, fairness, and trust in medical systems. These concerns include data bias, algorithmic and development bias, and interaction bias [40, 41]. However, such risks can be mitigated when AI is used under human supervision, creating a form of collaborative intelligence [12, 42]. In this study, all AI-generated responses were reviewed by clinical experts to minimize potential pitfalls and ensure responsible, accurate model recommendations.
ASPs can serve as specialty-specific clinical scenarios that are highly customizable to reflect a wide range of conditions and use cases. These virtual patients, powered by generative AI and LLMs, offer on-demand accessibility and can be used anytime, anywhere. Importantly, they present minimal ethical concerns regarding patient data privacy, as they are not based on real patient information [43, 44]. The novelty of this approach lies not in the rule content, which mirrors published algorithms, but in its ASP-driven simulation framework that allows reproducible testing without patient data.
However, the use of ASPs also raises important ethical considerations. Because they are generated from predefined assumptions, ASPs may reflect the constraints and biases of their simulation framework rather than the full complexity of real-world clinical variability. If their results are interpreted uncritically, they risk providing false reassurance by implying a level of accuracy or validation not supported by actual patient data. A specific concern is the potential misuse or misinterpretation of ASP-derived findings, such as citing them as evidence of clinically validated model performance or applying them directly in patient care. Such misrepresentation could foster unwarranted confidence in unverified prediction tools and lead to premature clinical adoption. There is also the danger of misinterpretation, especially if ASP-based findings are applied directly to care rather than understood as exploratory [45]. To mitigate this risk, all ASP-based findings in this study are explicitly identified as exploratory and intended solely for hypothesis generation and conceptual illustration, not as validated clinical evidence. These cautions underscore that ASPs should serve as a safe, ethical environment for testing ideas, not as a substitute for empirical data or clinical judgment [40].
This proof-of-concept study has several important limitations. First, it relied entirely on ASPs rather than real clinical data. While this approach enabled safe and reproducible testing, it does not capture the full complexity and heterogeneity of real patient populations. Second, because the simulated cohort was not anchored to epidemiologic or registry-based data, the resulting distributions of age, BMI, diabetes duration, and related metabolic variables may not accurately reflect real-world patterns. The plausibility of these distributions and inter-variable relationships has not yet been empirically verified. Third, although safeguards such as a fixed temperature setting were applied to minimize variability, prompts in large language models remain sensitive to small wording changes, and exact reproducibility cannot be guaranteed. Fourth, expert review of GPT-4 outputs was qualitative rather than quantitative; inter-rater reliability was not assessed, and verification was limited to confirming face validity. Fifth, the study did not include benchmarking against other large language models (e.g., GPT-5 or Gemini), which may differ in reasoning stability and output structure. Sixth, the synthetic cohort included 79% female profiles, slightly higher than the ~ 70% typically reported in bariatric surgery cohorts, likely reflecting random variation or inherent bias in GPT-based data generation. Finally, although expert oversight ensured clinical plausibility, it may also have introduced subjective bias, as human judgments are inherently influenced by prior experience.
These limitations underscore the exploratory nature of this work and highlight the need for future studies incorporating real-world data anchoring, quantitative expert validation, and multi-model benchmarking to confirm generalizability and robustness.
While ASPs offer a powerful and ethical way to explore clinical scenarios, they don’t fully capture the complexity and nuance of real-world patient data. Models that are overly tuned to synthetic patterns may struggle to generalize to actual clinical populations. This underscores the importance of thoughtful design and rigorous external validation when using ASPs and AI tools in healthcare [17, 38].
Despite these limitations, GPT-4 shows strong potential to synthesize complex clinical information and prediction model criteria, delivering personalized, human-like recommendations that align with existing evidence and clinical reasoning. This study helps bridge the gap between diabetes remission prediction models after MBS and the realities of clinical decision-making.
By combining ASP profiles with LLMs, we introduce a novel, structured approach to illustrate the selection process in existing prediction tools. This method not only can reveal where each model excels or falls short, but also may support a more personalized, patient-centered approach to model selection. Notably, our findings highlight that no single model delivers perfect accuracy across all patient types and MBS procedures. While combining multiple models might improve prediction precision, doing so can be impractical and overwhelming in fast-paced clinical environments.
It is important to note that, because this study relied exclusively on AI-simulated patients, no conclusions can be drawn about the model’s accuracy in real-world settings. In addition, even with overlapping baseline characteristics, surgical type and data completeness dictate model assignment. This highlights that model choice is procedure-sensitive, not purely patient-sensitive.
The observed selection frequencies are scenario-dependent and serve only as proof-of-concept outputs.
The AI-driven framework presented here addresses this challenge by simulating diverse clinical scenarios and evaluating each model’s performance in a consistent, scalable manner. This approach can help inform the development of more robust, inclusive prediction tools and ultimately support better, more tailored care for people living with T2DM. I.
Importantly, this framework should be interpreted as exploratory. While ASPs demonstrate feasibility for structured, reproducible testing of diabetes remission prediction models, their translational value will depend on validation against real patient cohorts. Future work incorporating GPT-5 and other LLMs, as well as systematic assessment of inter-rater reliability, will be critical to establish robustness and clinical relevance.
Conclusion
This proof-of-concept study demonstrates that LLM-generated simulated patients can be used to illustrate and refine a decision-support algorithm for selecting among validated diabetes remission prediction models. The approach is reproducible, transparent, and ethically safe for early-stage testing. Future work should validate this framework with real patient data and assess its impact on clinical decision-making.
Supplementary Information
Acknowledgements
Not applicable.
Author contributions
M.K., M.M., and R.V. were involved in the conceptualization, design, and conduct of the study, as well as the analysis of the results. M.K. and R.V.C. wrote the first draft of the manuscript. M.K., M.M., and R.V. edited, reviewed, and approved the final version of the manuscript. M.K. and R.V.C. are the guarantors of this work and, as such, had full access to all the data in the study and take responsibility for the integrity of the data and the accuracy of the data analysis.
Funding
No Funding.
Data availability
No datasets were generated or analysed during the current study.
Declarations
Ethics approval and consent to participate
As no primary data collection was performed, no formal ethical evaluation is required by our institutions.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Kirwan JP, Courcoulas AP, Cummings DE, Goldfine AB, Kashyap SR, Simonson DC, et al. Diabetes remission in the alliance of randomized trials of medicine versus metabolic surgery in type 2 diabetes (ARMMS-T2D). Diabetes Care. 2022;45(7):1574–83. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.De Luca M, Zese M, Bandini G, Chiappetta S, Iossa A, Merola G, et al. Metabolic bariatric surgery as a therapeutic option for patients with type 2 diabetes: a meta-analysis and network meta-analysis of randomized controlled trials. Diabetes Obes Metab. 2023;25(8):2362–73. [DOI] [PubMed] [Google Scholar]
- 3.Still CD, Wood GC, Benotti P, Petrick AT, Gabrielsen J, Strodel WE, et al. Preoperative prediction of type 2 diabetes remission after Roux-en-Y gastric bypass surgery: a retrospective cohort study. Lancet Diabetes Endocrinol. 2014;2(1):38–45. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Aron-Wisnewsky J, Sokolovska N, Liu Y, Comaneshter DS, Vinker S, Pecht T, et al. The advanced-Diarem score improves prediction of diabetes remission 1 year post-Roux-en-Y gastric bypass. Diabetologia. 2017;60(10):1892–902. [DOI] [PubMed] [Google Scholar]
- 5.Chen JC, Hsu NY, Lee WJ, Chen SC, Ser KH, Lee YC. Prediction of type 2 diabetes remission after metabolic surgery: a comparison of the individualized metabolic surgery score and the ABCD score. Surg Obes Relat Dis. 2018;14(5):640–5. [DOI] [PubMed] [Google Scholar]
- 6.Aminian A, Brethauer SA, Andalib A, Nowacki AS, Jimenez A, Corcelles R, et al. Individualized metabolic surgery score: procedure selection based on diabetes severity. Ann Surg. 2017;266(4):650–7. [DOI] [PubMed] [Google Scholar]
- 7.Pucci A, Tymoszuk U, Cheung WH, Makaronidis JM, Scholes S, Tharakan G, et al. Type 2 diabetes remission 2 years post Roux-en-Y gastric bypass and sleeve gastrectomy: the role of the weight loss and comparison of DiaRem and diabetter scores. Diabet Med. 2018;35(3):360–7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Ghusn W, Ma P, Vierkant RA, Mundi M, Fehervari M, Ikemiya K et al. The Diabetes Remission Index (DRI): A Novel Prognostic Calculator Model Predicting Diabetes Remission Before and After Metabolic Procedures. Ann Surg. 2025. [DOI] [PubMed]
- 9.Ghusn W, Hage K, Vierkant RA, Collazo-Clavell ML, Abu Dayyeh BK, Kellogg TA, et al. Type-2 diabetes mellitus remission prediction models after Roux-en-Y gastric bypass and sleeve gastrectomy based on disease severity scores. Diabetes Res Clin Pract. 2024;208:111091. [DOI] [PubMed] [Google Scholar]
- 10.Singh P, Adderley NJ, Hazlehurst J, Price M, Tahrani AA, Nirantharakumar K, et al. Prognostic models for predicting remission of diabetes following bariatric surgery: a systematic review and meta-analysis. Diabetes Care. 2021;44(11):2626–41. [DOI] [PubMed] [Google Scholar]
- 11.Kermansaravi M, Kermansaravi A, Kroh M, Shikora SA, Cohen RV. Collaborative intelligence in metabolic and bariatric surgery: integrating human expertise and artificial intelligence for better outcomes. Obes Surg. 2025;35(8):2787–9. 10.1007/s11695-025-07999-y. [DOI] [PubMed] [Google Scholar]
- 12.Bhatt AB, Bae J. Collaborative intelligence to catalyze the digital transformation of healthcare. NPJ Digit Med. 2023;6(1):177. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Brügge E, Ricchizzi S, Arenbeck M, Keller MN, Schur L, Stummer W, et al. Large language models improve clinical decision making of medical students through patient simulation and structured feedback: a randomized controlled trial. BMC Med Educ. 2024;24(1):1391. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Moingeon P, Chenel M, Rousseau C, Voisin E, Guedj M. Virtual patients, digital twins and causal disease models: paving the ground for in silico clinical trials. Drug Discov Today. 2023;28(7):103605. [DOI] [PubMed] [Google Scholar]
- 15.Wang H, Arulraj T, Ippolito A, Popel AS. From virtual patients to digital twins in immuno-oncology: lessons learned from mechanistic quantitative systems pharmacology modeling. NPJ Digit Med. 2024;7(1):189. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Hsieh VC-R, Liu M-Y, Lin H-C. AI-enabled clinical decision support system modeling for the prediction of cirrhosis complications. IRBM. 2024;45(5):100854. [Google Scholar]
- 17.Eertink JJ, Heymans MW, Zwezerijnen GJC, Zijlstra JM, de Vet HCW, Boellaard R. External validation: a simulation study to compare cross-validation versus holdout or external testing to assess the performance of clinical prediction models using PET data from DLBCL patients. EJNMMI Res. 2022;12(1):58. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Lee S, Kim DW, Oh NE, Lee H, Park S, Yon DK, et al. External validation of an artificial intelligence model using clinical variables, including ICD-10 codes, for predicting in-hospital mortality among trauma patients: a multicenter retrospective cohort study. Sci Rep. 2025;15(1):1100. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Dave T, Athaluri SA, Singh S. ChatGPT in medicine: an overview of its applications, advantages, limitations, future prospects, and ethical considerations. Front Artif Intell. 2023;6:1169595. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Kermansaravi M, Shahmiri SS, Khalaj A, Jalali SM, Amini M, Alamdari NM, et al. The first web-based Iranian national obesity and metabolic surgery database (INOSD). Obes Surg. 2022;32(6):2083–6. [DOI] [PubMed] [Google Scholar]
- 21.Wang GF, Yan YX, Xu N, Yin D, Hui Y, Zhang JP, et al. Predictive factors of type 2 diabetes mellitus remission following bariatric surgery: a meta-analysis. Obes Surg. 2015;25(2):199–208. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Singh P, Adderley NJ, Subramanian A, Gokhale K, Hazlehurst J, Singhal R, et al. Glycemic outcomes in patients with type 2 diabetes after bariatric surgery compared with routine care: a population-based, real-world cohort study in the United Kingdom. Surg Obes Relat Dis. 2022;18(12):1366–76. [DOI] [PubMed] [Google Scholar]
- 23.Ribaric G, Buchwald JN, McGlennon TW. Diabetes and weight in comparative studies of bariatric surgery vs conventional medical therapy: a systematic review and meta-analysis. Obes Surg. 2014;24(3):437–55. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Plaeke P, Beunis A, Ruppert M, De Man JG, De Winter BY, Hubens G, Review. Performance Comparison, and validation of models predicting type 2 diabetes remission after bariatric surgery in a Western European population. Obes Surg. 2021;31(4):1549–60. [DOI] [PubMed] [Google Scholar]
- 25.Zhu H, Guo P, Zhao Y, Wu X, Wang B, Yang H, et al. Prediction model of diabetes remission at 1-year after sleeve gastrectomy and comparison with other models. Obes Surg. 2025;35(1):249–56. [DOI] [PubMed] [Google Scholar]
- 26.Tharakan G, Scott R, Szepietowski O, Miras AD, Blakemore AI, Purkayastha S, et al. Limitations of the Diarem score in predicting remission of diabetes following Roux-en-Y gastric bypass (RYGB) in an ethnically diverse population from a single institution in the UK. Obes Surg. 2017;27(3):782–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Saberdoust F, Salehabadi G, Sheykholeslamy S, Noroozi E, Moradi M, Pazouki A, et al. Diagnostic value of Advanced-DiaRem for predicting diabetic remission after one anastomosis gastric bypass/Minigastric bypass. Obes Surg. 2024;34(9):3467–74. [DOI] [PubMed] [Google Scholar]
- 28.Debédat J, Sokolovska N, Coupaye M, Panunzi S, Chakaroun R, Genser L, et al. Long-term relapse of type 2 diabetes after Roux-en-Y gastric bypass: prediction and clinical relevance. Diabetes Care. 2018;41(10):2086–95. [DOI] [PubMed] [Google Scholar]
- 29.Dicker D, Golan R, Aron-Wisnewsky J, Zucker JD, Sokolowska N, Comaneshter DS, et al. Prediction of long-term diabetes remission after RYGB, sleeve gastrectomy, and adjustable gastric banding using DiaRem and Advanced-DiaRem scores. Obes Surg. 2019;29(3):796–804. [DOI] [PubMed] [Google Scholar]
- 30.Sjöholm K, Carlsson LMS, Taube M, le Roux CW, Svensson PA, Peltonen M. Comparison of preoperative remission scores and diabetes duration alone as predictors of durable type 2 diabetes remission and risk of diabetes complications after bariatric surgery: a post hoc analysis of participants from the Swedish obese subjects study. Diabetes Care. 2020;43(11):2804–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Fatima F, Hjelmesæth J, Hertel JK, Svanevik M, Sandbu R, Småstuen MC, et al. Validation of Ad-DiaRem and ABCD diabetes remission prediction scores at 1-year after Roux-en-Y gastric bypass and sleeve gastrectomy in the randomized controlled Oseberg trial. Obes Surg. 2022;32(3):801–9. [DOI] [PubMed] [Google Scholar]
- 32.Ohta M, Seki Y, Ohyama T, Bai R, Kim SH, Oshiro T, et al. Prediction of long-term diabetes remission after metabolic surgery in obese East Asian patients: a comparison between ABCD and IMS scores. Obes Surg. 2021;31(4):1485–95. [DOI] [PubMed] [Google Scholar]
- 33.Ahuja A, Tantia O, Chaudhuri T, Khanna S, Seetharamaiah S, Majumdar K, et al. Predicting remission of diabetes post metabolic surgery: a comparison of ABCD, DiaRem, and DRS scores. Obes Surg. 2018;28(7):2025–31. [DOI] [PubMed] [Google Scholar]
- 34.de Abreu Sesconetto L, da Silva RBR, Galletti RP, Agareno GA, Colonno BB, de Sousa JHB, et al. Scores for predicting diabetes remission in bariatric surgery: a systematic review and meta-analysis. Obes Surg. 2023;33(2):600–10. [DOI] [PubMed] [Google Scholar]
- 35.Saarinen I, Grönroos S, Hurme S, Peterli R, Helmiö M, Bueter M, et al. Validation of the individualized metabolic surgery score for bariatric procedure selection in the merged data of two randomized clinical trials (SLEEVEPASS and SM-BOSS). Surg Obes Relat Dis. 2023;19(5):522–9. [DOI] [PubMed] [Google Scholar]
- 36.Baldane S, Celik M, Korez MK, Baldane EG, Yilmaz H, Abusoglu S, et al. Comparison of scoring systems for predicting remission of type 2 diabetes in sleeve gastrectomy patients. Rom J Intern Med. 2022;60(4):235–43. [DOI] [PubMed] [Google Scholar]
- 37.Karpińska IA, Choma J, Wysocki M, Dudek A, Małczak P, Szopa M, et al. External validation of predictive scores for diabetes remission after metabolic surgery. Langenbecks Arch Surg. 2022;407(1):131–41. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Cai Y, Cai YQ, Tang LY, Wang YH, Gong M, Jing TC, et al. Artificial intelligence in the risk prediction models of cardiovascular disease and development of an independent validation screening tool: a systematic review. BMC Med. 2024;22(1):56. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Tsopra R, Fernandez X, Luchinat C, Alberghina L, Lehrach H, Vanoni M, et al. A framework for validating AI in precision medicine: considerations from the European ITFoC consortium. BMC Med Inform Decis Mak. 2021;21(1):274. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Zhang J, Zhang ZM. Ethics and governance of trustworthy medical artificial intelligence. BMC Med Inform Decis Mak. 2023;23(1):7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Rajpurkar P, Chen E, Banerjee O, Topol EJ. AI in health and medicine. Nat Med. 2022;28(1):31–8. [DOI] [PubMed] [Google Scholar]
- 42.Gichoya JW, Thomas K, Celi LA, Safdar N, Banerjee I, Banja JD, et al. AI pitfalls and what not to do: mitigating bias in AI. Br J Radiol. 2023;96(1150):20230023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Sardesai N, Russo P, Martin J, Sardesai A. Utilizing generative conversational artificial intelligence to create simulated patient encounters: a pilot study for anaesthesia training. Postgrad Med J. 2024;100(1182):237–41. [DOI] [PubMed] [Google Scholar]
- 44.Yamamoto A, Koda M, Ogawa H, Miyoshi T, Maeda Y, Otsuka F, et al. Enhancing medical interview skills through AI-simulated patient interactions: nonrandomized controlled trial. JMIR Med Educ. 2024;10:e58753. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Draghi B, Wang Z, Myles P, Tucker A. Identifying and handling data bias within primary healthcare data using synthetic data generators. Heliyon. 2024;10(2):e24164. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
No datasets were generated or analysed during the current study.


