Abstract
Background
Current diagnostic criteria for type 2 diabetes (T2D) capture disease heterogeneity poorly, and do not reliably predict progression, complications, or treatment response. The phenotypic clustering model proposed by Ahlqvist et al. identified five T2D subtypes using six clinical variables at diagnosis, each associated with distinct metabolic profiles and complication risks. Although this framework has been replicated in several cohorts, evidence in Mediterranean populations is lacking.
Methods
We conducted a prospective cohort study in Catalonia (Northeast Spain) including adults with newly diagnosed T2D recruited between March 2022 and January 2026. Using baseline data, we evaluated the Ahlqvist clustering approach. Autoantibody-positive individuals were classified as severe autoimmune diabetes (SAID), and sex-stratified k-means clustering (k = 4) was applied to autoantibody-negative participants. Cluster separation and stability were assessed using principal component analysis and silhouette analyses.
Results
A final total number of 991 individuals with newly diagnosed T2D were included in the analysis. Autoantibodies were present in 67 subjects (6.8%), thereby being classified as SAID. Among the remaining 924 participants, sex-stratified k-means clustering (k = 4) identified clusters with metabolic profiles consistent with the classical subtypes: mild age-related diabetes (MARD, n = 326), severe insulin-resistant diabetes (SIRD, n = 241), mild obesity-related diabetes (MOD, n = 206), and severe insulin-deficient diabetes (SIDD, n = 151). However, cluster separation was modest, and bootstrap stability was limited (Jaccard 0.555–0.718). In an unconstrained analysis, apart from the autoimmune diabetes group, silhouette optimisation identified three clusters as the most internally optimal structure, corresponding broadly to obesity/insulin-resistant (C1, n = 347), insulin-deficient (C2, n = 186), and age-related (C3, n = 391) phenotypes. Stability was substantially higher for the three-cluster solution (Jaccard 0.799–0.863). Concordance between the Ahlqvist and data-driven models was moderate (ARI = 0.473), with MOD individuals distributed across the other clusters.
Conclusions
The Ahlqvist clustering architecture could be approximated in this Mediterranean cohort at diagnosis, but internal stability of the five-cluster solution was limited. In this population, a four-cluster structure showed substantially better internal validity. These findings support the feasibility of phenotypic subclassification of T2D while underscoring the importance of evaluating population-specific cluster structures and their clinical relevance in longitudinal studies.
Trial registration: NCT05333718.
Graphical abstract
Supplementary Information
The online version contains supplementary material available at 10.1186/s12933-026-03198-w.
Keywords: Type 2 diabetes, Diabetes subgroups, Phenotypic clustering, Precision medicine, Disease heterogeneity, Cardiovascular risk, Primary healthcare, Epidemiology, Mediterranean population
Research insights
What is currently known about this topic?
Type 2 diabetes is heterogeneous and clustering identifies metabolic subtypes.
Ahlqvist clusters have been replicated in several populations.
Cluster distribution and phenotypes vary between populations.
What is the key research question?
Are Ahlqvist T2D clusters reproducible at diagnosis in Mediterranean populations?
What is new?
This is the first evaluation of Ahlqvist clustering in Mediterranean newly diagnosed T2D.
Five Ahlqvist-like phenotypes were identified but showed modest stability.
A four-cluster solution showed substantially better internal validity.
How might this study influence clinical practice?
Early clustering may help identify diabetes subgroups at onset of type 2 diabetes.
Background
Type 2 diabetes (T2D) is a heterogeneous metabolic disorder characterised by substantial variation in clinical presentation, pathophysiology, and risk of complications. However, traditional diagnostic criteria do not adequately capture this heterogeneity and provide limited information for predicting disease progression or therapeutic response [1, 2].
A major advance in phenotypic subclassification was reported by Ahlqvist et al. [3], who applied data-driven clustering using six clinical variables measured at diagnosis (age at diagnosis, glycated haemoglobin [HbA1c], body mass index [BMI], islet autoantibody status, and estimates of insulin resistance and β-cell function) to define five distinct T2D subtypes. These subtypes differed in metabolic characteristics at diagnosis, as well as in disease progression and risk of complications. Since this initial report, these clusters have been replicated in multiple cohorts across Europe, Asia, and the Americas, although differences in cluster prevalence and separation have been observed across populations [2, 4–10].
A recent systematic review evaluating 51 studies of subclassification strategies examined both simple clinical approaches and more complex machine-learning-based strategies [11]. The review concluded that the Ahlqvist phenotypic clusters remain the most consistently replicated classification framework. However, the reproducibility and clinical relevance of these clusters appear to depend partly on the characteristics of the underlying populations, highlighting the need to evaluate this framework across diverse ancestries and settings.
Regional differences in genetic background, lifestyle patterns, environmental exposures, and cardiometabolic risk profiles may contribute to differences in the phenotypic heterogeneity of T2D, underscoring the need for population-specific validation of clustering approaches [2, 11, 12]. Despite increasing interest in T2D subclassification, evidence regarding the applicability of the Ahlqvist phenotypic classification in Mediterranean populations remains limited.
Therefore, the primary aim of this study was to identify and characterize clinically relevant phenotypic clusters of newly diagnosed T2D in our Mediterranean population using the same clinical variables and analytical approach proposed by Ahlqvist et al. By examining cluster structure at the time of diagnosis, this study aimed to extend evidence on population-specific heterogeneity in T2D and to inform future research on precision approaches to diabetes classification across different populations. Understanding early metabolic subtypes may also help clarify differences in long-term cardiovascular risk among individuals with newly diagnosed T2D.
Methods
Study design and setting
This was a prospective, observational cohort study conducted in Catalonia, Northeast Spain [13]. The present analysis represents of the main objective of the study, i.e. the initial phenotypic characterization of a newly established cohort of individuals with newly diagnosed diabetes. This was also designed to enable future longitudinal and multi-omics investigations of diabetes heterogeneity.
Recruitment started on March 4, 2022 and ended January 31, 2026. Participants were recruited from 28 healthcare centres, including two reference hospitals and 26 primary healthcare centres from different public healthcare districts of Barcelona and Lleida. The present analysis was a cross-sectional assessment based on baseline data extracted on January 31, 2026 to reproduce the T2D clustering approach described by Ahlqvist et al. [3] at the time of diagnosis.
Study population
Eligible participants were adults (≥ 18 years) with newly diagnosed T2D according to the American Diabetes Association [1] criteria, with diagnosis documented within the previous three months according to medical records. Individuals with other types of diabetes, including type 1 diabetes, maturity-onset diabetes of the young (MODY), gestational diabetes, or diabetes secondary to other causes, were excluded.
Sample size considerations
Because clustering is an exploratory data-driven method, no formal power calculation was performed. A target sample size of 1200 participants was selected based on the expected incidence of T2D in the study population and on sample sizes used in previous clustering studies applying similar methodologies [14].
Clinical and laboratory assessment
A detailed description of the study design has already been published [13]. At inclusion, investigators conducted a standardized clinical assessment and recorded data through electronic case report forms. Fasting peripheral blood and urine samples were collected within two weeks of the baseline visit.
Main variables
The six variables included in the clustering analysis were: age at diagnosis, HbA1c, BMI, islet autoantibody status (glutamic acid decarboxylase 65 [GAD65] and/or insulinoma-associated antigen-2 [IA-2] antibodies), and estimates of insulin resistance and β-cell function derived from the homeostatic model assessment (HOMA2-IR and HOMA2-β, respectively). These variables were selected to replicate the original clustering approach described by Ahlqvist et al. [3], enabling direct comparison of diabetes subgroups between populations.
Body weight and height were measured using standardized procedures at participating centres, and BMI was calculated as weight in kilograms divided by height in meters squared (kg/m2). HbA1c at time of diagnosis was measured using standard clinical laboratory methods in all participating centres. Diabetes-related autoantibodies were measured using commercially available enzyme-linked immunosorbent assay (ELISA) kits for GAD65 and IA-2 autoantibodies as described in the study protocol [13]. C-peptide concentrations were measured using a high-sensitivity direct chemiluminescent immunoassay (LIASON XL “Flash” chemiluminescence technology, CLIA analyser) with a compatible assay kit. HOMA2-IR and HOMA2-β were calculated using fasting plasma glucose and C-peptide concentrations and through the HOMA-2 calculator developed by the Diabetes Trials Unit, University of Oxford, under licensed use [15]. These indices were used as measures of insulin resistance and β-cell function, respectively. Ancestry was assessed based on self-reported ethnicity.
Other baseline collected variables can be reviewed in the published study protocol [13].
Statistical analysis
Data preprocessing
Complete-case analysis was applied to exclude participants with missing values in clustering variables: age at diabetes onset, BMI, HbA1c, HOMA2-B, HOMA2-IR, autoantibody status, or sex. Implausible physiological values were excluded based on predefined plausibility thresholds to improve data quality: BMI ≤ 10 or ≥ 80 kg/m2 and non-positive HOMA2-B or HOMA2-IR values (incompatible with log-transformation). HOMA2-B and HOMA2-IR were natural log-transformed to reduce right-skewness and stabilise variance. Clustering was based on five continuous variables: age at diabetes onset (years), BMI (kg/m2), HbA1c (mmol/mol), log(HOMA2-B) (%), and log(HOMA2-IR). These were standardised to mean 0 and SD 1 prior to clustering. Sex was not included as a clustering variable but was used to stratify the Ahlqvist replication analysis and bootstrap stability assessment [3].
Participants testing positive for GAD65 autoantibodies (GADA) and/or IA2 autoantibodies were classified a priori as severe autoimmune diabetes (SAID). Autoantibody-negative participants (negative for both GADA and IA2) with known sex formed the primary cohort for data-driven clustering.
Analyses were performed in R version 4.4.2 using the following packages: readxl, cluster, fpc, networkD3, htmlwidgets, and htmltools [16].
Clustering procedures
Ahlqvist replication
To replicate the Ahlqvist subtypes, sex-stratified k-means clustering with a fixed number of clusters k = 4 was performed in the autoantibody-negative cohort. Within each sex-specific subset, the five continuous variables were standardised separately (mean 0, SD 1) prior to clustering. K-means clustering used the kmeansruns function (fpc package) with Euclidean distance, 100 random initialisations, and a maximum of 100 iterations (fixed random seed 123 for reproducibility). Cluster labels were assigned post hoc based on the dominant clinical characteristics of each cluster, consistent with the Ahlqvist framework: the cluster characterised by highest HbA1c/lowest HOMA2-B corresponded to Severe Insulin-Deficient Diabetes (SIDD); highest HOMA2-IR to Severe Insulin-Resistant Diabetes (SIRD); highest age at onset to Mild Age-Related Diabetes (MARD); and the remainder (typically higher BMI, intermediate HbA1c) to Mild Obesity-Related Diabetes (MOD). Sex-specific assignments were merged, with SAID appended as the fifth cluster. Sex was not included as a clustering variable but was used for subsequent sex-stratified analyses.
Data-driven optimal-k clustering
As a secondary exploratory analysis, k-means clustering was performed in the pooled autoantibody-negative cohort, without sex stratification. Variables were standardised globally (mean 0, SD 1) prior to clustering. Models were fitted for k = 2 to k = 6 using Euclidean distance, 100 random initialisations, and a maximum of 100 iterations per model (fixed random seed 123). The optimal number of clusters (kopt) was defined as the value of k that maximised the mean silhouette width. A final k-means model was then fitted using kopt; resulting clusters were labelled C1 to Ckopt. Sex was not included as a clustering variable but was used for subsequent sex-stratified analyses.
Cluster validation
Global and cluster-specific mean silhouette widths were computed using Euclidean distances on standardised features (silhouette function, cluster package). Cluster stability was assessed by sex using non-parametric bootstrap resampling (fpc package clusterboot procedure). For each sex and clustering solution (Ahlqvist k = 4, fixed seed 123; optimal-k, fixed seed 42), 2000 bootstrap samples with replacement were generated. K-means was refitted and Jaccard similarity indices computed between original and bootstrap assignments. Mean Jaccard values > 0.75 were considered indicative of acceptable stability. For the optimal-k solution, bootstrap resampling applied the global scaling parameters derived from the pooled autoantibody-negative cohort.
Principal component analysis (PCA) was applied to visualise cluster separation of standardised clustering variables. PC1-PC2 biplots were generated for the full cohort and stratified by sex. Boxplots summarised clinical variables by cluster (global and sex-stratified). Ahlqvist-like and optimal-k assignments were merged at the individual level using unique record identifiers. Concordance between the two classification schemes was assessed using contingency tables with row-wise percentages and quantified using the Adjusted Rand Index (ARI). An alluvial diagram was used to visualise individual-level patient flows between the two clustering solutions, with link widths proportional to patient numbers.
Results
Cohort characteristics
After exclusion of screen failures and participants with missing values in clustering variables, 991 participants with newly diagnosed T2D from were included in the analysis (Supplementary Fig. 1). Among participants of non-European ancestry, the most represented groups included individuals categorised as Latin American (n = 127, 12.8%), Hindustani (n = 40, 4%), and North Africa and the Middle East (n = 34, 3.4%). Baseline characteristics in the overall population and stratified by sex are detailed in Supplementary Table 1.
Islet autoantibodies (GAD65 and/or IA2) were detected in 6.8% of participants (n = 67), who were classified as SAID. The remaining 924 autoantibody-negative participants were then used for clustering analysis.
Replication of the Ahlqvist clustering framework
Sex-stratified k-means clustering with a fixed number of clusters (k = 4) identified four data-driven subtypes among autoantibody-negative participants. The largest subgroup was mild age-related diabetes (MARD, n = 326), followed by severe insulin-resistant diabetes (SIRD, n = 241), mild obesity-related diabetes (MOD, n = 206), and severe insulin-deficient diabetes (SIDD, n = 151). The distribution of participants across clusters in the cohort, stratified by sex, is shown in Fig. 1.
Fig. 1.
Subject distribution in the study cohort following Ahlqvist clusters. A, Total study participants. B, Analysis by sex. Abbreviations: MARD mild age-related diabetes, MOD mild obesity-related diabetes, SAID severe autoimmune diabetes, SIDD severe insulin-deficient diabetes, SIRD severe insulin-resistant diabetes
The clinical profiles of the clusters were consistent with the original Ahlqvist classification framework. The SIDD cluster was characterised by markedly elevated HbA1c (mean 106.2 mmol/mol) and reduced ß-cell function (HOMA2-B 78.2%). The SIRD cluster showed the highest insulin resistance (HOMA2-IR 4.05) and BMI (37.2 kg/m2). MARD included participants with the oldest age at diabetes onset (mean 68.5 years) and a near-normal metabolic profile, whereas MOD was characterised by intermediate HbA1c levels and BMI with relatively preserved ß-cell function. The distribution of main clustering across subgroups is shown in Fig. 2.
Fig. 2.
Distribution of main variables in the study cohort following Ahlqvist cluster subclassification A, Total study participants. B, Analysis by sex. Abbreviations: BMI body mass index, HbA1c glycated haemoglobin, HOMA2-B Homeostasis Model Assessment 2 of β-cell function, HOMA2-IR Homeostasis Model Assessment 2 of Insulin Resistance, MARD mild age-related diabetes, MOD mild obesity-related diabetes, SAID severe autoimmune diabetes SIDD severe insulin-deficient diabetes, SIRD severe insulin-resistant diabetes
Bootstrap stability analysis indicated poor cluster stability of the four-cluster solution across both sexes. Mean Jaccard similarity indices were 0.718, 0.623, 0.585, and 0.555 across clusters, with none exceeding the predefined stability threshold of 0.75.
Data-driven optimal k clustering
In the unconstrained clustering analysis excluding participants with positive antibodies, the optimal number of clusters determined by silhouette analysis was three. Mean silhouette widths across models were: 0.191 (k = 2); 0.207 (k = 3); 0.174 (k = 4); 0.192 (k = 5), and 0.175 (k = 6). The final model yielded three data-driven clusters (C1-C3) in addition to the predefined SAID group (Fig. 3).
Fig. 3.
Participant cluster distribution in the unconstrained analysis A, Total study participants. B, Analysis by sex. Abbreviations: C1 cluster 1, C2 cluster 2, C3 cluster 3, SAID severe autoimmune diabetes
The distribution of the main clustering variables across these clusters in the total population and stratified by sex is shown in Fig. 4.
Fig. 4.
Cluster distribution according to each variable in the unconstrained analysis. A, Total study participants. B, Analysis by sex. Abbreviations: BMI body mass index, C1 cluster 1, C2 cluster 2, C3 cluster 3, HbA1c glycated haemoglobin, HOMA2-B Homeostasis Model Assessment 2 of β-cell function, HOMA2-IR Homeostasis Model Assessment 2 of Insulin Resistance, SAID severe autoimmune diabetes
These resulting clusters preserved the main metabolic axes observed in the forced solution: an obesity/insulin-resistant phenotype (C1, n = 347), an insulin-deficient phenotype (C2, n = 186), and an age-related phenotype (C3, n = 391). Principal component analysis visualising cluster separation for both solutions are shown in Supplementary Figs. 2 and 3. Bootstrap stability was substantially higher for the three-cluster solution than in the four-cluster (Ahlqvist) model. Mean Jaccard similarity indices were 0.863 for C1, 0.851 for C2, and 0.799 for C3, all exceeding the stability threshold of 0.75.
Concordance between clustering approaches
Overall concordance between the Ahlqvist replication and the data-driven optimal-k solution was moderate (ARI = 0.473), indicating partial agreement beyond chance. Row-wise contingency analysis revealed a clear correspondence between the three data-driven clusters and four Ahlqvist subtypes, after excluding the group with autoimmune diabetes. Among subjects classified as SIDD in the Ahlqvist framework, 141 were assigned to C2 (93.4%), 8 to C1 (5.3%), and 2 to C3 (1.3%) in the data-driven solution. Similarly, among those classified as SIRD, 230 were assigned to C1 (95.4%), 7 to C3 (2.9%), and 4 to C2 (1.7%). Among individuals classified as MARD, 265 were assigned to C3 (81.3%), 58 to C1 (17.8%), and 3 to C2 (0.9%). Individuals classified as MOD were distributed across all three clusters, with 117 assigned to C3 (56.8%), 51 to C1 (24.8%), and 38 to C2 (18.4%). SAID was perfectly preserved across both solutions (100%), as expected given its a priori classification. Individual-level patient flows between clustering solutions are shown in the alluvial diagram (Supplementary Fig. 4).
Discussion
To the best of our knowledge, this study provides the first evaluation of the diabetes subgroup classification proposed by Ahlqvist and colleagues in a Mediterranean cohort of newly diagnosed T2D. In our population, the forced k-means clustering approach identified groups with metabolic profiles consistent with the four classical non-autoimmune phenotypes (SIDD, SIRD, MOD, and MARD), supporting the general transferability of this framework beyond the original Scandinavian setting. However, our findings suggest that the boundaries between these subgroups may be less clearly defined in this population.
First, although the forced k = 4 solution reproduced recognisable phenotypes, cluster separation was modest and bootstrap stability was low, with none of the clusters reaching the commonly accepted stability threshold (Jaccard > 0.75). This observation suggests that the four-cluster structure may not represent the most robust classification of metabolic heterogeneity in this Mediterranean cohort. Some notable discrepancies were observed while reproducing the Ahlqvist framework. BMI was not as markedly high in the MOD cluster, suggesting that this group may not be primarily driven by obesity in this population. Similarly, HbA1c levels in the SIRD cluster were not as high as expected when compared with the original clusters, indicating a less pronounced severity of hyperglycaemia at the time of diagnosis. Taken together, these findings further suggest less clear phenotypic separation between the original clusters in this cohort. Similar observations have been reported in other replication studies. For example, analyses of the multiple European cohorts within the IMI-RHAPSODY consortium showed that cluster proportions and boundaries may vary depending on the availability of metabolic variables and the handling of missing data [7].
The SAID cluster in our cohort also differed from that described in the original Ahlqvist study. Specifically, in our study, the SAID group had higher BMI, lower HbA1c and higher HOMA2-β than those reported in the Swedish study. One possible explanation is that individuals with a clear clinical diagnosis of type 1 diabetes were excluded at baseline in our study. As follow-up data are not yet available, the clinical trajectory of individuals within the SAID group in our study will be determined in the future, although it is conceivable that they will resemble those with latent autoimmune diabetes in adults.
When the clustering analysis was allowed to determine the optimal number of groups, a three-cluster solution emerged as the most internally consistent model, with higher silhouette width and markedly improved bootstrap stability. The detailed mapping between clustering approaches further highlights differences in the underlying structure of the data. While SIDD, SIRD, and MARD showed a degree of correspondence with the data-driven clusters, the MOD group was distributed across all three clusters, without a clear dominant allocation. This dispersion suggests that, in this population, MOD may not represent a distinct clear phenotype in this population, but rather a metabolic spectrum overlapping with other subgroups.
These findings are broadly consistent with observations from other European cohorts. In the German Diabetes Study, Zaharia et al. [4] reproduced the Ahlqvist clusters and reported MARD as the most prevalent subgroup, followed by MOD, SIDD and SIRD. Similar distributions were observed in the Danish DD2 cohort and in the Norwegian HUNT population study, supporting the external validity of the framework across Northern European populations [17, 18]. However, several replication studies have also noted substantial overlap between clusters and variability in cluster prevalence across cohorts, reinforcing the view that the metabolic heterogeneity of T2D may not always resolve into the same clearly discrete categories [11, 19, 20].
Population-specific factors may further contribute to variation in cluster structure. Studies in multi-ethnic cohorts such as MASALA and MESA have shown that prevalence insulin-deficient phenotypes are more common among South Asian individuals than among populations of European ancestry [21]. Replication studies conducted in Asian populations have similarly reported differences in cluster distribution [6, 22, 23]. Analyses from Mexican and Middle Eastern cohorts have also identified different distributions of insulin-resistant and obesity-related phenotypes, highlighting the influence of genetic background and environmental exposures on metabolic heterogeneity [24]. The distribution observed in our Mediterranean cohort, where SIRD represented a relatively prominent phenotype and MOD lacked clear separation, may therefore reflect regional differences in obesity patterns, insulin resistance, and diabetes presentation.
Taken together, these findings suggest that while the principal metabolic axes underlying diabetes heterogeneity appear to be largely conserved across populations, cluster-based subclassification may capture dominant metabolic tendencies rather than strictly discrete disease entities.
From a clinical perspective, subclassification of T2D has been proposed as a step toward precision medicine. In the original ANDIS study [3], the SIDD subgroup was associated with a higher risk of diabetic retinopathy, whereas SIRD showed increased susceptibility to diabetic kidney disease. Subsequent studies have confirmed associations between a given cluster and the complications’ risk; however, the clinical added value of clustering remains debated. For example, the population-based HUNT study reported that although cluster membership was associated with diabetes-related complications, it did not substantially outperform traditional clinical risk factors in predicting outcomes [18]. Importantly, as cardiovascular disease remains the leading cause of morbidity and mortality in T2D, early identification of metabolic subtypes may help to better characterise heterogeneous cardiovascular risk profiles, including the development of atherosclerotic disease. Because clustering in the present study was performed close to the time of the diagnosis, the COPERNICAN cohort offers a unique opportunity to examine whether these metabolic phenotypes are associated with distinct cardiovascular risk and other outcomes, and whether cluster-based stratification provides additional prognostic information beyond traditional risk factors.
A major strength of the present work is the prospective design. As specified in the current study protocol, this cohort was recruited in Catalonia specifically to characterize T2D subgroups at diagnosis in a Mediterranean region using the core phenotypic variables of the Ahlqvist model. In contrast to many replication studies [7, 17, 21], that have relied on secondary analyses of existing cohorts, the present study was designed explicitly for external evaluation of this clustering framework. Moreover, the clustering methodology closely followed the original Ahlqvist approach and included sex-stratified standardization of the clustering variables.
The timing of phenotyping is another important strength. Clustering was performed close to the time of diagnosis, consistent with the original ANDIS framework and with some subsequent studies such as the German Diabetes Study and the Danish DD2 cohort [3, 4, 17]. However, diagnosis-time clustering studies remain relatively uncommon: in a recent systematic review by Misra et al. [11], only 11 of 22 direct replication studies were conducted in individuals with new-onset diabetes. Moreover, even among cohorts described as newly diagnosed, phenotyping has not always occurred before treatment initiation; in DD2, for example, 85% of participants had already started glucose-lowering therapy at enrolment [17]. This distinction is important because cluster characteristics are not static and may evolve over time with disease progression duration or treatment. For instance, subgrouping in MASALA and MESA was performed on average 5.7 years after diagnosis, and in some RHAPSODY cohorts pretreatment diagnosis-time data were not available [7, 21].
Several limitations should also be considered. First, clustering relied on a limited set of metabolic variables (those pre-specified in the study design). In particular, the use of HOMA2-derived indices of insulin resistance and β-cell function represents an additional limitation, as these measures are not routinely available in clinical practice. While their inclusion allows direct comparison with the original Ahlqvist framework, this clustering approach is not yet readily applicable in many current clinical settings. Future analyses within this cohort are planned to evaluate whether other routinely available clinical variables might provide an improved clustering in our population. Importantly, this study represents the first report from an ongoing prospective cohort in which more extensive biological data have already been collected. Metabolomic profiling and genomic analyses are planned, which will also enable future studies to explore the molecular mechanisms underlying the metabolic subtypes identified here. Second, although the cohort will be followed longitudinally, no complete follow-up data are available. Without longitudinal outcome data, the clinical utility of clustering for risk prediction remains limited. As follow-up data accumulate, future analyses will address specific clinical aspects, such as the response to the different treatments, the evolution of individuals in each cluster, or the development of diabetes-related complications.
Conclusions
In conclusion, this study shows that the Ahlqvist clustering framework can be approximated in a Mediterranean cohort at the time of T2D diagnosis, while also suggesting that metabolic heterogeneity in this population might be better captured by a four-cluster structure. Although the clusters identified in our study showed metabolic characteristics consistent with those reported in other populations, some differences in subtype prevalence and glycaemic severity were observed. The partial overlap between clusters underscores the continuous nature of metabolic variation in diabetes and suggests that cluster-based subclassification may reflect dominant metabolic tendencies rather than strictly discrete disease entities. However, such subclassification provides a useful framework for understanding disease heterogeneity and may help guide future research into precision approaches to diabetes classification and management. Longitudinal studies integrating clinical outcomes, lifestyle factors, and molecular biomarkers will be essential to determine the clinical and prognostic relevance of these clusters throughout different populations.
Supplementary Information
Acknowledgements
COPERNICAN Research Group: Berta Fernandez-Camins, Bogdan Vlacho, Albert Canudas, Marta Ortega, Minerva Granado-Casas, Alexandre Perera-LLuna, Maria Antentas, Mieria Malcampo, Alejandro Boluda-Sanson, Yesmina El Khattabi Ofkir, Josep Franch-Nadal, Dídac Mauricio, Idoia Genua Trullos, Inka Miñambres, Natalia Riera, Antonia Valsera Robles, M Eulalia Pradas Sala, Norma Ballestar Palacios, Olga Vazquez Soler, Cristina Ledesma Serrano, Ana Cámara Caño, Teresa Villuendas, Begona Ribas Lopez, Ana Maria Pedro Pijoan, Eva Gonzalez Platas, Eva Garcia, Marta Hernández, Raquel Martí, Eva Serra Llavall, Montserrat Torra Solé, Maite Puig Solé, Ana Isabel Areny Tallaví, Marc Olivart, Neus Miro Vallve, Àngels Molló, Marta Puig Fito, Jesús Pujol Salud, Cristina Garcia Serrano,Eva Miquel Fernàndez, Lidia Aran Sole, Antonieta Lafarga, Cristina Dominguez Amador, Sandra Guerrero, Cristina Farras Salles, Maria Boldú Franque, Lorena Arnay, Mireia Marin, Laura Turo Pere, Anna Maria Torné Cortés, Laia Llubes Arrià, Cecilia Bañeres Argiles, Assumpció Florensa Flix, Paula Garcia, Manuel Mata Cases, Carla Muñoz, Isabel Bobe, Elena Hernández Boluda, María Begoña Cañas Diaz, Xavier Peligros Palma, Francisco Cegri Lombando, Francesc Xavier Cos Claramunt, Anna Martinez Sanchez, Gabriel Cuatrecasas, Anna Amorós, Blanca Simon, Anna Massana Raurich, Maria Isabel Prieto Fernández, Maria Olga Alvarez Fernández, Robert Cabanes Gómez, Ascensión Niubo Lopez, Marta Crespo Boixasa, Maria Miñana Castellanos, Assumpció Altaba Barcelo, Estel Felez Carroba, Carme Olius Peyrats, Eva Marsol Prieto, Eva Ferris Gallart, Mari Carmen Ruiz Salido, Alicia Mostazo Muntané, Marta Tàrrega Rubio, Maria Lluïsa Calonge Carbonell, Roser Ros Barnadas, Elena Artal Traveria, Zulema Martí Oltra, Carolina Lapena Estella, Olga Poveda Jovellar, Ana Carrera Rodriguez, Sonia Aldea Gomez, Xavier Herraiz Fabregat, Elisabet Llop Lozano, Ignacio Alborch Simó, Laura Güell Espigol, Pilar Pedro Benages, Javier Parramon Nuñez, Paloma Prats, Berta Tió, Carlos Arias, Pelayo Martinez, Cristian Llacer Pinos, Roser Ferrer Costa, Luz Maria Cruz Carlo, Carlos Brotons, Maria Teresa Vilella Moreno, Nuria López, Angela Marcela Colquechambi Castillo, Jordi Hoyo Sánchez, Joan Capellades Llopart, Marta Selvi Blasco. FP expresses gratitude for the support received through the Serra Húnter Programme, funded by the Government of Catalonia (Ref. UPC-LE-231-037).
Abbreviations
- ARI
Adjusted rand index
- ELISA
Enzyme-linked immunosorbent assay
- GAD65
Glutamic acid decarboxylase 65
- GADA
Glutamic acid decarboxylase antibodies
- HOMA-2
Homeostasis model assessment 2
- HOMA2-B
Homeostasis model assessment 2 of β-cell function
- HOMA2-IR
Homeostasis model assessment 2 of insulin resistance
- IA-2
Insulinoma-associated antigen-2
- MOD
Mild obesity-related diabetes
- MODY
Maturity-onset diabetes of the young
- PCA
Principal component analysis
- SAID
Severe autoimmune diabetes
- SIDD
Severe insulin-deficient diabetes
- SIRD
Severe insulin-resistant diabetes
- T2D
Type 2 diabetes
Author contributions
DM and JF-N contributed to the conception and design of the study. DM is the scientific coordinator of the project and designed the hypothesis and objectives. MO and JF-N are local coordinators for the project for primary care centres from Barcelona and Lleida. BV, BF-C and MG-C contributed to the study supervision and execution. MC, BF-C, and COPERNICAN Research Group are responsible for data collection. MIR-L, FP and AP-LL conducted the analysis. BF-C developed the draft manuscript. All authors critically revised the manuscript and approved the submitted version. DM and JF-N share the senior authorship. DM is the guarantor.
Funding
This research received financial support from the Spanish Ministry of Economy and Competitiveness (www.mineco.gob.es) PID2021-122952OB-I00, DPI2017-89827-R and Instituto de Salud Carlos III, Ministry of Health (FIS PI21/01163). This work has been financed by a PERIS grant from the Primary Care Oriented Research Projects program with the reference number SLT021/21/000030 granted by the Generalitat de Catalunya (2022–2024). CIBER for Diabetes and Associated Metabolic Diseases (CIBERDEM; leading group CB15/00071) and the Networking Biomedical Research Centre in the subject area of Bioengineering, Biomaterials and Nanomedicine (CIBER-BBN), are initiatives of Instituto de Salud Carlos III (ISCIII). This research was also partially supported by an unrestricted grant from Menarini Group. The sponsor did not have any roles in the study design, collection, analysis, interpretation of data, writing of the report or the decision to submit the article for publication.
Data availability
The datasets used and/or analysed during the current study are available from the corresponding author on reasonable request.
Declarations
Ethics approval
The present study complies with all the ethical requirements and with the protection of the participant subjects as it complies with the legal ethical Spanish regulation (Real Decreto 957/2020 and the Biomedical Investigation Law 14/2007, 3 July) and the Declaration of Helsinki in its last revision. This protocol was evaluated by the Ethics Committee (CEI) before the inclusion of the patients (Hospital de la Santa Creu i Sant Pau Ethics Committee, ref. HSCSP 21/369 (OBS) dated 17 July 2021.; Ethics Committee for drug research (CEIm) at IDIAP Jordi Gol, CEIm ref. 21/266-P dated 15 December 2021; University Hospital of Bellvitge Ethics Committee for Research, ref. PR438/21 (CSI 21/86) dated 09 February 2022). The study protocol was registered on ClinicalTrials.gov with Trial Registration Number NCT05333718.
Consent for publication
Not applicable.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.American Diabetes Association Professional Practice Committee. 6. Glycemic targets: standards of medical care in diabetes—2022. Diabetes Care. 2022;45:S83-96. 10.2337/dc22-S006. [DOI] [PubMed] [Google Scholar]
- 2.Deutsch AJ, Ahlqvist E, Udler MS. Phenotypic and genetic classification of diabetes. Diabetologia. 2022;65(11):1758–69. 10.1007/s00125-022-05769-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Ahlqvist E, Storm P, Käräjämäki A, Martinell M, Dorkhan M, Carlsson A, et al. Novel subgroups of adult-onset diabetes and their association with outcomes: a data-driven cluster analysis of six variables. Lancet Diabetes Endocrinol. 2018;6(5):361–9. 10.1016/S2213-8587(18)30051-2. [DOI] [PubMed] [Google Scholar]
- 4.Zaharia OP, Strassburger K, Strom A, Bönhof GJ, Karusheva Y, Antoniou S, et al. Risk of diabetes-associated diseases in subgroups of patients with recent-onset diabetes: a 5-year follow-up study. Lancet Diabetes Endocrinol. 2019;7(9):684–94. 10.1016/S2213-8587(19)30187-1. [DOI] [PubMed] [Google Scholar]
- 5.Bello-Chavolla OY, Bahena-López JP, Vargas-Vázquez A, Antonio-Villa NE, Márquez-Salinas A, Fermín-Martínez CA, et al. Clinical characterization of data-driven diabetes subgroups in Mexicans using a reproducible machine learning approach. BMJ Open Diabetes Res Care. 2020;8(1):e001550. 10.1136/bmjdrc-2020-001550. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Anjana RM, Baskar V, Nair ATN, Jebarani S, Siddiqui MK, Pradeepa R, et al. Novel subgroups of type 2 diabetes and their association with microvascular outcomes in an Asian Indian population: a data-driven cluster analysis: the INSPIRED study. BMJ Open Diabetes Res Care. 2020;8(1):e001506. 10.1136/bmjdrc-2020-001506. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Slieker RC, Donnelly LA, Fitipaldi H, Bouland GA, Giordano GN, kerlund M, et al. Replication and cross-validation of type 2 diabetes subtypes based on clinical variables: an IMI-RHAPSODY study. Diabetologia. 2021;64(9):1982–9. 10.1007/s00125-021-05490-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Tanabe H, Saito H, Kudo A, Machii N, Hirai H, Maimaituxun G, et al. Factors associated with risk of diabetic complications in novel cluster-based diabetes subgroups: a Japanese retrospective cohort study. J Clin Med. 2020;9(7):2083. 10.3390/jcm9072083. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Zou X, Zhou X, Zhu Z, Ji L. Novel subgroups of patients with adult-onset diabetes in Chinese and US populations. Lancet Diabetes Endocrinol. 2019;7(1):9–11. 10.1016/S2213-8587(18)30316-4. [DOI] [PubMed] [Google Scholar]
- 10.Dennis JM, Shields BM, Henley WE, Jones AG, Hattersley AT. Disease progression and treatment response in data-driven subgroups of type 2 diabetes compared with models based on simple clinical features: an analysis using clinical trial data. Lancet Diabetes Endocrinol. 2019;7(6):442–51. 10.1016/S2213-8587(19)30087-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Misra S, Wagner R, Ozkan B, Schön M, Sevilla-Gonzalez M, Prystupa K, et al. Precision subclassification of type 2 diabetes: a systematic review. Commun Med. 2023;3(1):138. 10.1038/s43856-023-00360-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Suzuki K, Hatzikotoulas K, Southam L, Taylor HJ, Yin X, Lorenz KM, et al. Genetic drivers of heterogeneity in type 2 diabetes pathophysiology. Nature. 2024;627(8003):347–57. 10.1038/s41586-024-07019-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Fernandez-Camins B, Vlacho B, Canudas A, Ortega M, Granado-Casas M, Perera-LLuna A, et al. Characterisation of type 2 diabetes subgroups at diagnosis: the COPERNICAN prospective observational cohort study protocol. BMJ Open. 2024;14(12):e083825. 10.1136/bmjopen-2023-083825. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Taurbekova B, Sarsenov R, Yaqoob MM, Atageldiyeva K, Semenova Y, Fazli S, et al. Cluster Analysis in Diabetes Research: A Systematic Review Enhanced by a Cross-Sectional Study. J Clin Med. 2025;14(10):3588. 10.3390/jcm14103588. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Diabetes Trials Unit, University of Oxford. HOMA2 Calculator, version [Version, 2.2.3]. Available at: https://www.dtu.ox.ac.uk/homacalculator/.
- 16.R Core Team. R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria.
- 17.Christensen DH, Nicolaisen SK, Ahlqvist E, Stidsen JV, Nielsen JS, Hojlund K, et al. Type 2 diabetes classification: a data-driven cluster study of the danish centre for strategic research in type 2 diabetes (DD2) cohort. BMJ Open Diabetes Res Care. 2022;10(2):e002731. 10.1136/bmjdrc-2021-002731. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Bjarkø VV, Haug EB, Langhammer A, Ruiz PLD, Carlsson S, Birkeland KI, et al. Clinical utility of novel diabetes subgroups in predicting vascular complications and mortality: up to 25 years of follow-up of the HUNT Study. BMJ Open Diabetes Res Care. 2024;12(6):e004493. 10.1136/bmjdrc-2024-004493. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Dsouza SM, Sulaiman F, Abdul F, Mulla F, Ahmed FS, AlSharhan M, et al. Aetiological clustering of newly diagnosed type 2 diabetes using machine learning: a retrospective cross-sectional study in Dubai, UAE. BMJ Open. 2025;15(11):e101284. 10.1136/bmjopen-2025-101284. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Wagner R, Heni M, Tabák AG, Machann J, Schick F, Randrianarisoa E, et al. Pathophysiology-based subphenotyping of individuals at elevated risk for type 2 diabetes. Nat Med. 2021;27(1):49–57. 10.1038/s41591-020-1116-9. [DOI] [PubMed] [Google Scholar]
- 21.Bancks MP, Bertoni AG, Carnethon M, Chen H, Cotch MF, Gujral UP, et al. Association of diabetes subgroups with race/ethnicity, risk factor burden and complications: the MASALA and MESA studies. J Clin Endocrinol Metab. 2021;106(5):e2106–15. 10.1210/clinem/dgaa962. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Xiong X, Yang Y, Wei L, Xiao Y, Li L, Sun L. Identification of two novel subgroups in patients with diabetes mellitus and their association with clinical outcomes: a two-step cluster analysis. J Diabetes Investig. 2021;12(8):1346–58. 10.1111/jdi.13494. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Prasad RB, Asplund O, Shukla SR, Wagh R, Kunte P, Bhat D, et al. Subgroups of patients with young-onset type 2 diabetes in India reveal insulin deficiency as a major driver. Diabetologia. 2022;65(1):65–78. 10.1007/s00125-021-05543-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Al-Thani NM, Zaghlool SB, Abou-Samra AB, Suhre K, Albagha OME. Subtyping of Type 2 Diabetes from a large Middle Eastern Biobank: Implications for Precision Medicine [Internet]. Genetic and Genomic Medicine; 2024 [cited 2026 Mar 15]. Available from: http://medrxiv.org/lookup/doi/10.1101/2024.12.26.24319530 [DOI] [PMC free article] [PubMed]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The datasets used and/or analysed during the current study are available from the corresponding author on reasonable request.





