Abstract
Background
Renal fibrosis (RF) is a key pathological hallmark and prognostic indicator of chronic kidney disease (CKD). Accurate evaluation of RF is critical for risk stratification and therapeutic decision-making, yet the current assessment relies mainly on renal biopsy, which has several limitations. This study aimed to develop a two-stage artificial intelligence framework that integrates multimodal MRI (native T1 mapping, ADC, and T2* mapping) and clinical indicators for the noninvasive assessment of RF in patients with CKD.
Methods
This prospective study included 152 patients with biopsy-proven CKD (RF 1: no RF, 34 patients; RF 2: mild RF, 69 patients; and RF 3: moderate to severe RF, 49 patients). The dataset was randomly partitioned into training and test cohorts at a 2:1 ratio. A two-stage model combining MobileNetV2-SE-based deep learning features with clinical indicators was developed for RF classification. Two binary tasks were performed: RF presence (RF 1 vs. RF 2 and RF 3) and severity (RF 2 vs. RF 3). Nested cross-validation was applied for model development and hyperparameter tuning, and bootstrapping was used to assess performance robustness. Model performance was evaluated with the area under the curve (AUC), calibration curves, decision curve analysis (DCA), and SHapley Additive exPlanations (SHAP) visualization.
Results
Compared with single-modality models, the multimodal deep learning model (DL-combine), which is based exclusively on native T1 mapping, ADC, and T2* mapping, demonstrated favourable and stable performance (test AUC: 0.930). Among the 14 classifiers, XGBoost performed numerically better in terms of RF presence (mean AUCs: 0.986, 0.887; accuracy: 0.947, 0.829), whereas ExtraTree performed better in terms of RF severity assessment (mean AUCs: 0.935, 0.883; accuracy: 0.886, 0.848). Calibration curves and DCA confirmed robust predictive reliability and clinical utility. SHAP analysis highlighted the relative contributions of the DL-sign and eGFR.
Conclusion
This two-stage multimodal MRI-based framework provides accurate and interpretable assessment of RF across different stages of CKD, supporting noninvasive risk stratification and complementary clinical decision-making.
Supplementary Information
The online version contains supplementary material available at 10.1186/s12882-026-04869-2.
Keywords: Multimodal magnetic resonance imaging, Deep learning, Renal fibrosis, Machine learning, Chronic kidney disease
Background
Chronic kidney disease (CKD) has emerged as a global health care problem, with an estimated 700 million people living with any stage of CKD [1, 2]. Renal fibrosis (RF), a pathological hallmark of CKD progression, is often irreversible and closely associated with renal prognosis [3]. While renal biopsy remains the gold standard for the diagnosis of RF, its invasiveness limits its clinical utility [4].
Functional magnetic resonance imaging (MRI) techniques provide noninvasive assessments of CKD and RF [5, 6]. Diffusion-weighted imaging (DWI) measures the mobility of water molecules through apparent diffusion coefficient (ADC) values, with more severe RF correlating with restricted diffusion and lower ADC [7]. T1 mapping, particularly non-contrast native T1 mapping, detects increasing cortical T1 values and decreasing differences in the corticomedullary region as fibrosis progresses [8, 9]. Similarly, T2* mapping reflects renal tissue hypoxia through R2* values, showing elevated levels in cortical and medullary regions with worsening kidney damage [10].
Conventional medical imaging is associated with operator bias and limited feature extraction. Radiomics improves objectivity by quantifying subtle features but still relies on manual engineering [11]. Deep learning (DL) overcomes this by automatically extracting complex patterns from raw images [12] and has shown great promise in tumour diagnosis and treatment [13]. However, its application in CKD is still in the early stages. Recent studies suggest that DL models based on single-modality MRI can identify CKD patients [14]. However, these models overlook the complementary pathological insights provided by sequences such as T1 mapping, ADC, and T2* mapping, which reflect distinct fibrotic mechanisms. [15]. Multimodal MRI fusion could address these limitations, yet no studies have explored combining multimodal MRI with DL for renal fibrosis assessment in patients with CKD.
In this study, we propose a two-stage, clinically oriented framework for noninvasive RF assessment in patients with CKD, with the aims of evaluating the utility of multimodal MRI-based deep learning for capturing fibrosis-related imaging information, examining the added value of multimodal integration over single-sequence models, and exploring whether a deep learning-derived imaging signature can provide a stable and interpretable representation for subsequent machine learning-based fibrosis assessment. Multimodal MRI-based deep learning models were developed using native T1 mapping, ADC, and T2* mapping to capture complementary structural and microstructural features, followed by integration of the derived DL signs with routine clinical indicators. This integrative two-stage framework is intended to support noninvasive fibrosis assessment and risk stratification across different stages of CKD, particularly in clinical settings where biopsy is constrained by eligibility or feasibility concerns.
Methods
Subjects
This single-centre, prospective study was approved by the ethics committee of our institution (JD-LK-2022–060-01). All patients underwent MRI voluntarily and signed informed consent. A prospective analysis of 216 CKD patients between September 2021 and November 2024 revealed that these patients were scheduled for renal biopsy and agreed to undergo renal MRI. The inclusion criteria included (1) meeting the clinical diagnostic criteria for CKD [16], and (2) performing multimodal MRI of the kidney within 3 days prior to scheduled renal biopsy. However, 64 patients were excluded because (1) 3 patients had claustrophobia during the MRI (2) 26 patients could not complete all the sequence scans and the MRI images were incomplete (3) 17 patients with severe image artifacts and poor image quality could not perform the post-processing, and (4) 18 patients ultimately did not consent to renal biopsy. Finally, a total of 152 CKD patients were enrolled. The inclusion and exclusion criteria are shown in Fig. 1.
Fig. 1.
Flowchart for the inclusion and exclusion of patients
Multimodal MRI examination and preprocessing
MRI of both kidneys was performed using a Siemens Prisma 3.0T magnetic resonance scanner with an 18-channel body coil. Multimodal MRI sequences included native T1 mapping, DWI and T2* mapping imaging. MRI images of both kidneys were acquired in the coronal plane. The image acquisition details are summarized in Supplementary Table 1.
ROIs of the right renal cortex were manually delineated by a radiologist with 7 years of genitourinary experience on coronal native T1 mapping images. The right renal cortex was selected to ensure anatomical consistency with the biopsy site. T2* and ADC volumes were aligned to the T1 mapping space using an advanced normalization tools (ANTs)-based multistep registration strategy. Additional preprocessing included denoising, spatial resampling, and field-of-view cropping. A rigid-affine registration served as the primary alignment approach, and a lightweight symmetric normalization (SyN) refinement was selectively applied when minor residual mismatches remained after affine alignment. All the registered images underwent manual quality control to ensure accurate multimodal correspondence. [17].
Clinical data
Clinical data such as age, sex, height, weight, body mass index (BMI), blood pressure, blood glucose, serum creatinine (Scr), blood urea nitrogen (BUN) and 24-hour urinary protein (24h-UPRO) were collected. The estimated glomerular filtration rate (eGFR) was calculated according to the Chronic Kidney Disease Epidemiology Collaboration (CKD-EPI) formula [18].
Data splitting and augmentation
The dataset was randomly split into training and test sets at a 2:1 ratio. In each round, the model was trained on two folds and evaluated on the remaining fold, which served as the test data for that specific iteration. The performance metrics were calculated for each round and subsequently averaged across the three rounds. To address class imbalance, negative samples in the training set were duplicated, and weighted loss was applied. To augment the training dataset, random flipping and rotation were used.
Performance-sample size analysis
To evaluate whether the available sample size was sufficient to support stable model performance, a performance-sample size curve analysis was conducted. We incrementally increased the training sample size for two research objectives: “RF presence” and “RF severity,” and calculated the area under the receiver operating characteristic curve (AUC) on a fixed, independent test set.
Deep learning modelling
MobileNetV2-SE, a lightweight convolutional network with a squeeze-and-excitation (SE) module, was used to enhance feature extraction efficiency [19]. Multimodal input channels were concatenated and passed through a multihead self-attention module, followed by a multilayer perceptron. Binary cross-entropy loss was optimized using the Adam optimizer (mini-batch size = 8, learning rate = 0.001), with batch normalization, dropout, and L2 regularization to mitigate overfitting. A softmax classifier generated probabilistic outputs.
Category probabilities were further integrated into a composite DL-sign score using the following formula:
[\mathrm{DL\text{-}sign}=\sum_{i = 0}^{n-1} v_i\cdot p_i]
![]() |
The composite DL-sign score is a scalar value. It is mathematically defined as the expected value of the categorical outcome and is calculated by summing the products of each category’s assigned value (vi) and its corresponding predicted probability (pi).(n: The total number of categories. Indexing ranges from (0) to (n-1). i: The category index (the i-th category). pi: The predicted probability of a sample belonging to the i-th category (from the model’s final softmax output). Value range: ([0,1]) Constraint:
)
To ensure the reliability of the composite DL-sign score, probability calibration was applied to the model outputs prior to DL-sign construction. Temperature scaling was performed on the logits of the model output on the test set (learning only a single temperature parameter T, estimated by minimizing the NLL/cross-entropy objective function), and
was used to generate calibrated multiclass probabilities for calculating the DL-sign in the testing phase. The test set was not involved in fitting any calibration parameters to avoid information leakage. To evaluate the calibration effect and the stability of the DL-sign, a supplementary analysis was conducted: DL-signs derived twice from the same batch of samples were compared, and their consistency was quantified using Spearman’s rank correlation coefficient.
An ablation analysis was performed by systematically removing or replacing selected architectural components, including the SE module and transformer-based fusion, while keeping all other experimental conditions unchanged. MM-MNetV2-MLP was consisted of a MobileNetV2 backbone (without SE), simple fusion, and an MLP head; MM-MNetV2SE-MLP incorporated an SE module while keeping all the other components unchanged; and MM-MNetV2-Trans replaced the simple fusion with a transformer fusion module but without using the SE module. MobileNetV2-SE represents the complete proposed framework and was used as the primary model for subsequent analyses, including DL-sign derivation. Macro-averaged performance metrics were calculated by averaging the one-vs-rest performance of each class.
Machine learning modelling
Logistic regression revealed significant clinical indicators (p < 0.05), which along with the DL-sign, formed a model to assess RF presence and severity. Fourteen machine learning models were used, including logistic regression (LR), naive Bayes (NB), SVM variants, decision tree (DT), random forest, ExtraTree, XGBoost, AdaBoost, multi-layer perceptron (MLP), gradient boosting machines (GBM), and light GBM. The full workflow is summarized in Fig. 2.
Fig. 2.
The full workflow. MobileNetV2-SE-based DL models extract DL-sign features from multimodal MRI images (native T1 mapping, ADC and T2* mapping). The DL-sign combined with selected clinical indicators (eGFR, Scr, BUN) were input into 14 machine learning classifiers for evaluation of the presence and severity of RF in patients with CKD
Model stability was evaluated using three-fold cross-validation. The entire dataset was randomly partitioned into three folds of comparable size and distribution. In each round, two folds were used to train the model, and the remaining single fold served exclusively as the held-out set for testing within that round. This process was repeated three times so that each fold was used as the test set exactly once. The performance metrics obtained from the three rounds were averaged to produce the final aggregated estimate. Multimodal MRI samples with different RF stages are shown in Figs. 3A–D.
Fig. 3.
Multimodal MRI images (native T1 mapping, ADC and T2* mapping) of different degrees of renal fibrosis (RF) in CKD patients, and using the shapley additive explanations (SHAP) to visualize the interaction contribution of features in the optimal model. (A) RF 1 patients (no fibrosis); (B) RF 2 patients(mild fibrosis); (C) RF 3 Patients(moderate fibrosis); (D) RF 3 patients (severe fibrosis); (E, F) SHAP summary plots show that DL-sign contributes the highest SHAP value to XGBoost and ExtraTree models, followed by eGFR
To further assess the robustness of the main findings and to reduce potential model selection bias, we additionally performed nested cross-validation (nested CV) for representative models corresponding to the two primary clinical tasks (RF presence and RF severity). Specifically, a three-fold outer cross-validation loop was used exclusively for unbiased performance estimation. Within each outer training fold, a three-fold inner cross-validation loop was applied for hyperparameter tuning. Model optimization was strictly confined to the inner loop, whereas the outer loop served solely for independent performance evaluation, thereby reducing the risk of information leakage and optimism bias. Model interpretability was enhanced using SHAP (SHapley Additive exPlanations), highlighting how the DL-sign interacts with clinical biomarkers (eGFR, Scr, BUN) [20].
Detailed pseudo code is provided via a link in the Supplementary Materials.
Renal histopathology
An ultrasound-guided renal biopsy was conducted within 3 days after the MRI by an experienced nephrologist. Patients were positioned prone with a sandbag under the abdomen to minimize kidney movement. Typically, the lower pole of the right kidney is the preferred puncture site. [21]. Following standard histopathological procedures, the degree of fibrosis in kidney biopsy specimens was assessed by Masson’s staining. According to the degree of fibrosis, the RF group was divided into a no RF group (referred as “RF 1”, no fibrosis), a mild RF group (“RF 2”, fibrosis proportion ≤ 25%), and a moderate to severe RF group (“RF 3”, fibrosis proportion > 25%) [22].
Statistical analysis
Statistical analyses were conducted with SPSS 25.0 (IBM Corp.) and MedCalc 15.2.2 (MedCalc Software Ltd.). The MobileNetV2-SE network was developed in Python 3.8 using PyTorch 2.0. Continuous variables with a normal distribution were analysed using ANOVA and are reported as the mean ± standard deviation. Non-normal variables are presented as median (interquartile range) and were analysed with the Kruskal-Wallis test. Categorical variables were compared using the chi-squared test. Model discrimination performance was evaluated primarily using the AUC. A bootstrap-based approach was applied to estimate AUC differences and corresponding confidence intervals in the independent test set. Model calibration and clinical utility were assessed using calibration curves and decision curve analysis (DCA), respectively. Threshold-dependent performance metrics, including sensitivity, specificity, accuracy, positive predictive value (PPV), negative predictive value (NPV), and F1 score, were also reported to provide a comprehensive evaluation beyond the AUC alone. All tests were two-tailed with a significance threshold of p < 0.05.
To evaluate sample size sufficiency, a performance-sample size curve analysis was performed.
Results
Patient baseline characteristics
According to the degree of fibrosis, the 152 patients with CKD included in this study were divided into the RF 1 group (no fibrosis, n = 34), RF 2 group (mild fibrosis, n = 69), and RF 3 group (moderate to severe fibrosis, n = 49), as shown in Table 1.
Table 1.
Basic characteristics
| RF 1 (n = 34) |
RF 2 (n = 69) |
RF 3 (n = 49) |
Statistics | P | |
|---|---|---|---|---|---|
| Age (years) | 44 ± 13 | 49 ± 15 | 47 ± 13 | 1.430 | 0.243 |
| Gender (male, %) | 14(41.2%) | 33(47.8%) | 30(61.2%) | 3.633 | 0.163 |
| Height (m) | 1.64 ± 0.09 | 1.61(1.59, 1.70) | 1.67 ± 0.08 | 4.286 | 0.117 |
| BMI (kg/m2) | 24.97 ± 3.18 | 24.26 ± 3.69 | 24.53 ± 3.34 | 0.475 | 0.623 |
|
Blood pressure (hypertension, %) |
21(61.8%) | 40(58.0%) | 32(65.3%) | 0.655 | 0.721 |
| Blood glucose (mmol/L) | 4.71(4.40, 5.06) | 4.72(4.46, 5.31) | 5.00(4.47, 5.47) | 3.881 | 0.144 |
| eGFR (mL/min/1.73 m2) | 109.39 ± 19.96 | 90.20 ± 28.26 | 59.23 ± 22.19 | 44.662 | <0.001* |
| Scr (μmol/L) | 61(48, 72.5) | 74(56, 96) | 110(100, 139) | 58.265 | <0.001* |
| BUN (mmol/L) | 4.80(3.50, 5.95) | 5.30(4.10, 6.70) | 6.90(6.00,9.40) | 33.864 | <0.001* |
| 24 h-UPRO (g/24 h) | 2.92(1.37, 5.84) | 2.81(1.51, 5.24) | 2.00(0.96,4.02) | 3.006 | 0.222 |
RF, renal fibrosis; RF 1, no RF (0% fibrosis); RF 2, mild RF (≤25% fibrosis); RF 3, moderate to severe RF (>25% fibrosis); CKD, chronic kidney disease; BMI, body mass index; eGFR, estimated glomerular filtration rate; Scr, serum creatinine; BUN, blood urea nitrogen; 24 h-UPRO, 24-hour urinary protein; Hypertension was defined as systolic/diastolic blood pressure ≥ 140/90 mmHg; values are mean with standard deviation or median with lower and upper quartile in parentheses or number with percentage in parentheses; *statistically significant
Patient age, sex, height, weight, BMI, blood pressure, blood glucose, and 24 h-UPRO did not significantly differ among the three RF groups (p > 0.05). The eGFR, Scr and BUN values of CKD patients were significantly different among the three RF groups (p < 0.05). With the worsening of fibrosis, the eGFR levels gradually decreased, while the SCr and BUN levels gradually increased.
Ablation analysis
The contributions of individual components were evaluated through ablation experiments (Table 2).
Table 2.
Ablation analysis of model variants for RF classification (macro-averaged metrics on the test set)
| Model | Fusion | Val AUC (macro) | Val ACC (macro) | Val F1 (macro) | Notes |
|---|---|---|---|---|---|
| MobileNetV2-SE | Full fusion | 0.880 | 0.843 | 0.763 | Best overall, used for downstream signature |
| MM-MNetV2-MLP | Concat | 0.674 | 0.660 | 0.597 | No SE, no Transformer |
| MM-MNetV2-Trans | Transformer | 0.593 | 0.608 | 0.555 | Transformer fusion only (no SE) |
| MM-MNetV2SE-MLP | Concat | 0.637 | 0.647 | 0.501 | SE only (no Transformer) |
RF, renal fibrosis; SE: Squeeze-and-Excitation; MLP: Multi-Layer Perceptron; MM-MNetV2-Trans, MobileNetV2 (without SE) + Transformer fusion; MM-MNetV2SE-MLP, MobileNetV2 + SE + simple fusion + MLP head (no Transformer); MM-MNetV2-MLP, MobileNetV2 (without SE) + simple fusion + MLP head (no Transformer); MobileNetV2-SE (Full model), complete model used as the main baseline; AUC, area under the curve; ACC, accuracy
Compared with the MM-MNetV2-MLP model, incorporating the SE module (MM-MNetV2SE-MLP) did not improve the generalization performance and resulted in reduced macro AUC (0.637 vs. 0.674), macro accuracy (0.647 vs. 0.660), and macro F1 score (0.501 vs. 0.597). Similarly, replacing the simple fusion with transformer fusion (MM-MNetV2-Trans) led to a further decrease in the macro AUC (0.593 vs. 0.674), macro accuracy (0.608 vs. 0.660), and macro F1 score (0.555 vs. 0.597). Overall, the complete framework (MobileNetV2-SE) achieved the best performance among all ablation settings. Removal or modification of individual components consistently degraded performance; therefore, MobileNetV2-SE was selected to derive the DL-sign for subsequent machine learning models. Detailed performance metrics are provided in Supplementary Table 6.
Optimization of multimodal MRI-based deep learning models
The native T1 mapping, ADC, and T2* mapping images were input into the MobileNetV2-SE architecture to construct four DL models (DL-native T1 mapping, DL-ADC, DL-T2* mapping and DL-combine), which can be used to discriminate RF 1, RF 2, and RF 3 in CKD patients. Comparisons of the RF diagnostic performance of these MRI-based DL models are shown in the training and test cohorts (Table 3).
Table 3.
Comparisons of diagnostic performance of deep learning models based on multimodal MRI sequences for renal fibrosis
| AUC (95%CI) | SEN | SPE | ACC | PPV | NPV | F1-score | |
|---|---|---|---|---|---|---|---|
| DL-combine | |||||||
| RF 1 | |||||||
| Training | 0.934 (0.875,0.977) | 0.826 | 0.910 | 0.891 | 0.731 | 0.947 | 0.776 |
| Test | 0.930 (0.837,0.988) | 0.818 | 0.900 | 0.882 | 0.692 | 0.947 | 0.750 |
| RF 2 | |||||||
| Training | 0.883 (0.822,0.948) | 0.932 | 0.789 | 0.851 | 0.774 | 0.938 | 0.845 |
| Test | 0.827 (0.700,0.910) | 0.720 | 0.846 | 0.784 | 0.818 | 0.759 | 0.766 |
| RF 3 | |||||||
| Training | 0.925 (0.879,0.966) | 0.765 | 0.925 | 0.871 | 0.839 | 0.886 | 0.800 |
| Test | 0.882 (0.768,0.963) | 0.800 | 0.889 | 0.863 | 0.750 | 0.914 | 0.774 |
| DL-native T1mapping | |||||||
| RF 1 | |||||||
| Training | 0.902 (0.840,0.950) | 0.913 | 0.756 | 0.792 | 0.525 | 0.967 | 0.667 |
| Test | 0.777 (0.643,0.907) | 0.909 | 0.650 | 0.706 | 0.417 | 0.963 | 0.571 |
| RF 2 | |||||||
| Training | 0.766 (0.670,0.848) | 0.864 | 0.579 | 0.703 | 0.613 | 0.846 | 0.717 |
| Test | 0.588 (0.449,0.745) | 0.760 | 0.462 | 0.608 | 0.576 | 0.667 | 0.655 |
| RF 3 | |||||||
| Training | 0.873 (0.798,0.938) | 0.882 | 0.716 | 0.772 | 0.612 | 0.923 | 0.723 |
| Test | 0.657 (0.502,0.820) | 0.733 | 0.639 | 0.667 | 0.458 | 0.852 | 0.564 |
| DL-ADC | |||||||
| RF 1 | |||||||
| Training | 0.941 (0.897,0.974) | 0.783 | 0.936 | 0.901 | 0.783 | 0.936 | 0.783 |
| Test | 0.727 (0.547,0.862) | 0.727 | 0.675 | 0.686 | 0.381 | 0.900 | 0.500 |
| RF 2 | |||||||
| Training | 0.572 (0.439,0.697) | 0.909 | 0.421 | 0.634 | 0.548 | 0.857 | 0.684 |
| Test | 0.512 (0.362,0.685) | 0.960 | 0.154 | 0.549 | 0.522 | 0.800 | 0.676 |
| RF 3 | |||||||
| Training | 0.847 (0.743,0.918) | 1.000 | 0.701 | 0.802 | 0.630 | 1.000 | 0.773 |
| Test | 0.541 (0.408,0.741) | 0.667 | 0.556 | 0.588 | 0.385 | 0.800 | 0.488 |
| DL-T2*mapping | |||||||
| RF 1 | |||||||
| Training | 0.871 (0.785,0.927) | 0.913 | 0.718 | 0.762 | 0.488 | 0.966 | 0.636 |
| Test | 0.723 (0.574,0.855) | 0.909 | 0.500 | 0.588 | 0.333 | 0.952 | 0.488 |
| RF 2 | |||||||
| Training | 0.741 (0.630,0.842) | 0.682 | 0.737 | 0.713 | 0.667 | 0.750 | 0.674 |
| Test | 0.597 (0.455,0.738) | 0.640 | 0.577 | 0.608 | 0.593 | 0.625 | 0.615 |
| RF 3 | |||||||
| Training | 0.803 (0.703,0.887) | 0.765 | 0.821 | 0.802 | 0.684 | 0.873 | 0.722 |
| Test | 0.648 (0.495,0.815) | 0.600 | 0.778 | 0.725 | 0.529 | 0.824 | 0.562 |
RF, renal fibrosis; DL-combine, deep learning model based on the combination of native T1 mapping, ADC and T2* mapping images; DL-native T1 mapping, deep learning model based on native T1 mapping images; DL-ADC, deep learning model based on ADC images; DL- T2* mapping, deep learning model based on T2* mapping images; RF 1, no RF (0% fibrosis); RF 2, mild RF (≤25% fibrosis); RF 3, moderate to severe RF (>25% fibrosis); AUC, area under the curve; SEN, sensitivity; SPE, specificity; ACC, accuracy; PPV, positive predictive value; NPV, negative predictive value; CI, confidence interval
When the three RF groups were compared, the AUCs (0.930, 0.827, and 0.882), accuracy (0.882, 0.784, and 0.863) and F1 score (0.750, 0.766, and 0.774) were greater in the DL-combine group than in the other groups in the test cohort. Additionally, the overall performance of DL-combine was consistently favourable in the training cohort. Thus, compared with the three single models, the combined DL model demonstrated improved discriminative ability for differentiating RF in patients with CKD. The weighted probability of the DL-combine model (DL-sign) was obtained. Spearman’s rank correlation analysis demonstrated a strong monotonic association between the DL-sign values before and after probability calibration (ρ=0.963, p = 5.3 × 10-8 7), indicating the high stability of the DL-sign rankings across the calibration procedures (Supplementary Figure 1).
RF evaluation of models integrating deep learning with clinical indicators
The performance-sample size curve analysis demonstrated that, at the current sample size, the generalization performance of both models essentially converged (Supplementary Figures 2 and 3).
In terms of identifying the presence of RF (RF 1 vs. RF 2 and RF 3), compared with the other machine learning algorithms, the XGBoost model demonstrated a numerically favourable overall diagnostic performance, with a mean AUC and a mean ACC of 0.986 and 0.947,respectively, in the training cohort and 0.887 and 0.829,respectively, in the test cohort (Table 4, Supplementary Table 2).
Table 4.
Cross-validation results of the optimal machine learning classification model (XGBoost) for identifying the presence of renal fibrosis (RF 1 vs. RF 2 and RF 3)
| Fold number | Training cohort | Test cohort | ||
|---|---|---|---|---|
| AUC | ACC | AUC | ACC | |
| Fold 1 | 0.980 | 0.931 | 0.911 | 0.843 |
| Fold 2 | 0.991 | 0.950 | 0.895 | 0.765 |
| Fold 3 | 0.987 | 0.961 | 0.855 | 0.880 |
| Mean | 0.986 | 0.947 | 0.887 | 0.829 |
AUC, area under the curve; ACC, accuracy
The confusion matrix revealed that most of the no RF and RF cases were correctly classified, with only a small number of false-negative and false positive predictions (Figs. 4A–C)
Fig. 4.
Confusion matrices for renal fibrosis classification. (A-C) XGBoost model for distinguishing non-fibrosis from fibrosis; (D-F) ExtraTree model for distinguishing mild fibrosis from moderate to severe fibrosis
For identifying RF severity (RF 2 vs. RF 3), the Extratree model showed better overall diagnostic performance, with a mean AUC and a mean ACC of 0.935 and 0.886, respectively, in the training cohort and 0.883 and 0.848, respectively, in the test cohort (Table 5, Supplementary Table 3).
Table 5.
Cross-validation results of the optimal machine learning classification model (ExtraTree) for identifying the severity of renal fibrosis (RF 2 vs. RF 3)
| Fold number | Training cohort | Test cohort | ||
|---|---|---|---|---|
| AUC | ACC | AUC | ACC | |
| Fold 1 | 0.941 | 0.897 | 0.852 | 0.800 |
| Fold 2 | 0.926 | 0.861 | 0.918 | 0.872 |
| Fold 3 | 0.939 | 0.899 | 0.880 | 0.872 |
| Mean | 0.935 | 0.886 | 0.883 | 0.848 |
ACC, accuracy; AUC, area under the curve
Within the nested CV framework, the XGBoost model showed relatively stable performance across different outer test folds for RF presence classification. The AUC values in the outer testing sets were distributed mainly between 0.85 and 0.95, with consistent accuracy and F1-scores across folds (Table 6).
Table 6.
Nested cross-validation results of the optimal machine learning classification model (XGBoost) for identifying the presence of renal fibrosis (RF 1 vs. RF 2 and RF 3)
| Source | Model | ACC | AUC | SEN | SPE | NPV | PPV | F1 |
|---|---|---|---|---|---|---|---|---|
| inner-train-fold0 | XGBoost | 0.896 | 0.958 | 0.887 | 0.929 | 0.684 | 0.979 | 0.931 |
| inner-test-fold0 | XGBoost | 0.765 | 0.856 | 0.680 | 1 | 0.529 | 1 | 0.810 |
| outer-test-fold1 | XGBoost | 0.922 | 0.951 | 0.950 | 0.818 | 0.818 | 0.950 | 0.950 |
| inner-train-fold1 | XGBoost | 0.851 | 0.910 | 0.808 | 1 | 0.600 | 1 | 0.894 |
| inner-test-fold1 | XGBoost | 0.882 | 0.894 | 0.923 | 0.750 | 0.750 | 0.923 | 0.923 |
| outer-test-fold0 | XGBoost | 0.902 | 0.918 | 0.950 | 0.727 | 0.800 | 0.927 | 0.938 |
| inner-train-fold2 | XGBoost | 0.779 | 0.912 | 0.706 | 1 | 0.531 | 1 | 0.828 |
| inner-test-fold2 | XGBoost | 0.909 | 0.929 | 0.926 | 0.833 | 0.714 | 0.962 | 0.943 |
| outer-test-fold0 | XGBoost | 0.902 | 0.924 | 0.925 | 0.818 | 0.750 | 0.949 | 0.937 |
| inner-train-fold0 | XGBoost | 0.851 | 0.952 | 0.811 | 1 | 0.583 | 1 | 0.896 |
| inner-test-fold0 | XGBoost | 0.912 | 0.926 | 0.929 | 0.833 | 0.714 | 0.963 | 0.945 |
| outer-test-fold0 | XGBoost | 0.784 | 0.899 | 0.730 | 0.929 | 0.565 | 0.964 | 0.831 |
| inner-train-fold1 | XGBoost | 0.762 | 0.925 | 0.714 | 1 | 0.407 | 1 | 0.833 |
| inner-test-fold1 | XGBoost | 0.912 | 0.953 | 0.92 | 0.889 | 0.800 | 0.958 | 0.939 |
| outer-test-fold1 | XGBoost | 0.725 | 0.867 | 0.622 | 1 | 0.500 | 1 | 0.767 |
| inner-train-fold2 | XGBoost | 0.912 | 0.942 | 0.925 | 0.867 | 0.765 | 0.961 | 0.942 |
| inner-test-fold2 | XGBoost | 0.606 | 0.811 | 0.536 | 1 | 0.278 | 1 | 0.698 |
| outer-test-fold1 | XGBoost | 0.824 | 0.864 | 0.838 | 0.786 | 0.647 | 0.912 | 0.873 |
| inner-train-fold0 | XGBoost | 0.912 | 0.938 | 0.925 | 0.867 | 0.765 | 0.961 | 0.942 |
| inner-test-fold0 | XGBoost | 0.735 | 0.824 | 0.625 | 1 | 0.527 | 1 | 0.769 |
| outer-test-fold2 | XGBoost | 0.720 | 0.854 | 0.659 | 1 | 0.391 | 1 | 0.794 |
| inner-train-fold1 | XGBoost | 0.897 | 0.964 | 0.863 | 1 | 0.708 | 1 | 0.926 |
| inner-test-fold1 | XGBoost | 0.735 | 0.815 | 0.731 | 0.750 | 0.462 | 0.905 | 0.809 |
| outer-test-fold2 | XGBoost | 0.860 | 0.923 | 0.854 | 0.889 | 0.571 | 0.972 | 0.909 |
| inner-train-fold2 | XGBoost | 0.779 | 0.887 | 0.740 | 0.889 | 0.552 | 0.949 | 0.831 |
| inner-test-fold2 | XGBoost | 0.971 | 0.989 | 0.963 | 1 | 0.875 | 1 | 0.981 |
| outer-test-fold2 | XGBoost | 0.800 | 0.940 | 0.756 | 1 | 0.474 | 1 | 0.861 |
AUC, area under the curve; ACC, accuracy; SEN, sensitivity; SPE, specificity; NPV, negative predictive value; PPV, positive predictive value
For the task of renal fibrosis severity stratification, the ExtraTrees model demonstrated stable classification performance across repeated conventional cross-validation. Across different test folds, the model achieved a moderate-to-high AUC and balanced sensitivity and specificity (Table 7).
Table 7.
Nested cross-validation results of the optimal machine learning classification model (ExtraTree) for identifying the severity of renal fibrosis (RF 2 vs. RF 3)
| Source | Model | ACC | AUC | Sensitivity | Specificity | NPV | PPV | F1 |
|---|---|---|---|---|---|---|---|---|
| inner-train-fold0 | ExtraTree | 0.865 | 0.912 | 0.889 | 0.853 | 0.935 | 0.762 | 0.821 |
| inner-test-fold0 | ExtraTree | 0.885 | 0.950 | 1 | 0.813 | 1 | 0.769 | 0.870 |
| outer-test-fold1 | ExtraTree | 0.800 | 0.886 | 0.952 | 0.632 | 0.923 | 0.741 | 0.833 |
| inner-train-fold1 | ExtraTree | 0.942 | 0.981 | 0.889 | 0.971 | 0.943 | 0.941 | 0.914 |
| inner-test-fold1 | ExtraTree | 0.846 | 0.853 | 0.8 | 0.875 | 0.875 | 0.800 | 0.800 |
| outer-test-fold0 | ExtraTree | 0.775 | 0.858 | 0.810 | 0.737 | 0.778 | 0.773 | 0.791 |
| inner-train-fold2 | ExtraTree | 0.942 | 0.958 | 0.850 | 1 | 0.914 | 1 | 0.919 |
| inner-test-fold2 | ExtraTree | 0.692 | 0.781 | 1 | 0.556 | 1 | 0.5 | 0.667 |
| outer-test-fold0 | ExtraTree | 0.800 | 0.853 | 0.857 | 0.737 | 0.824 | 0.783 | 0.818 |
| inner-train-fold0 | ExtraTree | 0.904 | 0.944 | 0.958 | 0.857 | 0.960 | 0.852 | 0.902 |
| inner-test-fold0 | ExtraTree | 0.815 | 0.843 | 0.692 | 0.929 | 0.765 | 0.900 | 0.783 |
| outer-test-fold0 | ExtraTree | 0.872 | 0.951 | 1 | 0.815 | 1 | 0.706 | 0.828 |
| inner-train-fold1 | ExtraTree | 0.868 | 0.924 | 0.783 | 0.933 | 0.848 | 0.900 | 0.837 |
| inner-test-fold1 | ExtraTree | 0.769 | 0.821 | 0.857 | 0.667 | 0.800 | 0.750 | 0.800 |
| outer-test-fold1 | ExtraTree | 0.821 | 0.923 | 0.917 | 0.778 | 0.954 | 0.647 | 0.759 |
| inner-train-fold2 | ExtraTree | 0.887 | 0.945 | 0.889 | 0.885 | 0.885 | 0.889 | 0.889 |
| inner-test-fold2 | ExtraTree | 0.885 | 0.872 | 0.900 | 0.875 | 0.933 | 0.818 | 0.857 |
| outer-test-fold1 | ExtraTree | 0.897 | 0.923 | 0.917 | 0.889 | 0.96 | 0.786 | 0.846 |
| inner-train-fold0 | ExtraTree | 0.942 | 0.980 | 0.947 | 0.940 | 0.969 | 0.900 | 0.923 |
| inner-test-fold0 | ExtraTree | 0.778 | 0.838 | 0.643 | 0.923 | 0.706 | 0.900 | 0.750 |
| outer-test-fold2 | ExtraTree | 0.821 | 0.856 | 0.8125 | 0.826 | 0.864 | 0.765 | 0.788 |
| inner-train-fold1 | ExtraTree | 0.868 | 0.95 | 0.913 | 0.833 | 0.926 | 0.808 | 0.857 |
| inner-test-fold1 | ExtraTree | 0.808 | 0.875 | 0.9 | 0.750 | 0.923 | 0.692 | 0.783 |
| outer-test-fold2 | ExtraTree | 0.846 | 0.871 | 0.8125 | 0.870 | 0.870 | 0.813 | 0.813 |
| inner-train-fold2 | ExtraTree | 0.868 | 0.933 | 0.917 | 0.828 | 0.923 | 0.815 | 0.863 |
| inner-test-fold2 | ExtraTree | 0.923 | 0.974 | 1 | 0.882 | 1 | 0.818 | 0.900 |
| outer-test-fold2 | ExtraTree | 0.821 | 0.889 | 0.813 | 0.826 | 0.864 | 0.765 | 0.788 |
AUC, area under the curve; ACC, accuracy; SEN, sensitivity; SPE, specificity; NPV, negative predictive value;PPV, positive predictive value
The nested CV results yielded performance estimates that were generally consistent with those obtained from conventional cross-validation, supporting the robustness of the main findings under a more rigorous evaluation framework. The corresponding confusion matrix indicated that both mild and moderate to severe fibrosis cases were predominantly correctly identified, with limited misclassifications (Figs. 4D–F ).
As shown in the SHAP summary plot, the DL-sign emerged as the most influential feature, contributing the highest relative SHAP values to the XGBoost and ExtraTree models, followed by the eGFR (Figs. 3E, 3F).
ROC curves revealed that the XGBoost and ExtraTree models demonstrated strong diagnostic performance for RF evaluation, with AUC values greater than 0.85 for the test cohort (Fig. 5).
Fig. 5.
Receiver operating characteristic curves of the optimal machine learning model for assessment of renal fibrosis (RF) in training and test cohorts. (A-C) XGBoost model for identifying the presence of RF (RF 1 vs. RF 2 and RF 3); (D-F) ExtraTree model for identifying the severity of RF (RF 2 vs. RF 3)
The calibration curves revealed that both models were generally close to the ideal line in the training and test cohorts but with some biases in the local range (Fig. 6).
Fig. 6.
Calibration curves of the optimal machine learning classification model for assessment of renal fibrosis (RF) in training and test cohorts. (A-C) XGBoost model for identifying the presence of RF (RF 1 vs. RF 2 and RF 3); (D-F) ExtraTree model for identifying the severity of RF (RF 2 vs. RF 3)
DCA indicated that both models had net benefits in both the training and test cohorts (Fig. 7).
Fig. 7.
Decision curves of the optimal machine learning classification model for assessment of renal fibrosis (RF) in training and test cohorts. (A-C) XGBoost model for identifying the presence of RF (RF 1 vs. RF 2 and RF 3); (D-F) ExtraTree model for identifying severity of RF (RF 2 vs. RF 3)
Bootstrap resampling on the test set revealed that, for most model comparisons, the estimated AUC differences were small, with 95% confidence intervals overlapping zero, whereas a limited number of classifiers clearly exhibited lower discriminative performance. Detailed bootstrap results are provided in Supplementary Tables 4–5.
On the basis of the overall performance observed across these analyses, XGBoost and ExtraTrees were used for the subsequent evaluation of RF presence and RF severity, respectively.
Discussion
This study introduces a novel two-stage deep learning approach using multimodal MRI (native T1 mapping, ADC, and T2* mapping) for differentiating RF in patients with CKD, offering advantages over previous single-modality methods. The DL-combine model had the highest AUCs in both the training and test groups, outperforming the single-modality models in terms of the diagnosis of RF. To our knowledge, multimodal MRI-based DL frameworks for RF assessment in CKD have not been extensively investigated, as previous research focused primarily on MRI texture-based ML methods [23, 24]. Our previous studies demonstrated that native T1 mapping-based radiomic models accurately assess renal function and fibrosis in patients with CKD [17]. However, single-modality imaging only partially captures fibrosis features and can be affected by confounding factors such as inflammation and oedema, limiting its specificity. In this study, native T1 mapping was used to assess fibrosis and extracellular matrix changes, the ADC showed microstructural disruption from restricted water diffusion, and T2* mapping indicated hypoxia-induced iron deposition. Together, these methods provide a comprehensive view of fibrotic mechanisms.
We employed the MobileNetV2-SE architecture, which combines the lightweight MobileNetV2 network with a squeeze-and-excitation (SE) attention mechanism, owing to its favourable balance between computational efficiency and representational capacity. The SE module is designed to recalibrate channelwise feature responses and has been shown to be effective at enhancing discriminative features in various medical imaging tasks [19]. Lightweight and transfer learning-based convolutional architectures have demonstrated robust performance across diverse medical image classification problems, including mammographic breast cancer identification and chest radiograph analysis [25, 26]. These studies highlight the practical advantages of resource-efficient CNN frameworks in real-world clinical settings. However, our ablation analysis indicates that, under the current data setting, simply introducing SE modules or increasing architectural complexity does not necessarily translate into improved generalization performance. Instead, the observed performance gains are primarily attributable to the overall design of the proposed framework, including multimodal feature integration and the two-stage modelling strategy. In contrast to segmentation-focused models such as U-Net and its variants [27, 28], which are computationally intensive and less suited for real-time applications, the proposed framework adopts a classification-oriented and resource-efficient design. Various registration strategies have been developed and applied in multimodal medical imaging to facilitate image alignment, including both intensity-based and feature-based approaches [29]. In the present study, an intensity-based registration framework was adopted to ensure consistent voxel-wise correspondence across MRI sequences. The use of depthwise separable convolutions enables effective feature extraction with reduced computational cost. Moreover, compared with heavier architectures such as ResNet and DenseNet [30], this lightweight design is better suited for small-sample medical imaging scenarios, offering a practical balance between efficiency and robustness. Overall, these characteristics support the applicability of the proposed framework as a resource-efficient tool for renal fibrosis staging.
We further combined the DL-sign with clinical indicators (eGFR, Scr, BUN) to construct machine learning models to identify both the presence and severity of RF. XGBoost emerged as the optimal model for detecting the presence of RF, whereas ExtraTree excelled for evaluating the severity of RF. During model comparison, bootstrap-based AUC analyses indicated that the differences in discrimination between the selected optimal models and alternative classifiers were generally modest, with overlapping confidence intervals in most comparisons. These observations suggest that several candidate models achieved comparable AUC performance when similar feature representations were used. Importantly, model selection was not based on the AUC alone but was informed by a comprehensive assessment integrating multiple threshold-dependent metrics, including the F1-score, sensitivity, specificity, and predictive values, as well as model stability across cross-validation folds. From an integrated evaluation perspective, the selected optimal models consistently ranked among the top-performing classifiers across multiple performance metrics and exhibited limited fold-to-fold variability, supporting their robustness and reliability in practical RF assessment. This multi-dimensional evaluation strategy helps mitigate potential bias associated with reliance on a single performance metric. Similar multimodal MRI-based approaches have demonstrated clinical feasibility in CKD studies. A single-centre retrospective study combined a T2WI-based model with ADC and R2* values to assess renal function in patients with diabetic nephropathy [24]. Hua et al. employed an SVM model integrating T1 mapping and DWI to effectively evaluate CKD and RF [15]. In our study, the inclusion of clinical biomarkers in the multimodal MRI-based DL framework improved both diagnostic accuracy and clinical interpretability. While eGFR reflects glomerular filtration, Scr and BUN represent nitrogen metabolism. Their combination with DL-derived structural features offers a comprehensive perspective on both functional decline and tissue remodelling. Notably, DL-sign features from native T1 mapping, ADC, and T2* mapping consistently ranked highest in SHAP analyses, indicating the pivotal role of multimodal MRI-derived features in RF stratification and highlighting the complementary value of clinical markers. The consistent dominance of the DL-sign in both the XGBoost and ExtraTree models further supports its robustness and diagnostic relevance across machine learning frameworks.
This proposed two-stage framework has meaningful potential for clinical adoption. By integrating quantitative multiparametric MRI with DL-derived features, the approach could support more standardized and objective evaluation of RF within routine kidney MRI workflows. In practice, such a system may assist radiologists in improving the consistency of fibrosis assessment and could facilitate earlier identification of patients at risk of progressive CKD. However, several practical considerations may influence large-scale deployment. The availability of scanners remains heterogeneous across institutions, and the cost and accessibility of multiparametric MRI sequences may limit their use in low-resource settings. In addition, variations in acquisition protocols across centres may introduce challenges to model transferability, underscoring the need for harmonization strategies. In the future, federated learning represents a promising direction for enhancing multicenter generalizability while avoiding raw data sharing [31]. Further development of more efficient and streamlined inference pipelines may also enable real-time implementation in clinical systems. These future improvements have the potential to expand the applicability of the proposed framework in broader clinical environments.
This study had several limitations. First, although 152 biopsy-proven CKD patients were included, cases with advanced fibrosis remained relatively rare. In addition, the use of relatively broad fibrosis grading thresholds resulted in the combination of moderate and severe fibrosis into a single RF3 category, which may introduce intragroup heterogeneity. Second, the dataset was derived from a single centre, which may limit generalizability. Although a multicentre collaboration for CKD imaging and biopsy data collection has been initiated by our team, the external datasets are still in the early acquisition stage and were not yet suitable for model validation in the present study. Third, the moderate sample size constrained further analysis of the associations between DL-derived features and different CKD aetiologies, larger multicentre cohorts will be needed to identify aetiology-specific fibrosis signatures. Finally, we did not compare the proposed lightweight MobileNetV2-SE-based feature extractor with deeper architectures such as ResNet, EfficientNet, or the Swin Transformer. Given the current dataset size, deeper networks pose a high risk of overfitting, however, future studies with expanded multicentre datasets will incorporate such architectures to further evaluate model robustness and benchmark model performance comprehensively.
In conclusion, this study demonstrates that a multimodal MRI-based deep learning framework can effectively capture fibrosis-related imaging information and enable noninvasive assessment of RF in patients with CKD. By integrating complementary information from native T1 mapping, ADC, and T2* mapping, the proposed approach provides added value over single-sequence models for evaluating both the presence and severity of RF. Furthermore, the derived DL-sign serves as a stable and continuous imaging representation that can be readily combined with routine clinical indicators. Using this two-stage strategy, XGBoost was identified as the optimal classifier for detecting the presence of renal fibrosis, whereas ExtraTree showed superior performance for assessing fibrosis severity. Collectively, these findings highlight the potential clinical utility of combining multimodal MRI with clinical data to support noninvasive RF evaluation and personalized management in patients with CKD.
Electronic supplementary material
Below is the link to the electronic supplementary material.
Acknowledgements
We thank AJE Editing Service for editing this manuscript.
Abbreviation
- CKD
Chronic kidney disease
- RF
Renal fibrosis
- MRI
Magnetic resonance imaging
- DWI
Diffusion-weighted imaging
- ADC
Apparent diffusion coefficient
- ANTs
Advanced normalization tools
- SyN
Symmetric normalization
- DL
Deep learning
- BMI
Body mass index
- Scr
Serum creatinine
- BUN
Blood urea nitrogen
- 24 h-UPRO
24-hour urinary protein
- eGFR
Estimated glomerular filtration rate
- CKD-EPI
Chronic Kidney Disease Epidemiology Collaboration
- SE
Squeeze-and-Excitation
- ROI
Region of interest
- LR
Logistic regression
- NB
Naive bayes
- SVM
Support Vector Machine
- DT
Decision Tree
- ExtraTree
Extremely Randomized Trees
- XGBoost
eXtreme Gradient Boosting
- MLP
Multi-Layer Perceptron
- GBM
Gradient Boosting Machines
- AUC
Area under the curve
- ROC
Receiver operating characteristic
- DCA
Decision curve analysis
Author contributions
Xiaojing Li: Conceptualization, Formal analysis, Writing-Original Draft. Yirui Li: Visualization, Formal analysis, Writing- Original Draft. Qing Ma: Investigation, Data Curation. Yilin Xu: Resources. Ye Zhu: Investigation. Jing Zhang: Software. Junkang Shen: Validation. Wu Cai: Supervision. Zhen Jiang, Chaogang Wei: Project administration, Writing-Review & Editing.
Funding
This study was financially supported by the Suzhou Medical College-QiLu Medical Research Program of Soochow University (24QL200214); the Suzhou Science and Technology Healthcare Innovation Project (SYW2025047); the Suzhou Science and Education Strong Health Project (MSXM2025013); the National Natural Science Foundation of China (81801754); the Project of State Key Laboratory of Radiation Medicine and Protection, Soochow University (GZK12025016).
Declarations
Data availability
The datasets generated and analysed during the current study are not publicly available due to ethical obligations to protect patient confidentiality and the dataset’s integral role in an ongoing longitudinal research program, but are available from the corresponding author on reasonable request.
Ethics approval and consent to participate
This study was performed in line with the principles of the Declaration of Helsinki. Approval was granted by the Ethics Committee of The Second Affiliated Hospital of Soochow University. (Approval number: JD-LK-2022–060-01).Written informed consent was obtained from all individual participants included in the study.
Consent for publication
All patients signed informed consent regarding publishing their data and photographs.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Xiaojing Li, Yirui Li and Qing Ma contributed equally to this work as co-first authors.
Contributor Information
Chaogang Wei, Email: weichaogang1122@163.com.
Zhen Jiang, Email: jiangzhen0416@suda.edu.cn.
References
- 1.Stewart S, Kalra PA, Blakeman T, Kontopantelis E, Cranmer-Gordon H, Sinha S. Chronic kidney disease: detect, diagnose, disclose-a UK primary care perspective of barriers and enablers to effective kidney care. BMC Med. 2024;22:331. 10.1186/s12916-024-03555-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Francis A, Harhay MN, Ong A, et al. Chronic kidney disease and the global public health agenda: an international consensus. Nat Rev Nephrol. 2024;20:473–85. 10.1038/s41581-024-00820-6. [DOI] [PubMed] [Google Scholar]
- 3.Panizo S, Martínez-Arias L, Alonso-Montes C, et al. Fibrosis in chronic kidney disease: pathogenesis and consequences. Int J Mol Sci. 2021;22:408. 10.3390/ijms22010408. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Schnuelle P. Renal biopsy for diagnosis in kidney disease: indication, technique, and safety. J Clin Med. 2023;12:6424. 10.3390/jcm12196424. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Jiang B, Liu F, Fu H, Mao J. Advances in imaging techniques to assess kidney fibrosis. Ren Fail. 2023;45:2171887. 10.1080/0886022X.2023.2171887. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Li J, An C, Kang L, Mitch WE, Wang Y. Recent Advances in magnetic resonance imaging assessment of renal fibrosis. Adv Chronic Kidney Dis. 2017;24:150–53. 10.1053/j.ackd.2017.03.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Berchtold L, Crowe LA, Combescure C, et al. Diffusion-magnetic resonance imaging predicts decline of kidney function in chronic kidney disease and in patients with a kidney allograft. Kidney Int. 2022;101:804–13. 10.1016/j.kint.2021.12.014. [DOI] [PubMed] [Google Scholar]
- 8.Berchtold L, Friedli I, Crowe LA, et al. Validation of the corticomedullary difference in magnetic resonance imaging-derived apparent diffusion coefficient for kidney fibrosis detection: a cross-sectional study. Nephrol Dial Transpl. 2020;35:937–45. 10.1093/ndt/gfy389. [DOI] [PubMed] [Google Scholar]
- 9.Wei CG, Zeng Y, Zhang R, et al. Native T1 mapping for non-invasive quantitative evaluation of renal function and renal fibrosis in patients with chronic kidney disease. Quant Imag Med Surg. 2023;13:5058–71. 10.21037/qims-22-1304. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Inoue T, Kozawa E, Okada H, et al. Noninvasive evaluation of kidney hypoxia and fibrosis using magnetic resonance imaging. J Am Soc Nephrol. 2011;22:1429–34. 10.1681/ASN.2010111143. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Guiot J, Vaidyanathan A, Deprez L, et al. A review in radiomics: making personalized medicine a reality via routine imaging. Med Res Rev. 2022;42:426–40. 10.1002/med.21846. [DOI] [PubMed] [Google Scholar]
- 12.Zhang M, Ye Z, Yuan E, et al. Imaging-based deep learning in kidney diseases: recent progress and future prospects. Insights Imag. 2024;15:50. 10.1186/s13244-024-01636-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Jiang X, Hu Z, Wang S, Zhang Y. Deep learning for medical image-based cancer diagnosis. Cancers (Basel). 2023;15:3608. 10.3390/cancers15143608. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Aslam I, Aamir F, Kassai M, et al. Validation of automatically measured T1 map cortico-medullary difference (ΔT1) for eGFR and fibrosis assessment in allograft kidneys. PLoS One. 2023;18:e0277277. 10.1371/journal.pone.0277277. [DOI] [PMC free article] [PubMed]
- 15.Hua C, Qiu L, Zhou L, et al. Value of multiparametric magnetic resonance imaging for evaluating chronic kidney disease and renal fibrosis. Eur Radiol. 2023;33:5211–21. 10.1007/s00330-023-09674-1. [DOI] [PubMed] [Google Scholar]
- 16.Stevens PE, Levin A. Kidney disease: improving global outcomes chronic kidney disease guideline development work group members (2013) evaluation and management of chronic kidney disease: synopsis of the kidney disease: improving global outcomes 2012 clinical practice guideline. Ann Intern Med. 158:825–30. 10.7326/0003-4819-158-11-201306040-00007. [DOI] [PubMed]
- 17.Wei C, Jin Z, Ma Q, et al. Native T1 mapping-based radiomics diagnosis of kidney function and renal fibrosis in chronic kidney disease. iScience. 2024;27:110493. 10.1016/j.isci.2024.110493. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Levey AS, Stevens LA, Schmid CH, et al. A new equation to estimate glomerular filtration rate. Ann Intern Med. 2009;150:604–12. 10.7326/0003-4819-150-9-200905050-00006. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Zhu Q, Zhuang H, Zhao M, Xu S, Meng R. A study on expression recognition based on improved mobilenetV2 network. Sci Rep. 2024;14:8121. 10.1038/s41598-024-58736-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Chen Z, Wang Y, Ying M, Su Z. Interpretable machine learning model integrating clinical and elastosonographic features to detect renal fibrosis in Asian patients with chronic kidney disease. J Nephrol. 2024;37:1027–39. 10.1007/s40620-023-01878-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Xu J, Wu X, Xu Y, et al. Acute kidney disease increases the risk of post-kidney biopsy bleeding complications. Kidney Blood Press Res. 2020;45:873–82. 10.1159/000509443. [DOI] [PubMed] [Google Scholar]
- 22.Srivastava A, Palsson R, Kaze AD, et al. The prognostic value of histopathologic lesions in native kidney biopsy specimens: results from the Boston kidney biopsy cohort study. J Am Soc Nephrol. 2018;29:2213–24. 10.1681/ASN.2017121260. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Mo X, Chen W, Chen S, et al. MRI texture-based machine learning models for the evaluation of renal function on different segmentations: a proof-of-concept study. Insights Imag. 2023;14:28. 10.1186/s13244-023-01370-4. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Chen W, Zhang L, Cai G, et al. Machine learning-based multimodal MRI texture analysis for assessing renal function and fibrosis in diabetic nephropathy: a retrospective study. Front Endocrinol (Lausanne). 2023;14:1050078. 10.3389/fendo.2023.1050078. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Patel RK, Choudhary A, Kumari N, Lamkuche HS. Pneumonia screening from radiology images using homomorphic transformation filter-based FAWT and customized VGG-16. Int J Imag Syst Technol. 2025;35:e70093. 10.1002/ima.70093.
- 26.Patel RK, Kashyap M. The study of various registration methods based on maximal stable extremal region and machine learning. Comput Methods Biomech Biomed Eng: Imag Visual. 2023;11(6):2508–15. 10.1080/21681163.2023.2243351. [Google Scholar]
- 27.Siddique N, Sidike P, Elkin C, Devabhaktuni V. U-net and its variants for medical image segmentation: a review of theory and applications. IEEE Access. 2021;9:82031–57. 10.1109/ACCESS.2021.3081920. [Google Scholar]
- 28.Liu J, Yildirim O, Akin O, Tian Y. AI-Driven robust kidney and renal mass segmentation and classification on 3D CT images. Bioeng (Basel). 2023;10:116. 10.3390/bioengineering10010116. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Deshpande S, Chouhan SS, Patel RK, Vishwakarma H. Transfer learning with ResNet50 for enhanced mammographic breast cancer identification. 2024 5th International Conference on Circuits, Control, Communication and Computing (I4C). IEEE; 2024.
- 30.Sharma N, Gupta S, Gupta D, et al. UMobileNetV2 model for semantic segmentation of gastrointestinal tract in MRI scans. PLoS One. 2024;19(5):e0302880. 10.1371/journal.pone.0302880. [DOI] [PMC free article] [PubMed]
- 31.Sheller MJ, Edwards B, Reina GA, Martin J, Pati S, Kotrotsou A, et al. Federated learning in medicine: facilitating multi-institutional collaborations without sharing patient data. Sci Rep. 2020;10(1):12598. 10.1038/s41598-020-69250-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The datasets generated and analysed during the current study are not publicly available due to ethical obligations to protect patient confidentiality and the dataset’s integral role in an ongoing longitudinal research program, but are available from the corresponding author on reasonable request.








