Abstract
Immunoglobulin A nephropathy (IgAN) is the most common primary glomerular disease worldwide, with a highly heterogeneous clinical course. Current biopsy-based risk assessment relies largely on structured histological scores, such as the Oxford MEST-C score, which summarize selected lesions but may omit information contained in routine narrative pathology reports. Here we show that large language model-based analysis of routine biopsy reports identifies reproducible pathological subtypes with prognostic relevance in IgAN. We studied 3,078 adults with primary IgAN from two hospitals in Wenzhou, China, using the First Affiliated Hospital cohort for subtype discovery and model development and the Second Affiliated Hospital cohort for independent external validation. Unsupervised clustering identified two subtypes, low-injury and high-injury, that differed in clinical presentation, histological severity and kidney outcomes. The high-injury subtype was independently associated with a higher risk of the composite kidney endpoint after adjustment for baseline clinical factors and the complete Oxford MEST-C score set (hazard ratio, 2.57; 95% confidence interval, 1.25-5.29). Interpretability analyses indicated that subtype separation was driven mainly by chronic tubulointerstitial injury and inflammatory burden. Exploratory analyses suggested heterogeneity in corticosteroid-associated outcome patterns across subtypes. These findings indicate that routine narrative biopsy reports contain prognostic information that complements structured pathological scoring.
Subject terms: IgA nephropathy, Kidney diseases
Kidney pathology reports contain rich information that structured assessment does not fully capture. Here, the authors show that large language models extract and standardize it to define reproducible IgA nephropathy subtypes that improve risk prediction and reveal differential treatment responses.
Introduction
Immunoglobulin A nephropathy (IgAN) is the most common type of primary glomerulonephritis worldwide and is characterized by marked heterogeneity in its clinical course, with a substantial proportion of patients progressing to kidney failure throughout their lifetime1. Accurate risk stratification and the rational selection of patients for immunosuppressive therapy, therefore, remain central and unresolved challenges in IgAN management2,3.
Renal biopsy is the cornerstone of diagnosis and prognostic assessment in IgAN4. The Oxford Classification (MEST-C) provides a standardized and reproducible framework for grading key histopathological lesions and has substantially advanced clinicopathological research5–7. Nevertheless, as an expert-defined categorical system, the MEST-C necessarily compresses a continuous and multidimensional spectrum of renal morphology into a limited set of discrete variables. In routine practice, kidney pathology reports contain rich narrative descriptions detailing lesion severity, spatial distribution, and contextual relationships among pathological features-information that is not fully captured by existing scoring systems8–10. The inability to systematically integrate this unstructured information represents a critical bottleneck in translating pathological data into precision medicine for IgAN.
This limitation is particularly consequential for immunosuppressive therapy. Although systemic glucocorticoids can reduce the risk of kidney disease progression in selected high-risk patients, their clinical benefit is heterogeneous and counterbalanced by substantial adverse effects11–15. Identifying patients whose underlying pathological activity is most likely to respond to immunomodulation remains a major unmet need16–18. Importantly, currently available tools, including the MEST-C score, have not been prospectively validated as predictive biomarkers of therapeutic response, underscoring the need for alternative approaches that more directly reflect disease activity at the tissue level2,3.
Recent advances in artificial intelligence, particularly large language models (LLMs), offer a unique opportunity to overcome these limitations19–22. Unlike conventional rule-based or feature-engineered natural language processing methods, LLMs can capture complex medical semantics and contextual relationships within unstructured clinical text, enabling scalable extraction of high-dimensional representations from narrative reports without the need for rigid, expert-defined ontologies23–25. This capability supports a fundamentally data-driven paradigm for pathological phenotyping, in which latent disease patterns can be discovered directly from the full textual content of biopsy reports rather than being constrained by predefined categories.
In this multicenter study, we developed and validated an LLM-based analytical framework that systematically converts narrative kidney pathology reports into high-dimensional pathological representations. Using unsupervised learning on these representations, we derived data-driven pathological subtypes that capture latent disease heterogeneity more thoroughly than do conventional scoring systems. We evaluated the clinical relevance of these subtypes by examining their associations with kidney outcomes and their relationship to outcome patterns in patients with corticosteroid exposure. By showing that clinically meaningful pathological subtypes can be derived directly from routine narrative biopsy reports, our study provides a framework for pathology-informed risk stratification and supports the potential of narrative pathology as a scalable source of prognostic information in IgAN.
Results
Study population and kidney pathology subtyping
The overall analytical workflow is shown in Fig. 1. In the derivation cohort from the First Affiliated Hospital, 8635 renal biopsies were screened, yielding 2920 patients with primary IgAN. After exclusion of duplicate records (n = 6), missing data (n = 1), and incident malignancies (n = 124), 2789 patients were included for model development and clinical analyses. An independent external validation cohort was assembled from the Second Affiliated Hospital, in which 289 patients with primary IgAN were included after 1205 biopsies were screened, and cases with incomplete data were excluded (n = 916).
Fig. 1. Overview of the LLM-based framework for pathological subtype discovery and validation in IgA nephropathy.

A Study design and analytical workflow. Narrative renal pathology reports on 3078 patients with IgA nephropathy in a two-center retrospective cohort were analyzed. The derivation cohort was used for embedding-based pathological subtype discovery, and an independent external cohort was used to evaluate subtype assignment and generalizability. The clinical relevance of the resulting low-injury and high-injury pathological subtypes was assessed through their associations with baseline clinicopathological features, kidney outcomes, and exploratory analyses of corticosteroid-associated outcome patterns. B Extraction, embedding, dimensionality-reduction, and clustering pipeline. Narrative pathology reports were processed using DeepSeek-V3 to perform structured extraction of 71 predefined pathological elements. Text from four pathology domains - glomerular, tubulointerstitial, vascular, and immunofluorescence descriptors - was flattened and encoded using Qwen3-Embedding 8B to generate four 4096-dimensional domain-specific embeddings. These embeddings were concatenated into a 16,384-dimensional report-level vector, reduced from 16,384 to 2500 and then to 300 dimensions using two-step principal component analysis, and used to derive the LLM-based pathological subtypes via unsupervised clustering. A supervised classifier was subsequently trained using the derivation cohort and applied to the external cohort for subtype prediction. Source data are provided as a Source Data file.
Unsupervised clustering analyses using multiple algorithms showed that partition-based methods, particularly k-means, generally achieved more favorable internal clustering performance than density-based approaches, whereas DBSCAN and HDBSCAN were less stable across parameter settings (Supplementary Fig. 1A, B). Based on the overall clustering performance, stability, and interpretability, a 2-cluster solution was selected for downstream analyses. The resulting subtypes were designated low-injury and high-injury according to their distinct clinicopathological profiles (Supplementary Fig. 1C).
We performed repeated subsampling-based consensus analyses and perturbation analyses to further assess the robustness of clustering. Across 30 iterations of 80% random subsampling, the selected 2-cluster solution showed excellent reproducibility relative to the reference clustering (mean ARI, 0.988 ± 0.006; mean NMI, 0.971 ± 0.011). The consensus matrix displayed two sharply separated blocks, with very high within-cluster consensus (0.994), negligible between-cluster consensus (0.0067), and a low PAC value (0.017), indicating a highly stable clustering structure. Mild noise injection produced near-identical assignments (mean ARI, 0.999 ± 0.002; mean NMI, 0.997 ± 0.004), supporting robustness to small perturbations in the embedding space.
To enable subtype assignment in external datasets, supervised classifiers were trained using the derivation cohort. An ensemble model displayed the best overall performance (Supplementary Table 1) and reliably classified patients in the independent validation cohort into the two pathological subtypes (Supplementary Fig. 1D).
Integrated interpretability analyses indicate that a biologically robust continuum of tubulointerstitial injury underlies the embedding-derived subtype axis
To define the pathological basis of the embedding-derived subtypes, we integrated teacher-student SHAP attribution, feature- and block-level counterfactual perturbation, encoded-level decomposition, and cross-method concordance analyses. To assess whether stable pathological interpretation accompanied the numerically stable clustering structure, we further examined the consistency of subtype-defining pathological drivers across these complementary attribution and perturbation approaches. Across these complementary approaches, the dominant signal consistently localized to the tubulointerstitial compartment, with more modest contributions from vascular lesions and relatively limited contributions from glomerular descriptors (Fig. 2A–H).
Fig. 2. Map of the embedding-derived subtypes obtained from the integrated interpretability analysis to a tubulointerstitial injury continuum.

A mean absolute SHAP values for the top 12 raw pathological descriptors in the teacher-student model are displayed in a hierarchical pathology layout. B Mean absolute counterfactual support scores for grouped pathology blocks after replacement with opposite-subtype prototype values. C Joint evidence plot integrating raw feature SHAP importance with feature-level counterfactual support. Each point represents one feature; the x-axis denotes mean absolute SHAP value, the y-axis denotes swap mean absolute support, the point size denotes deletion mean absolute support, and the color denotes the pathological domain. D Encoded-level SHAP decomposition of the degree of tubular atrophy after one-hot encoding. E–G Distributions of feature categories across the low-injury and high-injury subtypes for the degree of tubular atrophy (E), degree of interstitial fibrosis (F), and degree of inflammatory cell infiltration (G). H Cross-method concordance heat map for the top-ranked features. Values were normalized within the method for SHAP, deletion, and swap analyses. In clinical terms, the high-injury subtype corresponds to pathology-report descriptors of more advanced tubular atrophy, interstitial fibrosis, and inflammatory cell infiltration, whereas the low-injury subtype corresponds to descriptors of absent or milder chronic tubulointerstitial injury and less prominent inflammatory burden. Source data are provided as a Source Data file.
In the hierarchical SHAP analysis of the raw features, the strongest contributors were the degree of tubular atrophy, degree of interstitial fibrosis, and degree of inflammatory cell infiltration, whereas glomerular descriptors, including mesangial hypercellularity and global sclerosis, contributed less strongly overall (Fig. 2A). Counterfactual analyses supported the same interpretation. At the block level, tubulointerstitial chronic injury showed the highest support scores, followed by features of the inflammatory burden, whereas vascular changes contributed more modestly and immunofluorescence-related features contributed minimally (Fig. 2B). At the individual-feature level, the degree of tubular atrophy, degree of interstitial fibrosis, and degree of inflammatory cell infiltration showed the strongest convergent evidence across SHAP ranking, single-feature deletion, and opposite-prototype swapping (Fig. 2C). Thus, the key pathological drivers were not restricted to a single interpretability method but were repeatedly identified across independent analytical views.
Clinically, these model-derived drivers corresponded to recognizable patterns in routine pathology reports. The low-injury subtype was characterized by absent or mild chronic tubulointerstitial injury, limited tubular atrophy and interstitial fibrosis, and focal or mild inflammatory infiltration. In contrast, the high-injury subtype was associated with reports describing moderate-to-severe tubular atrophy and interstitial fibrosis accompanied by broader inflammatory distribution and more prominent inflammatory cell infiltration that was often lymphocyte-rich. Thus, the embedding-derived axis could be translated into a simplified pathological pattern of chronic tubulointerstitial remodeling with superimposed inflammatory burden.
Encoded-level decomposition further indicated that subtype separation was driven primarily by graded shifts in lesion severity, rather than by the mere presence or absence of lesions. For the degree of tubular atrophy, the largest contributions arose from ordinal severity categories after one-hot encoding (Fig. 2D). Consistently, subtype-specific distributions of the values for the degree of tubular atrophy, degree of interstitial fibrosis, and degree of inflammatory cell infiltration showed an enrichment of more advanced lesion categories in the high-injury subtype and milder categories in the low-injury subtype (Fig. 2E–G). The cross-method concordance analysis likewise identified tubular atrophy, interstitial fibrosis, and inflammatory cell infiltration as the most robust and reproducible drivers of the subtype assignment. Notably, although glomerular and vascular descriptors contributed less strongly overall, several features had relatively high support in deletion-based analyses, including capillary loop abnormalities, arteriolosclerosis description, and intimal fibrosis severity. These results suggest that the omission of these descriptors can shift the subtype assignment, indicating that they also provide important discriminatory information for the pathological stratification of IgAN (Fig. 2H).
Collectively, these analyses indicate that the latent axis captured by pathology-report embedding reflects a clinically recognizable and biologically robust continuum of tubulointerstitial injury accompanied by the inflammatory burden rather than an opaque clustering pattern in high-dimensional space. Together with the numerical resampling stability analyses, the findings suggest that the embedding-derived subtypes are stable not only in their cluster assignment but also in their underlying pathological interpretation. This back-mapping from the embedding space to specific pathological descriptors strengthens the biological plausibility and the clinical interpretability of the embedding-derived subtypes and supports their potential utility for prognostic stratification and the exploratory assessment of treatment-response heterogeneity.
Baseline characteristics
Baseline clinical and pathological characteristics differed markedly between the two pathological subtypes (Table 1). Compared with patients with the low-injury subtype (n = 1562), patients with the high-injury subtype (n = 1227) exhibited higher proteinuria levels, lower eGFRs, a higher incidence of hypertension, and more severe histological injury, as assessed by the Oxford MEST-C score (all P < 0.001). A high-risk clinical classification and use of corticosteroids or other immunosuppressive therapies were also more frequent in the high-injury group.
Table 1.
Baseline characteristics of IgAN patients by LLM classification
| Characteristic | Overall | Low-injury | High-injury | P value |
|---|---|---|---|---|
| N | 2789 | 1562 | 1227 | |
| Age (years) | 41.0 (12.9) | 39.4 (12.8) | 43.0 (12.9) | 1.16 × 10−13 |
| Female | 1543 (55.3) | 943 (60.4) | 600 (48.9) | 1.85 × 10−9 |
| Systolic blood pressure (mmHg) | 122.7 (14.3) | 119.1 (13.4) | 127.3 (14.2) | 5.44 × 10−45 |
| Diastolic blood pressure (mmHg) | 76.7 (9.8) | 74.3 (8.9) | 79.6 (10.0) | 1.05 × 10−39 |
| Hypertension | 421 (15.1) | 168 (10.8) | 253 (20.6) | 7.53 × 10−13 |
| Diabetes | 102 (3.7) | 38 (2.4) | 64 (5.2) | 0.000154 |
| High-risk IgAN | 2275 (81.8) | 1189 (76.4) | 1086 (88.6) | 2.33 × 10−16 |
| Serum albumin (g/L) | 38.1 (5.0) | 38.9 (4.9) | 37.0 (4.9) | 1.05 × 10−22 |
| Serum creatinine (μmol/L) | 90.8 [70.1, 119.3] | 76.4 [62.0, 93.8] | 116.8 [92.4, 152.0] | 2.59 × 10−192 |
| Serum uric acid (μmol/L) | 369.9 (84.8) | 342.7 (76.6) | 404.5 (82.0) | 4.14 × 10−85 |
| Hemoglobin (g/L) | 128.8 (18.1) | 131.0 (17.0) | 126.0 (19.1) | 6.93 × 10−13 |
| 24-H urine protein (g/24 h) | 1.2 [0.6, 2.4] | 0.9 [0.5, 1.7] | 1.7 [0.9, 3.1] | 9.75 × 10−50 |
| eGFR (ml/min/1.73 m2) | 81.4 [58.4, 105.7] | 99.2 [79.5, 114.7] | 60.1 [43.5, 77.2] | <2.2 × 10−16 |
| eGFR category | 1.31 × 10−185 | |||
| G1 | 1138 (40.8) | 979 (62.7) | 159 (13.0) | |
| G2 | 904 (32.4) | 446 (28.6) | 458 (37.3) | |
| G3a-G3b | 593 (21.3) | 119 (7.6) | 474 (38.6) | |
| G4 | 98 (3.5) | 14 (0.9) | 84 (6.8) | |
| G5 | 56 (2.0) | 4 (0.3) | 52 (4.2) | |
| MEST-C M1 | 1184 (42.9) | 447 (28.8) | 737 (61.2) | 5.57 × 10−65 |
| MEST-C E1 | 1327 (47.8) | 729 (46.9) | 598 (48.9) | 0.324 |
| MEST-C S1 | 2202 (79.3) | 1094 (70.4) | 1108 (90.6) | 1.84 × 10−38 |
| MEST-C T score | <2.2 × 10−16 | |||
| 1 | 907 (32.7) | 152 (9.8) | 755 (61.7) | |
| 2 | 293 (10.6) | 5 (0.3) | 288 (23.5) | |
| MEST-C C1 | 1417 (51.0) | 740 (47.6) | 677 (55.4) | 0.0000650 |
| Corticosteroid therapy | 1434 (51.4) | 774 (49.6) | 660 (53.8) | 0.0289 |
| RAS blockade therapy | 2362 (84.7) | 1346 (86.2) | 1016 (82.8) | 0.0164 |
| Mycophenolate mofetil | 255 (9.1) | 102 (6.5) | 153 (12.5) | 9.51 × 10−8 |
| Cyclophosphamide | 258 (9.3) | 105 (6.7) | 153 (12.5) | 2.84 × 10−7 |
| Tacrolimus | 82 (2.9) | 51 (3.3) | 31 (2.5) | 0.302 |
| Azathioprine | 3 (0.1) | 3 (0.2) | 0 (0.0) | 0.340 |
| Leflunomide | 60 (2.2) | 30 (1.9) | 30 (2.4) | 0.415 |
| Immunosuppressive therapy | 574 (20.6) | 256 (16.4) | 318 (25.9) | 8.76 × 10−10 |
Data are shown as mean (SD), median (IQR), or n (%). P-values were calculated using two-sided Student’s t-tests, Mann-Whitney U-tests, χ² tests, or Fisher’s exact tests, as appropriate, without adjustment for multiple comparisons. Exact P-values are shown in the table. MEST-C, Oxford classification of IgA nephropathy; RAS, renin-angiotensin system. Immunosuppressive therapy included mycophenolate mofetil, cyclophosphamide, tacrolimus, azathioprine, or leflunomide initiated within 1 year after biopsy. Source data are provided as a Source Data file.
Importantly, substantial differences in both clinical manifestations and MEST-C lesion severity persisted between pathological subtypes even within the same clinical risk categories (Supplementary Fig. 2), which indicated that the LLM-derived pathological subtypes captured the pathological heterogeneity beyond the conventional clinical risk stratification.
These findings were replicated in the external validation cohort (Supplementary Table 2), in which patients assigned to the high-injury subtype exhibited significantly worse kidney function, a higher incidence of hypertension, and more advanced MEST-C lesions compared with those with the low-injury subtype (Supplementary Fig. 3).
Pathological subtypes and composite kidney endpoint-free survival
Kaplan-Meier analyses revealed significant differences in composite kidney endpoint-free survival between patients with different pathological subtypes in the derivation cohort (P < 0.001; Supplementary Fig. 4A). Although the proteinuria-based clinical risk stratification was also associated with kidney outcomes in the derivation cohort (P < 0.001; Supplementary Fig. 4B), it yielded a highly imbalanced distribution across risk categories, with 507 patients classified as low risk and 2275 as high risk. External validation confirmed that the high-injury subtype was consistently associated with worse composite kidney endpoint-free survival than was the low-injury subtype (P < 0.001; Supplementary Fig. 4C), whereas the clinical risk stratification failed to distinguish outcomes in the validation cohort (P = 0.83; Supplementary Fig. 4D). Together, these findings indicate that proteinuria-based clinical risk stratification showed inconsistent discriminatory performance across cohorts and was sensitive to the underlying distribution of patients across risk categories, particularly when risk groups were highly imbalanced.
Concordance, added resolution, and stability of LLM-derived pathological subtypes relative to the Oxford MEST-C classification
To examine the relationships between LLM-derived pathological subtypes and established histologic grading, we first evaluated the associations between LLM-derived pathological subtypes and individual Oxford MEST-C components. In multivariable logistic regression models adjusted for all MEST-C variables, pathological subtypes were most strongly associated with the tubulointerstitial component (T score), with markedly elevated odds ratios for both T1 and T2 lesions (T1: OR, 33.37; T2: OR, 363.79; both P < 0.001; Supplementary Table 3). Significant associations were also observed for the M and S components, whereas E and C were not independently associated. These findings were replicated in the external validation cohort, in which the T score remained the dominant histologic correlate of pathological subtypes (Supplementary Table 4). Collectively, these results supported strong concordance between LLM-derived pathological subtypes and established markers of chronic structural kidney damage.
In contrast, clinical risk stratification showed broader but weaker associations across multiple MEST-C components, including M, E, S, and T, while remaining unassociated with the C score. This diffuse pattern suggested that clinical risk stratification reflected overall disease severity but lacked specificity for discrete renal morphology patterns.
We next assessed whether LLM-derived pathological subtypes provide additional discriminatory resolution beyond MEST-C patterns. In terms of identical MEST-C combinations, pathological subtypes exhibited substantially greater separation between the low-injury and high-injury groups than did clinical risk stratification between the low-risk and high-risk groups (Fig. 3A). Quantitative similarity analyses using Euclidean distances between comprehensive pathological embeddings further confirmed these findings, with pathological subtypes displaying consistently greater between-group separation and moderate-to-large effect sizes (Cohen’s d) across most MEST-C patterns (Fig. 3B). In contrast, the effect sizes for clinical risk stratification were predominantly small to moderate. These results indicated that, even within the same MEST-C category, LLM-derived pathological subtypes captured latent morphological heterogeneity that was not reflected by conventional grading systems.
Fig. 3. Relationships of MEST-C patterns with pathological subtypes and clinical risk stratification.

A Distribution of pathological subtypes and clinical risk categories across Oxford MEST-C patterns. The bar plots show the proportion of patients classified as low-injury or high-injury according to pathological subtype and as low-risk or high-risk according to clinical risk stratification within each MEST-C pattern. B Quantitative similarity analysis of MEST-C patterns based on Euclidean distances between MEST-C pattern embeddings and comprehensive renal pathology representations derived from full biopsy reports. Standardized distances are shown for pathological subtype-based and clinical risk stratification-based grouping. In the forest plots, central markers represent Cohen’s d point estimates for the between-group difference within each MEST-C pattern, and horizontal error bars represent the corresponding 95% confidence intervals. For the clinical risk comparison, distances were compared between low-risk patients (n = 507) and high-risk patients (n = 2275) in the complete follow-up cohort. For the pathological subtype comparison, distances were compared between low-injury patients (n = 1562) and high-injury patients (n = 1227). The unit of study was an individual patient with one baseline kidney biopsy report. These n values represent independent biological replicates from different patients; no technical replicates were used. Each patient contributed one derived distance value for each MEST-C pattern evaluated. Statistical comparisons were performed using two-sided Student’s t-tests. No separate control group was used because this was an observational comparison between predefined clinical risk strata and model-derived pathological subtypes. Source data are provided as a Source Data file.
In additional model-comparison analyses, the LLM-derived pathological subtypes provided measurable but modest incremental prognostic information beyond that provided by clinical variables and MEST-C. Adding the subtype to the clinical + MEST-C model improved model fit, as reflected by lower AIC and BIC and a significant likelihood ratio test, whereas gains in Harrell’s C-index and decision-curve net benefit were small (Fig. 4).
Fig. 4. Incremental prognostic value of the LLM-derived pathological subtype beyond clinical variables and the Oxford MEST-C score.

Decision curve analyses at 36 months (A) and 60 months (B) comparing the results obtained using clinical variables alone, clinical variables plus the Oxford MEST-C score, clinical variables plus the LLM-derived pathological subtype, and clinical variables plus both the Oxford MEST-C score and the LLM-derived subtype. C Summary of model-comparison analyses, including model fit indices, likelihood ratio tests, Harrell’s C-index, and the ΔC-index. Nested Cox models were compared using two-sided likelihood ratio tests, with no adjustment for multiple comparisons. The likelihood ratio test statistics were as follows: clinical versus clinical + MEST-C, χ2 = 28.158, df = 6, P = 8.77 × 10−5; clinical versus clinical + LLM-derived subtype, χ2 = 21.956, df = 1, P = 2.79 × 10−6; clinical + MEST-C versus clinical + MEST-C + LLM-derived subtype, χ2 = 9.058, df = 1, P = 0.00262; and clinical + LLM-derived subtype versus clinical + MEST-C + LLM-derived subtype, χ2 = 15.260, df = 6, P = 0.0183. For C-index and ΔC-index estimates, central markers represent point estimates, and horizontal error bars represent 95% confidence intervals estimated using 1000 patient-level bootstrap resamples. All analyses in this figure were performed in the complete-case prognostic analysis cohort of 2576 independent patients, including 305 patients who developed the composite kidney endpoint. The unit of study was an individual patient; each patient contributed one baseline clinical-pathological record and one follow-up outcome record. These n values represent independent biological replicates from different patients; no technical replicates were used. Bootstrap resamples were computational resampling iterations and were not biological or technical replicates. Likelihood ratio tests were used to compare nested Cox models. No experimental control group was used because the analysis was an observational prognostic model comparison; the clinical model served as the reference model. Source data are provided as a Source Data file.
Finally, we evaluated the stability and reproducibility of pathological subtype classification across institutions. Using the LLM embedding distance-based classification, interhospital agreement for pathological subtypes ranged from substantial to almost perfect (κ = 0.61–0.94; all P < 0.001), with the highest concordance observed for subtype classification between centers (Fig. 5). In contrast, agreement based on the observed distribution-based classification was poor, with κ values ranging from −0.61 to 0.21, including multiple comparisons indicating worse-than-chance agreement. Importantly, these patterns of enhanced resolution and interhospital stability for LLM-derived pathological subtypes were independently reproduced in the external validation cohort (Supplementary Fig. 5), underscoring the robustness and generalizability of the morphology-informed classification framework.
Fig. 5. Agreement between clinical and pathological subtype classifications across hospitals and modeling strategies.

Heatmap of κ statistics illustrating agreement between clinical risk stratification and LLM-derived pathological subtype classification across two independent hospitals under two modeling strategies: observed distribution-based classification and LLM embedding distance-based classification. Each cell reports the κ value with its 95% confidence interval. Two-sided tests of H₀: κ = 0 were performed using asymptotic standard errors; no adjustment for multiple comparisons was applied. Background colors indicate the strength of agreement according to standard κ interpretation categories. Asterisks denote levels of statistical significance (*P < 0.05, **P < 0.01, ***P < 0.001). Exact P-values are provided in the Source Data. Greater agreement was consistently observed for LLM embedding distance-based classification, particularly for pathological subtype classification across hospitals. Source data are provided as a Source Data file.
Interactions between pathological subtypes and immunosuppressive therapy
In analyses stratified by proteinuria-based clinical risk group, neither exposure to corticosteroids nor exposure to therapy with pooled immunosuppressive drugs showed evidence of interaction with clinical risk status (Fig. 6A–D). Kaplan-Meier curves demonstrated broadly similar treatment-associated outcome patterns across low- and high-risk groups, and the corresponding interaction terms in the subgroup Cox models were not significant either for corticosteroid therapy (P for interaction = 0.466) or for pooled immunosuppressive therapy (P for interaction = 0.992).
Fig. 6. Associations of corticosteroid and pooled immunosuppressive therapy exposure with kidney outcomes according to clinical risk strata and pathological subtype.

Kaplan-Meier curves show composite kidney endpoint-free survival according to treatment exposure, stratified by proteinuria-based clinical risk group (A, B) and pathological subtype (E, F). Panels A and B compare treated (+) and untreated (−) patients within low- and high-risk groups for corticosteroid therapy and pooled immunosuppressive therapy, respectively. Panels E, F show the corresponding comparisons within the low-injury and high-injury pathological subtypes. The number of patients at risk is shown below each Kaplan-Meier panel. Forest plots summarize subgroup-specific hazard ratios (HRs) and 95% confidence intervals for corticosteroid therapy (C, G) and pooled immunosuppressive therapy (D, H). Central markers indicate HR point estimates from Cox proportional hazards models, and horizontal bars indicate 95% confidence intervals. Clinical risk group did not significantly modify the association of corticosteroid therapy (C; P for interaction = 0.466) or pooled immunosuppressive therapy (D; P for interaction = 0.992) with the composite endpoint. Pathological subtype significantly modified the association of corticosteroid therapy with the composite endpoint (G; P for interaction = 0.006), whereas the interaction with pooled immunosuppressive therapy was not statistically significant (H; P for interaction = 0.091). All analyses were performed at the patient level; each patient contributed one independent observation, and no technical replicates were used. P-values in the Kaplan-Meier panels were derived from log-rank tests. Subgroup-specific HRs and interaction P-values in (C, D, G, H) were estimated using Cox proportional hazards models. Interaction P-values were calculated using two-sided Wald tests of the treatment-by-stratum interaction term. No adjustment for multiple comparisons was applied. Exact P-values are provided in the Source Data. Source data are provided as a Source Data file.
In contrast, the pathological subtype significantly modified the association between corticosteroid exposure and kidney outcome (Fig. 6E, G). In subgroup Cox analyses, corticosteroid exposure was associated with a lower risk of the composite kidney endpoint in the high-injury subtype (HR, 0.77; 95% CI, 0.59–0.99), whereas the association was in the opposite direction in the low-injury subtype (HR, 1.59; 95% CI, 1.01–2.52), with a significant interaction between subtype and corticosteroid exposure (P for interaction = 0.006). These divergent association patterns were also evident in the Kaplan-Meier curves. Taken together, these findings suggest that the embedding-derived pathological subtype may capture heterogeneity in corticosteroid-associated outcome patterns that is not reflected in conventional clinical risk stratification, although residual confounding by indication cannot be excluded.
For pooled immunosuppressive therapy, no statistically significant interaction with pathological subtype was observed (Fig. 6F, H; P for interaction = 0.091). However, the subgroup-specific estimates were directionally similar to those observed for corticosteroids; the high-injury subtype was associated with lower risk (HR, 0.74; 95% CI, 0.54–1.02), whereas there was no clear association in the low-injury subtype (HR, 1.24; 95% CI, 0.71–2.19). These results suggest that the treatment-associated heterogeneity was most clearly evident for corticosteroid exposure in the present dataset.
Inverse probability-weighted analyses
We fitted IPTW Cox proportional hazards models for the composite kidney endpoint to reduce potential treatment selection bias (Table 2). In the corticosteroid analysis, corticosteroid use was not associated with the overall kidney outcomes in either the base model or the fully adjusted model (base model: HR, 1.17, 95% CI 0.71–1.91, P = 0.54; fully adjusted model: HR, 0.86, 95% CI 0.50–1.48, P = 0.59). In contrast, the pathological subtype showed a strong main effect, with the high-injury subtype associated with a substantially higher risk than the low-injury subtype in the base model (HR, 9.98, 95% CI 6.53–15.27, P < 0.001). This association was attenuated but remained significant after full adjustment (HR, 2.57, 95% CI 1.25–5.29, P = 0.01), indicating that part of the excess risk was captured by conventional clinicopathological variables included in the model, but that the subtype retained independent prognostic value beyond these factors.
Table 2.
IPTW-adjusted Cox models for the composite kidney endpoint according to the pathological subtype and use of immunosuppressive therapy
| Variable | Base Model | Fully Adjusted Model | ||
|---|---|---|---|---|
| HR (95% CI) | P-value | HR (95% CI) | P-value | |
| Steroids: Main Effect | ||||
| High-injury subtype | 9.98 (6.53–15.27) | 2.70 × 10−26 | 2.57 (1.25–5.29) | 0.0105 |
| Steroids (Yes) | 1.17 (0.71–1.91) | 0.542 | 0.86 (0.50–1.48) | 0.589 |
| Age, years | — | — | 1.00 (0.99–1.01) | 0.533 |
| Female | — | — | 0.72 (0.56–0.93) | 0.0132 |
| Urine protein excretion, g/24 h | — | — | 1.23 (1.18–1.28) | 2.42 × 10−23 |
| MEST-C, C1 | — | — | 0.87 (0.67–1.14) | 0.316 |
| MEST-C, E1 | — | — | 1.12 (0.82–1.54) | 0.478 |
| MEST-C, M1 | — | — | 1.26 (0.94–1.67) | 0.119 |
| MEST-C, S1 | — | — | 1.29 (0.86–1.94) | 0.220 |
| MEST-C, T1 | — | — | 1.89 (1.04–3.42) | 0.0353 |
| MEST-C, T2 | — | — | 5.14 (2.77–9.51) | 1.94 × 10−7 |
| Steroid Interaction Effect | ||||
| High-injury subtype × steroid use | 0.40 (0.22–0.70) | 0.00157 | 0.51 (0.28–0.95) | 0.0339 |
| Immunosuppressants: Main Effect | ||||
| High-injury subtype | 7.10 (5.26–9.59) | 2.17 × 10−37 | 2.49 (1.32–4.70) | 0.00502 |
| Immunosuppressants (Yes) | 0.73 (0.38–.39) | 0.333 | 0.60 (0.32–1.13) | 0.114 |
| Age, years | — | — | 1.00 (0.99–1.01) | 0.616 |
| Female | — | — | 0.92 (0.69–1.24) | 0.596 |
| Urine protein excretion, g/24 h | — | — | 1.20 (1.14–1.25) | 2.01 × 10−15 |
| MEST-C, C1 | — | — | 1.08 (0.79–1.46) | 0.635 |
| MEST-C, E1 | — | — | 0.92 (0.67–1.28) | 0.634 |
| MEST-C, M1 | — | — | 1.10 (0.79–1.52) | 0.573 |
| MEST-C, S1 | — | — | 1.43 (0.88–2.32) | 0.153 |
| MEST-C, T1 | — | — | 1.74 (0.86–3.53) | 0.122 |
| MEST-C, T2 | — | — | 3.90 (1.88–8.09) | 0.000251 |
| Immunosuppressants: Interaction Effect | ||||
| High-injury subtype × immunosuppressant use | 0.68 (0.32–1.44) | 0.313 | 0.85 (0.40–1.80) | 0.663 |
Hazard ratios, 95% confidence intervals, and P-values were estimated from IPTW-weighted Cox proportional hazards models for the composite kidney endpoint. Fully adjusted models included age, sex, baseline proteinuria, and Oxford MEST-C scores (M, E, S, T, and C). P-values were calculated using two-sided Wald tests of Cox model coefficients, including treatment-by-subtype interaction terms. No adjustment for multiple comparisons was applied. Exact P-values are reported where available. CI confidence interval, HR hazard ratio, IPTW inverse probability of treatment weighting, MEST-C, Oxford classification for IgA nephropathy. Source data are provided as a Source Data file.
Notably, a significant interaction between the pathological subtype and corticosteroid therapy was observed in both the base and fully adjusted IPTW models (base model: HR for interaction, 0.40, 95% CI 0.22–0.70, P = 0.002; fully adjusted model: HR, 0.51, 95% CI 0.28–0.95, P = 0.03). A complementary 1:1 propensity score-matched analysis yielded directionally consistent findings (HR for interaction, 0.516; 95% CI 0.279–0.958; P = 0.036). These results are consistent with potential heterogeneity in corticosteroid-associated outcome patterns across pathological subtypes, although residual confounding by indication cannot be excluded. In the fully adjusted model, the relative hazard associated with the high-injury subtype was attenuated in steroid-treated patients compared with untreated patients, suggesting a possible subtype-dependent difference in the observed treatment association.
In contrast, for pooled immunosuppressive therapy, no significant interaction with the pathological subtype was detected in either the base model or the fully adjusted model (base model: HR for interaction, 0.68, 95% CI 0.32–1.44, P = 0.31; fully adjusted model: HR, 0.85, 95% CI 0.40–1.80, P = 0.66). Although the high-injury subtype remained independently associated with a worse outcome in this analysis, these results suggest that the effect modification was specific to corticosteroid therapy rather than to immunosuppressive treatment overall.
Discussion
In this multicenter retrospective study, we show that clinically meaningful pathological subtypes of IgAN can be derived directly from routine narrative kidney biopsy reports using an LLM-based framework. Unlike the Oxford MEST-C classification, which relies on predefined semiquantitative lesion scores, this data-driven approach captures latent pathological semantics embedded in free-text reports, reflecting integrative assessments routinely made by renal pathologists but not fully represented by structured scoring systems. Through unsupervised clustering of LLM-derived pathological representations, we identified two reproducible subtypes, designated low-injury and high-injury subtypes, with distinct clinicopathological profiles, markedly different risks of kidney disease progression, and heterogeneous corticosteroid-associated outcome patterns.
The traditional pathological assessment of IgAN has largely relied on expert-defined categorical systems, most notably the Oxford MEST-C classification26. Although this framework has provided an important foundation for standardized risk stratification, it necessarily compresses a complex, continuous, and spatially heterogeneous spectrum of histological injury into a limited set of discrete variables5–7. Within this context, our findings suggest that LLM-derived pathological subtypes are biologically grounded in established histological lesions while also capturing dimensions of injury that may not be fully represented by categorical scoring alone. The integration of established histological metrics with SHAP-based interpretability analyses demonstrated clear biological coherence of the derived pathological subtypes, with the dominant separating signal mapping to canonical Oxford lesions, particularly tubular atrophy/interstitial fibrosis, and more modest contributions from mesangial hypercellularity and segmental sclerosis. These observations suggest that the subtype axis captures a higher-resolution representation of tissue injury centered on the tubulointerstitial compartment that has long been recognized as the most consistent histological correlate of the prognosis of IgAN27–31.
In addition to its correlation with established histological metrics, the subtype axis was also clinically interpretable. Across integrated interpretability analyses, the dominant separating signals arose from chronic tubulointerstitial injury with the superimposed inflammatory burden, in which tubular atrophy, interstitial fibrosis, and inflammatory cell infiltration constituted the principal discriminating features, whereas vascular lesions and selected glomerular descriptors provided additional but more modest information. This pattern supports the use of low-injury and high-injury terminology and suggests that the key heterogeneity captured by the model is organized around a whole-kidney injury phenotype rather than glomerular activity alone. In other words, the distinction between subtypes appears to reflect the combined burden of chronic tubulointerstitial remodeling and immune cell infiltration, rather than proliferative glomerular lesions in isolation. This interpretation is consistent with prior work showing that active tubulointerstitial inflammation in the non-scarred cortex provides prognostic information beyond established clinical and histological indicators in patients with IgAN, supporting the view that fibrosis scores alone do not fully capture the tissue-level disease burden32.
This additional pathological resolution may be clinically meaningful. The pathological subtype remained independently associated with kidney outcomes after adjustment for baseline clinical risk factors and Oxford MEST-C scores, indicating that routine narrative biopsy reports contain prognostically relevant information in addition to that captured by structured scoring alone. Moreover, even among patients with the same MEST-C pattern, LLM-derived pathological subtypes distinguished substantial heterogeneity in the overall renal morphology and outcome risk that was not apparent from the conventional clinical stratification. Taken together, these findings suggest that narrative biopsy reports preserve contextual information on lesion severity, spatial distribution, and cross-compartment co-occurrence that is only partially represented by categorical scoring systems, but can be systematically recovered through language-based representation learning.
The stability and reproducibility of this morphology-informed stratification were further supported by agreement analyses across institutions. Risk propensity estimates derived from LLM embedding-based similarity showed higher interhospital and intermethod agreement than those based on observed MEST-C score distributions, indicating that LLM-derived pathological representations may provide a more stable and transferable abstraction of renal morphology. In this sense, the framework mitigates variation in the descriptive emphasis while preserving the contextual richness inherent to free-text reports, offering a scalable approach for pathology-informed stratification across centers. This point is particularly relevant to treatment stratification. Current decisions in the treatment of IgAN remain driven largely by proteinuria, the estimated glomerular filtration rate, and overall clinical risk, and the available prognostic tools, including the International IgAN Prediction Tool, were developed to estimate the progression risk rather than to select a specific therapy2. The KDIGO 2025 guidelines similarly emphasize biopsy scoring and clinical risk assessments for prognostication but do not recommend these tools to determine the likely benefit of a particular immunosuppressive regime33. Against this background, the interaction between pathological subtype and corticosteroid exposure is an observation of potential clinical relevance, suggesting that the derived classification may capture heterogeneity in corticosteroid-associated outcome patterns that is not fully reflected by conventional clinicopathological assessment.
Given the retrospective nature of these analyses, these findings should be interpreted with caution. In IPTW-adjusted models and subsequent sensitivity analyses, corticosteroid exposure was associated with more favorable kidney outcome patterns in patients with the high-injury subtype, whereas no clear favorable association was observed in patients with the low-injury subtype. These observations are consistent with the treatment heterogeneity reported in prior clinical studies, including the TESTING trial, in which corticosteroids improved kidney outcomes at the population level, while conventional histological scoring did not clearly delineate treatment-responsive pathological strata11,14. As noted by Cook et al., the future value of kidney biopsy classification of IgAN may lie not only in prognostic stratification but also in identifying pathological features that are informative for the therapeutic response, particularly by distinguishing potentially reversible injury from irreversible injury34. In this context, active tubulointerstitial injury may represent a treatment-relevant component of the disease burden that is not fully captured by the current categorical scoring frameworks. Rather than suggesting that the subtype assignment should guide treatment in current practice, our findings raise the possibility that the pathological substrate of steroid-associated heterogeneity extends beyond active glomerular lesions to include the immune and injury milieu of the tubulointerstitial compartment. This interpretation is biologically plausible. Although IgAN is fundamentally a glomerular immune complex-mediated disease, therapeutic decision-making has traditionally focused more heavily on active glomerular lesions, such as endocapillary hypercellularity and crescents. Previous studies suggest that patients with histologically active lesions may show different responses to immunosuppressive therapy35. However, glucocorticoids exert broad effects on lymphocyte- and monocyte/macrophage-mediated inflammatory pathways36–38, suggesting that the tubulointerstitial inflammatory burden may represent a complementary pathological substrate of treatment heterogeneity beyond glomerular descriptors alone. Our findings, therefore, support the idea that closer attention should be paid to the tubulointerstitial compartment in future prospective studies evaluating pathology-informed treatment-response heterogeneity in IgAN.
An additional implication of this work concerns the role of LLMs in kidney pathology research. In IgAN, much of the clinically meaningful morphological information in routine biopsy reports exists in narrative form, including lesion severity, distribution, co-occurrence, and contextual qualifiers that are difficult to encode manually and are only partially represented by categorical scores. Recent studies have shown that large language models and natural language processing approaches can extract structured clinical information from free-text medical documents, including kidney biopsy pathology reports, at scale19,22,39. By explicitly structuring feature extraction across key anatomical compartments and integrating these representations into a unified embedding space, our framework preserves expert-defined pathological knowledge while enabling clinically meaningful inference from routine pathology text without reannotation or additional imaging infrastructure.
Collectively, these findings position LLM-derived pathological analysis not as a replacement for established histological systems or expert renal pathological interpretation but as a principled extension that preserves expert-defined knowledge while enabling reproducible quantification of clinically relevant information embedded in routine narrative reports. The framework offers higher discriminatory resolution than does MEST-C alone and complements existing clinical and pathological risk models. However, direct model-comparison analyses indicated that the incremental gain in individualized prediction and decision-analytic utility beyond that provided by clinical variables and MEST-C was modest. Therefore, the LLM-derived pathological subtype should be interpreted as a complementary pathology-informed stratifier and hypothesis-generating tool rather than as a replacement for established risk models or careful expert review of complete pathology reports. Future studies in which this framework is directly compared with blinded expert reinterpretation of full biopsy reports are needed to define its added clinical value in routine practice.
Several limitations should be considered. First, the retrospective design inherently restricts causal inference, and the treatment-related findings should be regarded as exploratory and hypothesis-generating. Therefore, the observed subtype-by-corticosteroid interaction should not be used to guide treatment decisions without its validation in prospective cohorts or randomized trial datasets. Second, although the external cohort supports a degree of transportability, both cohorts arose from a broadly similar regional context, and broader validation across more diverse populations, reporting conventions, and languages will be required. Third, the interpretability of LLM-derived representations, although substantially improved by the additional analyses, remains incomplete and still depends on the completeness and quality of the source reports. Finally, electron microscopy data were unavailable in this study, and steroid treatment protocols were not fully standardized across centers. In addition, we did not directly compare the LLM-derived pathological subtypes framework with blinded expert reinterpretation of complete pathology reports. Therefore, although the framework provides a scalable and reproducible representation of narrative pathology semantics, how far its clinical value extends beyond meticulous expert review remains to be established.
Overall, our results suggest that LLM-based analysis of narrative kidney biopsy reports can recover a clinically and pathologically interpretable tubulointerstitial injury axis that complements existing pathological and clinical risk frameworks in IgAN. By showing that routine narrative reports contain prognostic information that extends beyond structured scoring and may relate to heterogeneity in corticosteroid-associated outcome patterns, this work supports further investigation of pathology-informed, biologically grounded risk stratification frameworks in IgAN.
Methods
Ethics statement
This retrospective study was conducted in accordance with the Declaration of Helsinki and was approved by the Ethics Committee of the First Affiliated Hospital of Wenzhou Medical University (approval no. KY2024-R337) and the Ethics Committee of the Second Affiliated Hospital of Wenzhou Medical University (approval no. 2025-K436-01). Both ethics committees waived the requirement for informed consent because: (i) the study involved a retrospective secondary analysis of existing clinical records and pathology reports; (ii) all clinical data and narrative pathology reports were deidentified before they were accessed by the research team and were linked solely using study-specific identifiers; (iii) the investigators, including the corresponding authors, had no access to the linkage key or any other information that could be used to identify individual patients; (iv) the study involved no intervention, additional examination, or further contact with patients and posed no more than minimal risk; and (v) obtaining individual informed consent from all patients included in these long-term historical cohorts was impracticable. The ethics committees determined that the waiver would not adversely affect the rights, privacy, or welfare of the participants.
Study population
This multicenter retrospective study was conducted at the First Affiliated Hospital of Wenzhou Medical University (January 2011-March 2025) and the Second Affiliated Hospital of Wenzhou Medical University (July 2012-February 2025). Pathology reports were provided by the Department of Renal Pathology at the First Affiliated Hospital and by KingMed Diagnostics Center (Guangzhou KingMed Diagnostics Group Co., Ltd.) for the Second Affiliated Hospital. Patients with primary IgAN from the First Affiliated Hospital formed the model development cohort, whereas those from the Second Affiliated Hospital constituted an independent external validation cohort that was not included in the model training. Pathological subtype classification derived from the development cohort was subsequently analyzed for clinical significance using follow-up data.
For both cohorts, the inclusion criteria were as follows: (1) availability of a complete initial kidney biopsy report and (2) age greater than 14 years at the time of biopsy. The exclusion criteria were secondary IgAN, kidney transplant diagnosis, membranous nephropathy or minimal change disease with IgA deposition, concurrent malignancies, or missing critical data.
Predictors and outcome definitions
Baseline data at biopsy included demographic characteristics, including age and sex, comorbidities (hypertension and diabetes), blood pressure, serum creatinine level, and 24-h urinary protein level (or spot protein-to-creatinine ratio). Sex was recorded from routine clinical records at the time of biopsy as a baseline demographic variable; gender identity was not routinely collected and was therefore not analyzed. The estimated glomerular filtration rate (eGFR) was calculated using the Chronic Kidney Disease Epidemiology Collaboration (CKD-EPI) equation. Histopathologic Oxford MEST-C scores were extracted from reports. The medications used included renin-angiotensin system blockers, corticosteroids, and pooled immunosuppressive therapy (mycophenolate mofetil, cyclophosphamide, tacrolimus, azathioprine, and leflunomide). Early immunosuppressive therapy was defined as the initiation of immunosuppressants within 1 year after biopsy. In accordance with KDIGO guidelines33, patients were classified as high risk under proteinuria-based clinical risk stratification if 24-h protein excretion was ≥0.5 g at the time of biopsy. Sex- and gender-based analyses were not prespecified in this retrospective study.
The primary outcome was a composite kidney endpoint, defined as any of the following: sustained eGFR <15 mL/min/1.73 m2, initiation of maintenance dialysis, kidney transplantation, or a confirmed ≥40% decrease in eGFR from baseline.
LLM-based pathological feature extraction and safeguards
Pathological features were extracted from narrative kidney biopsy reports using a structured LLM. Specifically, the DeepSeek-V3 model40 was accessed via its API to perform automated extraction. A total of 71 discrete pathological elements across six domains were identified through iterative review with our nephropathologists (D.L., B.C. and Y.L.), including Oxford MEST-C components; detailed descriptions of glomerular, tubulointerstitial, and vascular lesions; immunofluorescence findings; and a summary of overall inflammatory activity (Supplementary Fig. 6). These elements were selected to cover the large majority of features routinely documented in clinical kidney biopsy reports and broadly align with the Oxford classification framework. Because the framework operates on source text reports, the systematic analysis was limited to features that were explicitly documented in those reports. Accordingly, pathological findings that were not routinely described in historical reports could not be robustly evaluated in the present study.
Textual descriptions of four histopathological domains (glomerular, tubulointerstitial, vascular, and immunofluorescence staining) were flattened and encoded using the Qwen3-Embedding model (8B parameter version, quantized to Q4_K_M)41 via the Ollama framework (version 0.18.0) to generate a unified numerical representation. These vectors were concatenated to form a single 16,384-dimensional feature vector for each biopsy report. The MEST-C score and inflammatory activity summary were not separately embedded, as they were derived from and inherently encapsulated within the corresponding histopathological descriptions.
The LLM and embedding model were applied in a frozen configuration without task‑specific fine‑tuning. Flattened features, which were derived from the extraction of pathology report texts, served as the model input. Unsupervised clustering was performed solely on the LLM‑derived embeddings and was conducted independently of downstream clinical analyses.
For external validation, the same pretrained models and identical processing pipeline were applied without retraining or recalibration, ensuring that the pathological subtype assignment in the validation cohort was fully independent of the clinical outcomes.
Unsupervised clustering and supervised machine learning
We applied a two-step principal component analysis (PCA) procedure to reduce the dimensionality of the high-dimensional concatenated embedding vectors. Each biopsy report was represented by a 16,384-dimensional embedding vector, which was first projected into an intermediate 2500-dimensional space and then further reduced to a final 300-dimensional representation using PCA with a randomized singular value decomposition solver. The first PCA step served as a high-retention compression step to reduce redundancy and collinearity in the original embedding space, whereas the second PCA step generated a compact representation for the downstream clustering analysis. No clinical outcomes or downstream labels were used in either PCA step. We evaluated several unsupervised clustering algorithms, including k-means, spectral clustering, hierarchical clustering, DBSCAN, and HDBSCAN. The primary clustering solution was selected based on the overall interpretability and internal clustering performance, with silhouette analysis used as an initial model selection criterion.
We performed complementary stability analyses to assess the robustness of the selected solution. First, we conducted a subsampling-based consensus analysis using 30 repeated iterations of 80% random subsampling. For each iteration, PCA and clustering were repeated within the subsample, and agreement with the reference solution was quantified using ARI and NMI. We additionally summarized within-cluster consensus, between-cluster consensus, and the proportion of ambiguous clustering (PAC; defined over the consensus interval of 0.1–0.9). Then, we assessed the robustness of the results to perturbations by adding mild Gaussian noise to the embedding representation (noise standard deviation ratio 0.01) across 30 iterations and recalculating the cluster assignments, which again were summarized using ARI and NMI relative to the reference clustering. These analyses were performed to determine whether the observed two-cluster structure was reproducible across resampling, robust to mild perturbations, and not solely dependent on a single clustering algorithm.
For pathological subtype prediction model development (First Affiliated Hospital), we evaluated eight machine learning algorithms: logistic regression, random forest, support vector machine (SVM), k-nearest neighbors, light gradient boosting machine (LightGBM), neural network, extreme gradient boosting (XGBoost), and naïve Bayes. The three highest-performing models were combined into an ensemble classifier. Model performance was assessed using accuracy, the area under the receiver operating characteristic curve (AUC), and five-fold cross-validation. The final ensemble model was applied to predict pathological subtypes in an independent external validation cohort (Second Affiliated Hospital).
Integrated interpretability analysis of embedding-derived pathological subtypes
Interpretability analyses were performed post hoc to explain, rather than derive, the embedding-derived subtype assignments. Precomputed subtype labels were aligned to report embeddings and structured pathology JSON files using unique case identifiers. For each pathology report, section-specific text was extracted from the structured JSON representation and embedded independently using the Qwen3-embedding:latest model; section embeddings were then concatenated to generate the full-report representation used for the subtype analysis.
We implemented a teacher-student framework in which subtype labels served as the teacher signal and a random forest classifier trained on structured pathological features served as the student model to translate the embedding-derived subtypes into clinically interpretable pathological variables. Numeric variables were median-imputed. Categorical variables were pooled for infrequent levels, encoded with an explicit missing category, and one-hot encoded using a ColumnTransformer42. SHAP values were computed with TreeExplainer43 and summarized at both the encoded feature and raw feature levels by aggregating one-hot encoded levels back to their parent pathological descriptors.
Counterfactual perturbation analyses were then performed in the embedding space to quantify the contributions of individual descriptors and grouped pathology blocks to subtype separation. We evaluated three perturbation schemes: single-feature deletion, opposite-subtype prototype swapping, and grouped block replacement. Perturbation effects were quantified as the change in projection along the subtype separation axis, defined as the vector connecting the mean embeddings of the two subtypes. Subtype-specific prototypes were constructed from the most frequent non-empty value of each section-feature pair within each subtype. For key ordinal features, encoded-level SHAP decomposition and subtype-specific value distributions were further examined to determine whether subtype separation was driven by the presence/absence of lesions or by graded severity. Finally, feature rankings from SHAP, deletion, and swap analyses were normalized and integrated in a cross-method concordance heat map.
To assess whether numerical clustering stability was accompanied by stable pathological interpretation, we evaluated the consistency of subtype-defining pathological drivers across teacher-student SHAP attribution, encoded-level decomposition, feature- and block-level counterfactual perturbation, and cross-method concordance analyses.
Quantitative association of Oxford MEST-C patterns with clinical risk stratification and pathological subtypes
We investigated the quantitative relationships between the Oxford MEST-C histologic patterns, clinical risk stratification, and LLM-derived pathological subtypes. All observed combinations of MEST-C components were enumerated in the study cohort. Multivariable logistic regression models were applied to evaluate associations between individual MEST-C components (M, E, S, T, and C), as well as composite MEST-C patterns, and (1) clinical risk categories or (2) pathological subtypes. Effect estimates are reported as odds ratios with 95% confidence intervals. The distributions of patients across MEST-C combinations were summarized according to both clinical risk stratification and pathological subtype classification.
To quantify morphological similarity, each MEST-C pattern was encoded as a numerical embedding using Qwen3-Embedding. In parallel, comprehensive renal pathology representations reflecting overall renal morphology were derived from full biopsy reports using structured LLM-derived feature extraction, yielding high-dimensional pathological feature embedding vectors (16,384 dimensions). Pairwise Euclidean distances between MEST-C pattern embeddings and pathology feature embedding vectors were calculated, with shorter distances indicating greater morphological similarity. Distance distributions were stratified by clinical risk group and pathological subtype and visualized using density ridge plots to assess the alignment of MEST-C patterns with low- and high-risk phenotypes.
Agreement analysis of risk propensity across MEST-C patterns, hospitals, and stratification methods
We evaluated the robustness and reproducibility of morphology-based risk stratification across two independent cohorts (derivation and external validation) and two stratification approaches (pathological subtype-based and clinical risk-based). For each MEST-C pattern, risk propensity was quantified using two complementary metrics: (1) the proportion of patients classified as high risk and (2) the Euclidean distance between pattern-specific centroids and high-risk pathological centroids within the multidimensional morphology space.
Two strategies for MEST-C pattern construction were compared: expert-defined (“ground-truth”) patterns and embedding-based similarity-derived patterns. A total of twelve predefined pairwise comparisons were performed across three domains: (1) comparisons of stratification methods within the same hospital, (2) comparisons of MEST-C construction strategies within the same hospital, and (3) comparisons of the same MEST-C condition and stratification method across hospitals. Agreement was quantified using Fleiss’ κ statistics with 95% confidence intervals and interpreted according to the Landis-Koch criteria (slight, fair, moderate, substantial, and almost perfect agreement).
Associations of pathological subtypes with clinical outcomes and immunosuppressive therapy
We evaluated differences in clinical outcomes across pathological subtypes and clinical categories. We assessed whether the effect of immunosuppressive treatment (corticosteroids and other immunosuppressants) differed across these subgroups by conducting interaction analyses within a Cox proportional hazards framework.
Potential confounding in treatment assignment was addressed using inverse probability of treatment weighting (IPTW). Propensity scores were estimated using multivariable logistic regression models that included age, sex, baseline estimated glomerular filtration rate (eGFR), baseline 24‑h proteinuria, and MEST‑C scores as covariates. Stabilized weights were calculated from the average propensity scores across imputed datasets, followed by trimming at the 1st and 99th percentiles to limit extreme weights.
Within the IPTW-weighted sample, we fitted two sets of Cox models for each treatment (corticosteroids and other immunosuppressants). First, a base model included the main effects of the pathological subtypes and treatment, as well as their interaction term. Second, a fully adjusted model additionally incorporated age, sex, baseline 24‑h proteinuria, and MEST-C scores as covariates. These models allowed us to formally test for interactions by evaluating the statistical significance of the interaction term and to estimate subgroup‑specific hazard ratios.
In an independent external validation cohort, we performed survival analyses using the derived pathological subtypes and further characterized differences in baseline clinical characteristics across subtypes. We also examined the relationships between these subtypes and standardized MEST-C patterns.
Statistical analysis
Continuous variables are presented as means (standard deviations) or medians (interquartile ranges), as appropriate, and categorical variables are presented as counts (percentages). Between-group comparisons were performed using Student’s t-test or the Mann-Whitney U-test for continuous variables and the χ2-test or Fisher’s exact test for categorical variables, depending on data distribution and expected cell counts.
Pairwise agreement between classification strategies was assessed using Fleiss’ kappa coefficient with asymptotic standard errors. Time-to-event outcomes were evaluated using Kaplan-Meier survival curves and compared with the log-rank test. Multivariable Cox proportional hazards models were fitted to assess associations with kidney disease progression and to evaluate interaction effects. The proportional hazards assumption was assessed using Schoenfeld residuals and was not violated.
To assess the incremental prognostic value of LLM-derived pathological subtypes beyond the Oxford MEST-C score, we compared prespecified Cox models including clinical variables alone, clinical variables plus MEST-C, clinical variables plus LLM-derived pathological subtypes, and clinical variables plus both MEST-C and LLM-derived pathological subtypes. Model fit was evaluated using the Akaike information criterion, Bayesian information criterion, and likelihood ratio tests. Discrimination was assessed using Harrell’s C-index. Changes in the C-index (ΔC-index) were estimated using 1000 patient-level bootstrap resamples, with 95% confidence intervals derived from the 2.5th and 97.5th percentiles of the bootstrap distribution. Time-specific decision curve analyses were performed at 36 and 60 months.
Multivariable models were constructed based on clinical relevance and prior knowledge. Sex was recorded as a baseline demographic variable and included as a covariate in multivariable models to account for potential confounding. Because this retrospective study was not designed to test sex-specific hypotheses, sex-disaggregated analyses were not prespecified. Missing data were infrequent and were handled using complete-case analysis; sensitivity analyses yielded consistent results. All statistical tests were two-sided, and P < 0.05 was considered statistically significant. Statistical analyses were performed using R version 4.4.3 and Python version 3.13.
Reporting summary
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Supplementary information
Source data
Acknowledgments
We acknowledge the International IgA Nephropathy Network for its foundational contributions to the study of IgA nephropathy. We thank the staff of the nephrology departments of the First and Second Affiliated Hospitals of Wenzhou Medical University for their contributions to clinical care, data curation, and long-term patient follow-up. We also acknowledge the developers of publicly available large language model-based tools, including Qwen and DeepSeek, which enabled the scalable analysis of narrative pathology reports in this study.
Author contributions
Ji Zhang conceived and designed the study. Ji Zhang and J.L. developed the large language model-based analytical framework and performed the statistical analyses. Ji Zhang, J.L., L.J., and Jiatian Zhang contributed to data acquisition, curation, and preprocessing. Q.Z., S.C., D.L., F.H., B.C., Y.L., and Z.S. contributed to clinical data collection, pathological interpretation, and patient follow-up. Ji Zhang drafted the manuscript. J.L., L.J., and Jiatian Zhang contributed to data visualization and figure preparation. Y.Z., C.C., and M.P. provided critical supervision, interpreted the results, and revised the manuscript for important intellectual content. All authors reviewed and approved the final manuscript.
Peer review
Peer review information
Nature Communications thanks Haotian Lin and the other anonymous reviewer(s) for their contribution to the peer review of this work. A peer review file is available.
Funding
This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.
Data availability
The source data underlying the figures and tables are provided with this paper. The deidentified individual-level clinical data analyzed in this study are available under restricted access because they are subject to patient privacy protections, ethics oversight, and institutional data-governance requirements. Researchers may request access to these data by submitting a scientifically justified proposal to the corresponding authors. The proposal should specify the research purpose, study design, requested data elements, intended use, and proposed data-security measures. All requests will be reviewed by the relevant institutional ethics committee and the data-holding institution in accordance with applicable ethical, privacy, data-security, and institutional requirements. Applicants will normally receive a decision within approximately six weeks after submission of a complete request. Before data release, approved applicants will be required to enter into a data-use agreement restricting use of the data to the approved purpose, requiring secure data storage and processing, prohibiting any attempt to reidentify participants, and prohibiting unauthorized transfer or disclosure of the data to third parties. The corresponding authors will serve solely as points of contact and will facilitate communication between applicants and the relevant institutional bodies. They will not exercise personal or discretionary authority over access decisions. Access will not be conditional on authorship, collaboration with the study investigators, or any other non-objective requirement. Access to the data will be available for a period of 10 years following the publication date of this article. Source data are provided with this paper.
Code availability
The source code used for large language model–based feature extraction, unsupervised clustering, pathological subtype classification, model validation, and clinical association analyses is publicly available under the MIT License at GitHub (https://github.com/zhji0426/LLM-for-pathological-subtypes). A permanent, archived version of the code used in this study is available through Zenodo at https://doi.org/10.5281/zenodo.20777820. A user-friendly web interface for predicting pathological subtypes from narrative kidney pathology reports is available at http://www.igan123.cn and will remain publicly accessible after publication, subject to temporary interruptions for routine maintenance or technical updates. All analyses were conducted using publicly available software, and the repository contains the scripts required to reproduce the principal analyses and results reported in this study.
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
These authors contributed equally: Ji Zhang, Jiadan Lu.
Contributor Information
Yu Zheng, Email: zheng_yu2@hotmail.com.
Chaosheng Chen, Email: wzccs8@163.com.
Min Pan, Email: pm_1234@163.com.
Supplementary information
The online version contains supplementary material available at https://doi.org/10.1038/s41467-026-76326-5.
References
- 1.Floege, J., Bernier-Jean, A., Barratt, J. & Rovin, B. Treatment of patients with IgA nephropathy: a call for a new paradigm. Kidney Int.107, 640–651 (2025). [DOI] [PubMed] [Google Scholar]
- 2.Barbour, S. J. et al. Evaluating a new international risk-prediction tool in IgA nephropathy. JAMA Intern. Med.179, 942–952 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Rovin, B. H. et al. KDIGO 2021 clinical practice guideline for the management of glomerular diseases. Kidney Int.100, S1–S276 (2021). [DOI] [PubMed] [Google Scholar]
- 4.Cattran, D. C., Floege, J. & Coppo, R. Evaluating progression risk in patients with immunoglobulin A nephropathy. Kidney Int. Rep.8, 2515–2528 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Working Group of the International IgAN et al. The Oxford classification of IgA nephropathy: rationale, clinicopathological correlations, and classification. Kidney Int.76, 534–545 (2009). [DOI] [PubMed] [Google Scholar]
- 6.Roberts, I. S. et al. The Oxford classification of IgA nephropathy: pathology definitions, correlations, and reproducibility. Kidney Int.76, 546–556 (2009). [DOI] [PubMed] [Google Scholar]
- 7.Trimarchi, H. et al. Oxford classification of IgA nephropathy 2016: an update from the IgA Nephropathy Classification Working Group. Kidney Int.91, 1014–1021 (2017). [DOI] [PubMed] [Google Scholar]
- 8.Lee, S. M. IgA nephropathy: morphologic predictors of progressive renal disease. Hum. Pathol.13, 314–322 (1982). [DOI] [PubMed] [Google Scholar]
- 9.Haas, M. Histologic subclassification of IgA nephropathy a clinicopathologic study of 244 cases. Am. J. Kidney Dis.29, 829–842 (1997). [DOI] [PubMed] [Google Scholar]
- 10.SM. Kurt, L. ee Prognostic indicators of progressive renal disease in IgA nephropathy emergence of a new histologic grading system. Am. J. Kidney Dis.29, 953–958 (1997). [DOI] [PubMed] [Google Scholar]
- 11.Lv, J. et al. Effect of oral methylprednisolone on clinical outcomes in patients with IgA nephropathy: the TESTING Randomized Clinical Trial. JAMA318, 432–442 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Rauen, T. et al. Intensive supportive care plus Immunosuppression in IgA nephropathy. N. Engl. J. Med.373, 2225–2236 (2015). [DOI] [PubMed] [Google Scholar]
- 13.Rauen, T. et al. After ten years of follow-up, no difference between supportive care plus immunosuppression and supportive care alone in IgA nephropathy. Kidney Int.98, 1044–1052 (2020). [DOI] [PubMed] [Google Scholar]
- 14.Lv, J. et al. Effect of oral methylprednisolone on decline in kidney function or kidney failure in patients with IgA nephropathy: the TESTING Randomized Clinical Trial. JAMA327, 1888–1898 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Chen, P. & Lv, J. Low-dose glucocorticoids, mycophenolate mofetil and hydroxychloroquine in IgA nephropathy, an update of current clinical trials. Nephrology (Carlton)29, 25–29 (2024). [DOI] [PubMed] [Google Scholar]
- 16.Rauen, T., Eitner, F., Fitzner, C. & Floege, J. Con: STOP immunosuppression in IgA nephropathy. Nephrol. Dial. Transplant.31, 1771–1774 (2016). [DOI] [PubMed] [Google Scholar]
- 17.Schena, F. P. & Manno, C. Intensive supportive care plus immunosuppression in IgA nephropathy. N Engl J Med374, 992 (2016). [DOI] [PubMed] [Google Scholar]
- 18.Rovin, B. H. et al. Executive summary of the KDIGO 2021 Guideline for the Management of Glomerular Diseases. Kidney Int.100, 753–779 (2021). [DOI] [PubMed] [Google Scholar]
- 19.Lee, D. et al. Using large language models to automate data extraction from surgical pathology reports: retrospective cohort study. JMIR Form Res.9, e64544 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Di Palma, L., Darvizeh, F., Alì, M. & Fazzini, D. Structured transformation of unstructured prostate MRI reports using large language models. Tomography11 10.3390/tomography11060069 (2025). [DOI] [PMC free article] [PubMed]
- 21.Alzaid, E., Pergola, G., Evans, H., Snead, D. & Minhas, F. Large multimodal model-based standardisation of pathology reports with confidence and its prognostic significance. J .Pathol. Clin. Res.10, e70010 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Choi, H. S., Song, J. Y., Shin, K. H., Chang, J. H. & Jang, B. S. Developing prompts from a large language model for extracting clinical information from pathology and ultrasound reports in breast cancer. Radiat. Oncol. J.41, 209–216 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Toal, M., Hill, C., Quinn, M., O’Neill, C. & Maxwell, A. P. Large language models’ clinical decision-making on when to perform a kidney biopsy: comparative study. J. Med. Internet Res.27, e73603 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Iqbal, U. et al. Impact of large language model (ChatGPT) in healthcare: an umbrella review and evidence synthesis. J. Biomed. Sci.32, 45 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Luo, I. et al. Leveraging large language models to extract smoking history from clinical notes for lung cancer surveillance. NPJ Digit. Med.8, 731 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Soares, M. F. & Roberts, I. S. IgA nephropathy: an update. Curr. Opin. Nephrol. Hypertens.26, 165–171 (2017). [DOI] [PubMed] [Google Scholar]
- 27.Lv, J. et al. Evaluation of the Oxford classification of IgA nephropathy: a systematic review and meta-analysis. Am. J. Kidney Dis.62, 891–899 (2013). [DOI] [PubMed] [Google Scholar]
- 28.Coppo, R. et al. Validation of the Oxford classification of IgA nephropathy in cohorts with different presentations and treatments. Kidney Int.86, 828–836 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 29.Bellur, S. S., Troyanov, S., Vorobyeva, O., Coppo, R. & Roberts, I. S. D. Evidence from the large VALIGA cohort validates the subclassification of focal segmental glomerulosclerosis in IgA nephropathy. Kidney Int.105, 1279–1290 (2024). [DOI] [PubMed] [Google Scholar]
- 30.Howie, A. J. & Lalayiannis, A. D. Systematic review of the Oxford classification of IgA nephropathy: reproducibility and prognostic value. Kidney3604, 1103–1111 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Roberts, I. S. Pathology of IgA nephropathy. Nat. Rev. Nephrol.10, 445–454 (2014). [DOI] [PubMed] [Google Scholar]
- 32.Rankin, A. J. et al. Assessment of active tubulointerstitial nephritis in non-scarred renal cortex improves prediction of renal outcomes in patients with IgA nephropathy. Clin. Kidney J.12, 348–354 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Floege, J. et al. Executive summary of the KDIGO 2025 Clinical Practice Guideline for the Management of Immunoglobulin A Nephropathy (IgAN) and Immunoglobulin A Vasculitis (IgAV). Kidney Int.108, 548–554 (2025). [DOI] [PubMed] [Google Scholar]
- 34.Cook, H. T. Interpretation of renal biopsies in IgA nephropathy. Contrib. Nephrol.157, 44–49 (2007). [DOI] [PubMed] [Google Scholar]
- 35.Hou, J. H. et al. Mycophenolate mofetil combined with prednisone versus full-dose prednisone in IgA nephropathy with active proliferative lesions: a randomized controlled trial. Am. J. Kidney Dis.69, 788–795 (2017). [DOI] [PubMed] [Google Scholar]
- 36.Shimba, A. & Ikuta, K. Control of immunity by glucocorticoids in health and disease. Semin. Immunopathol.42, 669–680 (2020). [DOI] [PubMed] [Google Scholar]
- 37.Ehrchen, J. M., Roth, J. & Barczyk-Kahlert, K. More than suppression: glucocorticoid action on monocytes and macrophages. Front. Immunol.10, 2028 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Cain, D. W. & Cidlowski, J. A. Immune regulation by glucocorticoids. Nat. Rev. Immunol.17, 233–247 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Woźnicki, P. et al. Automatic structuring of radiology reports with on-premise open-source large language models. Eur. Radiol.35, 2018–2029 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Guo, D. et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature645, 633–638 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Zhang, Y. et al. Qwen3 embedding: advancing text embedding and reranking through foundation models. Preprint at arXiv 10.48550/arXiv.2506.05176 (2025). [DOI]
- 42.Pedregosa, F. et al. Scikit-learn: machine learning in Python. J. Mach. Learn. Res.12, 2825–2830 (2011). [Google Scholar]
- 43.Lundberg, S. M. & Lee, S.-I. A unified approach to interpreting model predictions. Advances in neural information processing systems30 (2017).
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The source data underlying the figures and tables are provided with this paper. The deidentified individual-level clinical data analyzed in this study are available under restricted access because they are subject to patient privacy protections, ethics oversight, and institutional data-governance requirements. Researchers may request access to these data by submitting a scientifically justified proposal to the corresponding authors. The proposal should specify the research purpose, study design, requested data elements, intended use, and proposed data-security measures. All requests will be reviewed by the relevant institutional ethics committee and the data-holding institution in accordance with applicable ethical, privacy, data-security, and institutional requirements. Applicants will normally receive a decision within approximately six weeks after submission of a complete request. Before data release, approved applicants will be required to enter into a data-use agreement restricting use of the data to the approved purpose, requiring secure data storage and processing, prohibiting any attempt to reidentify participants, and prohibiting unauthorized transfer or disclosure of the data to third parties. The corresponding authors will serve solely as points of contact and will facilitate communication between applicants and the relevant institutional bodies. They will not exercise personal or discretionary authority over access decisions. Access will not be conditional on authorship, collaboration with the study investigators, or any other non-objective requirement. Access to the data will be available for a period of 10 years following the publication date of this article. Source data are provided with this paper.
The source code used for large language model–based feature extraction, unsupervised clustering, pathological subtype classification, model validation, and clinical association analyses is publicly available under the MIT License at GitHub (https://github.com/zhji0426/LLM-for-pathological-subtypes). A permanent, archived version of the code used in this study is available through Zenodo at https://doi.org/10.5281/zenodo.20777820. A user-friendly web interface for predicting pathological subtypes from narrative kidney pathology reports is available at http://www.igan123.cn and will remain publicly accessible after publication, subject to temporary interruptions for routine maintenance or technical updates. All analyses were conducted using publicly available software, and the repository contains the scripts required to reproduce the principal analyses and results reported in this study.
