Skip to main content
PLOS One logoLink to PLOS One
. 2026 May 6;21(5):e0344600. doi: 10.1371/journal.pone.0344600

Robust disease prognosis via diagnostic knowledge preservation: A sequential learning approach

Haresh Rengaraj Rajamohan 1,*, Yanqi Xu 1, Weicheng Zhu 1, Richard Kijowski 2, Kyunghyun Cho 1, Krzysztof J Geras 1,3, Narges Razavian 3,4, Cem M Deniz 3,5
Editor: Edward Hoffer6
PMCID: PMC13148697  PMID: 42090385

Abstract

Accurate disease prognosis is essential for patient care but is often hindered by the scarcity of longitudinal data. This study explores deep learning training strategies that utilize large, accessible diagnostic datasets to pretrain models aimed at predicting future disease progression in knee osteoarthritis (OA), Alzheimer’s disease (AD), and breast cancer (BC). While diagnostic pretraining improves prognostic task performance, naive fine-tuning for prognosis can cause ‘catastrophic forgetting,’ where the model’s original diagnostic accuracy degrades, a significant patient safety concern in real-world settings. To address this, we propose a sequential learning strategy with experience replay. We used cohorts with knee radiographs, brain MRIs, and digital mammograms to predict 4-year structural worsening in OA, 2-year cognitive decline in AD, and 5-year cancer diagnosis in BC. Our results showed that diagnostic pretraining on larger datasets improved prognosis model performance compared to standard baselines, boosting both the Area Under the Receiver Operating Characteristic curve (AUROC) (e.g., Knee OA external: 0.770 vs 0.747; Breast Cancer: 0.874 vs 0.848) and the Area Under the Precision-Recall Curve (AUPRC) (e.g., Alzheimer’s Disease: 0.752 vs 0.683). Additionally, a sequential learning approach with experience replay achieved prognostic performance comparable to dedicated single-task models (e.g., Breast Cancer AUROC 0.876 vs 0.874) while also preserving diagnostic ability. This method maintained high diagnostic accuracy (e.g., Breast Cancer Balanced Accuracy 50.4% vs 50.9% for a dedicated diagnostic model), unlike simpler multitask methods prone to catastrophic forgetting (e.g., 37.7%). Our findings show that leveraging large diagnostic datasets is a reliable and data-efficient way to enhance prognostic models while maintaining essential diagnostic skills.

1. Introduction

Progressive diseases present significant challenges to healthcare systems worldwide. Their gradual onset and potential for irreversible damage make early intervention, guided by accurate prognosis, crucial for optimizing patient outcomes, reducing healthcare costs, and improving quality of life [1–3]. Key examples include knee osteoarthritis (OA), Alzheimer’s disease (AD), and breast cancer (BC) – conditions representing diverse physiological systems but sharing the critical need for early identification and reliable progression prediction [1,3–5].

Deep learning (DL) [6] offers powerful tools for automating medical image analysis and has shown considerable promise for both diagnosis and prognosis tasks across various diseases. Studies have demonstrated DL’s effectiveness in diagnosing OA severity [7], detecting AD from neuroimaging [8–11], and screening for BC [12,13]. Similarly, DL models have been developed to predict disease progression, such as total knee replacement risk in OA [14–18], progression to AD [19,20], and long-term BC outcomes [21,22].

However, developing robust prognosis models faces a significant hurdle: the scarcity of longitudinal data required to track disease progression over time. This contrasts with diagnosis, where larger cross-sectional datasets are often more readily available. As a result, recent advances in medical imaging AI – across knee osteoarthritis, Alzheimer’s disease, and oncology have largely focused on improving diagnostic accuracy and disease severity assessment, reflecting the scale and availability of diagnostic datasets rather than long-horizon progression labels [23–25]. To mitigate this data scarcity, previous works have explored multitask learning (MTL) as additional regularization, where a single model is trained to perform several related tasks simultaneously, aiming to improve generalization and performance through shared representations [13,14,16,21,22]. For instance, MTL combining diagnosis and prognosis has shown benefits in OA [14,16], and pretraining on BI-RADS has been shown to improve current cancer diagnosis performance [13]. Yet these MTL studies typically rely on relatively small prognosis datasets and do not fully capitalize on the abundance of diagnostic data, even as recent medical foundation-model approaches increasingly leverage large-scale diagnostic pretraining with limited labeled data [26,27]. This study therefore hypothesizes that sequential learning – diagnosis pretraining on large, unbiased datasets followed by prognostic training, achieves a favorable balance of prognostic and diagnostic performance.

A key challenge arises, however, when building a single model that learns tasks sequentially. When a model pretrained on a data-rich task (e.g., diagnosis) is subsequently fine-tuned on a new, data-scarce task (e.g., prognosis), it often suffers from catastrophic forgetting, a significant degradation of its ability to perform the original task [28]. In the context of healthcare, this phenomenon is not merely a technical limitation but a critical issue for real-world model maintenance and patient safety. A prognostic model that compromises diagnostic accuracy could lead to missed disease during routine follow-up, undermining its clinical utility and trust. A primary strategy to combat this phenomenon is experience replay (or rehearsal), where the model is periodically retrained on stored examples from the original task while learning the new one, thereby refreshing and preserving its previously acquired knowledge [29].

This work demonstrates that pretraining models on larger diagnostic datasets significantly improves prognostic performance, generalization to external cohorts, and discrimination within patient subgroups compared to standard initialization techniques across knee OA, Alzheimer’s disease and breast cancer. We also identify that a sequential multitask learning approach with experience replay yields the best overall results, matching the prognostic performance of dedicated single-task models while effectively preserving diagnostic capabilities and outperforming simpler joint training strategies.

2. Materials and methods

2.1. Ethics statement

This retrospective study involved the analysis of fully de-identified data from three publicly available research datasets (OAI, MOST, and ADNI) and one institutional dataset (NYU Langone Health). The Osteoarthritis Initiative (OAI) (ClinicalTrials.gov identifier: NCT00080171) was approved by the Institutional Review Boards (IRB) at the UCSF Coordinating Center (approval #10–00532) and all clinical sites, including Memorial Hospital of Rhode Island, Ohio State University, University of Pittsburgh, and University of Maryland/Johns Hopkins University. Ethical approval for the Multicenter Osteoarthritis Study (MOST) was obtained from the IRBs at Boston University (H-32956), University of Alabama at Birmingham (IRB-000329007), University of California San Francisco (301480), and University of Iowa (201511711); approval for secondary data analysis was granted by the University of Florida (IRB202201899). Data for the Alzheimer’s Disease Neuroimaging Initiative (ADNI ClinicalTrials.gov identifier: NCT00106899) were obtained in accordance with the Declaration of Helsinki and approved by the IRBs of all participating institutions. All participants in these public cohorts provided written informed consent. The Breast Cancer dataset was sourced from New York University Langone Health under a protocol approved by the NYU Langone Health IRB (IRB00010481, S18-00712). This dataset consisted of fully de-identified data and was approved for retrospective analysis with a waiver of informed consent, in compliance with HIPAA regulations. As this study involved only the retrospective analysis of fully de-identified data, it did not require new participant recruitment or additional IRB approval beyond the existing protocols cited above.

2.2. Prediction task definitions

We investigate the following clinical measures for knee OA, Alzheimer’s disease and breast cancer:

2.2.1. Knee Osteoarthritis (OA).

Diagnosis: Classification based on Kellgren-Lawrence Grade (KLG) [30] – A severity scale (0–4) assigned by radiologists based on radiographic features including joint space narrowing, bone spurs, and sclerosis.

Prognosis: Prediction of disease worsening over a 4-year horizon [31], defined as:

  • Structural Incidence for early-stage OA (KLG 0–1): Progression to KLG ≥ 2 or undergoing a Total Knee Replacement (TKR).

  • Structural Progression for radiographic OA (KLG ≥ 2): An increase in KLG or undergoing a TKR.

2.2.2. Alzheimer’s Disease (AD).

Diagnosis: Classification of cognitive status using structural brain MRI, based on the diagnostic criteria established by the Alzheimer’s Disease Neuroimaging Initiative (ADNI) study protocol:

  • Cognitive Normal (CN): Participants with no significant impairment in cognitive functions or activities of daily living.

  • Mild Cognitive Impairment (MCI): Participants with a subjective memory concern, objective memory loss, and preserved general cognition and functional performance.

  • Alzheimer’s Disease (AD): Participants meeting the NINCDS/ADRDA criteria for probable AD.

Prognosis: Prediction of cognitive decline over a 2-year horizon:

  • CN → MCI: Onset of cognitive impairment.

  • MCI → AD or CN → AD: Progression to confirmed disease.

2.2.3. Breast Cancer (BC).

BI-RADS (Diagnosis): Classification system for breast imaging findings:

  • Grade 0: Incomplete assessment; additional imaging evaluation is needed.

  • Grade 1: Negative; no abnormalities found.

  • Grade 2: Benign finding; no malignant characteristics present.

Prognosis: Prediction of a patient having a confirmed breast cancer diagnosis within 5 years of the scan date.

2.3. Datasets and cohorts

2.3.1. Knee osteoarthritis.

Data were sourced from the Osteoarthritis Initiative (OAI) [32] for model development and the Multicenter Osteoarthritis Study (MOST) [33] for external validation. The OAI is a long-term, multicenter observational study of 4,796 participants focused on knee OA biomarkers. We utilized bilateral posterior-anterior fixed flexion knee radiographs.

  • Diagnosis Cohort (OAI): Comprised 47,041 radiographs from 4,508 subjects with KLG assessments.

  • Prognosis Cohort (OAI): To predict structural worsening, a 4 year follow up from the baseline visit was used. This timeframe was selected to balance observing meaningful change and maximizing cohort size, as longer horizons led to significant participant drop-off. Further, knees with KLG 4 or a TKR at baseline were excluded. Two distinct prognosis tasks and their corresponding cohorts were defined based on baseline disease severity:

    • Structural Incidence: This task focused on predicting the onset of radiographic OA in knees with baseline KLG 0–1. The cohort included 420 patients who progressed (defined as reaching KLG ≥ 2 or undergoing a TKR) and 2,439 non-progressing controls.

    • Structural Progression: This task focused on predicting the worsening of established OA in knees with baseline KLG 2–3. The cohort included 571 patients who progressed (defined as an increase in KLG or undergoing a TKR) and 1,677 non-progressing controls.

  • MOST Validation Cohort: Corresponding cohorts were created from the MOST dataset using a 5-year follow-up due to data availability. The incidence cohort included 505 progressors and 1,395 controls. The progression cohort included 562 progressors and 617 controls.

2.3.2. Alzheimer’s disease.

Data were obtained from the ADNI database [25], a longitudinal multicenter study aimed at validating biomarkers for AD progression. We utilized T1-weighted brain MRI scans.

  • Diagnosis Cohort: Consisted of 2,723 scans from 662 patients classified as CN, MCI or AD.

  • Prognosis Cohort: To predict cognitive decline within a 24-month follow-up, 914 scans from 365 unique patients were used. Progression was defined as a change from CN to MCI/AD or MCI to AD. To augment the dataset size, MRI scans from both baseline and available intermediate follow-up visits were utilized as distinct input time points; however, scans from patients already classified as AD were excluded. To prevent data leakage and ensure the model learned generalizable prognostic features rather than patient-specific anatomy, all scans from a single patient were strictly kept within the same data partition (training, validation, or test) during all experiments. The cohort included 148 progressing patients (contributing 343 scans) and 217 stable controls (contributing 571 scans).

2.3.3. Breast cancer.

We utilized a curated dataset of digital screening mammograms from NYU Langone Health [34], enhanced with longitudinal follow-up cancer labels up to five years post-imaging.

  • Diagnosis Cohort: Contained 56,733 examinations from 33,681 patients with radiologist-assigned BI-RADS grades 0, 1, or 2.

  • Prognosis Cohort: To facilitate robust model development, a balanced cohort was constructed. This included 3,000 examinations from 2,319 patients who were diagnosed with breast cancer within 5 years (but more than 130 days after the index scan) and 3,000 control examinations from 2,837 patients who remained cancer-free for the 5-year follow-up period. The mean age for patients contributing case examinations was 61.5 ± 11.4 years, while the mean age for controls was 56.4 ± 10.8 years.

To prevent data leakage, strict patient-level splits were used to ensure disjoint sets of individuals across all training, validation, and test sets, and no information from future visits was used as input features. Detailed demographic and clinical characteristics for all patient cohorts are provided in S1-S4 Tables.

2.4. Approach

Our work investigates two key approaches: utilizing diagnosis tasks for model pretraining and proposing an effective multitask learning approach.

2.4.1. Diagnosis pretraining.

We compare diagnosis-based initialization against modality-specific baselines (Fig 1). We use modality-appropriate initializations: OA is 2D and moderate-resolution, where ImageNet pretraining is standard; AD uses 3D brain MRI (no ImageNet analogue), and BC requires very high-resolution mammography where running large ImageNet backbones at native resolution is not tractable, so we use random initialization similar to previous works [8,12]. Concretely, prognosis models are initialized from either (i) diagnosis-pretrained weights learned on the corresponding diagnostic task, (ii) ImageNet weights (OA only), or (iii) random weights (AD, BC).

Fig 1. Illustrates the proposed approaches (a) single task prognosis with diagnostic pretraining and (b) Sequential learning with experience replay.

Fig 1

2.4.2. Sequential learning with experience replay.

To mitigate catastrophic forgetting [28,29] in multitask settings, we used a two-phase approach (Fig 1).

  1. Phase 1 (Diagnosis Pretraining): A model is trained on the large diagnosis cohort.

  2. Phase 2 (Prognosis training with Diagnosis Replay): The model is fine-tuned using a multitask objective. At each step, a batch is randomly sampled with equal probability from either the prognosis cohort (to update on the prognosis task) or the diagnosis cohort (to “replay” and refresh diagnosis knowledge). This allows the model to learn the new task while preserving its original capabilities.

2.4.3. Comparative multitask approaches.

For comparison, we evaluated three alternative strategies:

  • Single-cohort MT: An approach similar to that used by Tiulpin et al. [14] and Leung et al. [16], where the model is trained on both diagnosis and prognosis tasks using only the data available in the smaller prognosis cohort.

  • Diagnosis-pretrained MT: Initializes with diagnostic pretraining but then fine-tunes solely on the prognosis cohort’s diagnosis and prognosis labels, without experience replay

  • Concurrent MT: Trains simultaneously on both diagnosis and prognosis cohorts from a baseline initialization, without any pretraining.

2.5. Experimental setup

2.5.1. Model architectures and preprocessing.

  • OA: Input radiographs were normalized, and knee joints were extracted prior to analysis. ResNet-34 with attention [35,36] architecture was used on extracted knee joints.

  • AD: Input MRI scans underwent standard bias correction and spatial normalization. A custom 3D-CNN [8] was used on processed T1 MRIs.

  • BC: The four standard mammographic views were individually processed and cropped to a uniform resolution. A custom multi-view CNN [12] was used on the processed mammogram views.

2.5.2. Training protocol.

All models were trained using the Adam optimizer [37] coupled with a cosine annealing learning rate scheduler to ensure stable convergence. For the highly imbalanced breast cancer BI-RADS task, a weighted cross-entropy loss was employed to give appropriate importance to rare classes. Condition-specific data augmentation techniques (including random cropping, flips, and rotations) were applied during training to enhance model robustness.

For tasks with smaller sample sizes (all prognosis cohorts and the AD diagnosis cohort), we employed a 5-fold cross-validation scheme with strict, non-overlapping patient-level splits. For each of the 5 iterations, one fold was designated as the held-out test set. Of the remaining four folds, three were used for training the model, and one was used as the validation set for hyperparameter tuning and model checkpointing via early stopping. This process was repeated five times, ensuring that every patient served as part of a test set exactly once. The final reported performance metrics are the mean and standard deviation across the results from the five test folds.

For multitask learning with experience replay, tasks were sampled with equal probability (p = 0.5). Crucially, to prioritize the more challenging prognostic task, model checkpoints were saved based on the epoch achieving the highest prognosis AUROC on the validation set, rather than on the total loss

For each disease, the diagnosis and prognosis cohorts were drawn from the same underlying patient populations (OAI, ADNI, and NYU dataset). As such, patients in the longitudinal prognosis cohort were also represented in the larger cross-sectional diagnosis cohort. To prevent any data leakage between the pretraining and fine-tuning stages, a strict patient-level splitting procedure was enforced. For each of the 5 cross-validation folds of the prognosis task, the patient IDs assigned to the test set were identified first. These testset patients were then completely excluded from the training set of the large diagnosis model used for pretraining. This ‘test set holdout’ strategy ensures that our prognosis models were always evaluated on patients that the entire training pipeline, including the pretraining stage, had never seen, providing an unbiased estimate of performance.

2.5.3. Evaluation and analysis.

This study was prepared and reported in accordance with the TRIPOD (Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis) guideline wherever applicable. The pre-specified primary endpoint for all prognostic tasks was the Area Under the Receiver Operating Characteristic curve (AUROC), which measures the model’s overall ability to discriminate between patients who experience disease progression and those who do not.

Secondary endpoints for prognosis included the Area Under the Precision-Recall Curve (AUPRC), which provides additional insight into model performance, especially in cohorts with class imbalance. Sensitivity and specificity were computed additionally for breast cancer.

For the threshold-dependent metrics of sensitivity and specificity, the following procedure was used for each of the 5 cross-validation folds: first, an optimal decision threshold was determined by maximizing the Youden Index [38] on the validation set for that fold. This fold-specific threshold was then applied to the predictions on the corresponding test set to calculate sensitivity and specificity. The final reported metrics are the mean and standard deviation of the five resulting sensitivity and specificity values.

For diagnosis, we report balanced accuracy for the five-class knee OA task, following [16] to mitigate class imbalance. For the multi-class AD and BC tasks, we additionally report macro-AUROC and micro-AUROC, consistent with prior work. Macro-AUROC is the unweighted mean of one-vs-rest AUROCs across classes, giving each class equal weight. Micro-AUROC pools predictions and labels across all classes to compute a single AUROC, effectively weighting classes by their prevalence. Balanced accuracy is the mean of per-class recall (sensitivity).

To formally compare the primary approaches, ROC curves of ensembled cross-validation predictions were compared using the DeLong test [39]. For every patient in the test set, the predictive scores (probabilities) from the five respective cross-validation models were averaged. This resulted in a single set of ensembled predictions for the entire test cohort, from which a single, more stable ROC curve was generated for the statistical comparison. This practice of averaging predictions is a standard ensembling technique to reduce variance and provide a more robust estimate of model performance.

Stratified subgroup analyses were performed on pre-defined clinical subgroups: structural incidence (KL 0–1) vs. progression (KL 2–3) for Knee OA, and cognitive transition groups (e.g., CN vs. MCI) for Alzheimer’s Disease. These subgroup analyses were conducted to verify that the models were learning genuine prognostic signals rather than simply using baseline disease severity as a proxy for progression risk. This analysis was not powered for formal statistical tests of subgroup interaction. Statistical significance testing was focused on our primary hypothesis comparing the baseline and diagnosis-pretrained models.

3. Results

The experimental results are presented in two main parts. The first subsection evaluates the impact of initializing prognosis models with weights pretrained on large diagnostic datasets, comparing their performance against standard baselines for single-task prognosis. The second subsection compares four multitask learning strategies on both prognosis and diagnosis.

3.1. Impact of diagnosis pretraining

Initializing prognosis models with weights pretrained on diagnostic tasks with full diagnosis cohorts yielded improvements over baseline initializations (ImageNet for Knee OA, random for AD/BC), as detailed in Table 1.

Table 1. Comparison of prognosis prediction performance using models with baseline initialization versus diagnostic pretraining across three diseases. Knee OA models used ImageNet pretraining as baseline; AD and Breast Cancer models used random initialization as baseline.

Disease Dataset Metric Baseline Init. Diagnosis Pretrained Init.
Knee Osteoarthritis OAI (Internal) AUROC 0.744 ± 0.014 0.738 ± 0.005
AUPRC 0.452 ± 0.014 0.448 ± 0.008
MOST (External) AUROC 0.747 ± 0.012 0.770 ± 0.006
AUPRC 0.564 ± 0.011 0.602 ± 0.008
Alzheimer’s Disease ADNI AUROC 0.807 ± 0.025 0.840 ± 0.008
AUPRC 0.683 ± 0.030 0.752 ± 0.012
Sensitivity 60.5% ± 12.0% 78.1% ± 8.2%
Specificity 82.5% ± 6.1% 72.6% ± 11.8%
Breast Cancer NYU AUROC 0.848 ± 0.007 0.874 ± 0.002
AUPRC 0.873 ± 0.012 0.900 ± 0.001
Sensitivity 67.6% ± 2.1% 69.9% ± 1.8%
Specificity 92.0% ± 3.5% 93.4% ± 2.7%

*Note: AUROC: Area Under the Receiver Operating Characteristic curve; AUPRC: Area Under the Precision-Recall curve. Values are mean ± standard deviation. Boldface indicates the best-performing approach for each metric.

For Knee Osteoarthritis, while performance on the internal OAI dataset was comparable, diagnosis-pretrained models demonstrated superior generalization on the external MOST dataset (e.g., AUROC 0.770 vs 0.747) and exhibited lower variance, suggesting more stable performance. This improved generalization on the external MOST dataset was statistically significant when comparing model ensembles via DeLong’s test (AUROC 0.781 for diagnosis-pretrained vs. 0.768 for ImageNet-pretrained, p = 0.006, n = 4251). Additionally, this performance benefit on the external MOST dataset was consistent across both early-stage (Incidence) and established (Progression) disease subgroups (Fig 2a).

Fig 2. Comparison of baseline initialization versus diagnostic pretraining for prognosis prediction performance (AUROC) within specific disease subgroups.

Fig 2

(a) Knee OA subgroups include Incidence (progression from KL 0-1) and Progression (worsening from KL 2-3) cohorts evaluated on the internal OAI and external MOST datasets. (b) Alzheimer’s Disease subgroups represent different cognitive transitions within the ADNI dataset (CN: Cognitively Normal, MCI: Mild Cognitive Impairment; 0: Stable status over 2 years, 1: Progressing status over 2 years). Diagnostic pretraining generally improves or maintains performance with better stability, showing particular benefit in the external MOST and all AD sub cohorts. Full numerical results are available in S5 and S6 Tables.

For Alzheimer’s Disease, diagnostic pretraining significantly enhanced AD progression prediction across all metrics, notably increasing AUROC (0.840 vs 0.807) and AUPRC (0.752 vs 0.683). Benefits were observed across different cognitive subgroups (Fig 2b). However, the difference in overall AUROC between model ensembles on the whole test set was not statistically significant via DeLong’s test (0.846 for diagnosis-pretrained vs. 0.823 for random initialized, p = 0.137, n = 236).

For Breast Cancer, diagnostic pretraining substantially improved 5-year prognosis prediction (AUROC 0.874 vs 0.848). This boost extended to precision-recall (AUPRC 0.900 vs 0.873) and clinical utility metrics (sensitivity/specificity), alongside improved stability (lower std. dev.). This performance increase was statistically significant when comparing model ensembles (AUROC 0.876 for diagnostic-pretrained vs. 0.854 for random initialized, p < 0.001, n = 1735).

Across all three diseases, these findings reinforce that leveraging diagnostic knowledge provides a valuable initialization for prognostic tasks, likely by helping the model learn representations of relevant anatomical and pathological features that precede clinically evident progression.

3.2. Multitask learning performance

Four multitask learning strategies were further compared, evaluating performance on both the prognosis task (Table 2) and the diagnosis task (Table 3).

Table 2. Prognosis prediction performance across different multitask learning approaches. ‘Ref (Diag. Pret)’ refers to the single-task prognosis model initialized with diagnostic pretraining (from Table 1). ‘Single-cohort MT’ uses only prognosis cohort data for multitask training. ‘Concurrent MT’ trains on diagnosis and prognosis cohorts simultaneously from baseline initialization. ‘Diag-pret MT’ performs diagnostic pretraining then trains multitask on the prognosis cohort. ‘Seq Learn w/ Replay’ initializes with diagnosis pretraining, trains multitask on prognosis cohort and uses experience replay with the diagnosis cohort.

Disease Dataset Metric Single-task prognosis (diagnosis-pretrained init.) Single-cohort MT Concurrent MT Diag-pret MT Seq Learn w/ Replay
Knee Osteoarthritis OAI (Internal) AUROC 0.738 ± 0.005 0.739 ± 0.017 0.743 ± 0.011 0.748 ± 0.005 0.750 ± 0.006
AUPRC 0.448 ± 0.008 0.444 ± 0.015 0.438 ± 0.011 0.458 ± 0.005 0.456 ± 0.006
MOST (External) AUROC 0.770 ± 0.006 0.756 ± 0.027 0.765 ± 0.005 0.763 ± 0.007 0.764 ± 0.007
AUPRC 0.602 ± 0.008 0.564 ± 0.020 0.589 ± 0.004 0.599 ± 0.008 0.594 ± 0.014
Alzheimer’s Disease ADNI AUROC 0.840 ± 0.008 0.798 ± 0.014 0.828 ± 0.016 0.848 ± 0.008 0.851 ± 0.009
AUPRC 0.752 ± 0.012 0.653 ± 0.027 0.716 ± 0.021 0.754 ± 0.011 0.754 ± 0.005
Breast Cancer NYU AUROC 0.874 ± 0.002 0.864 ± 0.002 0.860 ± 0.007 0.878 ± 0.003 0.876 ± 0.003
AUPRC 0.900 ± 0.001 0.894 ± 0.002 0.889 ± 0.008 0.903 ± 0.003 0.903 ± 0.002
Sensitivity 69.9% ± 1.8% 66.1% ± 1.6% 67.4% ± 2.6% 69.0% ± 2.5% 67.7% ± 1.7%
Specificity 93.4% ± 2.7% 97.1% ± 1.0% 94.6% ± 4.3% 95.5% ± 1.8% 97.4% ± 1.2%

*Note: Values are mean ± standard deviation. Boldface indicates the best-performing approach for each metric.

Table 3. Diagnosis task performance across different multitask learning approaches. ‘Reference’ refers to a model trained only on the diagnosis task. Performance is measured by Balanced Accuracy (Bal Acc) for Knee OA; and Macro AUROC, Micro AUROC, and Bal Acc for Alzheimer’s Disease (AD) and Breast Cancer (BC). Other approach definitions are as in Table 2.

Disease Condition Dataset Metric Reference Single-cohort MT Concurrent MT Diag-pret MT Seq Learn w/ Replay
Knee Osteoarthritis OAI (Internal) Bal Acc (%) 75.1% 51.9% ± 1.3% 73.4% ± 2.4% 56.6% ± 0.7% 75.1% ± 1.1%
MOST (External) Bal Acc (%) 65.6% 47.0% ± 3.0% 67.3% ± 1.2% 48.8% ± 1.7% 67.5% ± 2.6%
Alzheimer’s Disease ADNI Macro AUROC 0.752 ± 0.012 0.530 ± 0.011 0.741 ± 0.013 0.532 ± 0.026 0.758 ± 0.005
Micro AUROC 0.760 ± 0.016 0.581 ± 0.007 0.746 ± 0.007 0.618 ± 0.015 0.763 ± 0.011
Bal Acc (%) 56.3% ± 4.2% 35.3% ± 3.2% 52.7% ± 3.0% 38.1% ± 2.9% 53.0% ± 2.8%
Breast Cancer NYU Macro AUROC 0.7 0.577 ± 0.007 0.571 ± 0.007 0.657 ± 0.004 0.697 ± 0.002
Micro AUROC 0.724 0.687 ± 0.002 0.655 ± 0.013 0.726 ± 0.012 0.735 ± 0.012
Bal Acc (%) 50.9% 37.7% ± 1.2% 41.3% ± 0.4% 45.2% ± 0.4% 50.4% ± 0.6%

*Note: Values are reported as mean ± standard deviation where available from cross-validation, otherwise as single run results. Boldface indicates the best-performing approach for each metric.

For Knee Osteoarthritis, Sequential Learning with Experience Replay and Diagnosis-pretrained MT demonstrated the strongest prognostic performance (e.g., OAI AUROC 0.75, MOST AUROC 0.764), matching the dedicated single-task reference model. On the diagnosis task, Sequential Learning with Experience Replay maintained robust performance (OAI Bal Acc 75.1%), while other methods showed substantial degradation.

For Alzheimer’s Disease, Sequential Learning with Experience Replay and Diagnosis-pretrained MT again achieved the highest prognosis performance (AUROC 0.851 and 0.848). For diagnosis, Sequential Learning with Experience Replay preserved strong performance (macro AUROC 0.758), whereas Single-cohort MT and Diagnosis-pretrained MT suffered significant degradation (macro AUROC 0.53).

For Breast Cancer, Sequential Learning with Experience Replay and Diagnosis-pretrained MT yielded the best prognosis performance (AUROC 0.876 and 0.878 respectively). Critically, Sequential Learning with Experience Replay maintained diagnostic metrics nearly identical to the single-task reference model (Bal Acc 50.4% vs 50.9%), while all other multitask methods showed poor to degraded diagnostic performance.

The robustness of the Sequential Learning with Experience Replay approach was further confirmed in the subgroup analyses (Fig 3). Across the various Knee OA and Alzheimer’s Disease subgroups, this method consistently achieved prognostic performance that was comparable or superior to the dedicated single-task reference model. In contrast, simpler strategies like ‘Single-cohort MT’ often showed a notable degradation in performance. These results highlight Sequential Learning with Experience Replay as an optimal multitask approach, effectively preventing catastrophic forgetting while achieving top-tier prognostic performance.

Fig 3. Performance comparison of various multitask learning approaches for prognosis prediction (AUROC) within specific disease subgroups.

Fig 3

(a) Knee OA subgroups (OAI/MOST Incidence and Progression) and (b) Alzheimer’s Disease subgroups (ADNI CN/MCI transitions) are shown. Strategies compared include a single-task reference model (‘Ref (Diag. Pret)’); Multitask training using only the prognosis cohort (‘Single-cohort MT’); Concurrent training on diagnosis and prognosis cohorts (‘Concurrent MT’); Multitask training on the prognosis cohort after diagnostic pretraining (‘Diag-pret MT’); and Sequential Learning with Experience Replay (‘Seq Learn w/ Replay’). Sequential Learning with Replay demonstrates robust performance across most subgroups, often matching or exceeding the dedicated single-task reference model. Full numerical results, including standard deviations, are available in S7 and S8 Tables.

4. Discussion

4.1. Principal findings

The experiments demonstrate two main findings across the studied conditions: (1) diagnostic pretraining improves prognosis prediction and provides stability, and (2) a sequential learning approach with experience replay performed best among the multitask strategies evaluated.

Diagnostic-pretrained initialization improved prognostic discrimination and generalization to an external cohort and across demographic/clinical subgroups. The advantage remained after controlling for diagnostic severity (Fig 2), indicating that pretraining helps models capture relevant prognostic features beyond simply using diagnosis severity as a proxy for risk.

Recent work on generalist and multimodal medical AI suggests that representations learned from large-scale diagnostic data can serve as effective foundations for a wide range of downstream diagnostic classification tasks, particularly in settings with limited labeled data [24,40]. Our results provide empirical evidence that this paradigm extends specifically to disease prognosis, where longitudinal labels are especially scarce and difficult to collect. By leveraging large, cross-sectional diagnostic cohorts, our approach improves prognostic discrimination and stability while preserving diagnostic competence, addressing a key gap between foundation-model principles and real-world longitudinal prediction tasks.

Furthermore, our evaluation of joint diagnosis-prognosis approaches showed that sequential learning with experience replay performed best. Across knee OA and AD subgroups, it reached prognostic performance comparable to dedicated, diagnosis-pretrained single-task models, while crucially, also preserving diagnostic performance. This dual success effectively mitigates the critical issue of catastrophic forgetting (Table 3, Fig 4).

Fig 4. Confusion matrices for the diagnosis task across different multitask learning approaches.

Fig 4

(a) Knee OA KLG prediction, (b) Alzheimer’s diagnosis, and (c) Breast Cancer BI-RADS prediction. The matrices show that methods like ‘Diagnosis Pretrained MT’ forget how to classify rare classes (e.g., KLG 4, AD, BI-RADS 0) after prognosis fine-tuning. In contrast, ‘Seq learning w/ Replay’ retains knowledge across all classes, demonstrating its effectiveness at mitigating catastrophic forgetting.

In contrast, the ‘Single-cohort MT’ and ‘Diagnosis-pretrained MT’ variants were evaluated to represent simpler strategies from prior works and serve as a crucial baseline. Their performance starkly illustrates the severity of catastrophic forgetting. Lacking access to the full, unbiased diagnostic dataset during fine-tuning, these approaches suffered a sharp degradation in diagnostic ability. They effectively forgot knowledge from pretraining, including how to classify advanced disease stages in knee OA/AD and the rare class BI-RADS 0 in breast cancer (Fig 4).

This predictable decay in the baseline models highlights the essential value of the replay mechanism. By continually refreshing knowledge from the original task, experience replay preserves diagnostic competence – including for rare classes, while the model learns the new prognosis objective. The consistent success of this strategy across three physiologically distinct conditions (musculoskeletal, neurological, and oncological) suggests its broad applicability for other progressive diseases where longitudinal data is limited.

This work improves upon prior multitask learning strategies by demonstrating that the Single-cohort MT approach, similar to that used by Tiulpin et al. [14] and Leung et al. [16], is suboptimal for knee OA prognosis. Similarly in breast cancer, this study extends the findings of Wu et al. [13], showing that pretraining on diagnosis BI-RADS labels benefits not just breast cancer detection but also long-term prognosis.

4.2. Limitations

Several limitations warrant discussion. First, our knee OA and AD datasets derive from standardized research protocols, potentially lacking the variability of routine clinical practice and therefore affecting generalizability to diverse care settings. Second, the inherent class imbalance (few progressors versus many stable patients) in progressive disease studies affects all conditions studied. While we employed balanced sampling and specific evaluation metrics (AUPRC) to mitigate this, the imbalance remains a challenge for model deployment.

Third, a practical consideration for the sequential learning with experience replay strategy is its reliance on access to the original diagnosis dataset, which may not always be available when models are shared. It is important to note, however, that this limitation applies only to the multitask fine-tuning stage. The initial performance gains from using a diagnosis-pretrained model for standard fine-tuning – our first key finding remains an applicable benefit of our work, even without access to the original data for replay.

Finally, while experience replay effectively mitigates catastrophic forgetting, it increases the computational cost of training compared to simple fine-tuning, as the model must process ‘replay’ batches alongside new data.

4.3. Clinical implications

From a clinical perspective, these findings are relevant. A single, unified model that supports both diagnosis and prognosis mirrors routine workflow, assessing current status and planning forward. Such a model can aid counseling by presenting a fuller view – from present severity to future risk and can support personalization. For example, flagging patients at high risk of rapid OA progression for earlier intervention, or using screening mammograms to guide breast-cancer surveillance intensity. Because the sequential learning strategy preserves diagnostic competence, the model remains a dependable diagnostic aid while adding prognostic value, critical for trust and clinical adoption. In addition, producing diagnosis and prognosis together enables cross-verification at the point of care: clinicians can check whether current findings align with predicted risk and more closely review discordant cases, which may support trust and uptake. We view this as a usability/trust benefit rather than a validated safety outcome. Future prospective studies are needed to quantify its impact on decision quality. Additional future work should test these models on larger, more heterogeneous datasets from diverse care settings to assess generalizability beyond the research cohorts used here.

4.4. Conclusion

In conclusion, our work demonstrates an effective approach to extracting more value from limited longitudinal datasets by leveraging larger diagnosis datasets for prognostic model training.

Supporting information

S1 Fig. Comparative multitask learning approaches.

Visualizations of (a) Single Cohort MT, (b) Diagnosis Pretrained MT, and (c) Concurrent MT.

(TIF)

pone.0344600.s001.tif (641.3KB, tif)
S1 Table. Patient Characteristics in Structural Progression OAI and MOST Cohorts.

(DOCX)

pone.0344600.s002.docx (15.4KB, docx)
S2 Table. Patient Characteristics in Structural Incidence OAI and MOST Cohorts.

(DOCX)

pone.0344600.s003.docx (15.6KB, docx)
S3 Table. Demographic Characteristics and Cognitive Status Distribution across the ADNI Prognosis Cohort.

(DOCX)

pone.0344600.s004.docx (14.5KB, docx)
S4 Table. Breast cancer Patient distribution in the Prognosis Cohort.

(DOCX)

pone.0344600.s005.docx (14.5KB, docx)
S5 Table. Comparison of AUROC performance for models trained on progression and incidence, using different initializations.

(DOCX)

pone.0344600.s006.docx (14.2KB, docx)
S6 Table. Detailed AUROC analysis comparing model performance across cognitive status subgroups.

(DOCX)

pone.0344600.s007.docx (14.1KB, docx)
S7 Table. Detailed metrics for various training strategies for progression and incidence prediction tasks on OAI and MOST Datasets.

(DOCX)

pone.0344600.s008.docx (14.7KB, docx)
S8 Table. Detailed progression prediction performance (AUROC) across cognitive status subgroups for multitask methods.

(DOCX)

pone.0344600.s009.docx (16.5KB, docx)

Acknowledgments

We gratefully acknowledge data provision from the Osteoarthritis Initiative (OAI), a public-private partnership (this manuscript does not necessarily reflect the opinions of OAI investigators); the Multicenter Osteoarthritis Study (MOST); and the Alzheimer’s Disease Neuroimaging Initiative (ADNI). A full list of ADNI contributors is available at adni.loni.usc.edu. We thank the investigators of these studies for their contributions.

Data Availability

- Osteoarthritis Initiative (OAI): OAI data are publicly available through the NIMH Data Archive (NDA). Access requires creating an NDA account, agreeing to the OAI data access terms, and requesting the desired collections through the NDA portal. Full instructions are provided at https://nda.nih.gov/oai. - Multicenter Osteoarthritis Study (MOST): MOST data are publicly available through the NIA Aging Research Biobank. Access requires creating a Biobank account and submitting a data request in accordance with the Biobank’s terms and conditions. More information is available at https://agingresearchbiobank.nia.nih.gov/. - Alzheimer’s Disease Neuroimaging Initiative (ADNI): ADNI data are publicly available from the LONI Image & Data Archive upon completion of a web application and acceptance of the ADNI Data Use Agreement. Data can be accessed at https://adni.loni.usc.edu/. - Breast Cancer Cohort (NYU Langone Health): This institutional dataset contains protected health information and cannot be made publicly available. De-identified data may be shared upon reasonable request, pending approval by NYU Langone’s data governance committee. Inquiries can be directed to Krzysztof Geras (k.j.geras@nyu.edu). - Code and Model Availability: All code for preprocessing and model training sufficient to reproduce the analyses reported in this study are available at https://github.com/denizlab/diag-to-prog-replay.

Funding Statement

This study was supported by the National Institute of Arthritis and Musculoskeletal and Skin Diseases in the form of a grant awarded to CMD (R01-AR074453) and the National Institute of Arthritis and Musculoskeletal and Skin Diseases in the form of a salary for HRR. The specific roles of these authors are articulated in the ‘author contributions’ section. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

References

  • 1.Cross M, Smith E, Hoy D, Nolte S, Ackerman I, Fransen M, et al. The global burden of hip and knee osteoarthritis: estimates from the global burden of disease 2010 study. Annals of the Rheumatic Diseases. 2014;73(7):1323–30. [DOI] [PubMed] [Google Scholar]
  • 2.GBD 2021 Osteoarthritis Collaborators. Global, regional, and national burden of osteoarthritis, 1990-2020 and projections to 2050: A systematic analysis for the Global Burden of Disease Study 2021. Lancet Rheumatol. 2023;5(9):e508–22. doi: 10.1016/S2665-9913(23)00163-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Rasmussen J, Langerman H. Alzheimer’s disease - Why we need early diagnosis. Degener Neurol Neuromuscul Dis. 2019;9:123–30. doi: 10.2147/DNND.S228939 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Arnold M, Morgan E, Rumgay H, Mafra A, Singh D, Laversanne M, et al. Current and future burden of breast cancer: Global statistics for 2020 and 2040. The Breast. 2022;66:15–23. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Allemani C, Matsuda T, Di Carlo V, Harewood R, Matz M, Nikšić M, et al. Global surveillance of trends in cancer survival 2000–14 (CONCORD-3): Analysis of individual records for 37 513 025 patients diagnosed with one of 18 cancers from 322 population-based registries in 71 countries. The Lancet. 2018;391(10125):1023–75. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 6.Lecun Y, Bottou L, Bengio Y, Haffner P. Gradient-based learning applied to document recognition. Proc IEEE. 1998;86(11):2278–324. doi: 10.1109/5.726791 [DOI] [Google Scholar]
  • 7.Tiulpin A, Thevenot J, Rahtu E, Lehenkari P, Saarakkala S. Automatic knee osteoarthritis diagnosis from plain radiographs: A deep learning-based approach. Scientific Reports. 2018;8(1):1727. doi: 10.1038/s41598-018-35557-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 8.Liu S, Masurkar AV, Rusinek H, Chen J, Zhang B, Zhu W. Generalizable deep learning model for early Alzheimer’s disease detection from structural MRIs. Scientific Reports. 2022;12(1):17106. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Liu S, Yadav C, Fernandez-Granda C, Razavian N. On the design of convolutional neural networks for automatic detection of Alzheimer’s disease. Machine learning for health workshop, 2020. 184–201. [Google Scholar]
  • 10.Castellano G, Esposito A, Lella E, Montanaro G, Vessio G. Automated detection of Alzheimer’s disease: A multi-modal approach with 3D MRI and amyloid PET. Sci Rep. 2024;14(1):5210. doi: 10.1038/s41598-024-56001-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 11.Suk HI, Lee SW, Shen D, Alzheimer’s Disease Neuroimaging Initiative. Hierarchical feature representation and multimodal fusion with deep learning for AD/MCI diagnosis. NeuroImage. 2014;101:569–82. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Geras KJ, Wolfson S, Shen Y, Wu N, Kim S, Kim E. High-resolution breast cancer screening with multi-view deep convolutional neural networks. 2017. doi: arXiv:1703.07047 [Google Scholar]
  • 13.Wu N, Phang J, Park J, Shen Y, Huang Z, Zorin M, et al. Deep neural networks improve radiologists’ performance in breast cancer screening. IEEE Trans Med Imaging. 2020;39(4):1184–94. doi: 10.1109/TMI.2019.2945514 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.Leung K, Zhang B, Tan J, Shen Y, Geras KJ, Babb JS, et al. Prediction of total knee replacement and diagnosis of osteoarthritis by using deep learning on knee radiographs: Data from the osteoarthritis initiative. Radiology. 2020;296(3):584–93. doi: 10.1148/radiol.2020192091 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 15.Tolpadi AA, Lee JJ, Pedoia V, Majumdar S. Deep learning predicts total knee replacement from magnetic resonance images. Sci Rep. 2020;10(1):6371. doi: 10.1038/s41598-020-63395-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Tiulpin A, Klein S, Bierma-Zeinstra SMA, Thevenot J, Rahtu E, Meurs J, et al. Multimodal machine learning-based knee osteoarthritis progression prediction from plain radiographs and clinical data. Sci Rep. 2019;9(1):20038. doi: 10.1038/s41598-019-56527-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Schiratti J-B, Dubois R, Herent P, Cahané D, Dachary J, Clozel T, et al. A deep learning method for predicting knee osteoarthritis radiographic progression from MRI. Arthritis Res Ther. 2021;23(1):262. doi: 10.1186/s13075-021-02634-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Rajamohan HR, Wang T, Leung K, Chang G, Cho K, Kijowski R, et al. Prediction of total knee replacement using deep learning analysis of knee MRI. Sci Rep. 2023;13(1):6922. doi: 10.1038/s41598-023-33934-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 19.Lee G, Nho K, Kang B, Sohn K-A, Kim D, for Alzheimer’s Disease Neuroimaging Initiative. Predicting Alzheimer’s disease progression using multi-modal deep learning approach. Sci Rep. 2019;9(1):1952. doi: 10.1038/s41598-018-37769-z [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 20.Spasov S, Passamonti L, Duggento A, Lio P, Toschi N, Alzheimer’s Disease Neuroimaging Initiative. A parameter-efficient deep learning approach to predict conversion from mild cognitive impairment to Alzheimer’s disease. Neuroimage. 2019;189:276–87. [DOI] [PubMed] [Google Scholar]
  • 21.Sun D, Wang M, Li A. A Multimodal deep neural network for human breast cancer prognosis prediction by integrating multi-dimensional data. IEEE/ACM Trans Comput Biol Bioinform. 2019;16(3):841–50. doi: 10.1109/TCBB.2018.2806438 [DOI] [PubMed] [Google Scholar]
  • 22.Huang Z, Zhang X, Ju Y, Zhang G, Chang W, Song H, et al. Explainable breast cancer molecular expression prediction using multi-task deep-learning based on 3D whole breast ultrasound. Insights Imaging. 2024;15(1):227. doi: 10.1186/s13244-024-01810-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.Harkey MS, Costello KE, Mehta B, Wen C, Malfait AM, Madry H. Artificial intelligence in osteoarthritis research: summary of the 2025 OARSI pre-congress workshop. Osteoarthritis and Cartilage Open. 2025;1:100687. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 24.Moor M, Banerjee O, Abad ZSH, Krumholz HM, Leskovec J, Topol EJ, et al. Foundation models for generalist medical artificial intelligence. Nature. 2023;616(7956):259–65. doi: 10.1038/s41586-023-05881-4 [DOI] [PubMed] [Google Scholar]
  • 25.Weiner MW, Kanoria S, Miller MJ, Aisen PS, Beckett LA, Conti C, et al. Overview of Alzheimer’s disease neuroimaging initiative and future clinical trials. Alzheimers Dement. 2025;21(1):e14321. doi: 10.1002/alz.14321 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 26.Akinci D’Antonoli T, Bluethgen C, Cuocolo R, Klontzas ME, Ponsiglione A, Kocak B. Foundation models for radiology: fundamentals, applications, opportunities, challenges, risks, and prospects. Diagn Interv Radiol. 2025;:10.4274/dir.2025.253445. doi: 10.4274/dir.2025.253445 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Ma D, Pang J, Gotway MB, Liang J. A fully open AI foundation model applied to chest radiography. Nature. 2025;643(8071):488–98. doi: 10.1038/s41586-025-09079-8 [DOI] [PubMed] [Google Scholar]
  • 28.Robins A. Catastrophic forgetting, rehearsal and pseudorehearsal. Connection Science. 1995;7(2):123–46. doi: 10.1080/09540099550039318 [DOI] [Google Scholar]
  • 29.Chaudhry A, Rohrbach M, Elhoseiny M, Ajanthan T, Dokania PK, Torr PH. On tiny episodic memories in continual learning. 2019. doi: 10.48550/arXiv.1902.10486 [DOI] [Google Scholar]
  • 30.Kellgren JH, Lawrence JS. Radiological assessment of osteo-arthrosis. Ann Rheum Dis. 1957;16(4):494–502. doi: 10.1136/ard.16.4.494 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Lee S, Kwon Y, Lee N, Bae K-J, Kim J, Park S, et al. The Prevalence of osteoarthritis and risk factors in the korean population: The Sixth Korea National Health and Nutrition Examination Survey (VI-1, 2013). Korean J Fam Med. 2019;40(3):171–5. doi: 10.4082/kjfm.17.0090 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Nevitt M, Felson D, Lester G. The osteoarthritis initiative. Protocol for the cohort study. 2006;1(2).
  • 33.Segal NA, Nevitt MC, Gross KD, Hietpas J, Glass NA, Lewis CE. The Multicenter Osteoarthritis Study (MOST): Opportunities for Rehabilitation Research. PM & R: The Journal of Injury, Function, and Rehabilitation. 2013;5(8):10–6. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 34.Wu N, Phang J, Park J, Shen Y, Kim SG, Heacock L. The NYU breast cancer screening dataset v1. 0. New York, NY, USA: New York Univ. 2019. [Google Scholar]
  • 35.He K, Zhang X, Ren S, Sun J. Deep Residual Learning for Image Recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016. 770–8. doi: 10.1109/cvpr.2016.90 [DOI] [Google Scholar]
  • 36.Woo S, Park J, Lee J-Y, Kweon IS. CBAM: Convolutional Block Attention Module. Lecture Notes in Computer Science. Springer International Publishing. 2018. 3–19. doi: 10.1007/978-3-030-01234-2_1 [DOI] [Google Scholar]
  • 37.Kingma DP. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980. 2014. [Google Scholar]
  • 38.Youden WJ. Index for rating diagnostic tests. Cancer. 1950;3(1):32–5. doi: 10.1002/1097-0142(1950)3:1<32::aid-cncr2820030106>3.0.co;2-3 [DOI] [PubMed] [Google Scholar]
  • 39.DeLong ER, DeLong DM, Clarke-Pearson DL. Comparing the areas under two or more correlated receiver operating characteristic curves: A nonparametric approach. Biometrics. 1988;44(3):837–45. doi: 10.2307/2531595 [DOI] [PubMed] [Google Scholar]
  • 40.Acosta JN, Falcone GJ, Rajpurkar P, Topol EJ. Multimodal biomedical AI. Nat Med. 2022;28(9):1773–84. doi: 10.1038/s41591-022-01981-2 [DOI] [PubMed] [Google Scholar]

Decision Letter 0

Naveen Baskaran

10 Dec 2025

-->PONE-D-25-52607-->-->Robust Disease Prognosis via Diagnostic Knowledge Preservation: A Sequential Learning Approach-->-->PLOS One

Dear Dr. Rajamohan,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please improve the language proficiency to improve flow and clear information. Provide IRB details.

Please submit your revised manuscript by Jan 24 2026 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:-->

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

-->If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: https://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols. Additionally, PLOS ONE offers an option for publishing peer-reviewed Lab Protocol articles, which describe protocols hosted on protocols.io. Read more information on sharing protocols at https://plos.org/protocols?utm_medium=editorial-email&utm_source=authorletters&utm_campaign=protocols.

We look forward to receiving your revised manuscript.

Kind regards,

Naveen Baskaran, MD

Academic Editor

PLOS One

Journal Requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2. Please note that PLOS One has specific guidelines on code sharing for submissions in which author-generated code underpins the findings in the manuscript. In these cases, we expect all author-generated code to be made available without restrictions upon publication of the work. Please review our guidelines at https://journals.plos.org/plosone/s/materials-and-software-sharing#loc-sharing-code and ensure that your code is shared in a way that follows best practice and facilitates reproducibility and reuse.

3. Thank you for stating the following in the Acknowledgments Section of your manuscript:

“We gratefully acknowledge data provision from the Osteoarthritis Initiative (OAI), an NIH526 funded public-private partnership (this manuscript does not necessarily reflect the opinions of OAI investigators or funders); the Multicenter Osteoarthritis Study (MOST), sponsored by the NIH/National Institute on Aging; and the Alzheimer's Disease Neuroimaging Initiative (ADNI). ADNI is funded by NIH Grant U01 AG024904, DOD ADNI (W81XWH-12-2-0012), and numerous public and private contributions, with a full list of contributors available at adni.loni.usc.edu.”

We note that you have provided funding information that is not currently declared in your Funding Statement. However, funding information should not appear in the Acknowledgments section or other areas of your manuscript. We will only publish funding information present in the Funding Statement section of the online submission form.

Please remove any funding-related text from the manuscript and let us know how you would like to update your Funding Statement. Currently, your Funding Statement reads as follows:

“This work was supported by the National Institutes of Health (NIH) under grant R01-AR074453”

Please include your amended statements within your cover letter; we will change the online submission form on your behalf.

4. We note that the grant information you provided in the ‘Funding Information’ and ‘Financial Disclosure’ sections do not match.

When you resubmit, please ensure that you provide the correct grant numbers for the awards you received for your study in the ‘Funding Information’ section.

5. Thank you for stating the following financial disclosure:

“This work was supported by the National Institutes of Health (NIH) under grant R01-AR074453”

Please state what role the funders took in the study.  If the funders had no role, please state: "The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript."

If this statement is not correct you must amend it as needed.

Please include this amended Role of Funder statement in your cover letter; we will change the online submission form on your behalf.

6. When completing the data availability statement of the submission form, you indicated that you will make your data available on acceptance. We strongly recommend all authors decide on a data sharing plan before acceptance, as the process can be lengthy and hold up publication timelines. Please note that, though access restrictions are acceptable now, your entire data will need to be made freely accessible if your manuscript is accepted for publication. This policy applies to all data except where public deposition would breach compliance with the protocol approved by your research ethics board. If you are unable to adhere to our open data policy, please kindly revise your statement to explain your reasoning and we will seek the editor's input on an exemption. Please be assured that, once you have provided your new statement, the assessment of your exemption will not hold up the peer review process.

7. If the reviewer comments include a recommendation to cite specific previously published works, please review and evaluate these publications to determine whether they are relevant and should be cited. There is no requirement to cite these works unless the editor has indicated otherwise.

8. Please review your reference list to ensure that it is complete and correct. If you have cited papers that have been retracted, please include the rationale for doing so in the manuscript text, or remove these references and replace them with relevant current references. Any changes to the reference list should be mentioned in the rebuttal letter that accompanies your revised manuscript. If you need to cite a retracted article, indicate the article’s retracted status in the References list and also include a citation and full reference for the retraction notice.

[Note: HTML markup is below. Please do not edit.]

Reviewers' comments:

Reviewer's Responses to Questions

-->Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented. -->

Reviewer #1: Yes

Reviewer #2: Yes

**********

-->2. Has the statistical analysis been performed appropriately and rigorously? -->

Reviewer #1: Yes

Reviewer #2: Yes

**********

-->3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.-->

Reviewer #1: Yes

Reviewer #2: Yes

**********

-->4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.-->

Reviewer #1: Yes

Reviewer #2: Yes

**********

-->5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)-->

Reviewer #1: This manuscript is already well written and provides important insights into robust disease prognosis via diagnostic knowledge preservation. I have a few suggestions that may help improve the manuscript:

1. Line 91: Please provide the IRB institution and IRB number for this study.

2. Line 112: The abbreviation for ADNI should be introduced here, not at line 159.

3. Lines 162–163: Since the abbreviations for Cognitive Normal (CN), Mild Cognitive Impairment (MCI), and Alzheimer’s Disease (AD) have already been defined earlier, please use the abbreviations directly in this section.

4. Discussion: The manuscript would benefit from a dedicated Limitations section, as several inherent limitations appear to be present.

5. A thorough English language review is recommended to further improve clarity, flow, and readability.

Reviewer #2: A very convincing paper with sound findings & elaboration on the significance of the findings. A more recent literature would strengthened the juatification s of the findings. Critical argumenta are still lacking on some parts of discussions.

**********

-->6. PLOS authors have the option to publish the peer review history of their article (what does this mean?). If published, this will include your full peer review and any attached files.

If you choose “no”, your identity will remain anonymous but your review may still be made public.

Do you want your identity to be public for this peer review? For information about this choice, including consent withdrawal, please see our Privacy Policy.-->

Reviewer #1: No

Reviewer #2: No

**********

[NOTE: If reviewer comments were submitted as an attachment file, they will be attached to this email and accessible via the submission site. Please log into your account, locate the manuscript record, and check for the action link "View Attachments". If this link does not appear, there are no attachment files.]

To ensure your figures meet our technical requirements, please review our figure guidelines: https://journals.plos.org/plosone/s/figures

You may also use PLOS’s free figure tool, NAAS, to help you prepare publication quality figures: https://journals.plos.org/plosone/s/figures#loc-tools-for-figure-preparation.

NAAS will assess whether your figures meet our technical requirements by comparing each figure against our figure specifications.

-->

PLoS One. 2026 May 6;21(5):e0344600. doi: 10.1371/journal.pone.0344600.r002

Author response to Decision Letter 1


6 Jan 2026

Dear Dr. Baskaran and Reviewers,

We thank the Academic Editor and the Reviewers for their thoughtful comments and constructive feedback on our manuscript. We are pleased that the reviewers found the study technically sound.

We have revised the manuscript to address the specific points raised, particularly regarding the inclusion of specific IRB details, the clarification of abbreviations, and the addition of a dedicated Limitations section. Our point-by-point responses are detailed below.

Reviewer #1

Comment: This manuscript is already well written and provides important insights into robust disease prognosis via diagnostic knowledge preservation

Response: We thank Reviewer 1 for their positive assessment and we appreciate their helpful feedback.

Comment 1: Line 91: Please provide the IRB institution and IRB number for this study.

Response: We thank the reviewer for this important request. We have updated the Ethics Statement (Section 2.1) to explicitly state that this is a retrospective analysis of fully de-identified data which did not require a separate IRB approval. We have added the available IRB approval numbers, ClinicalTrials.gov identifiers for the public datasets (OAI, MOST,ADNI) and the specific IRB number for the NYU breast cancer dataset.

Changes in the manuscript: “This retrospective study involved the analysis of fully de-identified data from three publicly available research datasets (OAI, MOST, and ADNI) and one institutional dataset (NYU Langone Health). The Osteoarthritis Initiative (OAI) (ClinicalTrials.gov identifier: NCT00080171) was approved by the Institutional Review Boards (IRB) at the UCSF Coordinating Center (approval #10-00532) and all clinical sites, including Memorial Hospital of Rhode Island, Ohio State University, University of Pittsburgh, and University of Maryland/Johns Hopkins University. Ethical approval for the Multicenter Osteoarthritis Study (MOST) was obtained from the IRBs at Boston University (H-32956), University of Alabama at Birmingham (IRB-000329007), University of California San Francisco (301480), and University of Iowa (201511711); approval for secondary data analysis was granted by the University of Florida (IRB202201899). Data for the Alzheimer’s Disease Neuroimaging Initiative (ADNI ClinicalTrials.gov identifier: NCT00106899) were obtained in accordance with the Declaration of Helsinki and approved by the IRBs of all participating sites. All participants in these public cohorts provided written informed consent. The Breast Cancer dataset was sourced from New York University Langone Health under a protocol approved by the NYU Langone Health IRB (IRB00010481, S18-00712). This dataset consisted of fully de-identified data and was approved for retrospective analysis with a waiver of informed consent, in compliance with HIPAA regulations. As this study involved only the retrospective analysis of fully de-identified data, it did not require new participant recruitment or additional IRB approval beyond the existing protocols cited above.“

Comment 2: Line 112: The abbreviation for ADNI should be introduced here, not at line 159.

Response: We have corrected this oversight. The abbreviation "(ADNI)" is now introduced at the first mention of the "Alzheimer’s Disease Neuroimaging Initiative" in the text (Section 2.1).

Comment 3: Lines 162–163: Since the abbreviations for Cognitive Normal (CN), Mild Cognitive Impairment (MCI), and Alzheimer’s Disease (AD) have already been defined earlier, please use the abbreviations directly in this section.

Response: We have updated Section 2.2.2 to use the abbreviations CN, MCI, and AD directly, as they were defined previously.

Comment 4: Discussion: The manuscript would benefit from a dedicated Limitations section, as several inherent limitations appear to be present.

Response: We agree that a transparent discussion of limitations is vital. We have added a dedicated subsection 4.2. Limitations within the Discussion section and expanded the previously mentioned Limitations.

Changes in the manuscript: The updated limitations subsection is as follows:

“Several limitations warrant discussion. First, our knee OA and AD datasets derive from standardized research protocols, while the breast cancer dataset originates from a single institution. These data sources may lack the variability of routine clinical practice and therefore affect generalizability to diverse care settings. Second, the inherent class imbalance (few progressors versus many stable patients) in progressive disease studies affects all conditions studied. While we employed balanced sampling and specific evaluation metrics (AUPRC) to mitigate this, the imbalance remains a challenge for model deployment.

Third, a practical consideration for the sequential learning with experience replay strategy is its reliance on access to the original diagnosis dataset, which may not always be available when models are shared. It is important to note, however, that this limitation applies only to the multitask fine-tuning stage. The initial performance gains from using a diagnosis-pretrained model for standard fine-tuning - our first key finding remains an applicable benefit of our work, even without access to the original data for replay.

Finally, while experience replay effectively mitigates catastrophic forgetting, it increases the computational cost of training compared to simple fine-tuning, as the model must process 'replay' batches alongside new data.”

Comment 5: A thorough English language review is recommended to further improve clarity, flow, and readability.

Response: We have carefully reviewed the entire manuscript and edited the text to improve clarity, flow, and readability as requested. We have tightened the phrasing in the Abstract and Introduction to ensure that key technical concepts are defined clearly. We have also corrected minor grammatical inconsistencies, improved sentence transitions, and refined the vocabulary throughout the Methods and Results sections to ensure a professional academic tone.

Changes in the manuscript: Below are a few representative examples of these revisions:

Section

Original Text

Revised Text

Abstract

"...hindered by the lack of long-term data."

"...hindered by the scarcity of longitudinal data."

Introduction

"Progressive diseases pose significant challenges... Their gradual onset and potential for irreversible damage..."

"Progressive diseases present significant challenges... Due to their gradual onset and potential for irreversible damage..."

Methods

"All models were trained using the Adam optimizer [33]. A cosine annealing learning rate scheduler was used..."

"All models were trained using the Adam optimizer [33], coupled with a cosine annealing learning rate scheduler..."

Results

"...weights pretrained on diagnostic tasks with full diagnosis cohorts**,** yielded improvements..."

"...weights pretrained on diagnostic tasks with full diagnosis cohorts yielded improvements..."

Results

"This performance benefit on the external MOST dataset"

"Additionally, this performance benefit on the external MOST dataset"

Discussion

"...exhibited lower variance, indicating more stable performance."

"...exhibited lower variance, suggesting more stable performance."

Discussion

"...mirrors routine workflow, assess current status, then plan forward."

"...mirrors routine workflow: assessing current status and planning forward."

Reviewer #2:

Comment: A very convincing paper with sound findings & elaboration on the significance of the findings.

Response: We thank Reviewer 2 for their positive assessment and we appreciate their encouraging feedback.

Comment 1:A more recent literature would strengthen the justifications of the findings. Critical arguments are still lacking in some parts of discussions.

Response: We agree that connecting our findings to the latest developments in medical AI strengthens the manuscript. We have added targeted, up-to-date citations in both the Introduction and Discussion to contextualize our approach relative to recent work on diagnostic-scale datasets and foundation models. Specifically, we now cite recent OA-focused AI literature and dataset-driven trends across progressive diseases [23–25], and we reference recent reviews and examples of medical foundation models that leverage large-scale diagnostic pretraining for improved downstream performance under limited labeled data [26,27]. In the Discussion (Section 4.1), we additionally connect our empirical findings to the broader paradigm of foundation models and multimodal AI [24,40], and we explicitly argue that our results extend these ideas to longitudinal prognosis, where labels are scarce and temporal prediction is challenging. Finally, to strengthen critical argumentation, we added a dedicated Limitations subsection (Section 4.2) that discusses key practical trade-offs, including generalizability, class imbalance, reliance on access to diagnostic data for replay, and the additional computational cost of experience replay.

Changes in the manuscript:

Introduction:

“However, developing robust prognosis models faces a significant hurdle: the scarcity of longitudinal data required to track disease progression over time. This contrasts with diagnosis, where larger cross-sectional datasets are often more readily available. As a result, recent advances in medical imaging AI - across knee osteoarthritis, Alzheimer’s disease, and oncology have largely focused on improving diagnostic accuracy and disease severity assessment, reflecting the scale and availability of diagnostic datasets rather than long-horizon progression labels [23-25]. To mitigate this data scarcity, previous works have explored multitask learning (MTL) as additional regularization, where a single model is trained to perform several related tasks simultaneously, aiming to improve generalization and performance through shared representations [13, 14, 16, 21, 22]. For instance, MTL combining diagnosis and prognosis has shown benefits in OA [14, 16], and pretraining on BI-RADS has been shown to improve current cancer diagnosis performance [13]. Yet these MTL studies typically rely on relatively small prognosis datasets and do not fully capitalize on the abundance of diagnostic data, even as recent medical foundation-model approaches increasingly leverage large-scale diagnostic pretraining with limited labeled data [26, 27]. This study therefore hypothesizes that sequential learning - diagnosis pretraining on large, unbiased datasets followed by prognostic training, achieves a favorable balance of prognostic and diagnostic performance.”

Discussion

“Recent work on generalist and multimodal medical AI suggests that representations learned from large-scale diagnostic data can serve as effective foundations for a wide range of downstream diagnostic and classification tasks, particularly in settings with limited labeled data [24,40]. Our results provide empirical evidence that this paradigm extends specifically to disease prognosis, where longitudinal labels are especially scarce and difficult to collect. By leveraging large, cross-sectional diagnostic cohorts, our approach improves prognostic discrimination and stability while preserving diagnostic competence, addressing a key gap between foundation-model principles and real-world longitudinal prediction tasks. ”

New Papers Cited:

[23] Harkey MS, Costello KE, Mehta B, Wen C, Malfait AM, Madry H, et al. Artificial intelligence in osteoarthritis research: summary of the 2025 OARSI pre-congress workshop. Osteoarthritis and Cartilage Open. 2025;100687.

[24] Moor M, Banerjee O, Abad ZS, Krumholz HM, Leskovec J, Topol EJ, et al. Foundation models for generalist medical artificial intelligence. Nature. 2023 Apr 13;616(7956):259-65.

[25] Weiner MW, Kanoria S, Miller MJ, Aisen PS, Beckett LA, Conti C, Diaz A, Flenniken D, Green RC, Harvey DJ, Jack Jr CR. Overview of Alzheimer's Disease Neuroimaging Initiative and future clinical trials. Alzheimer's & Dementia. 2025 Jan;21(1):e14321.

[26] D’Antonoli TA, Bluethgen C, Cuocolo R, Klontzas ME, Ponsiglione A, Kocak B. Foundation models for radiology: fundamentals, applications, opportunities, challenges, risks, and prospects. Diagnostic and Interventional Radiology. 2025.

[27] Ma, DongAo, et al. "A fully open AI foundation model applied to chest radiography." Nature (2025): 1-11.

[40] Acosta JN, Falcone GJ, Rajpurkar P, Topol EJ. Multimodal biomedical AI. Nature medicine. 2022 Sep;28(9):1773-84.

Further, the updated Limitations subsection is as follows:

“Several limitations warrant discussion. First, our knee OA and AD datasets derive from standardized research protocols, while the breast cancer dataset originates from a single institution. These data sources may lack the variability of routine clinical practice and therefore affect generalizability to diverse care settings. Second, the inherent class imbalance (few progressors versus many stable patients) in progressive disease studies affects all conditions studied. While we employed balanced sampling and specific evaluation metrics (AUPRC) to mitigate this, the imbalance remains a challenge for model deployment.

Third, a practical consideration for the sequential learning with experience replay strategy is its reliance on access to the original diagnosis dataset, which may not always be available when models are shared. It is important to note, however, that this limitation applies only to the multitask fine-tuning stage. The initial performance gains from using a diagnosis-pretrained model for standard fine-tuning - our first key finding remains an applicable benefit of our work, even without access to the original data for replay.

Finally, while experience replay effectively mitigates catastrophic forgetting, it increases the computational cost of training compared to simple fine-tuning, as the model must process 'replay' batches alongside new data.”

Attachment

Submitted filename: Response to reviewers.docx

pone.0344600.s011.docx (19.3KB, docx)

Decision Letter 1

Edward Hoffer

23 Feb 2026

Robust disease prognosis via diagnostic knowledge preservation: A sequential learning approach

PONE-D-25-52607R1

Dear Dr. Rajamohan,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice will be generated when your article is formally accepted. Please note, if your institution has a publishing partnership with PLOS and your article meets the relevant criteria, all or part of your publication costs will be covered. Please make sure your user information is up-to-date by logging into Editorial Manager at Editorial Manager® and clicking the ‘Update My Information' link at the top of the page. For questions related to billing, please contact billing support.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Edward Hoffer, MD

Academic Editor

PLOS One

Additional Editor Comments (optional):

Reviewers' comments:

Acceptance letter

Edward Hoffer

PONE-D-25-52607R1

PLOS One

Dear Dr. Rajamohan,

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS One. Congratulations! Your manuscript is now being handed over to our production team.

At this stage, our production department will prepare your paper for publication. This includes ensuring the following:

* All references, tables, and figures are properly cited

* All relevant supporting information is included in the manuscript submission,

* There are no issues that prevent the paper from being properly typeset

You will receive further instructions from the production team, including instructions on how to review your proof when it is ready. Please keep in mind that we are working through a large volume of accepted articles, so please give us a few days to review your paper and let you know the next and final steps.

Lastly, if your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

You will receive an invoice from PLOS for your publication fee after your manuscript has reached the completed accept phase. If you receive an email requesting payment before acceptance or for any other service, this may be a phishing scheme. Learn how to identify phishing emails and protect your accounts at https://explore.plos.org/phishing.

If we can help with anything else, please email us at customercare@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. Edward Hoffer

Academic Editor

PLOS One

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 Fig. Comparative multitask learning approaches.

    Visualizations of (a) Single Cohort MT, (b) Diagnosis Pretrained MT, and (c) Concurrent MT.

    (TIF)

    pone.0344600.s001.tif (641.3KB, tif)
    S1 Table. Patient Characteristics in Structural Progression OAI and MOST Cohorts.

    (DOCX)

    pone.0344600.s002.docx (15.4KB, docx)
    S2 Table. Patient Characteristics in Structural Incidence OAI and MOST Cohorts.

    (DOCX)

    pone.0344600.s003.docx (15.6KB, docx)
    S3 Table. Demographic Characteristics and Cognitive Status Distribution across the ADNI Prognosis Cohort.

    (DOCX)

    pone.0344600.s004.docx (14.5KB, docx)
    S4 Table. Breast cancer Patient distribution in the Prognosis Cohort.

    (DOCX)

    pone.0344600.s005.docx (14.5KB, docx)
    S5 Table. Comparison of AUROC performance for models trained on progression and incidence, using different initializations.

    (DOCX)

    pone.0344600.s006.docx (14.2KB, docx)
    S6 Table. Detailed AUROC analysis comparing model performance across cognitive status subgroups.

    (DOCX)

    pone.0344600.s007.docx (14.1KB, docx)
    S7 Table. Detailed metrics for various training strategies for progression and incidence prediction tasks on OAI and MOST Datasets.

    (DOCX)

    pone.0344600.s008.docx (14.7KB, docx)
    S8 Table. Detailed progression prediction performance (AUROC) across cognitive status subgroups for multitask methods.

    (DOCX)

    pone.0344600.s009.docx (16.5KB, docx)
    Attachment

    Submitted filename: Response to reviewers.docx

    pone.0344600.s011.docx (19.3KB, docx)

    Data Availability Statement

    - Osteoarthritis Initiative (OAI): OAI data are publicly available through the NIMH Data Archive (NDA). Access requires creating an NDA account, agreeing to the OAI data access terms, and requesting the desired collections through the NDA portal. Full instructions are provided at https://nda.nih.gov/oai. - Multicenter Osteoarthritis Study (MOST): MOST data are publicly available through the NIA Aging Research Biobank. Access requires creating a Biobank account and submitting a data request in accordance with the Biobank’s terms and conditions. More information is available at https://agingresearchbiobank.nia.nih.gov/. - Alzheimer’s Disease Neuroimaging Initiative (ADNI): ADNI data are publicly available from the LONI Image & Data Archive upon completion of a web application and acceptance of the ADNI Data Use Agreement. Data can be accessed at https://adni.loni.usc.edu/. - Breast Cancer Cohort (NYU Langone Health): This institutional dataset contains protected health information and cannot be made publicly available. De-identified data may be shared upon reasonable request, pending approval by NYU Langone’s data governance committee. Inquiries can be directed to Krzysztof Geras (k.j.geras@nyu.edu). - Code and Model Availability: All code for preprocessing and model training sufficient to reproduce the analyses reported in this study are available at https://github.com/denizlab/diag-to-prog-replay.


    Articles from PLOS One are provided here courtesy of PLOS

    RESOURCES