Skip to main content
Nature Communications logoLink to Nature Communications
. 2025 Oct 24;16:9423. doi: 10.1038/s41467-025-64469-w

Modality-projection universal model for comprehensive full-body medical imaging segmentation

Yixin Chen 1,#, Lin Gao 2,#, Yajuan Gao 3,4, Rui Wang 5, Jingge Lian 3, Xiangxi Meng 6, Yanhua Duan 2, Leiying Chai 2, Hongbin Han 1,3,4, Zhaoping Cheng 2,, Zhaoheng Xie 1,7,
PMCID: PMC12552709  PMID: 41136445

Abstract

The integration of deep learning in medical imaging has significantly advanced diagnostic, therapeutic, and research outcomes. However, applying universal models across multiple modalities remains challenging due to inherent inter-modality variability. Here we present the Modality Projection Universal Model (MPUM), trained on 861 subjects, which dynamically adapts to diverse imaging modalities through a modality-projection strategy. MPUM achieves state-of-the-art, whole-body organ segmentation, providing rapid localization for computer-aided diagnosis and precise anatomical quantification to support clinical decision-making. A controller-based convolutional layer further enables saliency map visualization, enhancing model interpretability for clinical use. Beyond segmentation, MPUM reveals metabolic correlations along the brain-body axis and between distinct brain regions, providing insights into systemic and physiological interactions from a whole-body perspective. Here we show that this universal framework accelerates diagnosis, facilitates large-scale imaging analysis, and bridges anatomical and metabolic information, enabling discovery of cross-organ disease mechanisms and advancing integrative brain-body research.

Subject terms: Translational research, Diagnostic markers


Applying universal AI models across multiple modalities remains challenging. Here, the authors present a model that enables multi-modality whole-body segmentation for computer-aided diagnosis, revealing metabolic correlations across brain, body, and brain–body systems.

Introduction

Universal models, characterized by their ability to generalize across diverse tasks without fine-tuning, have emerged as a powerful framework in many fields. By leveraging shared representations1, these models offer unparalleled adaptability across a wide range of applications. In medical imaging, universal models have gained attention for their potential to generalize across various anatomical regions, imaging modalities, and clinical tasks2. However, significant challenges arise due to intrinsic modality differences and the complexity of anatomical structures involved. In response, several large-scale datasets, such as TotalSegmentator3,4 and Dense Anatomical Prediction (DAP)5, various universal challenges like BodyMaps246 and the Universal Lesion Segmentation7, as well as related models like MedSAM2, CDUM8, PCNet9, TotalSegmentator3, STUNet10, and LUCIDA11 have collectively advanced the development of universal medical imaging models. These models support efficient multi-task processing by enabling accurate organ identification, clinical diagnostics, report generation, and treatment monitoring, while also reducing the time and effort required for manual annotation. This brings substantial benefits to both clinical practice and research.

Similar to universal models, foundation models represent an alternative approach to developing generalizable AI in medical imaging. Foundation models are typically pre-trained using self-supervised learning on large volumes of unannotated or weakly annotated data1215, whereas universal models are trained directly on annotated datasets spanning multiple tasks. Despite these differences, both approaches aim to generate high-dimensional representations that can generalize across various tasks. However, foundation models face inherent limitations, particularly in multi-category segmentation tasks, where the binary nature of contrastive learning is less effective. Additionally, foundation models typically require task-specific fine-tuning, which compromises their practicality as truly versatile tools in medical applications. In contrast, universal models demand large annotated datasets, which is a major bottleneck since high-quality labels in medical imaging are labor-intensive16,17. Despite these limitations, exploring universal models is crucial because they offer the potential for truly versatile tools in medical applications without the need for task-specific fine-tuning8,18,19, capable of handling multiple tasks simultaneously and streamlining clinical workflows.

With the efforts of the communities, annotated medical multi-task datasets have become increasingly mature. For example, the TotalSeg dataset3 includes segmentation annotations for 104 types of CT anatomical structures from 1204 unique subjects. The DAP dataset5 covers 133 types of anatomical structures with 533 CT scans. The CDUM dataset8 combines multiple single-task CT datasets to build a multi-task model capable of recognizing 25 organs and 6 types of tumors. The TotalSegMRI dataset4 covers 59 anatomical structures from 298 MR scans. Compared to single-task segmentation tasks, multi-task training not only scales dataset size but also leverages task synergy. This is evident in how shared boundaries among adjacent tissues lead to improved6,9,10,20.

Currently, research on cross-modal foundation/universal segmentation has followed four paradigms: prompt-driven models, structure-adaptive models, native 3D models, and few-shot/zero-shot models. Prompt-driven models like MedSAM2, adapt Segment-Anything with point- or box-based cues and can tackle more than 140 tasks, but their dependence on user interaction and diminished accuracy on irregular organs limit full automation. Structure-adaptive networks (e.g., SPADNet21 and UniSeg22) mix expert kernels at inference, yet their encoders receive modality information only after feature extraction and the publicly released models recognize just seven anatomies. Few- or zero-shot approaches typified by UniverSeg18 generalize to unseen organs through a 2D support set, although whole-body scans would require thousands of slices and the accuracy still lags behind fully supervised baselines. Native 3D foundations like TotalSegmentator3 and VISTA3D23 operate directly on voxels and cover more than 100 classes, but each is trained on a single modality and therefore inherits modality-specific bias. CDUM8 integrates text embeddings into segmentation models to capture anatomical relationships. STUNet10 is a scalable and transferable UNet model series, with sizes ranging from 14 million to 1.4 billion parameters. SAT19 is designed to segment a wide array of medical images using text prompts. PCNet9 utilizes prior category knowledge to guide the universal segmentation model in capturing inter-category relationships. Among the existing research on universal medical models, only SAT is trained as a multi-task model on multi-modality data, whereas others are focused on single-modality data. Although multi-modality datasets are larger and potentially more powerful, they introduce challenges like conflicting feature distributions and increased training instability.

To address these limitations, we propose a versatile medical segmentation model based on a multi-modality projection mechanism. This mechanism allows for the extraction of modality-specific features from a shared high-dimensional space, enabling generalization across different imaging modalities without fine-tuning. Each organ has a high-dimensional latent feature, which can be projected in different directions to representations. This multi-modality projection allows for a unified understanding across different imaging techniques. Our modality projection universal model (MPUM) was trained using data from 861 unique subjects. Figure 1 illustrates our key methodological innovations and experimental designs. Our proposed MPUM is designed with two key characteristics: precise brain segmentation and comprehensive whole-body segmentation. To demonstrate the clinical impact of MPUM, we focus on three case studies: identification (Case 1), diagnosis (Case 2), and analysis (Case 3).

  • Case 1 (Fig. 1c): Technical validation of segmentation performance, comparing MPUM to other advanced universal models.

  • Case 2 (Fig. 1d): MPUM’s role as a computer-aided diagnosis tool for intracranial hemorrhage (ICH) localization in CT scans.

  • Case 3 (Fig. 1e): MPUM as a comprehensive analysis tool. Both epilepsy and Alzheimer’s Disease (AD), which are marked by structural-metabolic inter-brain correlations, benefit from whole-body PET/CT analysis.

Fig. 1. Overview of the modality projection universal model (MPUM).

Fig. 1

a Training process of the MPUM leveraging data from three distinct modalities. b Comparison of two common multi-modality data training strategies with our proposed modality-projection strategy. c Application of the MPUM as an aided identification tool across three modalities (over 200 categories). d The MPUM is utilized as a computer-aided diagnosis (CAD) tool for precise localization of intracranial hemorrhage with CT scans. e Application of the MPUM as an aided analysis tool in identifying altered metabolic correlations in regions affected by epilepsy and AD. f, Additional experimental results, including t-SNE feature visualizations and saliency map analysis.

For ICH, our model addresses a critical challenge in the emergency room, where rapid diagnosis is vital, but radiologists may experience delays. MPUM enhances diagnostic efficiency by accurately identifying hemorrhages in CT scans, enabling quicker decision-making and timely intervention. Additionally, our team’s long-standing focus on the brain-body axis24, which regulates multiple physiological systems. MPUM supports comprehensive metabolic analysis in neurological conditions. By revealing brain-brain and brain-body metabolic associations, MPUM has the potential to uncover systemic biomarkers and deepen understanding of disease, such as epilepsy and Alzheimer’s Disease (AD). Overall, the MPUM has transformative potential for clinical workflows, reducing manual annotation and improving diagnostic accuracy.

Results

We developed a deep learning-based multi-modality universal segmentation model. In the identification tasks, the model demonstrates superior performance in anatomical structure identification, outperforming existing segmentation models in terms of Dice and surface Dice metrics. In aided diagnosis, it accurately detected and quantified intracranial hemorrhages in CT scans, significantly improving diagnostic accuracy among general practitioners in a clinical setting. For aided analysis, the model facilitated the identification of significant metabolic alterations in epilepsy and AD, revealing both brain-brain and brain-body associations across the body. Furthermore, this section includes interpretable insights through visualization of feature operators and saliency maps.

Identification: anatomical structure identification (Case 1)

We compared MPUM with state-of-the-art universal segmentation models, including CDUM8, PCNet9, and STUNet10, as well as the UNet25 with parameters adjusted to match the complexity of the other models. PCNet9, MPUM, and STUNet10 each have approximately 60M trainable parameters, whereas the classic 3D UNet contains around 19M. We scaled the classic 3D UNet by increasing the number of feature channels at each stage by a factor of 1.75 (e.g., 32 → 56, 64 → 112, …, 512 → 896). This brought its parameter count to approximately 59M. By normalizing model capacity, we ensure that the performance differences genuinely reflect innovations in architecture and design, rather than differences in model complexity. In addition, we evaluated three training strategies for multi-modality data: modality-mixed, modality-specific, and our proposed modality-projection strategy. To make a fair and comprehensive comparison, we applied a consistent training protocol across all models. Specifically, we adopted standardized preprocessing (e.g., resampling to 2 mm isotropic resolution), patch-based training with 128 × 128 × 128 input size, and augmentation strategies such as random Gaussian smoothing and contrast adjustment. All networks were optimized with the Adam optimizer (initial learning rate 3e − 4, weight decay 3e − 5), with the learning rate reduced by a factor of 10 whenever the validation loss plateaued. We used a batch size of 4 and trained for up to 200 epochs, employing early stopping based on validation performance to avoid overfitting. The training objective was a combination of categorical cross-entropy and soft Dice loss. Each network’s weights were initialized with the He-normal method and trained from scratch.

As shown in Fig. 2a, for MRI body segmentation, our projection strategy outperformed all other approaches, achieving the highest Dice score of 0.7751 and the highest surface Dice score of 0.5471. In comparison, the best-performing mixed strategy model, STUNet, achieved a Dice of 0.7560 and a surface Dice of 0.5100, while the best modality-specific strategy, STUNet, reached a Dice of 0.7627 and a surface Dice of 0.5211.

Fig. 2. Performance comparison of MPUM with advanced segmentation models.

Fig. 2

a Dice score and surface Dice metrics for CT (left), MR (middle), and PET (right) segmentation across UNet, CDUM, PCNet, and STUNet using three strategies: modality-mixed, modality-specific, and modality-projection (used by MPUM). b The impact of multi-modality training on brain and body segmentation tasks, reported in terms of Dice performance. Source data are provided as a Source Data file.

Similarly, in the CT body segmentation, the projection strategy again demonstrated superior performance, with a Dice score of 0.8517 and a surface Dice score of 0.8506. The closest competitor, the modality-specific STUNet, achieved a Dice of 0.8462 and a surface Dice of 0.8395, whereas the mixed strategy STUNet model recorded a Dice of 0.8394 and a surface Dice of 0.8288.

In the case of CT brain segmentation, the projection strategy continued to yield the highest scores, with a Dice of 0.7419 and a surface Dice of 0.6872. The highest scores among the other strategies were from the modality-specific PCNet, which achieved a Dice of 0.6540 and a surface Dice of 0.5982, and the mixed strategy STUNet, with a Dice of 0.6318 and a surface Dice of 0.5417.

In Fig. 2, we compare three multi-modality training strategies: (1) modality-mixed, where a single model is trained on all modalities with shared parameters; (2) modality-specific, where separate models are trained per modality; and (3) our proposed modality-projection strategy, where shared latent representation is dynamically projected into modality-specific convolutional kernels via a controller module. While the modality-mixed approach benefits from data scale, it suffers from conflicting optimization gradients across modalities. The modality-specific strategy avoids such conflict but misses cross-modality synergy. In contrast, our projection method resolves this tension by learning a shared latent representation and dynamically projecting it into modality-specific convolutional kernels using a controller module. As shown in Fig. 2a, this approach achieves the highest Dice and surfaceDice scores across all benchmarks (e.g., MRI body: 0.7751 vs. 0.7560/0.7627; CT brain: 0.7419 vs. 0.6318/0.6540), confirming its robustness and effectiveness.

We further benchmarked MPUM against two universal segmentation frameworks (TotalSegmentator and VISTA3D) on a single organ analysis in Supplementary Table 2. MPUM attains the highest Dice in the majority of the categories, with statistically significant gains (p < 0.05) in over 60 organs. It matches or slightly outperforms existing models on high-contrast targets like the liver and spleen, while showing notable advantages in challenging structures such as vessels (e.g., inferior vena cava: 0.903 vs. 0.893/0.872), musculoskeletal regions (e.g., gluteus maximus: 0.951 vs. 0.939/0.920), and small abdominal organs (e.g., colon: 0.904 vs. 0.861/0.834). For interactive segmentation frameworks like MedSAM2, manual box prompts segment some structures accurately (e.g., left kidney, right rib 10), but its segmentation of the liver and the left autochthon was less precise than the outputs of TotalSegmentator, VISTA3D, and MPUM (see Supplementary Fig. 6). For the spleen, its irregular shape forced the box prompt to include adjacent rib, causing MedSAM to mis-segment the rib and fail to isolate the spleen. For these reasons, we excluded SAM-based models from our automated benchmark. The details are shown in Supplementary Table 2.

Diagnosis: intracranial hemorrhage (Case 2)

Intracranial hemorrhage (ICH) is a life-threatening emergency where quick diagnosis and treatment decisions are critical for patient survival2628. ICH has a high early mortality, with up to 40% of patients dying within a year of the event29,30, and a significant proportion of survivors suffering from lasting functional impairments31. Quick and accurate identification of ICH is essential for timely medical intervention, which can significantly improve patient outcomes3234.

Recent studies highlight that precise ICH localization significantly influences management and outcomes. Brainstem hemorrhages carry the worst prognosis, as they are significantly associated with higher mortality35. Furthermore, patients with brainstem hemorrhages experience significantly worse health-related quality of life, particularly in the pain domain, compared to those with cerebellar hemorrhages35. Intraventricular extension with resulting hydrocephalus markedly worsens outcome: a meta-analysis reveals that patients with ICH, IVH, and hydrocephalus have substantially higher 30- and 90-day mortality rates compared to those with ICH alone36. Likewise, lobar (cortical) ICH is associated with higher early mortality than deep hemorrhages (26.7% vs. 16.5% in one cohort)37. Pooled analyses indicate early surgical evacuation of lobar ICH does not significantly improve functional outcome38. In contrast, basal ganglia or thalamic hemorrhages can benefit from surgery in selected cases: a large cohort study demonstrated that patients with good preoperative motor function who underwent hematoma evacuation had favorable neurological outcomes at 3 months39. Together, these findings highlight the significant prognostic and therapeutic implications of distinguishing between lobar, deep, cerebellar, brainstem, and ventricular ICH.

The goal of Case 2 is to enable the MPUM as a clinical aid for fast and accurate ICH diagnosis. The experimental setup for Case 2 consists of two parts: firstly, we tested 100 cases on the Instance2022 dataset40,41 and conducted statistical analyses on the volume of brain regions affected by ICH. Secondly, we validated the effectiveness of the MPUM with the validation of three senior radiologists and three junior radiologists on the in-house ICH dataset from emergency department.

Our MPUM framework was fine-tuned on the Instance2022 dataset to identify hemorrhagic areas from CT head scans. As depicted in Fig. 3a, the original pre-trained MPUM provides brain regions maps from CT scans, while the fine-tuned MPUM yields segmentation map for hemorrhages. By integrating these two predictions, we obtained precise aided diagnosis results. For instance, in the Fig. 3a, the model autonomously detected a hemorrhage involving 3868 mm3 in the left insula and 2241 mm3 in the left putamen. Precise quantification of ICH through CT scans plays a critical role in clinical decision-making and treatment planning33. Accurate volume measurements are essential for determining the extent of hemorrhage, monitoring its progression, and evaluating the risk of further complications like hematoma expansion, which is closely associated with worse outcomes32. We conducted a detailed analysis of hemorrhage volumes across various brain regions (see Supplementary Fig. 1). Some regions, such as the insula, show relatively higher average volumes, which aligns with clinical observations where certain brain areas are more prone to larger hemorrhages due to their vascular structures and the prevalence of small vessel disease32,42. This detailed quantification has proven crucial for developing more targeted therapies and enhancing diagnostic accuracy43.

Fig. 3. Performance of the MPUM framework as aided diagnosis tool.

Fig. 3

a MPUM assists in hemorrhage detection and brain region mapping from CT head scans, facilitating accurate diagnosis. The red-shaded region indicates an intracranial hemorrhage. b MPUM enhances diagnostic performance and supports less-experienced radiologists in real-world clinical settings. Source data are provided as a Source Data file.

Additionally, we collected 28 ICH CT scans, along with diagnostic reports, from the emergency department of an external medical center. The reports contain precise diagnostic results assisted by MRI imaging, which is regarded as ground truth. For the diagnostic accuracy test, three experienced radiologists with more than five years of experience and three general radiologists analyzed these cases, assessing the brain regions affected by ICH. Each radiologist should determine one or multiple hemorrhagic regions from the following categories: frontal lobe hemorrhage, temporal lobe hemorrhage, parietal lobe hemorrhage, occipital lobe hemorrhage, basal ganglia hemorrhage, cerebellar and brainstem hemorrhage, subarachnoid hemorrhage, subdural hemorrhage, and ventricular hemorrhage. We adopted a rigorous experimental design:

  • CT only. Both junior and senior radiologists reviewed raw CT volumes without any computer assistance. These results serve as baselines, with junior performance as the reference for measuring improvement and senior performance as the expert benchmark.

  • CT + lesion mask. Junior radiologists were provided with lesion masks generated by Viola44, the top-performing model from the 2022 ICH Segmentation Challenge. Comparing this setting with the CT-only baseline quantifies the benefit of automated hemorrhage detection.

  • CT + lesion mask + MPUM brain atlas. Junior radiologists received both the lesion mask and the MPUM-derived brain region map. This setting evaluates whether providing anatomical localization in addition to lesion detection further improves diagnostic accuracy by assisting spatial reasoning and regional interpretation.

As summarised in Fig. 3b, without any assistance, senior radiologists achieved accuracies of 100%, 92.9%, and 85.7%, whereas junior radiologists reached only 78.6%, 75.0%, and 60.7%. Most junior radiologists’ errors stemmed from incomplete identification of all hemorrhage regions and confusion between adjacent brain areas. Supplying an automated lesion mask raised junior performance (e.g., Junior-1 from 78.6% to 89.3%, with similar gains for Junior-2 and Junior-3). When the MPUM brain-region atlas was also provided, their accuracies climbed further to 96.4%, 85.7%, and 89.3%, underscoring the utility of precise regional localization and the diagnostic value of the MPUM framework.

The automated lesion mask directs radiologists to voxels with high hemorrhage probability, shortening search time and minimizing the risk of missing small or low-contrast foci, thereby reducing false-negatives. The MPUM atlas provides a detailed anatomical map that clearly delineates boundaries, such as the interface between the temporal lobe and the lateral ventricle. Junior doctors often struggle to identify these boundaries accurately. By linking each highlighted lesion voxel to a specific neuroanatomical region, the atlas reduces region-of-interest confusion (e.g. mistaking parietal hemorrhage for ventricular bleeding), thereby boosting localization accuracy and specificity. In short, the lesion mask addresses the question “Where is the bleed?”, while the atlas identifies its exact neuro-anatomical location. Their combination replicates the workflow of expert radiologists and explains the stepwise performance gains observed in the study.

Analysis: multi-organ metabolic associations (Case 3)

We assessed the efficacy of our universal model as a comprehensive analysis tool across neurological disorders, examining how epilepsy and AD affect metabolic associations. PET/CT enables simultaneous, non-invasive metabolic assessment of multiple organs, facilitating the study of systemic and inter-organ interactions45,46. This capability is crucial for elucidating complex diseases through metabolic connectivity studies that span both brain networks and whole-body physiology.

Our universal model facilitates rapid identification of regions of interest (ROIs) in human tissue structures, significantly reducing the manpower costs. We analyzed whole-body PET/CT data from a control group (n = 22) and a group of patients without active epilepsy episodes (n = 50). Utilizing our pre-trained universal model, we identified 215 ROIs in the CT scans (details in Supplementary Table 3). From 215 ROIs, we excluded 12 ROIs due to high background radioactivity, including the bladder, inferior vena cava, aorta, and pulmonary artery, leaving 203 effective ROIs. We defined metabolic associations between two different regions as one pair, such as the left kidney and liver. In total, we analyzed 20,503 pairs for metabolic associations, which comprised 3403 brain-brain pairs, 7140 body-body pairs, and 9960 brain-body pairs. Given that epilepsy is a brain-triggered disorder, our analysis focused on brain-related pairs, excluding the 7140 body-body pairs.

We examined whether the metabolic associations in each pair were significantly altered due to epilepsy. Among the 3,403 brain-brain pairs, 228 pairs showed significant changes in correlation due to epilepsy (p < 0.001). Notably, 108 pairs involved the ‘right anterior temporal lobe lateral part’, and 68 pairs involved the ‘right middle and inferior temporal gyrus’ (see Supplementary Fig. 2a, b). These findings suggest a strong link between the metabolic activity in these brain areas and epilepsy since epilepsy significantly alters the correlation. Figure 4a illustrates the metabolic associations of the ‘right Anterior temporal lobe lateral part’ with other brain regions. For the control group, there are high correlations between the ‘right anterior temporal lobe lateral part’ and other brain regions, whereas the patient group showed weak metabolic connections. Figure 4a also shows a similar phenomenon for the ‘right middle and inferior temporal gyrus’. Previous studies indicate that both regions are common sites for epileptic foci4749.

Fig. 4. Multi-organ metabolic association analysis for pediatric epilepsy based on the universal model.

Fig. 4

We analyzed the metabolic associations between the epilepsy patient group (n = 50) and the control group (n = 22), using Fisher Z-Transformation to calculate the significance of differences in Pearson correlation coefficients. a Schematic representation of the connectivity among brain regions associated with the right anterior temporal lobe lateral part and the right middle and inferior temporal gyrus. The left diagrams illustrate the strong metabolic connection within the control group. Notably, these correlations are statistically significantly reduced in the patient group (p < 0.001). b Altered metabolic connectivity between the Pallidum and Vertebrae T1-T12 affected by epilepsy.

To illustrate the broader applicability of our model in neurodegenerative disease, we conducted an additional analysis using brain 18F-FDG PET data from the ADNI cohort (97 cognitively normal [CN], 328 mild cognitive impairment [MCI], and 42 Alzheimer’s disease [AD] subjects). Due to the lack of publicly available whole-body PET/CT scans for AD patients, we utilized ADNI brain-only data. Following the same pipeline, we included age adjustment and FDR correction. As shown in Supplementary Fig. 8a, b, in the CN vs. MCI comparison, altered associations emerged in the middle and inferior temporal gyri and posterior superior temporal gyrus. These regions are well-known to distinguish MCI or early AD converters from healthy controls50. In the MCI vs. AD comparison (Supplementary Fig. 8c, d), the most notable changes were found in the inferior lateral parietal cortex and postcengbtral gyrus. These regions are consistent with prior 18F-FDG PET findings showing parietal hypometabolism as a hallmark of early AD50,51. These results demonstrate the potential of our model to uncover clinically relevant, disease-associated imaging biomarkers.

We also analyzed metabolic associations between brain regions and body across 9960 pairs in the epilepsy cohort. Interestingly, we identified significant changes in metabolic associations in 14 pairs (p < 0.001). As illustrated in Fig. 4b, a notable metabolic association change was observed between the ‘pallidum’ and ‘vertebrae T1 to T12’. This finding suggests a significant disruption in metabolic connectivity, potentially due to neuronal dysfunctions in epilepsy. These changes in metabolic connectivity, along with their correlation coefficients, are detailed in Supplementary Fig. 2c.

Furthermore, we conducted a significant analysis of metabolic changes within individual organs between the control and patient groups, as detailed in Supplementary Fig. 3. We discovered that epilepsy not only causes metabolic abnormalities in certain brain regions but also leads to significant metabolic changes in other parts of the body, such as in some bones and muscles. Epilepsy increases the cerebral metabolic rate of oxygen and ATP demand, leading to mitochondrial exhaustion, which might help substantiate changes in some brain regions, such as temporal lobe and thalamus52. Additionally, metabolic dysfunctions also affect various bodily functions, including muscle activity53.

Ablation study

As shown in Fig. 2b, we conducted a multi-modality ablation study by training identical segmentation models under four input settings (CT only, CT+PET, CT+MR, and CT+PET+MR) and evaluated both body and brain tasks. Compared to the CT-only baseline, adding PET or MR each yielded noticeable improvements in Dice and surface Dice, reflecting complementary anatomical and functional cues. Finally, the tri-modality model achieved the best performance overall, confirming that integrating multiple modalities provides synergistic improvements in segmentation accuracy.

We also conducted controlled experiments to assess the impact of training volume on model performance. As shown in Supplementary Fig. 7, the results indicate that segmentation accuracy improves steadily when increasing the dataset size from 30% to 70%, with gains beginning to plateau between 70% and 85%, and nearing saturation at 100%. This trend suggests diminishing returns beyond a certain data threshold. We attribute this plateau to the rich inter-organ supervision in our fully annotated, multi-modality dataset.

Visualization of feature operators

The modality-projection model incorporates a controller module that transforms features from the modality space into feature extraction operators. The distribution of these feature operators across different modalities, as visualized in Fig. 5, enhances the interpretability of the model.

Fig. 5. Comprehensive visualization of saliency maps and feature operators.

Fig. 5

The black region displays a progression of saliency maps from shallow to deep layers, encompassing nine cases: a, a body CT scan; b-e, brain CT scans; f & g, PET scans; h& i, MR scans. The bottom right subfigure shows the t-SNE visualization of convolutional kernel operators, capturing the feature extraction from shallow to deep layers.

Stage 1 (shallowest layer) extracts low-level semantic features (e.g., edges, textures, simple shapes), which are largely modality-independent. In the t-SNE plots, MR and PET are overlapped due to their similar soft-tissue contrast. Since CT is well-suited to capture bone edges and basic texture features, it still shows partial overlap with MR and PET modalities in shallow layers, reflecting shared low-level feature patterns across imaging techniques. In deep convolutional layers (stage 3 and stage 4), high-level semantic features become modality-specific. The t-SNE plots display distinct clustering of feature operators by modality, reflecting each modality’s unique advanced feature extraction. This analysis highlights the contrast between shared low-level universality and modality-specific high-level abstraction, reconciled by the modality-projection controller.

Saliency map

To illustrate the interpretability of our universal model, we visualized saliency maps across network layers. In contrast to traditional gradient-based CAM methods54 that highlight the saliency of the final decision layer, our approach provides layer-wise saliency throughout inference. As shown in Fig. 5, we plot the saliency maps for different regions under various modalities, showing systematic patterns tied to layer depth, modality, and organ category.

In the shallow layers of the network, saliency is primarily focused on the image’s contours and textures, while in the deeper layers, saliency increasingly concentrates on the ROI regions. As shown in Fig. 5, the saliency map of the shallowest layer (the leftmost column) highlights the edge information of the image. The subsequent layer focuses on the texture information over large areas of the image (second column from the left). In contrast, the deeper network layers progressively refine the ROI regions with increasing precision. This progression indicates that the shallow layers capture global edge and texture information, whereas the deeper layers leverage localized information to improve the delineation of target regions.

As shown in Fig. 5a, the task of identifying the right femur from a CT image is relatively straightforward due to the high density of bone structures. In contrast, Fig. 5b–e depict the identification of various brain regions, challenging tasks given the difficulty of distinguishing soft tissue areas from CT images. In Fig. 5a, the shallow layers recognize a small amount of contour information, which is sufficient to clearly identify the location of the femur. Conversely, Fig. 5b–e need recognize the overall contour of the skull for preliminary localization, followed by a focus on the subtle variations in the soft tissue across the entire brain region.

Figure 5 a–e represent cases from the CT modality, while Fig. 5f–g are from the PET modality and Fig. 5h–i are from the MR modality. CT saliency maps show continuous contouring, MR maps show soft-tissue emphasis, and PET maps focus on metabolic activity at shallow layers with deeper refinement around the ROI.

Discussion

Our study demonstrates that the proposed MPUM, employing a modality-projection training strategy, achieves robust segmentation performance. The experimental results highlight significant advantages of this strategy in multi-modality training. This model shows potential in medical imaging tasks3,8,9, extending the applications of traditional universal segmentation models. The experimental results highlight two key advantages of this approach: precise brain segmentation and comprehensive whole-body segmentation. Specifically, the model enhances diagnostic accuracy for ICH, supporting physicians in timely and accurate diagnosis. Additionally, it enables systematic metabolic analysis across both brain-brain and brain-body associations in epilepsy and AD.

Universal CT-based brain segmentation

Previous models have struggled with CT-based brain segmentation due to the low soft tissues contrast. Existing methods often rely on knowledge transfer, such as BraSEDA55, which uses GANs to transfer MR brain region knowledge to CT, and UNSB56, which transfers ventricle region knowledge via a diffusion Schrödinger bridge. However, these methods are limited by domain shifts and focus on only a few regions. MPUM addresses these issues by delivering high accuracy across 83 distinct brain regions from CT scans.

Additionally, accurate segmentation of brain ventricles in CT scans is critical for emergency procedures like ventriculostomy, used to treat conditions such as hydrocephalus, brain injury, and tumors56. Although MRI remains the gold standard for ICH detection, CT’s speed and availability make it vital in acute ICH care. MPUM enables precise ventricle segmentation from CT scans, including the third ventricle, lateral ventricles, and temporal horn. This capability offers a significant clinical advantage3234.

Furthermore, the key factors in predicting outcomes and treatment strategies for ICH are midline shift and hemorrhage volume57. Specifically, hemorrhage volume in the brainstem is crucial in determining whether conservative medical treatment or surgical intervention is required57. As shown in Case 2, MPUM is capable of assessing hemorrhage volume in brainstem, providing valuable information for clinical decision-making. While MPUM does not directly detect the midline, it infers midline location by mapping left and right cerebellum regions.

From brain imaging analysis to brain-body axis exploration

MPUM excels in whole-body segmentation, which supports our research on the brain-body axis—the bidirectional communication between the brain and body that underpins many physiological and psychological processes58,59. While epilepsy has traditionally been viewed as a brain-centric disorder, recent studies suggest it may also involve systemic metabolic effects beyond the brain53,60,61. This insight prompted us to explore whole-body metabolic changes in epilepsy.

Leveraging MPUM’s high-throughput capability, we systematically identified altered metabolic associations across both brain-brain and brain-body ROIs. Notably, significant metabolic changes were observed in brain regions such as the right anterior temporal lobe and right middle/inferior temporal gyrus, which are consistent with existing research on epileptic foci4749. Our whole-body analysis also revealed a statistically significant metabolic association in the epilepsy cohort between pallidum and vertebrae T1-T12. This result highlights MPUM’s utility in uncovering data driven biomarkers for neurological disorders through comprehensive, system level metabolic profiling.

Annotation density

The capacity and diversity of training data are vital for developing universal segmentation models. Large-scale works like MedSAM2 represent a significant milestone in this direction. However, increasing the number of training subjects must be carefully balanced with considerations of annotation density and label consistency.

High annotation density compensates for subject volume

While our dataset includes 861 subjects (533 CT, 533 PET, 328 MR), each CT scan is annotated with 214 structures, resulting in over 200,000 image-mask pairs across modalities. In contrast, many large-scale public datasets rely on partial, coarse, or prompt-derived annotations (e.g., MedSAM’s reliance on semi-automatic label generation). Supplementary Fig. 9 illustrates that even with 30% of the dataset, our model achieves strong segmentation performance, which improves until performance saturates around 85% of the training data. It demonstrates that our dataset contains rich and dense supervision, enabling effective training even with fewer subjects.

Challenges of mixing incompatible datasets

Most public datasets (e.g., BTCV62) annotate only a limited number of structures, leading to label granularity mismatch when combined with our 214-class target. Additionally, the partial annotation problem introduces conflicting supervision. For example, structures labeled as “background” in one dataset may be labeled as “foreground” in another, introducing conflicting supervisory signals63,64. Furthermore, sparse annotations lack the anatomical context necessary to model spatial dependencies between neighboring structures, which is a key feature captured by our densely labeled dataset.

Limitations and future directions

Despite the progress our model has made in clinical applications and medical analysis, several challenges remain. Integration into real-world workflows requires not only high accuracy but also seamless user experience, secure deployment, and compatibility with clinical platforms. Current inference runs fully automatically, but basic Python and Linux skills are required. Future versions will integrate with platforms such as 3D Slicer to provide plug-and-play functionality, removing the need for scripting and facilitating deployment in routine radiology workflows.

Here we envision MPUM as a reusable imaging module within broader clinical AI pipelines. Its structured outputs, including whole-body and brain ROI labels, volumetric measurements, and metabolic matrices, serve as standardized inputs for a wide spectrum of downstream tasks. These span rapid localization to support acute clinical workflows; standardized feature extraction for population-scale imaging studies, enabling cross-organ and cross-modality analysis; and automated reporting or decision-support systems, where MPUM outputs feed directly into multi-modal analytics and domain-specific large language models (LLMs). To translate this potential from prototype to bedside, future work will focus on addressing key technical challenges such as domain shift, continual learning, and secure deployment. We also aim to expand multi-task learning and integrate MPUM seamlessly into hospital infrastructure, ultimately enabling a stepwise transition toward AI-assisted, imaging-driven clinical care.

Methods

Study population

In Case 1, we employed the 18F-FDG PET/CT dataset20 alongside two MR datasets4,65. Specifically, we utilized 533 negative control subjects from the 18F-FDG PET/CT dataset20. The body region labels were sourced from DAP Atlas Dataset5, covering 142 distinct anatomical structures. DAP label dataset contains the annotations of 533 subjects’ scans. After excluding irrelevant labels such as “background” and “left to annotations”, we utilized 133 effective labels from the DAP dataset. As the DAP dataset focuses primarily on body regions, it lacked detailed brain annotations. To address this gap, we pre-segmented the brain regions in the PET scans using the MOOSE tool66,67, followed by refinement and correction by two experienced radiologists, resulting in 83 annotated brain regions. Consequently, the “brain” label from the DAP dataset was removed, finalizing 132 non-brain categories and 83 brain region categories. The MR dataset includes 298 MR body scans4 annotated with 43 body regions, and 30 MR brain scans, annotated with the same 83 brain regions65.

For Case 2, we used the Instance2022 dataset40,41, which includes 100 non-contrast head CT volumes from clinically diagnosed patients with various types of intracranial hemorrhage (ICH), such as subdural, epidural, intraventricular, intraparenchymal, and subarachnoid hemorrhage. These CT volumes were sourced from Peking University Shougang Hospital, China, and meticulously labeled by 10 radiologists with over five years of clinical experience. We fine-tuned our MPUM using the Instance2022 data, enabling it to accurately identify ICH. In addition, we obtained 28 ICH cases along with diagnostic reports from the emergency department of Peking University Third Hospital for further validation. The diagnostic reports include findings confirmed by MRI imaging.

Case 3 involved 50 pediatric epilepsy patients who were enrolled at Qianfoshan Hospital, Shandong, China, between August 3, 2021, and June 4, 2024. The cohort consisted of 27 females and 23 males, with ages ranging from 2 to 18 years (11.78 ± 5.94 years). The height of the participants ranged from 0.93 to 1.85 m (1.49 ± 0.236 m), and their weight ranged from 14.5 to 90 kg (46.22 ± 19.92 kg). The control group comprised 22 participants, including 17 males and 5 females, with data collected from August 18, 2020, to May 9, 2024.

Detailed demographic statistics for both cohorts are provided in Supplementary Table 1. Median ages do not differ significantly between groups (Epilepsy: 12 years [IQR 7–15] vs. Control: 13 years [IQR 7–14]; Mann–Whitney U, p = 0.782). Height (p = 0.586) and weight (p = 0.959) were likewise comparable, indicating an age- and size-matched control cohort. PET/CT scanning in very young patients was clinically justified. For example, a 2-year-old with drug-resistant epilepsy required whole-body PET to locate an epileptogenic zone ahead of surgery. Supplementary Table 1 summarizes seizure duration, frequency, and semiology (78% focal, 22% generalized). Screening and cohort inclusion details are illustrated in Supplementary Fig. 9.

The multiple datasets used in this study, along with their descriptions, the division of training and testing data, and download links, are detailed in the Supplementary Table 4.

Implements

All experiments were conducted using Python (v3.10.14). The following packages were used for data preprocessing, analysis, and visualization: PyTorch (v2.6.0), NumPy (v1.26.4), pandas (v1.3.5), SciPy (v1.10.1), scikit-learn (v1.5.2), matplotlib, seaborn (v0.13.2), tqdm (v4.66.5), einops (v0.7.0), SimpleITK (v2.2.1), pydicom (v2.2.2), dicom2nifti (v2.4.8), MONAI (v1.2.0), and nnUNet (v1.7.1).

Data preprocessing

For this study, all imaging modalities, including CT, PET, and MR scans, underwent linear interpolation to achieve isotropic voxel sizes with a 2 mm resolution. This standardization was essential to address variations in slice thickness and in-plane resolutions across different datasets12.

To ensure consistent and stable model training, we normalized the voxel values within these patches. For the CT patches, the voxel values were normalized between 0 and 1. MR images are first intensity-truncated at the 0.5th and 99.5th percentiles to suppress outliers, then linearly scaled to the [0, 1] range. PET data were first converted to Standardized Uptake Value (SUV) and then divided by 20 for normalization. These normalization steps were crucial for maintaining data consistency across the different imaging modalities.

Additionally, we employed data augmentation techniques, specifically RandGaussianSmooth and RandAdjustContrast, to enhance the diversity and robustness of our training dataset. The RandGaussianSmooth method involves applying Gaussian smoothing to the images with a standard deviation randomly chosen between 0.5 and 1.5. This technique helps in reducing noise and simulating various levels of blurriness. RandAdjustContrast randomly scaled intensity values between 0.5 and 1.5, simulating different contrast conditions across scanners and reconstruction settings.

To optimize training efficiency, we pre-cropped all CT, PET, and MR scans into fixed-size patches. Loading a 128 × 128 × 128 patch is significantly faster than reading the full image and cropping on the fly. This pre-processing step dramatically reduces the I/O time during training, allowing for more efficient use of computational resources. We have made the patch-wise multi-modality training dataset publicly available to facilitate further research and development.

Metrics

We employed two evaluation metrics: Dice and surface Dice. Dice measures the overlap between predicted and true segmentations. It is calculated as twice the area of overlap between the two segmentations divided by the total number of pixels in both segmentations, providing an overall accuracy of how well the two align. Surface Dice, on the other hand, is a more specific measure that focuses on the boundary accuracy of the segmentation. It assesses how closely the boundaries of the predicted segmentation conform to the true surface contours of the object being analyzed.

Modality projection principle

Motivation

In recent studies of universal models, such as SAT19 and CDUM8, the modality-mixed strategy is commonly employed during the training phase. This strategy mixes data from various modalities, enabling a single model to adapt to data from multiple modalities. As shown in Fig. 1b, the modality-mixed strategy results in a multi-modality universal model. Although this approach benefits from a larger training data set, it also encounters challenges related to feature interference among modalities. The feature extraction operators (weights and biases) must deal with imaging data from all three modalities. This may lead the model to prioritize modality-independent features, potentially sacrificing modality-specific details. In contrast, the modality-specific strategy involves training separate models for each modality. The performance of these models serves as the baseline for this study. As illustrated in Fig. 1b, there are three types of modality data—MR, PET, and CT—the modality-specific strategy could produce three modality-specific models: a PET model, a CT model, and an MR model. The drawbacks of training individual models for each modality are the limited amount of training data and the lack of feature collaboration between the multi-modality data. In the realm of multi-modality medical imaging, different imaging modalities provide complementary views of the same underlying biological tissues. Each modality captures specific aspects of tissue characteristics due to differences in imaging principles and physical interactions. To effectively integrate information across modalities, we propose a modality projection theory centered on the concept of high-dimensional latent features that represent the comprehensive properties of each tissue type.

Fundamental principles

The core of our theory is the assumption that for each tissue, there exists a high-dimensional latent feature vector TRdT, where dT denotes the dimensionality of the latent feature space. This latent feature encapsulates all intrinsic properties of the tissue that could be captured across various imaging modalities.

Each imaging modality provides a modality-specific projection of this latent feature into its own feature space. For a given modality m ∈ CT, MR, PET, we define a projection matrix PmRdT×dm, where dm is the dimensionality of the modality feature space. The modality-specific feature vector MmRdm is obtained by projecting the latent feature:

Mm=TPm 1

This projection models how each modality interprets the latent tissue representation, capturing modality-specific characteristics.

Inverse projection and reconstruction

The relationship between latent features and modality-specific features suggests that, under certain conditions, it is possible to reconstruct the latent features from the modality-specific representation using the inverse of the projection matrices:

T=MmPm1 2

However, each modality offers only a partial perspective of the latent tissue state. Reconstructing T from a single modality may miss critical information. Combining multiple modalities yields a more complete and accurate representation.

Extended projection with external models

Optimizing T and Pm simultaneously can lead to training instability, as the condition number of the system matrix may increase due to parameter interdependence. Changes in T propagate to Mm, impacting the eigenvalues of Pm and disrupting gradient flow.

To address this challenge, we propose incorporating external pre-trained models (e.g., CLIP68) to anchor projections using robust, stable feature embeddings. Let MiRdi represent the sub-space features from an external model i, with di being its feature dimensionality. Each external model has its projection matrix PiRdT×di:

Mi=TPi. 3

By leveraging these external features, we estimate the latent features by aggregating inversely projected sub-space features:

T=1Ni=1NMiPi1, 4

where N is the number of external models used. This approach anchors the latent features to stable, pre-trained representations, mitigating optimization instability. Additionally, the training datasets used by external pre-training models encompass a wide range of textual data, which facilitates the reconstruction of latent features.

Generalized modality projection framework

In this generalized framework, the latent features T serves as a central hub connected to various modality features Mm and external model features Mi through their respective projection matrices. This enables effective cross-modal and cross-domain knowledge integration.

Architecture of modality projection universal model(MPUM)

Building upon the modality projection principle, we introduce the MPUM, a comprehensive framework designed to effectively process and integrate multi-modality medical imaging data. The MPUM leverages the Modality Projection Controller (MPC) within a deep learning architecture to handle diverse imaging modalities—such as CT, MR, and PET—facilitating accurate and efficient image segmentation.

As illustrated in Supplementary Fig. 4a, the MPUM architecture begins with multi-modality patch inputs (e.g., CT, MR, PET), processed initially through a Head Layer, followed by a series of Dual-Branch Blocks with skip connections, and concluding with a Tail Layer to produce the segmentation output. The skip connections help preserve spatial information and integrate features from different layers, improving segmentation performance.

At the core of the MPUM is the Modality Projection Controller, illustrated in Supplementary Fig. 4b. The MPC is responsible for dynamically adjusting the feature extraction process based on the specific characteristics of the input modality.

The MPC process involves the following two steps. Firstly, the latent feature vector T is projected into the modality category space using the modality-specific projection matrix Pm, as shown in Eq. (1). Secondly, the modality features Mm are processed through a Multi-Layer Perceptron (MLP), serving as the Feature Operator Generator (FOG):

Km=FOG(Mm), 5

resulting in convolutional kernels Km tailored for each modality. This dynamic adjustment allows the MPUM to adapt its convolutional operations to the unique properties of each modality, thus improving feature extraction and segmentation accuracy.

As illustrated in Supplementary Fig. 4, we detail the operation of the controller-based convolution layer and clarify how the MLP output is shaped into a 3D kernel. Firstly, the input feature map is projected into a latent space via a 1 × 1 × 1 convolution, reducing its channel dimension to a predefined latent size L (e.g., L = 64). This step compresses the feature map into a latent representation while preserving spatial dimensions. The multi-layer perceptron (MLP) in the controller then outputs a flattened kernel of length L × 33, where 33 corresponds to the spatial dimensions of the kernel. Finally, the filtered feature map is projected back to the target output channel dimension using another 1 × 1 × 1 convolution, ensuring compatibility with downstream operations.

The MPC uses an MLP to generate convolution kernel parameters, a methodology widely validated in prior work. For instance, HyperNetworks69,70 pioneered the use of auxiliary networks to dynamically generate weights for primary network, demonstrating the feasibility of MLP-based parameter synthesis. Similarly, HyperConvolution71 extended this concept by leveraging MLPs to produce continuous convolution kernels. While our approach adopts this mechanism, our innovation lies in the use of modality-specific embeddings to guide kernel generation, encoding anatomical priors specific to multi-modal medical data.

The MPUM also incorporates a specialized Dual-Branch Block Structure, illustrated in Supplementary Fig. 4c, which consists of two parallel branches: traditional convolutional branch and controller-based convolutional branch. Traditional convolutional branch employs standard convolutional layers with fixed kernels to extract general features from the input data. It captures modality-invariant features. In contrast, the controller-based branch leverages convolutional kernels dynamically generated by MPC, enabling the network to adapt to specific imaging characteristics of each modality. These controller-driven convolutional layers are non-parametric and adjust their operations according to the modality type, allowing the model to extract highly specific anatomical and functional cues. Together, this dual-branch structure enhances both generalization and specialization.

Saliency map

Saliency maps are vital tools for visualizing the decision-making process of deep learning models. Traditional methods for generating saliency maps primarily rely on gradient-based techniques. Methods such as Vanilla Gradients72 and Grad-CAM54 generate saliency maps by calculating the gradients of the input image with respect to the model’s output, highlighting regions that significantly influence the model’s decisions.

As shown in Supplementary Fig. 4c, our modality projection block structure enables the generation of saliency maps across all layers of the network, not only the final decision layer. In our framework, the feature operator K from Eq. (5) is a tensor with dimensions RC×H×3×3×3, where C represents categories, H denotes the number of channels, and 3 × 3 × 3 specifies the kernel size. This structure enables interpretation of the model’s behavior across the entire network hierarchy.

Statistical analysis

We applied the Fisher Z-transformation to compare correlation coefficients between two independent groups. This transformation standardizes raw correlation coefficients into values approximating a normal distribution, enabling hypothesis testing between control and patient groups.

In our analysis, we analyzed metabolic data from organs a and b in a control group (22 individuals), calculating the mean standardized uptake value (SUV) for the regions of interest (ROI) within these organs. The metabolic data for organ a in the control group is denoted as Acontrol = {ann = 1, 2, …, 22}, representing the set of mean SUVs for each individual. Similarly, the metabolic data for organ b in the control group is represented as Bcontrol = {bnn = 1, 2, …, 22}. We calculated the Pearson correlation coefficient, which measures the strength and direction of a linear relationship between two variables:

ra,bgroup=i=1n(aia¯)(bib¯)i=1n(aia¯)2(bib¯)2,group{control,patient} 6

where ai and bi are the individual SUV for organs a and b in the control group, and a¯ and b¯ are the mean SUVs for organ a and b, respectively. We can simplify the Eq. (6) to ra,bgroup=pearson(Agroup,Bgroup). Similarly, replacing an organ variable with Age yields the Pearson correlations ra,Agegroup and rb,Agegroup.

To remove the linear effect of age, we calculated the partial correlation between organs a and b, conditional on age, using the closed-form solution:

ra,bAgegroup=ra,bgroupra,Agegrouprb,Agegroup1(ra,Agegroup)21(rb,Agegroup)2,group{control,patient} 7

Eq. (7) is equivalent to regressing a and b on age, extracting their residuals, and calculating the ordinary Pearson correlation between the residuals.

We then apply the Fisher Z-transformation to ra,bcontrol and ra,bpatient:

za,bAgecontrol=arctanhra,bAgecontrol 8
za,bAgepatient=arctanhra,bAgepatient 9

And computed the Z-score:

z=za,bAgecontrolza,bAgepatient1/(ncontrol3)+1/(npatient3), 10

where ncontrol and npatient are the sample sizes of control group and patient group, respectively. Finally, we calculate the p-value assuming a normal distribution:

p=2*(1cdf(abs(z))), 11

where abs(z) represent the absolute value of the Z-score and cdf stands for cumulative distribution function. The term 1 − cdf(abs(z)) represents the one-tailed p-value, and multiplying by 2 makes it a two-tailed test. Raw p values were adjusted by Benjamini-Hochberg FDR method.

In large-scale metabolic connectome studies like ours (>20,000 correlations), controlling for false positives is critical46,7375. Although Bonferroni Correction strictly controls the family-wise error rate (FWER), it would require an adjusted threshold of 0.01/20, 000 = 5 × 10−7 in our analysis, which is overly conservative. We therefore report both raw and FDR-adjusted p values. All results remain significant at α = 0.01 post-FDR, reinforcing their reliability (See Supplementary Fig. 2).

Ethical approval declaration

This study was approved by the Ethics Committee of The First Affiliated Hospital of Shandong First Medical University & Shandong Provincial Qianfoshan Hospital (Approval No. S917). All participants provided informed consent, and the study adhered to all applicable ethical standards and institutional guidelines. Participants received no financial remuneration.

Reporting summary

Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.

Supplementary information

Reporting Summary (200.1KB, pdf)

Source data

Source Data (13.2MB, xlsx)

Acknowledgements

Zhaoheng Xie discloses support for the research of this work from the Natural Science Foundation of China (62394311 [Z.X.], 62394310 [Z.X.]), Beijing Natural Science Foundation (Z210008 [Z.X.]), National Biomedical Imaging Facility Grant [Z.X.], and from the startup funds of Peking University Health Science Center [Z.X.]. Rui Wang discloses support for the research of this work from the Youth Fund of the National Natural Science Foundation of China (62301245 [R.W.]) and from the Guangdong Basic and Applied Basic Research Foundation (2022A1515110674 [R.W.]). Xiangxi Meng discloses support for this work from Peking University Clinical Medicine Plus X - Young Scholars Project. The authors gratefully acknowledge Prof. Lin Lu of Peking University Sixth Hospital, Prof. Jiang Yuwu of Peking University First Hospital, and B.E. Ruoyan Xu of Peking University Health Science Center for their valuable discussions and insightful suggestions.

Author contributions

Y. Chen conceived the study, designed the experiments, implemented the code, and drafted the manuscript. L. Gao performed data processing, statistical analysis, and interpretation of the experimental results. Y. Chen and L. Gao contributed equally to this work. Y. Gao collected the ICH dataset and carried out the corresponding data analysis. R. Wang and J. Lian supplied imaging data and contributed to manuscript revisions. X. Meng assisted with methodological development and contributed to manuscript revisions. Y. Duan, H. Han, and L. Chai curated the clinical data and contributed to manuscript revisions. Z. Cheng and Z. Xie jointly supervised the project, secured funding, and approved the final manuscript.

Peer review

Peer review information

Nature Communications thanks Yucheng Tang and the other anonymous reviewer(s) for their contribution to the peer review of this work. A peer review file is available.

Data availability

Publicly available datasets used in this study are listed below. Accession links are provided in Supplementary Table 4. 1. AutoPET 18F-FDG PET/CT (533 controls); 2. DAP Atlas body-region labels 3. TotalSegmentator-MRI 4. MRI-Brain 5. Instance2022 ICH head-CT challenge set The pediatric epilepsy PET/CT cohort (50 patients, 22 controls), acquired at Qianfoshan Hospital, Shandong, China, is not publicly available due to patient privacy regulations. Researchers interested in accessing the data should contact the corresponding author (Z. Cheng) with a formal data request and research purpose, which will be reviewed in accordance with institutional policies. Submit a formal request to the Data Access Committee via the corresponding author (Z. Cheng; [czpabc@163.com]), including: (i) institutional affiliation; (ii) a brief research proposal describing the intended analyses; and (iii) a data security plan. Receipt of a complete application will be acknowledged within 10 business days and an access decision is typically issued within 30 calendar days. Approved users must sign a DUA that (a) limits use to the approved research purpose; (b) forbids redistribution to third parties; (c) requires secure storage and access logging; (d) restricts public reporting to aggregate, non-identifying results. A template DUA is provided. Source data are provided with this paper.

Code availability

All code used for data processing and performance analysis, including well-trained model weights, is publicly available via GitHub at https://github.com/YixinChen-AI/MPUMunder the MIT licence. The code is archived on Zenodo with the 10.5281/zenodo.16730886. The inference time of MPUM when deployed on different graphics cards as shown in Supplementary Fig. 5.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Yixin Chen, Lin Gao.

Contributor Information

Zhaoping Cheng, Email: czpabc@163.com.

Zhaoheng Xie, Email: xiezhaoheng@pku.edu.cn.

Supplementary information

The online version contains supplementary material available at 10.1038/s41467-025-64469-w.

References

  • 1.Kirillov, A. et al. Segment anything. In Proceedings of the IEEE/CVF international conference on computer vision, 4015–4026 (2023).
  • 2.Ma, J. & Wang, B. Segment anything in medical images. Nat. Commun. 15, 654 (2024). [DOI] [PMC free article] [PubMed]
  • 3.Wasserthal, J. et al. TotalSegmentator: Robust segmentation of 104 anatomic structures in CT images. Radiology: Artif. Intell.5, e230024 (2023). [DOI] [PMC free article] [PubMed]
  • 4.D’Antonoli, T. A. et al. TotalSegmentator MRI: Sequence-Independent Segmentation of 59 Anatomical Structures in MR images. Radiology314, e241613 (2025). [DOI] [PubMed]
  • 5.Jaus, A. et al. Towards unifying anatomy segmentation: Automated generation of a full-body CT dataset via knowledge aggregation and anatomical guidelines. In International Conference on Image Processing, 41–47 (IEEE, 2024).
  • 6.Qu, C. et al. Abdomenatlas-8k: Annotating 8,000 CT volumes for multi-organ segmentation in three weeks. Advances in Neural Information Processing Systems36, (2024).
  • 7.de Grauw, M. et al. The ULS23 Challenge: A baseline model and benchmark dataset for 3D universal lesion segmentation in computed tomography. Med. Image Anal.102, 103525 (2025). [DOI] [PubMed]
  • 8.Liu, J. et al. Clip-driven universal model for organ segmentation and tumor detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 21152–21164 (2023).
  • 9.Chen, Y. et al. PCNet: Prior category network for CT universal segmentation model. IEEE Trans. Med. Imaging43, 3319–3330 (2024). [DOI] [PubMed]
  • 10.Huang, Z. et al. STU-Net: Scalable and transferable medical image segmentation models empowered by large-scale supervised pre-training https://arxiv.org/abs/2304.06716 (2023).
  • 11.Chen, Y. et al. LUCIDA: Low-dose universal-tissue CT image domain adaptation for medical segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 393–402 (Springer, 2024).
  • 12.Pai, S. et al. Foundation model for cancer imaging biomarkers. Nat. Mach. Intell.6, 354–367 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Chen, T., Kornblith, S., Norouzi, M. & Hinton, G. A simple framework for contrastive learning of visual representations. In International conference on machine learning, 1597–1607 (PMLR, 2020).
  • 14.Caron, M. et al. Unsupervised learning of visual features by contrasting cluster assignments. Adv. Neural Inf. Process. Syst.33, 9912–9924 (2020). [Google Scholar]
  • 15.Dwibedi, D., Aytar, Y., Tompson, J., Sermanet, P. & Zisserman, A. With a little help from my friends: Nearest-neighbor contrastive learning of visual representations. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 9588–9597 (2021).
  • 16.Upadhyay, A. K. & Bhandari, A. K. Advances in deep learning models for resolving medical image segmentation data scarcity problem: A topical review. Arch. Comput. Methods Eng.31, 1701–1719 (2024). [Google Scholar]
  • 17.Yang, J. Multi-task learning for medical foundation models. Nat. Comput. Sci.4, 473–474 (2024). [DOI] [PubMed]
  • 18.Butoi, V. I. et al. Universeg: Universal medical image segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 21438–21451 (2023).
  • 19.Zhao, Z. et al. One Model to Rule them All: Towards Universal Segmentation for Medical Images with Text Prompts https://arxiv.org/abs/2312.17183 (2024).
  • 20.Gatidis, S. et al. A whole-body FDG-PET/CT dataset with manually annotated tumor lesions. Sci. Data9, 601 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Luo, L. et al. Building Universal Foundation Models for Medical Image Analysis with Spatially Adaptive Networks. arXiv preprint arXiv:2312.07630 (2023).
  • 22.Ye, Y., Xie, Y., Zhang, J., Chen, Z. & Xia, Y. Uniseg: A prompt-driven universal segmentation model as well as a strong representation learner. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 508–518 (Springer, 2023).
  • 23.He, Y. et al. VISTA3D: A unified segmentation foundation model for 3D medical imaging. In Proceedings of the IEEE/CVF International Conference on Computer Vision and Pattern Recognition (2024).
  • 24.Lu, W., Duan, Y., Li, K., Qiu, J. & Cheng, Z. Glucose uptake and distribution across the human skeleton using state-of-the-art total-body PET/CT. Bone Res.11, 36 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 25.Ronneberger, O., Fischer, P. & Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18, 234–241 (Springer, 2015).
  • 26.Greenberg, S. M. et al. 2022 Guideline for the management of patients with spontaneous intracerebral hemorrhage: a guideline from the American Heart Association/American Stroke Association. Stroke53, e282–e361 (2022). [DOI] [PubMed] [Google Scholar]
  • 27.Hemphill III, J. C. et al. Guidelines for the management of spontaneous intracerebral hemorrhage: a guideline for healthcare professionals from the American Heart Association/American Stroke Association. Stroke46, 2032–2060 (2015). [DOI] [PubMed] [Google Scholar]
  • 28.Sanner, A. P., Grauhan, N. F., Brockmann, M. A., Othman, A. E. & Mukhopadhyay, A. Voxel Scene Graph for Intracranial Hemorrhage. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 519–529 (Springer, 2024).
  • 29.An, S. J., Kim, T. J. & Yoon, B.-W. Epidemiology, risk factors, and clinical features of intracerebral hemorrhage: an update. J. Stroke19, 3 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Pinho, J., Costa, A. S., Araújo, J. M., Amorim, J. M. & Ferreira, C. Intracerebral hemorrhage outcome: a comprehensive update. J. Neurol. Sci.398, 54–66 (2019). [DOI] [PubMed] [Google Scholar]
  • 31.Moon, J.-S. et al. Prehospital neurologic deterioration in patients with intracerebral hemorrhage. Crit. Care Med.36, 172–175 (2008). [DOI] [PubMed] [Google Scholar]
  • 32.Hillal, A., Ullberg, T., Ramgren, B. & Wassélius, J. Computed tomography in acute intracerebral hemorrhage: neuroimaging predictors of hematoma expansion and outcome. Insights Into imaging13, 180 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 33.Sage, A. & Badura, P. Intracranial hemorrhage detection in head CT using double-branch convolutional neural network, support vector machine, and random forest. Appl. Sci.10, 7577 (2020). [Google Scholar]
  • 34.Zimmerman, R. D., Maldjian, J. A., Brun, N., Horvath, B. & Skolnick, B. Radiologic estimation of hematoma volume in intracerebral hemorrhage trial by CT scan. Am. J. Neuroradiol.27, 666–670 (2006). [PMC free article] [PubMed] [Google Scholar]
  • 35.Chen, R. et al. Infratentorial intracerebral hemorrhage: relation of location to outcome. Stroke50, 1257–1259 (2019). [DOI] [PubMed] [Google Scholar]
  • 36.Wahjoepramono, P. O. P. et al. Hydrocephalus is an independent factor affecting morbidity and mortality of ICH patients: Systematic review and meta-analysis. World Neurosurg.: X19, 100194 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 37.Mendiola, J. M. F.-Pd, Arboix, A., García-Eroles, L. & Sánchez-López, M. J. Acute spontaneous lobar cerebral hemorrhages present a different clinical profile and a more severe early prognosis than deep subcortical intracerebral hemorrhages–a hospital-based stroke registry study. Biomedicines11, 223 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.Akram, M. J. et al. Surgical vs. Conservative management for lobar intracerebral hemorrhage, a meta-analysis of randomized controlled trials. Front. Neurol.12, 742959 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 39.Hazra, D., Chandy, G. M. & Ghosh, A. K. Surgical outcome of basal ganglia hemorrhage: A retrospective analysis of nearly 3,000 cases over 10 years. Asian J. Neurosurg.18, 742–750 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Li, X. et al. The state-of-the-art 3D anisotropic intracranial hemorrhage segmentation on non-contrast head CT: The INSTANCE challenge. arXiv preprint arXiv:2301.03281 (2023).
  • 41.Li, X. et al. Hematoma expansion context guided intracranial hemorrhage segmentation and uncertainty estimation. IEEE J. Biomed. Health Inform.26, 1140–1151 (2021). [DOI] [PubMed] [Google Scholar]
  • 42.Prakash, K. B., Zhou, S., Morgan, T. C., Hanley, D. F. & Nowinski, W. L. Segmentation and quantification of intra-ventricular/cerebral hemorrhage in CT scans by modified distance regularized level set evolution technique. Int. J. Comput. Assist. Radiol. Surg.7, 785–798 (2012). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 43.Kim, K. H., Koo, H.-W., Lee, B.-J., Yoon, S.-W. & Sohn, M.-J. Cerebral hemorrhage detection and localization with medical imaging for cerebrovascular disease diagnosis and treatment using explainable deep learning. J. Korean Phys. Soc.79, 321–327 (2021). [Google Scholar]
  • 44.Liu, Q. et al. Voxels intersecting along orthogonal levels attention u-net for intracerebral haemorrhage segmentation in head CT. In 2023 IEEE 20th International Symposium on Biomedical Imaging (ISBI), 1–5 (IEEE, 2023).
  • 45.Sundar, L. K. S., Hacker, M. & Beyer, T. Whole-body pet imaging: a catalyst for whole-person research? J. Nucl. Med.64, 197 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Sun, T. et al. Identifying the individual metabolic abnormities from a systemic perspective using whole-body PET imaging. Eur. J. Nucl. Med. Mol. Imaging49, 2994–3004 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 47.Roper, S. N. & Rhoton Jr, A. L. Surgical anatomy of the temporal lobe. Neurosurg. Clin. North Am.4, 223–231 (1993). [PubMed] [Google Scholar]
  • 48.Engel Jr, J. Introduction to temporal lobe epilepsy. Epilepsy Res.26, 141–150 (1996). [DOI] [PubMed] [Google Scholar]
  • 49.Engel Jr, J. Surgery for seizures. N. Engl. J. Med.334, 647–653 (1996). [DOI] [PubMed] [Google Scholar]
  • 50.Silverman, D. H. Brain 18F-FDG PET in the diagnosis of neurodegenerative dementias: comparison with perfusion SPECT and with clinical evaluations lacking nuclear imaging. J. Nucl. Med.45, 594–607 (2004). [PubMed] [Google Scholar]
  • 51.Wilson, H., Pagano, G. & Politis, M. Dementia spectrum disorders: lessons learnt from decades with PET research. J. Neural Transm.126, 233–251 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 52.Liotta, A. et al. Metabolic adaptation in epilepsy: From acute response to chronic impairment. Int. J. Mol. Sci.25, 9640 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 53.Fei, Y., Shi, R., Song, Z. & Wu, J. Metabolic control of epilepsy: a promising therapeutic target for epilepsy. Front. Neurol.11, 592514 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 54.Selvaraju, R. R. et al. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision, 618–626 (2017).
  • 55.Chen, Y. et al. Structure-Enhanced Unsupervised Domain Adaptation for CT Whole-Brain Segmentation. IEEE Transactions on Radiation and Plasma Medical Sciences (2024).
  • 56.Teimouri, R., Kersten-Oertel, M. & Xiao, Y. CT-based brain ventricle segmentation via diffusion Schrödinger Bridge without target domain ground truths. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 135–144 (Springer, 2024).
  • 57.Xu, X.-M., Zhang, H. & Meng, R.-L. Cranial midline shift is a predictor of the clinical prognosis of acute cerebral infarction patients undergoing emergency endovascular treatment. Sci. Rep.13, 21037 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 58.Li, S. et al. The mechanism of the gut-brain axis in regulating food intake. Nutrients15, 3728 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Sharp, H. E. C., Critchley, H. D. & Eccles, J. A. Connecting brain and body: Transdiagnostic relevance of connective tissue variants to neuropsychiatric symptom expression. World J. psychiatry11, 805 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 60.Yang, W. et al. A review of the pathogenesis of epilepsy based on the microbiota-gut-brain-axis theory. Front. Mol. Neurosci.17, 1454780 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 61.Zhu, H., Wang, W. & Li, Y. The interplay between microbiota and brain-gut axis in epilepsy treatment. Front. Pharmacol.15, 1276551 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 62.Miccai multi-atlas labeling beyond the cranial vault–workshop and challenge, author=Landman, Bennett and Xu, Zhoubing and Igelsias, Juan and Styner, Martin and Langerak, Thomas and Klein, Arno. In Proc. MICCAI multi-atlas labeling beyond cranial vault–workshop challenge, vol. 5, 12 (Munich, Germany, 2015).
  • 63.Koch, L. M. et al. Multi-atlas segmentation using partially annotated data: methods and annotation strategies. IEEE Trans. pattern Anal. Mach. Intell.40, 1683–1696 (2017). [DOI] [PubMed] [Google Scholar]
  • 64.Xie, Y., Zhang, J., Xia, Y. & Shen, C. Learning from partially labeled data for multi-organ and tumor segmentation. IEEE Trans. Pattern Anal. Mach. Intell.45, 14905–14919 (2023). [DOI] [PubMed] [Google Scholar]
  • 65.Serag, A. et al. Construction of a consistent high-definition spatio-temporal atlas of the developing brain using adaptive kernel regression. Neuroimage59, 2255–2265 (2012). [DOI] [PubMed] [Google Scholar]
  • 66.Sundar, L. K. S. et al. Fully automated, semantic segmentation of whole-body 18F-FDG PET/CT images based on data-centric artificial intelligence. J. Nucl. Med.63, 1941–1948 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 67.Isensee, F., Jaeger, P. F., Kohl, S. A., Petersen, J. & Maier-Hein, K. H. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nat. methods18, 203–211 (2021). [DOI] [PubMed] [Google Scholar]
  • 68.Radford, A. et al. Learning Transferable Visual Models From Natural Language Supervision https://arxiv.org/abs/2103.00020 (2021).
  • 69.Ha, D., Dai, A. & Le, Q. V. Hypernetworks. In International Conference on Learning Representations (2017).
  • 70.Chauhan, V. K., Zhou, J., Lu, P., Molaei, S. & Clifton, D. A. A brief review of hypernetworks in deep learning. Artif. Intell. Rev.57, 250 (2024). [Google Scholar]
  • 71.Ma, T., Dalca, A. V. & Sabuncu, M. R. Hyper-convolution networks for biomedical image segmentation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 1933–1942 (2022).
  • 72.Simonyan, K., Vedaldi, A. & Zisserman, A. Deep inside convolutional networks: Visualising image classification models and saliency maps. In International Conference on Learning Representations (2014).
  • 73.Sun, L. et al. Influence of diabetes mellitus on metabolic networks in lung cancer patients: an analysis using dynamic total-body PET/CT imaging. Eur. J. Nuclear Med. Mol. Imaging52, 2145–2156 (2025). [DOI] [PubMed]
  • 74.Lu, W., Song, T., Li, J., Zhang, Y. & Lu, J. Individual-specific metabolic network based on 18F-FDG PET revealing multi-level aberrant metabolisms in Parkinson’s disease. Hum. Brain Mapp.45, e70026 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 75.Wang, M. et al. Individual brain metabolic connectome indicator based on Kullback-Leibler divergence similarity estimation predicts progression from mild cognitive impairment to Alzheimer’s dementia. Eur. J. Nucl. Med. Mol. imaging47, 2753–2764 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Reporting Summary (200.1KB, pdf)
Source Data (13.2MB, xlsx)

Data Availability Statement

Publicly available datasets used in this study are listed below. Accession links are provided in Supplementary Table 4. 1. AutoPET 18F-FDG PET/CT (533 controls); 2. DAP Atlas body-region labels 3. TotalSegmentator-MRI 4. MRI-Brain 5. Instance2022 ICH head-CT challenge set The pediatric epilepsy PET/CT cohort (50 patients, 22 controls), acquired at Qianfoshan Hospital, Shandong, China, is not publicly available due to patient privacy regulations. Researchers interested in accessing the data should contact the corresponding author (Z. Cheng) with a formal data request and research purpose, which will be reviewed in accordance with institutional policies. Submit a formal request to the Data Access Committee via the corresponding author (Z. Cheng; [czpabc@163.com]), including: (i) institutional affiliation; (ii) a brief research proposal describing the intended analyses; and (iii) a data security plan. Receipt of a complete application will be acknowledged within 10 business days and an access decision is typically issued within 30 calendar days. Approved users must sign a DUA that (a) limits use to the approved research purpose; (b) forbids redistribution to third parties; (c) requires secure storage and access logging; (d) restricts public reporting to aggregate, non-identifying results. A template DUA is provided. Source data are provided with this paper.

All code used for data processing and performance analysis, including well-trained model weights, is publicly available via GitHub at https://github.com/YixinChen-AI/MPUMunder the MIT licence. The code is archived on Zenodo with the 10.5281/zenodo.16730886. The inference time of MPUM when deployed on different graphics cards as shown in Supplementary Fig. 5.


Articles from Nature Communications are provided here courtesy of Nature Publishing Group

RESOURCES