Abstract
Recent advances in vision–language models and LLMs have introduced contextual anatomical reasoning into brain MRI segmentation. However, the field still suffers from a fundamental limitation: the absence of a unified anatomical definition of the structures being segmented. Existing datasets rely on labels produced by heterogeneous manual workflows, often lacking explicit anatomical criteria or consistent annotation standards. As a result, models learn and evaluate within isolated labeling systems, limiting cross-model comparison and valid anatomical measurements. To address these challenges, we introduce NeuroLangSeg, a language-guided framework that enforces a consistent anatomical protocol for subcortical segmentation. A key component of the framework is an anatomical–linguistic evaluator that acts as a training discriminator, encouraging the model to produce outputs by assessing shape characteristics, protocol-defined spatial relationships, and age- and sex-adjusted volumetric norms. Building upon this constraint, NeuroLangSeg integrates a pretrained image encoder with protocol-aligned anatomical prompts and a masked pseudo-labeling strategy, enabling data-efficient and interpretable learning under limited supervision. Together, these components yield anatomically consistent segmentations and support subject-level reporting grounded in a unified anatomical standard. Evaluation across diverse MRI datasets—including comparisons with state-of-the-art models—shows that NeuroLangSeg achieves +4.1 DSC / +8.0 NSD in in-site settings and +3.6 DSC / +14.5 NSD in cross-site generalization over the average baseline, enabled by its LLM–visual integration, while delivering anatomically verifiable predictions suitable for both research and clinical use. GitHub: https://github.com/jlliu2001/SAT_MPL
Keywords: Anatomical Protocol, Language-Driven Segmentation, Anatomical–Linguistic Evaluation, Brain MRI
1. Introduction
Accurate segmentation of subcortical brain structures is fundamental to quantitative analysis and clinical assessment. Most regions such as the hippocampus, amygdala, and thalamus, enable detailed investigations of brain development, aging, and neuropathology, supporting downstream analyses of structure–function relationships and population-level biomarkers (Baribeau et al., 2019; Cruz et al., 2023; de Jong et al., 2008). Although manual delineation remains the gold standard for defining anatomical boundaries, it is time-consuming, labor-intensive, and dependent on expert knowledge.
While traditional neuroimaging pipelines such as FreeSurfer (Fischl, 2012), BrainSuite (Kim et al., 2024), ANTs (Avants et al., 2011), and FSL (Jenkinson et al., 2012) have been widely used for automated frameworks for structural analysis, their multi-stage registration and optimization procedures are computationally intensive and difficult to scale for large datasets or clinical workflows. In contrast, recent advances in deep learning have substantially improved medical image segmentation, allowing models to learn rich representations directly from MRI data and achieve high accuracy across diverse anatomical and clinical tasks (Billot et al., 2023; Guha Roy et al., 2019; Henschel et al., 2020, 2022; Estrada et al., 2023; Zhang et al., 2024). However, most existing approaches remain task-specific—trained for a single structure, cohort, or labeling rules—and demonstrate lower performance when deployed on heterogeneous datasets, limiting their generalization and clinical applicability.
To enhance flexibility and interpretability, large language models (LLMs) and vision-language models (VLMs) have been developed for medical image segmentation by coupling textual descriptions with visual representations (Ma et al., 2024; Fu et al., 2025; Su et al., 2025; Zhao et al., 2025). These multimodal models incorporate semantic context and enable adaptable segmentation across domains. Prompt-based VLMs extend this capability to open-vocabulary medical segmentation across organs and modalities (Ma et al., 2024; Zhao et al., 2025), while anatomical priors (e.g., shape templates and mesh constraints) further enhance spatial consistency (Su et al., 2025). In neuroimaging, emerging protocol-guided approaches encode hierarchical anatomical relationships, such as topology-based text generation to improve brain segmentation (Fu et al., 2025).
Despite these advances, subcortical segmentation still lacks clinically unified anatomical protocols and evaluation frameworks. First, current visual backbones are constrained by their training labels. Most large-scale datasets rely on FreeSurfer-derived masks because they are readily available (Fischl et al., 2002; Tae et al., 2008). Some models such as FastSurfer (Henschel et al., 2020, 2022; Estrada et al., 2023), SynthSeg (Billot et al., 2023), and QuickNAT (Guha Roy et al., 2019) largely reproduce or refine these outputs. However, FreeSurfer boundaries often diverge from expert manual labels (Morey et al., 2009; Schoemaker et al., 2016; Lerch et al., 2017), introducing systematic structural bias into both training and evaluation. Second, even when manual segmentations are available, there is still no clinically unified protocol: different experts and software tools apply different delineation rules—for example, outlining the amygdala or hippocampus with different boundaries—resulting in inconsistent ground-truth masks across datasets (Geuze et al., 2005; Yushkevich et al., 2015). Although recent vision–language models incorporate textual cues, they do not resolve this underlying protocol mismatch. Finally, current segmentation frameworks lack a standardized evaluation pipeline to assess anatomical accuracy from a clinical perspective. Conventional metrics based on overlap, such as the Dice coefficient, quantify geometric similarity but fail to capture the morphological integrity, topological consistency, or biological validity of the predicted structures (Babalola et al., 2009).
To address these challenges, we propose NeuroLangSeg, a language-guided subcortical segmentation framework with pseudo-supervision and anatomical–linguistic validation based on a consistent anatomical protocol from Neuromorphometrics, Inc. (Landman and Warfield, 2012). Our main contributions are: 1) Contextual anatomical prompts are encoded and fused with visual features, enabling flexible, prompt-conditioned segmentation across structures and cohorts. 2) A unified visual backbone combines large-scale 3D masked autoencoder pretraining, label-efficient pseudo-label refinement, and global–local stabilization to improve robustness across scanners, ages, and modalities. 3) Clinical anatomical protocols encoded by an LLM guide morphological and topological discriminators. During inference, the evaluator integrates morphological, topological, and BrainChart-normalized volumetric metrics (adjusted for age and sex) to assess anatomical consistency.
Together, these components make NeuroLangSeg a clinically aligned and explainable framework for subcortical segmentation, providing both high accuracy and anatomical validation. To our knowledge, it is the first model to unify language-guided learning, semi-supervised segmentation, and anatomical–linguistic evaluation. Experiments on healthy and clinical cohorts show strong generalization and anatomically consistent performance across diverse populations.
2. Method
We address subcortical segmentation across heterogeneous MRI cohorts that differ in annotation policies and lack a unified, protocol-driven evaluation standard (Figure 1). Each sample contains a 3D MRI volume and its corresponding segmentation map . A textual prompt describing the target anatomical structure is provided and converted into a semantic embedding , which is combined with visual features extracted from the MAE encoder to guide structure-specific prediction. The fused representation is passed to a segmentation decoder to produce structure-specific masks:
| (1) |
where is the visual encoder, is the visual decoder, is the text encoder, and is the query decoder. The segmentation head performs dot-product matching and projection to generate the final mask . Predictions are evaluated using morphological, topological, and volumetric metrics to assess anatomical consistency.
Figure 1:

Overview of NeuroLangSeg Segmentation and Evaluation
2.1. Language-Guided Prompt Encoding
Text Encoder:
To encode anatomical concepts and their associated positional knowledge, we adopt a BERT-based text encoder initialized from a biomedical language model (Zhao et al., 2025) and further adapted through supervised fine-tuning. The encoder maps heterogeneous textual descriptions, including structure names, morphological definitions, and pairwise spatial relations into a unified embedding space. This allows the resulting text embedding to capture structure-specific location cues and facilitates grounding of anatomical terms within the volumetric imaging space.
Query Decoder: To adapt the text-derived representation to each MRI volume, we employ a Transformer-based query decoder that fuses textual embeddings with multi-scale visual features. The text embedding acts as the query, and image features serve as keys and values. A stack of cross-attention decoder blocks refines the query by attending to anatomy-relevant visual cues, enabling inference of subject-specific variations. The query is matched with voxel-level features in the segmentation head, ensuring alignment between textual priors and the spatial context of the MRI scans.
2.2. Unified Visual Backbone for Label-Efficient Segmentation
We construct a unified visual backbone by combining 3D MAE pretraining, masked pseudo-label refinement, and global–local stabilization. This backbone provides a strong initialization for subsequent vision–language fine-tuning.
MAE pretraining.
A 3D Masked Autoencoder (MAE) (He et al., 2021) is first trained on large-scale MRI volumes to learn modality- and site-invariant features. Local patches and a downsampled global view are randomly masked and reconstructed using an MSE loss, yielding a pretrained visual encoder .
Masked Pseudo-Labeling (MPL).
To enable label-efficient domain adaptation, we adopt a 3D MPL teacher–student framework (Grill et al., 2020; Tarvainen and Valpola, 2017). We keep pretrained MAE encoder with a segmentation decoder to build segmentation . Given an input image and label from the source domain, the teacher model provides pseudo-labels for unlabeled target image and student model learns from masked source image and masked target image by minimizing the loss with weight :
| (2) |
Where is a compound segmentation loss that consists of cross-entropy and Dice loss (Zhang et al., 2024).
Global–Local Collaboration (GLC).
To stabilize pseudo-labels under domain shift, the GLC module (Zhang et al., 2024) fuses high-resolution local patches with global context extracted from the MAE encoder and regularizes their consistency. The full GLC formulation is provided in Appendix A. The visual backbone is pretrained with:
| (3) |
where is the loss of regular fully-supervised segmentation in source data and contains the global–local consistency terms (Appendix A). After these stages, the visual backbone is fine-tuned jointly with the language-guided module using only the supervised segmentation loss .
2.3. Anatomical–Linguistic Discriminator
2.3.1. Morphological Discriminator
Different subcortical structures exhibit distinct morphological variations, which serve as crucial reference points during manual annotation. Considering the shape characteristics of brain regions, the shape encoder employs an -equivariant convolutional neural network (Billot et al., 2024) to extract shape features invariant to rigid transformations, mapping 3D annotations into a compact shape embedding space.
The shape encoder is pretrained using a denoising autoencoder framework, mapping noisy inputs to embeddings, which are reconstructed by a decoder comprising transposed 3D convolutions with instance normalization. The reconstruction loss combines MSE and soft Dice loss. During the training of NeuroLangSeg, the pre-trained shape encoder is used to constrain the morphological features. In each training step, both the prediction and the ground truth are forwarded through the fixed to obtain their respective shape embeddings. The discrepancy between two embeddings is quantified using the MSE loss, which enforces the network to capture anatomically plausible shapes:
| (4) |
2.3.2. Topological Discriminator
In addition to shape characteristics, the spatial relationships among subcortical nuclei provide crucial cues for manual annotation. To extract these positional features, we used a LLM to parse natural-language descriptions in annotation protocols provided by Neuromorphometrics, Inc. We used the following prompt to extract anatomical rules into a JSON format: “Please extract the morphological features, relevant reference regions for manual annotation, and positional relationship descriptions… and convert them into a structured JSON description.” The LLM output identified 37 key anatomical pairs (15 left, 15 right, 7 cross-hemisphere) (shown in Appendix Table 4) and defined their relational types in a structured JSON format. For example, from the sentence “the hippocampus is posterior and inferior to the amygdala,” the LLM outputs structured JSON: hippocampus-amygdala: {relative_position: [−1, −1, 0], adjacency_ratio: 1, adjacency_vector: [1, 1,0]}. The discrete direction vector encodes the posterior–inferior offset under a standardized anatomical coordinate system (anterior, superior, right as positive). The adjacency ratio and vector denote whether two structures share a boundary and the dominant direction from one centroid toward the shared interface.
While the LLM identifies which relationships matter, the quantitative features are formalized by computing the statistics from the training set’s ground truths. Each structure pair is thus represented by a 7D relational feature , where is the continuous relative position, is the adjacency ratio, and is the adjacency-direction vector. To account for inter-subject variability in age and development, all relative position vectors are explicitly normalized based on the subject’s total brain volume before being processed by the discriminator. For each subject, anatomical pairs form the relational matrix , encoding the full anatomical topology. is the number of anatomical pairs. The MLP-based location encoder is pretrained in the task of reconstructing relative vectors extracted from annotation images of all subjects in the normal cohort. In NeuroLangSeg training, this fixed enforces topological consistency: for and , their relational matrices and are extracted and encoded as global location embeddings. A Mean Squared Error (MSE) loss minimizes the discrepancy between the two embeddings, constraining the network to preserve accurate anatomical relationships:
| (5) |
2.4. Total Loss
The total training objective of NeuroLangSeg integrates supervised segmentation with protocol-guided anatomical constraints. The supervised term, , combines binary cross-entropy and soft Dice losses to encourage both voxel-level accuracy and region-level overlap fidelity. Two auxiliary regularizers are used: a shape loss that penalizes deviations from protocol-defined morphological characteristics, and a location loss that constrains predictions to anatomically valid spatial neighborhoods derived from protocol-based adjacency rules. The anatomical–linguistic discriminators that define these protocol constraints are not optimized jointly with the segmentation model; they are trained once using manual labels and a fixed anatomical protocol and are frozen during segmentation training and evaluation. The overall loss is defined as:
| (6) |
where , and control the relative contributions of segmentation fidelity, morphological regularization, and anatomical location consistency.
3. Evaluation
3.1. Classical Metrics
We evaluate segmentation quality using Dice Similarity Coefficient (DSC) and Normalized Surface Distance (NSD) (Nikolov et al., 2021) against manual labels when available.
| (7) |
The DSC measures volumetric overlap between a prediction and the manual ground truth , and NSD evaluates boundary agreement within a tolerance . and denote tolerance bands around the prediction and ground-truth boundaries with .
3.2. Anatomical–Linguistic Evaluators
For large-scale or clinical datasets without manual annotations, we rely on three anatomical evaluators—morphological, topological, and volumetric.
The morphological evaluator assesses whether a predicted structure conforms to its anatomical shape. Using , we derive per-label shape priors by encoding the annotations of all healthy instances of each structure and averaging them into a prototype vector . During evaluation, is encoded by , and its embedding is compared with the prototype via cosine similarity. This similarity is reported as the shape-consistency score, reflecting the morphological correctness of the prediction.
| (8) |
The topology evaluator measures whether predicted regions preserve correct anatomical spatial relationships. We extracted the mean and standard deviation of relational features across all annotated subjects. is the number of anatomical pairs. During evaluation, relational features are extracted from the segmentation , normalized using and , and converted into a topological correctness score:
| (9) |
The volumetric evaluator checks whether predicted structure volumes align with population norms. For each structure, the predicted volume is converted to an age- and sex-adjusted BrainChart -score (Bethlehem et al., 2022; Rutherford et al., 2022), which we denote as . Volumes with are considered plausible:
| (10) |
4. Experiments
We evaluate NeuroLangSeg across three complementary settings: (1) in-site, (2) cross-site segmentation and generalization, and (3) clinical disease-cohort assessment. Segmentation accuracy (DSC, NSD) is reported wherever manual labels are available, while the three anatomical–linguistic evaluators (morphological, topological, volumetric) quantify anatomical robustness in both labeled and unlabeled datasets. Across all experiments, we compare NeuroLangSeg with four visual-only segmentation models (FastSurfer (Henschel et al., 2020, 2022), QuickNAT (Guha Roy et al., 2019), MAPSeg (Zhang et al., 2024), and nnU-Net (Isensee et al., 2021), as well as SAT (Zhao et al., 2025) as the vision-language baseline. FastSurfer (Henschel et al., 2022) baseline utilizes the latest VINNA architecture, which incorporates an internal augmentation strategy for resolution independence. Notably, nnU-Net and MAPSeg serve as the underlying backbones for both SAT and NeuroLangSeg to ensure a controlled comparison of linguistic integration. While methods like SynthSeg (Billot et al., 2023) are popular for domain-agnostic full-brain segmentation, they were excluded here as they rely on intensity simulations for whole-brain labels and are not directly applicable to our focus on protocol-specific subcortical structures and anatomical-linguistic alignment.
4.1. Dataset
MAE Pretraining:
We compile 11,948 unlabeled T1/T2 MRI scans spanning ages 1–100 years from nine publicly available datasets (e.g., ABCD (Casey et al., 2018) and HCP (Harms et al., 2018); full list in Appendix B). These scans contain no manual labels and are used solely for self-supervised MAE pretraining. Pseudo-supervised Fine-tuning: A total of 118 manually labeled T1-weighted subjects spanning ages 1–100 years are drawn from ADNI (Jack et al., 2008), CANDI (Kennedy et al., 2012), OASIS (Marcus et al., 2010), Colin (Holmes et al., 1998), and BCP (Howell et al., 2019) dataset. These subjects provide ground-truth annotations for supervised fine-tuning and in-site/cross-site segmentation evaluation. Clinical Cohorts: We additionally include two non–manually labeled clinical datasets—20 subjects from BrainTS (BraTS) (Li et al., 2023) tumor cohort and 30 subjects from ADNI (Jack et al., 2008) Alzheimer’s disease cohort—which are used exclusively to evaluate out-of-distribution anatomical generalization without manual ground truth.
4.2. Experimental Settings
In-Site Segmentation (Exp. 1):
We evaluate performance under matched training and testing conditions using the 118 manually labeled subjects. The dataset is randomly split 50% for training, 10% for validation, and 40% for testing.
Cross-Site Generalization (Exp. 2):
To quantify generalization under realistic domain shift, we use the ADNI and Colin datasets as an external test cohort. 12 ADNI subjects and one Colin subject are withheld from finetuning, thereby providing an independent evaluation of cross-site performance.
Disease–Cohort Assessment (Exp. 3):
To evaluate clinical robustness and out-of-distribution behavior, we apply NeuroLangSeg to the BraTS tumor dataset and the ADNI Alzheimer’s cohort. Since no manual labels are available, evaluation is performed using the morphological, topological, and volumetric anatomical–linguistic evaluators.
4.3. Results and Discussion
1. Segmentation Accuracy and Visualization:
Table 1 summarizes the in/cross-site DSC/NSD performance. In in-site evaluation, NeuroLangSeg obtains the highest average DSC (86.9%) and NSD (95.0%), exceeding the strongest baseline by +1.9% DSC and +2.2% NSD, with larger gains over the average baseline (+4.1% DSC, +8.0% NSD), particularly for small structures such as the amygdala and accumbens. Under cross-site evaluation, NeuroLangSeg again achieves the highest average DSC (84.0%) and NSD (93.6%). This yields +0.2% DSC and +1.4% NSD gain over the strongest baseline and substantial improvements over the average baseline (+3.6% DSC, +14.5% NSD).
Table 1:
Segmentation performance (DSC and NSD) across seven subcortical structures for in-site and cross-site generalization. Bold indicates the best performance.
| Method | DSC % | NSD % | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HIPP | AMG | CD | PT | PD | TM | AB | Avg | HIPP | AMG | CD | PT | PD | TM | AB | Avg | |
| In-site | ||||||||||||||||
| FastSurfer | 85.4 | 80.9 | 88.2 | 88.7 | 82.2 | 91.6 | 78.3 | 85.0 | 91.6 | 88.2 | 94.4 | 92.9 | 87.6 | 89.0 | 92.5 | 90.9 |
| QuickNAT | 77.5 | 59.5 | 80.7 | 83.4 | 72.7 | 87.6 | 59.3 | 74.4 | 78.2 | 43.2 | 79.4 | 82.5 | 61.0 | 74.6 | 61.9 | 68.7 |
| nnU-Net | 85.4 | 81.1 | 87.9 | 88.4 | 84.0 | 91.1 | 75.9 | 84.8 | 93.4 | 91.7 | 94.7 | 91.9 | 91.5 | 90.1 | 91.5 | 92.1 |
| MAPSeg | 85.7 | 79.8 | 88.3 | 88.7 | 84.5 | 91.5 | 76.8 | 85.0 | 91.9 | 85.5 | 95.0 | 93.2 | 88.6 | 88.8 | 91.3 | 90.6 |
| SAT | 85.7 | 81.2 | 87.4 | 88.3 | 84.4 | 90.9 | 77.1 | 85.0 | 92.6 | 91.7 | 95.2 | 95.1 | 91.3 | 89.6 | 93.8 | 92.8 |
| NeuroLangSeg | 87.6 | 84.0 | 88.9 | 89.3 | 86.2 | 92.0 | 80.4 | 86.9 | 95.3 | 94.9 | 97.2 | 96.1 | 93.5 | 92.8 | 95.7 | 95.0 |
| Cross-site | ||||||||||||||||
| FastSurfer | 77.2 | 69.1 | 85.4 | 86.2 | 75.5 | 87.8 | 70.1 | 78.8 | 67.2 | 62.1 | 76.4 | 72.7 | 65.1 | 63.7 | 69.6 | 68.1 |
| QuickNAT | 74.2 | 63.1 | 76.8 | 80.5 | 66.5 | 86.6 | 64.0 | 73.1 | 61.1 | 42.6 | 60.2 | 59.7 | 37.1 | 57.8 | 53.6 | 53.2 |
| nnU-Net | 84.6 | 79.3 | 87.7 | 86.9 | 80.5 | 89.4 | 76.3 | 83.5 | 92.8 | 90.3 | 95.1 | 88.9 | 88.0 | 86.4 | 93.2 | 90.7 |
| MAPSeg | 83.9 | 79.3 | 87.3 | 87.9 | 81.6 | 89.8 | 77.0 | 83.8 | 90.3 | 87.8 | 95.9 | 95.1 | 89.7 | 87.1 | 92.1 | 91.1 |
| SAT | 83.2 | 77.6 | 86.7 | 86.7 | 80.1 | 89.0 | 75.3 | 82.6 | 91.5 | 90.2 | 96.4 | 95.2 | 89.8 | 87.5 | 94.4 | 92.2 |
| NeuroLangSeg | 84.5 | 80.0 | 86.8 | 88.1 | 81.0 | 89.8 | 78.1 | 84.0 | 93.5 | 92.8 | 97.1 | 97.0 | 91.5 | 89.7 | 96.0 | 93.6 |
HIPP:Hippocampus, AMG:Amygdala, TM:Thalamus, CD:Caudate, PT:Putamen, PD:Pallidum, AB:Accumbens
Figure 2 shows in-site qualitative segmentation results. In the coronal view, several baselines enlarge or shrink the amygdala. In the sagittal view, manual annotations are inherently discontinuous because they were drawn primarily in the coronal plane. Models trained only on these labels tend to compensate incorrectly: methods such as SAT, MAPSeg, and QuickNAT often enlarge the structure, FastSurfer tends to shrink it, and nnU-Net frequently yields missing segments. In contrast, NeuroLangSeg produces anatomically consistent shapes across views without artificial expansion, collapse, or disappearance. Minor 1–2 pixel over-segmentation may still occur, which is a common and well-known behavior in deep learning–based segmentation methods.
Figure 2:

Qualitative comparisons. Coronal and sagittal planes, and zoomed-in regions of interest, respectively. Major segmentation errors are highlighted with red arrows. Ground-truth boundaries are indicated by dotted lines, while segmentations from different methods are shown as transparent overlays.
Table 2 reports ablation results on the in-site dataset. Removing all discriminators leads to a clear performance degradation in both Dice and NSD across all subcortical structures. Excluding either the shape or location discriminator generally reduces overall performance of segmentation accuracy compared with the full model.
Table 2:
Segmentation performance (DSC and NSD) across seven subcortical structures for ablation study. Bold indicates the best performance.
| Method | DSC % | NSD % | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| HIPP | AMG | CD | PT | PD | TM | AB | Avg | HIPP | AMG | CD | PT | PD | TM | AB | Avg | |
| In-site | ||||||||||||||||
| w/o discriminators | 86.2 | 81.9 | 87.5 | 88.0 | 84.5 | 91.0 | 77.5 | 85.2 | 93.8 | 92.8 | 95.9 | 94.8 | 91.2 | 89.8 | 94.0 | 93.2 |
| w/o morphological discriminator | 86.6 | 82.4 | 87.9 | 88.4 | 84.9 | 91.3 | 78.3 | 85.7 | 93.5 | 94.4 | 95.0 | 94.7 | 91.6 | 92.2 | 95.3 | 93.8 |
| w/o topological discriminator | 87.2 | 83.5 | 88.4 | 88.9 | 85.9 | 91.7 | 79.4 | 86.4 | 92.8 | 95.1 | 95.2 | 94.6 | 94.6 | 92.8 | 95.6 | 94.4 |
| NeuroLangSeg | 87.6 | 84.0 | 88.9 | 89.3 | 86.2 | 92.0 | 80.4 | 86.9 | 95.3 | 94.9 | 97.2 | 96.1 | 93.5 | 92.8 | 95.7 | 95.0 |
| Cross-site | ||||||||||||||||
| w/o discriminators | 84.1 | 79.3 | 86.5 | 87.5 | 80.1 | 89.3 | 77.5 | 83.5 | 92.6 | 91.9 | 96.7 | 96.6 | 91.3 | 89.0 | 95.7 | 93.4 |
| w/o morphological discriminator | 84.3 | 79.6 | 86.5 | 87.7 | 80.1 | 89.4 | 77.2 | 83.5 | 92.8 | 92.4 | 96.6 | 96.8 | 91.4 | 89.2 | 95.3 | 93.5 |
| w/o topological discriminator | 84.9 | 79.7 | 86.7 | 88.0 | 80.3 | 89.8 | 77.5 | 83.8 | 93.8 | 93.3 | 97.2 | 97.3 | 91.3 | 90.8 | 95.9 | 94.2 |
| NeuroLangSeg | 84.5 | 80.0 | 86.8 | 88.1 | 81.0 | 89.8 | 78.1 | 84.0 | 93.5 | 92.8 | 97.1 | 97.0 | 91.5 | 89.7 | 96.0 | 93.6 |
HIPP:Hippocampus, AMG:Amygdala, TM:Thalamus, CD:Caudate, PT:Putamen, PD:Pallidum, AB:Accumbens
2. Anatomical–Linguistic Evaluation and Clinical Generalization:
Table 3 reports the average anatomical–linguistic evaluator scores for in-site and cross-site cognitively normal (CN) subjects and clinical cohorts. For CN participants, NeuroLangSeg achieves high shape and location scores and volume scores close to 0 (within [−2, 2]), indicating good alignment with the anatomical protocol. Volume score is summarized using to properly capture deviation magnitude. Other segmentation methods also produce CN scores that cluster near the reference values, as expected for healthy controls; however, their morphological, topological, and volumetric metrics are consistently lower than those of NeuroLangSeg. As expected, ADNI (AD) and BraTS (tumor) cohorts show reduced scores due to pathology-related changes in morphology and spatial organization.
Table 3:
Anatomical–Linguistic Evaluators’ average scores for in-site, cross-site (CN), and clinical cohorts (AD and tumor).
| Score | In-site and Cross-site (CN) | ADNI (AD) | BraTS (Tumor) | |||||
|---|---|---|---|---|---|---|---|---|
| FastSurfer | QuickNAT | nnU-Net | MAPSeg | SAT | NeuroLangSeg | NeuroLangSeg | NeuroLangSeg | |
| Shape | 78.1 ± 1.7 | 78.9 ± 1.9 | 79.4 ± 1.3 | 79.0 ± 1.3 | 79.5 ± 1.4 | 79.2 ± 1.7 | 71.2 ± 1.4*** | 71.4 ± 1.4*** |
| Location | 87.9 ± 2.8 | 84.0 ± 2.5*** | 86.9 ± 4.2 | 88.3 ± 2.6 | 87.3 ± 2.9 | 87.9 ± 2.4 | 77.1 ± 2.7*** | 67.4 ± 3.5*** |
| Volume | 1.39 ± 0.51*** | 2.17 ± 0.32*** | 1.05 ± 0.34 | 1.02 ± 0.35 | 1.09 ± 0.35 | 0.94 ± 0.24 | 1.82 ± 0.49*** | 2.24 ± 0.53*** |
indicate statistically significant differences with p < 0.001. Stable CN scores indicate protocol-consistent anatomy, while significant deviations in AD and tumor reflect pathological changes.
A one-way ANOVA was used to assess whether the evaluators distinguish anatomical quality. Within CN subjects, evaluator scores from each method were compared to NeuroLangSeg. Most methods showed no significant difference, as all were tested on the same CN cohort. Only QuickNAT and FastSurfer differed significantly , consistent with their lower DSC/NSD performance. Using NeuroLangSeg across CN, AD, and tumor groups, all three evaluators showed significant differences , indicating stable scores in healthy controls and clear sensitivity to disease-related anatomical changes.
Figure 3 illustrates these patterns using violin plots of our method’s evaluator score distributions for CN, AD, and tumor participants. CN subjects cluster tightly around the reference values, whereas AD and tumor display volume outside [−2, 2], and reductions in shape and location scores. Slightly lower CN scores for the amygdala and pallidum arise because our manual labels are smaller than the FreeSurfer-derived volumes used in the BrainChart reference.
Figure 3:

Shape, location, and volume-score distributions of subcortical regions in Cognitively Normal (CN), Alzheimer’s Disease (AD), and tumor participants, stratified by sex.
5. Conclusion
We introduced NeuroLangSeg, a language-guided framework that unifies visual features with protocol-consistent anatomical reasoning for subcortical MRI segmentation. Through MAE pretraining, pseudo-supervised fine-tuning, and anatomical–linguistic evaluation, our method delivers accurate, consistent, and clinically interpretable segmentations across in-site, cross-site, and disease cohorts. Quantitative comparison with state-of-the-art models shows substantial gains, including +8.0 NSD in in-site evaluation and +14.5 NSD in cross-site generalization over the average baseline. ANOVA analyses further confirm that our anatomical–linguistic scores significantly distinguish healthy controls from pathological cases while remaining stable within CN subjects. By grounding segmentation in a standardized anatomical protocol, NeuroLangSeg advances robust, interpretable, and clinically aligned neuroimaging segmentation. The future work is to extend NeuroLangSeg to infant, pediatric, and fetal MRI, as well as to additional disease cohorts, to further assess robustness across developmental stages and pathological conditions.
Acknowledgments
This work was supported by NIH grants R00HD103912, R01MH133313 (Y.W.), and Neuromorphometrics, Inc.
Appendix A. Model Architecture
A.1. Text Encoder
Using text encoder , each text prompt including structure names, morphological definitions, and pairwise spatial relations is tokenized and processed through a multi-layer self-attention stack, producing a d-dimensional representation. Given a batch of paired textual descriptions , each anatomical structure’s concept is paired with , corresponding either to (1) a descriptive phrase capturing its anatomical morphology , or (2) relational statements describing its spatial relations relative to other structures . The encoder produces embedding pairs :
| (A1) |
To encourage semantically aligned anatomical concepts to share a nearby representation, we optimized using contrastive learning, with an InfoNCE contrastive loss.
| (A2) |
where is the temperature coefficient, and is the number of subcortical regions.
A.2. Query Decoder
The text embedding serves as the initial Query, while the visual feature extracted from the image encoder serves as the Keys and Values. A stack of decoder blocks, each consisting of cross-attention followed by feed-forward layers, progressively refines the representation:
| (A3) |
Through this cross-attention mechanism, the query decoder allows the text embedding to attend to anatomically relevant visual cues, enabling the model to infer subject-specific variations in location, orientation, and shape. The output serves as an image-conditioned anatomical query, which is subsequently matched against voxel-wise visual features in the segmentation head, using a dot-product operation to generate the final segmentation.
A.3. 3D Multi-Scale Masked Autoencoder (MAE)
Our 3D MAE uses 3D ResNet blocks (Zhang et al., 2024) instead of Vision Transformers. The encoder is composed of eight 3D ResNet blocks and we adopt an asymmetric architecture with a lightweight decoder, detailed in Supplementary Figure 4. During training, the model jointly learns from two input types—randomly sampled local patches and a downsampled version of the full volumetric scan , both resized to 963 voxels. To enable self-supervised learning, both and are partitioned into non-overlapping 3D patches and subjected to random masking. For , we use a patch size of 83 and mask 70% of the patches uniformly at random. For , which provides a broader field of view (FOV), we use a smaller patch size of 43 while maintaining the same masking ratio. The resulting masked inputs, denoted as and , are passed to the MAE, which is trained to reconstruct the original unmasked volumes using mean squared error loss computed only on the masked regions.
Figure 4:

Illustrations of MAE 3D ResNet Block and 3D architectures.
A.4. 3D Masked Pseudo-Labeling (MPL)
MAPSeg employs a 3D Masked Pseudo-Labeling (MPL) strategy based on a teacher–student architecture (Figure 5). The segmentation backbone combines the pretrained MAE encoder with a lightweight 3D decoder adapted from DeepLabV3. The decoder uses a 3D Atrous Spatial Pyramid Pooling (ASPP) module with multi-scale dilated convolutions to enlarge the effective receptive field. The student model processes both labeled source volumes and unlabeled target volumes, while the teacher model—an exponential moving average (EMA) of the student—produces stable pseudo-labels for the target domain.
To improve generalization, the student receives masked input volumes, following the same masking scheme used during MAE pretraining. This forces the model to rely on global context rather than local intensity alone. MPL integrates (i) supervised loss on source data and (ii) consistency loss between teacher and student predictions on target data. This yields a label-efficient adaptation mechanism that leverages MAE-learned priors while mitigating noisy pseudo-label propagation.
The teacher model’s parameters are updated during training via an exponential moving average (EMA) of the student model’s parameters (Tarvainen and Valpola, 2017):
| (A4) |
where and denote training iterations and is the EMA update weight. For models initialized from large-scale MAE pretraining, we set during the first 1,000 steps and afterwards. For models pretrained on small-scale source and target datasets (e.g., only dozens of scans), we set during the first 1,000 steps, during the next 2,000 steps, and for the remaining training. The teacher model is initialized with the student model’s parameters after a warm-up stage (e.g., 1,000 iterations) on the source-domain data.
A.5. 3D Global-Local Collaboration (GLC)
To improve pseudo-label stability under large domain shifts, we introduce a Global–Local Collaboration (GLC) module (Zhang et al., 2024). For each scan (Figure 5), we extract a high-resolution local patch and a downsampled global volume . The encoder produces local and global features:
| (A5) |
where is a binary mask and denotes cropping followed by interpolation to match spatial dimensions. The GLC module fuses the two feature streams by channel-wise concatenation:
| (A6) |
forming a unified 1024-dimensional latent representation processed by the ASPP head for segmentation. To provide global supervision, the model also predicts from the global view via
| (A7) |
To enforce alignment between local and global information, we impose a cosine similarity regularizer:
| (A8) |
The GLC losses for source and target data are:
| (A9) |
| (A10) |
During training, local 96 × 96 × 96 patches are randomly sampled to provide high-resolution detail while the global branch maintains coarse contextual awareness. At inference, predictions are generated with a sliding-window scheme (stride 80) to cover the full volume.
With the supervised loss
| (A11) |
the full training objective becomes:
| (A12) |
Figure 5:

Illustrations of MPL and GLC.
A.6. Morphological Discriminator
The morphological discriminator adopts the SE(3) -equivariant convolutional neural network (Billot et al., 2024) as the shape encoder to extract features. The shape encoder is pretrained using a denoising autoencoder framework, mapping noisy inputs to embeddings, which are reconstructed by a decoder comprising transposed 3D convolutions with instance normalization. For framework details, please refer to the supplementary Figure 6.
The reconstruction task is trained by minimizing the difference between the reconstructed image and the input image, and the loss uses a combination of multi-class Dice loss and MSE loss. For the input labeled image , the reconstructed labeled image is obtained after passing through the shape encoder and decoder. The loss is calculated as follows:
| (A13) |
Figure 6:

Illustrations of Morphological Discriminator.
A.7. Topological Discriminator
Topological Discriminator employs MLP as the location encoder , which is pretrained to reconstruct relation features from noisy inputs. For framework details, please refer to the supplementary Figure 7. For each subject, we extract relation vectors of 37 anatomical pairs from ground truth. is the number of anatomical pairs. The relation feature is composed of the continuous relative position , the adjacency ratio , and the adjacency-direction . is obtained by taking the difference between the centroids of structure and of structure . Adjacency ratio is the proportion of shared boundary voxels, and the adjacency-direction vector is defined as the difference between the centroid of the subset of structure ’s boundary voxels that are adjacent to and the centroid of structure as a whole. Here, is the subset of boundary voxels of structure that are in direct spatial contact with structure , and denotes the set of all boundary voxels belonging to anatomical structure . Then we generate n = 100 perturbation versions in the relation vectors by adding Gaussian noise . The encoder maps noisy relations to embeddings, which are reconstructed by a feedforward decoder. The reconstruction loss is:
| (A14) |
where denotes the reconstructed relation matrix.
The details of the parameter settings for the pre-training tasks of the shape discriminator and the topology discriminator are presented in Table 11.
Figure 7:

Illustrations of Topological Discriminator.
Appendix B. Dataset Description
We gather T1-weighted and T2-weighted MRI data across 1–100 years from 15 publicly available datasets. The detailed dataset information is as follows:
ABCD: Adolescent Brain Cognitive Development Study (Casey et al., 2018) is a large-scale, longitudinal neuroimaging and behavioral study tracking brain development and child health in over 10,000 U.S. children aged 9–10 years. Participants were enrolled at ages 9–10 and are being followed into their early 20s. We collected 2930 subjects with 3211 longitudinal scans spanning 9–17 years, with 3211 T1-weighted and 3209 T2-weighted images.
ABIDE-I: Autism Brain Imaging Data Exchange (Di Martino et al., 2014) is a cross-sectional multi-site initiative that shares resting-state fMRI and structural MRI data from individuals with autism and typically developing controls. We collected 1102 subjects/scans of T1-weighted images spanning 6 – 64 years.
ADHD-200: ADHD-200 Global Competition (Bellec et al., 2017) is a cross-sectional multi-site dataset sharing resting-state fMRI and structural MRI data to identify biomarkers of Attention Deficit Hyperactivity Disorder (ADHD). We collected 869 subjects/scans of T1-weighted images spanning 7 – 26 years.
ADNI: Alzheimer’s Disease Neuroimaging Initiative (Jack et al., 2008) is a longitudinal, multi-site study designed to develop clinical, imaging, genetic, and biochemical biomarkers for early detection and tracking of Alzheimer’s disease. We collected 50 subjects/scans of T1-weighted images spanning 60–96 years.
BCP: Baby Connectome Project (Howell et al., 2019) is a longitudinal neuroimaging study aiming to map early brain development and connectivity from infancy through early childhood. We collected 2126 subjects with 2444 longitudinal scans spanning 0–7 years, with 2406 T1-weighted and 2347 T2-weighted images.
HBN: Healthy Brain Network (Alexander et al., 2017) is a cross-sectional transdiagnostic pediatric study collecting neuroimaging, behavioral, cognitive, and genetic data to better understand mental health and learning disorders. We collected 1729 subjects/scans spanning 5 – 22 years, with 1698 T1-weighted and 562 T2-weighted images.
HCP-A: Human Connectome Project – Aging (Bookheimer et al., 2019) is a cross-sectional dataset focused on understanding brain connectivity and aging across the adult lifespan. We collected 725 subjects/scans spanning 36 – 100 years, with 725 T1-weighted and 725 T2-weighted images.
HCP-D: Human Connectome Project – Development (Somerville et al., 2018) is a cross-sectional study examining brain development and connectivity from childhood through young adulthood. We collected 652 subjects/scans spanning 6 – 22 years, with 652 T1-weighted and 652 T2-weighted images.
HCP-YA: Human Connectome Project – Development (Harms et al., 2018) is a cross-sectional study to map the healthy human connectome by collecting and freely distributing neuroimaging and behavioral data on 1,200 normal young adults, aged 22–35.
PING: Pediatric Imaging, Neurocognition, and Genetics (Jernigan et al., 2016) is a cross-sectional study designed to assess brain development and its genetic and environmental influences in children and adolescents. We collected 754 subjects/scans spanning 0 – 22 years, with 752 T1-weighted and 106 T2-weighted images.
CANDI: Child and Adolescent Neuro Development Initiative (Kennedy et al., 2012) includes structural MRI scans of children and adolescents, supporting research on brain development, psychiatric disorders, and neuroanatomical differences across diagnoses.
OASIS: Open Access Series of Imaging Studies (Marcus et al., 2010) provides structural brain MRI data across the adult lifespan, including individuals with and without Alzheimer’s disease, to support neurodegenerative and aging research.
COLIN: Colin27 Brain Atlas (Holmes et al., 1998) is a high-resolution MRI brain template created by averaging 27 T1-weighted scans of a single individual.
BraTS2023: The Brain Tumor Segmentation (BraTS) Challenge 2023 (Li et al., 2023) provides an expanded multi-site mpMRI dataset ( 4,500 cases) with expert tumor delineations across diverse populations and tumor types, enabling benchmarking of segmentation, missing-data handling, and cross-task generalizability.
Table 4:
List of the 37 anatomically relevant structure pairs used to construct relation vectors.
| Category | Structure pairs |
|---|---|
| Left intra-hemispheric (15) | (L-HIPP, L-AMG), (L-HIPP, L-TM), (L-HIPP, L-AB), (L-AMG, L-AB), (L-AMG, L-TM), (L-AMG, L-PT), (L-CD, L-PT), (L-CD, L-PD), (L-CD, L-AB), (L-CD, L-TM), (L-PT, L-PD), (L-PT, L-AB), (L-PT, L-TM), (L-PD, L-TM), (L-AB, L-TM) |
| Right intra-hemispheric (15) | (R-HIPP, R-AMG), (R-HIPP, R-TM), (R-HIPP, R-AB), (R-AMG, R-AB), (R-AMG, R-TM), (R-AMG, R-PT), (R-CD, R-PT), (R-CD, R-PD), (R-CD, R-AB), (R-CD, R-TM), (R-PT, R-PD), (R-PT, R-AB), (R-PT, R-TM), (R-PD, R-TM), (R-AB, R-TM) |
| Inter-hemispheric (7) | (L-HIPP, R-HIPP), (L-AMG, R-AMG), (L-CD, R-CD), (L-PT, R-PT), (L-PD, R-PD), (L-TM, R-TM), (L-AB, R-AB) |
We pretrain our models on a dataset of approximately 12,000 subjects, including both T1-weighted and T2-weighted scans. All datasets were preprocessed using N4 bias correction and skull stripping. Full dataset details are provided in Supplementary Table 5.
We use 118 manually labeled subjects from ADNI (Jack et al., 2008), CANDI (Kennedy et al., 2012), OASIS (Marcus et al., 2010), Colin (Holmes et al., 1998), and BCP-50 (Howell et al., 2019) (details in Supplementary Table 6). For robust model development, 50% of subjects are used for training, 10% for validation, and 40% are held out for testing. In our study, CANDI, Colin, and OASIS are single-site datasets. The ADNI manual dataset includes 30 subjects spanning 25 distinct sites. ADNI subjects are split such that 15 subjects from 12 sites are used for training, 3 subjects from 2 sites are used for validation, and 12 subjects from 11 different sites are used for testing. Colin and 12 ADNI testing subjects are reserved exclusively for cross-site inference. Two non–manually labeled clinical datasets—BraTS(Li et al., 2023) tumor cohort and ADNI (Jack et al., 2008)(Alzheimer’s disease) cohort (Table 7)—which are used exclusively to evaluate out-of-distribution anatomical generalization.
Table 5:
Dataset summary across age ranges for pretraining: modality, number of subjects, scans, and age span.
| Dataset | Modality | Subjects | T1 Scans | T2 Scans | Age (yrs) |
|---|---|---|---|---|---|
| ABCD | T1w, T2w | 2930 | 3211 | 3209 | 9–16 |
| ABIDE-I | T1w | 1102 | 1102 | – | 6–64 |
| ADHD-200 | T1w | 869 | 869 | – | 7–26 |
| BCP | T1w, T2w | 2126 | 5183 | 4303 | 1–7 |
| HBN | T1w, T2w | 1729 | 1684 | 560 | 6–64 |
| HCP-A | T1w, T2w | 725 | 725 | 725 | 36–100 |
| HCP-D | T1w, T2w | 652 | 652 | 652 | 6–21 |
| HCP-YA | T1w, T2w | 1061 | 10 | 660 | 22–35 |
| PING | T1w, T2w | 754 | 752 | 106 | 3–21 |
Table 6:
Dataset summary for subcortical segmentation: modality, number of subjects, scans, and age span.
| Dataset | Modality | Subjects | T1 Scans | T2 Scans | Age (yrs) |
|---|---|---|---|---|---|
| OASIS | T1w | 50 | 70 | – | 18–93 |
| ADNI | T1w | 29 | 30 | – | 71–88 |
| CANDI | T1w | 13 | 13 | – | 5–15 |
| Colin | T1w | 1 | 1 | – | 27 |
| BCP | T1w, T2w | 25 | 25 | 25 | 1–2 |
Appendix C. Baseline
C.1. FastSurfer
FastSurferCNN (Henschel et al., 2020) is a 2D fully convolutional architecture designed for fast whole-brain segmentation and serves as the segmentation module within the FastSurfer pipeline. The network follows an encoder–decoder design similar to QuickNAT but introduces several architectural improvements: competitive dense blocks (replacing concatenation with maxout operations) to encourage feature competition, unpooling layers for spatially accurate upsampling, and a wider contextual field to better capture neuroanatomical boundaries. FastSurferCNN predicts 2D segmentations for axial, coronal, and sagittal slices, which are combined through a multi-view aggregation strategy to produce the final 3D mask.
In its original formulation, FastSurfer is trained on FreeSurfer-derived labels for 95 anatomical structures, providing a high-speed alternative to traditional surface-based processing. Because the method relies on 2D slice-wise predictions, its performance can vary across views, especially for small or discontinuous subcortical structures.
For our study, we employ FastSurferCNN as a baseline and fine-tune it on our 7-class subcortical label set under the same training conditions as the other methods.
Table 7:
Dataset summary for clinical generalization: modality, number of subjects, scans, and age span.
| Dataset | Modality | Subjects | T1 Scans | T2 Scans | Age (yrs) |
|---|---|---|---|---|---|
| ADNI | T1w | 30 | 30 | – | 60–96 |
| BraTS | T1w | 20 | 20 | – | 50–85 |
C.2. QuickNAT
QuickNAT (Guha Roy et al., 2019) is a 2D fully convolutional framework that performs segmentation on individual slices rather than full 3D volumes. The method trains three independent F-CNNs on single coronal, axial, and sagittal slices, and fuses their outputs through a view-aggregation module to obtain the final 3D prediction. Each F-CNN adopts an encoder–decoder architecture with skip connections, unpooling layers, and dense connections to improve gradient flow and feature reuse. The network is optimized using a combination of multi-class Dice loss and weighted logistic loss to address class imbalance and enhance boundary delineation.
In its original formulation, QuickNAT is pre-trained using auxiliary labels generated by FreeSurfer and then fine-tuned on expert manual segmentations. This strategy leverages large-scale automated annotations while adapting to higher-quality ground truth. Because QuickNAT operates on single 2D slices without explicit 3D contextual modeling, its predictions may vary across views, particularly for small or irregularly shaped subcortical structures.
For our experiments, we fine-tune QuickNAT on our 7-class subcortical label set to serve as a baseline under consistent training and evaluation conditions.
C.3. nnU-Net
nnU-Net(Isensee et al., 2021) is a self-configuring segmentation framework that automatically adapts its network architecture, data preprocessing strategies, and training pipelines to the characteristics of a given dataset. Following the original design, we relied entirely on nnU-Net’s built-in mechanisms, including its automated determination of patch size, batch size, normalization scheme, deep supervision, and data augmentation policies.
For our experiments, we use the default 3D full-resolution configuration. During training, the model optimized the standard combination of Dice loss and cross-entropy loss under the framework’s predefined schedule, including the default learning rate, optimizer settings, and training epochs. After training, inference was performed using nnU-Net’s standard test-time augmentation and sliding-window strategy.
C.4. MAPSeg
MAPSeg (Zhang et al., 2024) is an unsupervised domain adaptation (UDA) framework designed for volumetric medical image segmentation. It integrates 3D masked autoencoding (MAE) with a masked pseudo-labeling (MPL) strategy and a global–local consistency (GLC) objective to improve robustness across heterogeneous imaging domains. The framework is self-supervised during pretraining through 3D MAE reconstruction, and subsequently refines pseudo-labels using MPL to adapt the model to new domains without requiring manual annotations. GLC further stabilizes training by enforcing consistency between global volumetric context and local structural details.
MAPSeg was originally proposed for centralized, federated, and test-time UDA settings, allowing models trained on one domain to generalize to unseen scanners or cohorts. Because MAPSeg operates directly on 3D volumes with domain-adaptive pseudo-labeling, it can handle cross-domain variations more effectively than purely supervised baselines.
For our experiments, we use MAPSeg’s domain-adapted 3D backbone and fine-tune it on our 7-class subcortical label set to serve as a UDA-based baseline under consistent training conditions.
C.5. SAT
SAT(Zhao et al., 2025) is a large-vocabulary medical image segmentation framework integrating an image encoder with a text-aware feature modulation module. The central design of SAT involves a text–image feature alignment mechanism and a hierarchical decoder with Mixture-of-Experts layers. The encoding process of SAT involves two encoders: a vision encoder responsible for extracting multi-scale 3D visual representations from the input medical volume, and a text encoder mapping natural language descriptions of target structures into embedding vectors. During the decoding stage, the segmentation decoder dynamically fuses these two modalities to produce structure-specific predictions.
In our experiments, we employed the SAT-nano variant built upon the nnU-Net vision backbone as recommended in the original paper. We fine-tuned the SAT-nano on our dataset using only the image domain and segmentation supervision corresponding to our task. No additional large-scale pretraining or external datasets were used.
Appendix D. Experiment Settings
D.1. Visual-Backbone
MAE Pretraining.
For MAE pretraining, we follow the training configurations listed in Table 8. Each mini-batch contains a randomly sampled local patch and a downsampled global scan . The masking patch size specified in Table 8 is applied only to ; for , the masking patch is always set to half the size due to its larger field of view. During the MAE stage, we apply random 3D affine transformations with isotropic scaling between 75–150% and rotation sampled from [−40°, 40°].
MPL-GLC.
For centralized UDA brain MRI segmentation, the detailed training configurations are provided in Table 9. Each mini-batch contains four patches: a local–global pair from the source domain and another pair from the target domain (each of size 963). During warm-up epochs, the model is trained exclusively on the source domain. Model selection is based on the validation Score, with a patience of 50 epochs.
Target-domain augmentation.
We apply a random 3D affine transformation with isotropic scaling of 70–130% and rotation sampled from [−30°, 30°].
Table 8:
MAE Pretraining Configurations
| config | value |
|---|---|
| patch size | 96 × 96 × 96 |
| local masking patch | 8 × 8 × 8 |
| global masking patch | 4 × 4 × 4 |
| masking ratio | 70% |
| optimizer | AdamW |
| learning rate | 2 × 10−4 |
| weight decay | 0.05 |
| momentum | |
| lr scheduler | cosine annealing |
| epochs | 300 |
| batch size | 4 |
| iters/epoch | 500 |
| aug. prob. | 0.35 |
| augmentation | random affine |
Table 9:
Fine-tuning Configurations
| config | value |
|---|---|
| patch size | 96 × 96 × 96 |
| local masking patch | 8 × 8 × 8 |
| global masking patch | 4 × 4 × 4 |
| masking ratio | 70% |
| optimizer | AdamW |
| learning rate | 1 × 10−4 |
| weight decay | 0.01 |
| momentum | , |
| lr scheduler | cosine WR |
| total epochs | 100 |
| warmup epochs | first 10 |
| early stop | 50 |
| batch size | 1 |
| iters/epoch | 100 |
| augmentation | random affine |
| source aug. | random bias field |
| target aug. | random gamma |
| source prediction weight | |
| EMA update weight | |
| auxiliary global loss weight | |
| cosine similarity weight |
Source-domain augmentation.
A stronger augmentation pipeline is used for the source domain, including random affine (70–140% scaling, [−30°, 30°] rotation), random bias field, and random gamma transformation .
Teacher–student update.
The teacher model is updated using an exponential moving average (EMA) of the student parameters :
| (A15) |
where denotes the training iteration and is the EMA decay rate.
EMA scheduling.
For models initialized from large-scale MAE pretraining, we set for the first 1,000 steps and 0.9999 thereafter. For models pretrained only on small-scale datasets (tens of scans), we use for the first 1,000 steps, 0.999 for the next 2,000 steps, and 0.9999 for the remaining iterations.
The teacher network is initialized using the student parameters after a warm-up stage (e.g., 1,000 iterations) trained solely on the source domain.
D.2. Language-Guided Prompt Encoding
The detailed training configurations of NeuroLangSeg are provided in Table 10.
Appendix E. Results
Figure 8 shows qualitative segmentation results across methods for cross-site. Table 12 reports anatomical–linguistic evaluator scores across seven subcortical structures for in-site and cross-site CN subjects and clinical cohorts. Figure 9 shows anatomical–linguistic evaluator score distributions of seven subcortical structures across methods.
Table 10:
NeuroLangSeg Configurations
| config | value |
|---|---|
| patch size | 96 × 96 × 96 |
| embedding dimension | 512 |
| text max query | 32 |
| optimizer | AdamW |
| learning rate | 1 × 10−4 |
| weight decay | 0.01 |
| momentum | , |
| lr scheduler | cosine annealing |
| epochs | 500 |
| batch size | 2 |
| iters/epoch | 500 |
| shape loss weight | 0.5 |
| location loss weight | 0.5 |
Table 11:
Discriminator Pretraining Configurations
| config | value |
|---|---|
| shape encoder | |
| shape embedding dimension | 128 |
| shape encoder crop size | 96 × 96 × 96 |
| optimizer | Adam |
| loss | Dice+MSE |
| learning rate | 1 × 10−4 |
| weight decay | 1 × 10−5 |
| lr scheduler | cosine annealing |
| epochs | 50 |
| batch size | 8 |
| augmentation | gaussian noise |
| noise perturbations | 100 |
| location encoder | |
| location encoder batch size | 32 |
| relation pairs | 37 |
| location embedding dimension | 64 |
| optimizer | Adam |
| loss | MSE |
| learning rate | 1 × 10−3 |
| weight decay | 1 × 10−4 |
| lr scheduler | cosine annealing |
| epochs | 50 |
| batch size | 32 |
| augmentation | gaussian noise |
| noise perturbations | 100 |
Table 12:
Anatomical–Linguistic Evaluators’ scores across seven subcortical structures for in-site, cross-site, and clinical generalization.
| Score | Methods | Seven subcortical structures | ||||||
|---|---|---|---|---|---|---|---|---|
| HIPP | AMG | CD | PT | PD | TM | AB | ||
| Shape | In-site and Cross-site (CN) | |||||||
| FastSurfer | 88.2±2.3 | 74.8±2.4 | 73.3±1.8 | 75.9±2.1 | 74.2±1.0 | 86.0±3.7 | 74.1±6.4 | |
| QuickNAT | 94.6 ± 1.6*** | 79.9 ± 3.2*** | 71.2 ± 1.8*** | 77.4±1.7 | 74.0±1.1 | 84.9±2.5 | 70.1 ± 8.4* | |
| nnU-Net | 89.7±1.9 | 76.5±1.2 | 74.4±1.8 | 77.6±1.6 | 73.6±0.9 | 87.3±3.2 | 76.6±5.4 | |
| MAPSeg | 89.3±2.1 | 76.7±1.1 | 73.4±1.7 | 75.5±2.3 | 74.0±0.9 | 86.3±3.1 | 77.7±5.3 | |
| SAT | 90.0±2.0 | 76.0±2.0 | 73.9±1.4 | 76.5±1.6 | 73.7±0.9 | 87.6±2.7 | 78.9±5.7 | |
| NeuroLangSeg | 90.0±2.3 | 75.1±2.7 | 74.1±1.9 | 76.8±1.8 | 74.3±1.0 | 86.9±2.9 | 77.2±6.7 | |
| Clinical generalization | ||||||||
| (ADNI) AD-NeuroLangSeg | 81.2 ±5.7 *** | 63.4 ± 3.4 *** | 70.9 ± 1.6 *** | 75.7 ± 0.8 ** | 75.4 ± 0.6 *** | 66.7 ± 3.4 *** | 65.5 ± 1.6 *** | |
| (BraTS) Tumor-NeuroLangSeg | 85.7 ± 3.7 *** | 64.6 ± 3.0 *** | 70.2 ± 2.3 *** | 77.6±0.6 | 74.2±0.9 | 64.7 ± 5.6 *** | 62.8 ± 1.5 *** | |
| Location | In-site and Cross-site (CN) | |||||||
| FastSurfer | 85.5±5.4 | 88.6±3.4 | 86.9±3.2 | 87.1±3.9 | 88.6 ± 2.6*** | 87.1±4.2 | 87.2±3.2 | |
| QuickNAT | 82.7 ± 3.7*** | 79.1 ± 9.5*** | 81.8 ± 4.6*** | 79.6 ± 3.3*** | 77.0 ± 10.2*** | 77.2 ± 8.6*** | 82.4±4.2 | |
| nnU-Net | 84.8±5.0 | 88.0±4.3 | 85.4±4.0 | 86.5±4.7 | 87.8±4.3 | 85.9 ±4.7 ** | 86.0±4.6 | |
| MAPSeg | 86.3±3.7 | 89.0±2.9 | 87.3±3.0 | 88.1±2.7 | 88.9 ± 2.6*** | 87.8±3.3 | 87.8±3.0 | |
| SAT | 85.8±3.7 | 88.2±3.2 | 85.8±3.4 | 86.3±3.5 | 87.2±3.5 | 87.4±3.2 | 86.4±3.9 | |
| NeuroLangSeg | 86.9±3.1 | 89.7±2.5 | 87.9±2.7 | 86.2±2.4 | 85.7±2.3 | 89.8±2.8 | 85.8±4.0 | |
| Clinical generalization | ||||||||
| (ADNI) AD-NeuroLangSeg | 69.3 ± 4.9 *** | 80.2 ± 5.1 *** | 70.6 ± 3.4*** | 74.7 ± 3.9*** | 74.3 ± 3.4*** | 74.0 ± 4.2*** | 73.0 ± 5.3 *** | |
| (BraTS) Tumor-NeuroLangSeg | 70.0 ± 5.7 *** | 73.9 ± 6.2 *** | 67.6 ± 9.0*** | 66.1 ± 6.9 *** | 67.0 ± 6.7*** | 70.5 ± 11.7*** | 62.7 ± 10.9*** | |
| Volume | In-site and Cross-site (CN) | |||||||
| FastSurfer | 1.46 ± 0.96 ** | 2.25 ± 1.05 *** | 1.40 ± 1.01 ** | 1.02±0.77 | 1.79±1.59 | 0.75±0.66 | 1.06±0.75 | |
| QuickNAT | 1.55 ± 1.10** | 3.56 ± 1.34 *** | 1.69 ± 1.13*** | 0.84±0.57 | 2.54 ± 0.94*** | 1.53 ± 0.79*** | 3.50 ± 1.02*** | |
| nnU-Net | 1.10±0.70 | 1.93 ± 0.89* | 0.82±0.60 | 0.75±0.54 | 1.29±0.76 | 0.67±0.52 | 0.79±0.61 | |
| MAPSeg | 1.01±0.68 | 1.40±0.68 | 1.13±0.80 | 0.94±0.65 | 1.28±0.82 | 0.70±0.55 | 0.71±0.49 | |
| SAT | 0.98±0.80 | 1.92 ± 0.90* | 1.07±0.69 | 0.68±0.54 | 1.36±1.03 | 0.74±0.60 | 0.86±0.60 | |
| NeuroLangSeg | 0.91±0.67 | 1.44±0.77 | 0.84±0.55 | 0.72±0.50 | 1.13±0.74 | 0.78±0.54 | 0.77±0.59 | |
| Clinical generalization | ||||||||
| (ADNI) AD-NeuroLangSeg | 2.30 ± 1.27*** | 2.88 ± 1.23 *** | 1.62 ± 1.02*** | 1.41 ± 0.84 *** | 2.80 ± 0.87*** | 0.81±0.54 | 0.93±0.78 | |
| (BraTS) Tumor-NeuroLangSeg | 2.12 ± 1.29 *** | 3.64 ± 1.97 *** | 1.20±1.41 | 1.59 ± 1.53*** | 3.58 ± 1.46 *** | 1.70 ± 1.87*** | 1.83 ± 1.71*** | |
p < 0.05;
p < 0.01;
p < 0.001; n.s., not significant, using ANOVA with Bonferroni correction for multiple comparisons.
HIPP:Hippocampus, AMG:Amygdala, TM:Thalamus, CD:Caudate, PT:Putamen, PD:Pallidum, AB:Accumbens
Figure 8:

Qualitative segmentation performance for cross-site. Coronal and sagittal views are shown, together with corresponding zoomed-in regions of interest. Major segmentation errors are highlighted with red arrows. Ground-truth boundaries are indicated by dotted lines, while segmentations from different methods are shown as transparent overlays.
Figure 9:

Shape, location, and Z-score distributions of seven subcortical structures compared to baselines. *** indicate statistically significant differences with p < 0.001 under Bonferroni correction.
References
- Alexander Lydia, Escalera Juan, Ai Leo, et al. An open resource for transdiagnostic research in pediatric mental health and learning disorders. Scientific Data, 4:170181, 2017. doi: 10.1038/sdata.2017.181. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Avants Brian B, Tustison Nicholas J, Song Gang, Cook Philip A, Klein Arno, and Gee James C. A reproducible evaluation of ANTs similarity metric performance in brain image registration. Neuroimage, 54(3):2033–2044, February 2011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Babalola Kolawole Oluwole, Patenaude Brian, Aljabar Paul, Schnabel Julia, Kennedy David, Crum William, Smith Stephen, Cootes Tim, Jenkinson Mark, and Rueckert Daniel. An evaluation of four automatic methods of segmenting the subcortical structures in the brain. NeuroImage, 47(4):1435–1447, 2009. ISSN 1053–8119. doi: 10.1016/j.neuroimage.2009.05.029. URL https://www.sciencedirect.com/science/article/pii/S1053811909005114. [DOI] [PubMed] [Google Scholar]
- Baribeau Danielle A, Dupuis Annie, Paton Tara A, Hammill Christopher, Scherer Stephen W, Schachar Russell J, Arnold Paul D, Szatmari Peter, Nicolson Rob, Georgiades Stelios, Crosbie Jennifer, Brian Jessica, Iaboni Alana, Kushki Azadeh, Lerch Jason P, and Anagnostou Evdokia. Structural neuroimaging correlates of social deficits are similar in autism spectrum disorder and attention-deficit/hyperactivity disorder: analysis from the POND network. Transl. Psychiatry, 9(1):72, February 2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bellec Pierre, Chu Carlton, Chouinard-Decorte François, Benhajali Yassine, Margulies Daniel S, and Craddock R Cameron. The neuro bureau ADHD-200 preprocessed repository. Neuroimage, 144(Pt B):275–286, January 2017. [DOI] [PubMed] [Google Scholar]
- Bethlehem RAI, Seidlitz J, White SR, Vogel JW, Anderson KM, Adamson C, Adler S, Alexopoulos GS, Anagnostou E, Areces-Gonzalez A, Astle DE, Auyeung B, Ayub M, Bae J, Ball G, Baron-Cohen S, Beare R, Bedford SA, Benegal V, Beyer F, Blangero J, Cábez M Blesa, Boardman JP, Borzage M, Bosch-Bayard JF, Bourke N, Calhoun VD, Chakravarty MM, Chen C, Chertavian C, Chetelat G, Chong YS, Cole JH, Corvin A, Costantino M, Courchesne E, Crivello F, Cropley VL, Crosbie J, Crossley N, Delarue M, Delorme R, Desrivieres S, Devenyi GA, Di Biase MA, Dolan R, Donald KA, Donohoe G, Dunlop K, Edwards AD, Elison JT, Ellis CT, Elman JA, Eyler L, Fair DA, Feczko E, Fletcher PC, Fonagy P, Franz CE, Galan-Garcia L, Gholipour A, Giedd J, Gilmore JH, Glahn DC, Goodyer IM, Grant PE, Groenewold NA, Gunning FM, Gur RE, Gur RC, Hammill CF, Hansson O, Hedden T, Heinz A, Henson RN, Heuer K, Hoare J, Holla B, Holmes AJ, Holt R, Huang H, Im K, Ipser J, Jack CR Jr, Jackowski AP, Jia T, Johnson KA, Jones PB, Jones DT, Kahn RS, Karlsson H, Karlsson L, Kawashima R, Kelley EA, Kern S, Kim KW, Kitzbichler MG, Kremen WS, Lalonde F, Landeau B, Lee S, Lerch J, Lewis JD, Li J, Liao W, Liston C, Lombardo MV, Lv J, Lynch C, Mallard TT, Marcelis M, Markello RD, Mathias SR, Mazoyer B, McGuire P, Meaney MJ, Mechelli A, Medic N, Misic B, Morgan SE, Mothersill D, Nigg J, Ong MQW, Ortinau C, Ossenkoppele R, Ouyang M, Palaniyappan L, Paly L, Pan PM, Pantelis C, Park MM, Paus T, Pausova Z, Paz-Linares D, Binette A Pichet, Pierce K, Qian X, Qiu J, Qiu A, Raznahan A, Rittman T, Rodrigue A, Rollins CK, Romero-Garcia R, Ronan L, Rosenberg MD, Rowitch DH, Salum GA, Satterthwaite TD, Schaare HL, Schachar RJ, Schultz AP, Schumann G, Schöll M, Sharp D, Shinohara RT, Skoog I, Smyser CD, Sperling RA, Stein DJ, Stolicyn A, Suckling J, Sullivan G, Taki Y, Thyreau B, Toro R, Traut N, Tsvetanov KA, Turk-Browne NB, Tuulari JJ, Tzourio C, Vachon-Presseau É, Valdes-Sosa MJ, Valdes-Sosa PA, Valk SL, van Amelsvoort T, Vandekar SN, Vasung L, Victoria LW, Villeneuve S, Villringer A, Vértes PE, Wagstyl K, Wang YS, Warfield SK, Warrier V, Westman E, Westwater ML, Whalley HC, Witte AV, Yang N, Yeo B, Yun H, Zalesky A, Zar HJ, Zettergren A, Zhou JH, Ziauddeen H, Zugman A, Zuo XN, 3R-BRAIN, AIBL, Alzheimer’s Disease Neuroimaging Initiative, Alzheimer’s Disease Repository Without Borders Investigators, CALM Team, Cam-CAN, CCNP, COBRE, cVEDA, ENIGMA Developmental Brain Age Working Group, Developing Human Connectome Project, FinnBrain, Harvard Aging Brain Study, IMAGEN, KNE96, Mayo Clinic Study of Aging, NSPN, POND, PREVENT-AD Research Group, VETSA, Bullmore ET, and Alexander-Bloch AF. Brain charts for the human lifespan. Nature, 604(7906):525–533, April 2022.35388223 [Google Scholar]
- Billot Benjamin, Greve Douglas N., Puonti Oula, Thielscher Axel, Van Leemput Koen, Fischl Bruce, Dalca Adrian V., and Iglesias Juan Eugenio. Synthseg: Segmentation of brain mri scans of any contrast and resolution without retraining. Medical Image Analysis, 86:102789, 2023. ISSN 1361–8415. doi: 10.1016/j.media.2023.102789. URL https://www.sciencedirect.com/science/article/pii/S1361841523000506. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Billot Benjamin, Dey Neel, Moyer Daniel, Hoffmann Malte, Turk Esra Abaci, Gagoski Borjan, Grant P Ellen, and Golland Polina. Se (3)-equivariant and noise-invariant 3d rigid motion tracking in brain mri. IEEE transactions on medical imaging, 43(11): 4029–4040, 2024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bookheimer Susan Y, Salat David H, Terpstra Melissa, Ances Beau M, Barch Deanna M, Buckner Randy L, Burgess Gregory C, Curtiss Sandra W, Diaz-Santos Mirella, Elam Jennifer Stine, Fischl Bruce, Greve Douglas N, Hagy Hannah A, Harms Michael P, Hatch Olivia M, Hedden Trey, Hodge Cynthia, Japardi Kevin C, Kuhn Taylor P, Ly Timothy K, Smith Stephen M, Somerville Leah H, Uğurbil Kâmil, van der Kouwe Andre, Van Essen David, Woods Roger P, and Yacoub Essa. The lifespan human connectome project in aging: An overview. Neuroimage, 185:335–348, January 2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Casey BJ, Cannonier Tariq, Conley May I, Cohen Alexandra O, Barch Deanna M, Heitzeg Mary M, Soules Mary E, Teslovich Theresa, Dellarco Danielle V, Garavan Hugh, Orr Catherine A, Wager Tor D, Banich Marie T, Speer Nicole K, Sutherland Matthew T, Riedel Michael C, Dick Anthony S, Bjork James M, Thomas Kathleen M, Chaarani Bader, Mejia Margie H, Hagler Donald J Jr, Cornejo M Daniela, Sicat Chelsea S, Harms Michael P, Dosenbach Nico U F, Rosenberg Monica, Earl Eric, Bartsch Hauke, Watts Richard, Polimeni Jonathan R, Kuperman Joshua M, Fair Damien A, Dale Anders M, and ABCD Imaging Acquisition Workgroup. The adolescent brain cognitive development (ABCD) study: Imaging acquisition across 21 sites. Dev. Cogn. Neurosci, 32:43–54, August 2018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Cruz K Guadalupe, Leow Yi Ning, Le Nhat Minh, Adam Elie, Huda Rafiq, and Sur Mriganka. Cortical-subcortical interactions in goal-directed behavior. Physiol. Rev, 103(1): 347–389, January 2023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- de Jong LW, van der Hiele K, Veer IM, Houwing JJ, Westendorp RGJ, Bollen ELEM, de Bruin PW, Middelkoop HAM, van Buchem MA, and van der Grond J. Strongly reduced volumes of putamen and thalamus in alzheimer’s disease: an MRI study. Brain, 131(Pt 12):3277–3285, December 2008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Martino A Di, Yan C-G, Li Q, Denio E, Castellanos FX, Alaerts K, Anderson JS, Assaf M, Bookheimer SY, Dapretto M, Deen B, Delmonte S, Dinstein I, Ertl-Wagner B, Fair DA, Gallagher L, Kennedy DP, Keown CL, Keysers C, Lainhart JE, Lord C, Luna B, Menon V, Minshew NJ, Monk CS, Mueller S, Müller R-A, Nebel MB, Nigg JT, O’Hearn K, Pelphrey KA, Peltier SJ, Rudie JD, Sunaert S, Thioux M, Tyszka JM, Uddin LQ, Verhoeven JS, Wenderoth N, Wiggins JL, Mostofsky SH, and Milham MP. The autism brain imaging data exchange: towards a large-scale evaluation of the intrinsic brain architecture in autism. Mol. Psychiatry, 19(6):659–667, June 2014. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Estrada Santiago, Kügler David, Bahrami Emad, Xu Peng, Mousa Dilshad, Breteler Monique M.B., Aziz N. Ahmad, and Reuter Martin. Fastsurfer-hypvinn: Automated sub-segmentation of the hypothalamus and adjacent structures on high-resolutional brain mri. Imaging Neuroscience, 1:1–32, November 2023. ISSN 2837–6056. doi: 10.1162/imag_a_00034. URL http://dx.doi.org/10.1162/imag_a_00034. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fischl Bruce. FreeSurfer. Neuroimage, 62(2):774–781, August 2012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Fischl Bruce, Salat David H., Busa Evelina, Albert Marilyn, Dieterich Megan, Haselgrove Christian, van der Kouwe Andre, Killiany Ron, Kennedy David, Klaveness Shuna, Montillo Albert, Makris Nikos, Rosen Bruce, and Dale Anders M.. Whole brain segmentation: Automated labeling of neuroanatomical structures in the human brain. Neuron, 33(3): 341–355, 2002. ISSN 0896–6273. doi: 10.1016/S0896-6273(02)00569-X. URL https://www.sciencedirect.com/science/article/pii/S089662730200569X. [DOI] [PubMed] [Google Scholar]
- Fu Qiang, Xia Xinyuan, and Hong Yi. Brain mri segmentation with language-driven detection and context-aware descriptions. In ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5, 2025. doi: 10.1109/ICASSP49660.2025.10887845. [DOI] [Google Scholar]
- Geuze E, Vermetten E, and Bremner JD. MR-based in vivo hippocampal volumetrics: 1. review of methodologies currently employed. Mol. Psychiatry, 10(2):147–159, February 2005. [DOI] [PubMed] [Google Scholar]
- Grill Jean-Bastien, Strub Florian, Altché Florent, Tallec Corentin, Richemond Pierre H., Buchatskaya Elena, Doersch Carl, Pires Bernardo Avila, Guo Zhaohan Daniel, Azar Mohammad Gheshlaghi, Piot Bilal, Kavukcuoglu Koray, Munos Rémi, and Valko Michal. Bootstrap your own latent: A new approach to self-supervised learning, 2020.
- Roy Abhijit Guha, Conjeti Sailesh, Navab Nassir, and Wachinger Christian. Quicknat: A fully convolutional network for quick and accurate segmentation of neuroanatomy. NeuroImage, 186:713–727, 2019. ISSN 1053–8119. doi: 10.1016/j.neuroimage.2018.11.042. URL https://www.sciencedirect.com/science/article/pii/S1053811918321232. [DOI] [PubMed] [Google Scholar]
- Harms Michael P, Somerville Leah H, Ances Beau M, Andersson Jesper, Barch Deanna M, Bastiani Matteo, Bookheimer Susan Y, Brown Timothy B, Buckner Randy L, Burgess Gregory C, Coalson Timothy S, Chappell Michael A, Dapretto Mirella, Douaud Gwenaëlle, Fischl Bruce, Glasser Matthew F, Greve Douglas N, Hodge Cynthia, Jamison Keith W, Jbabdi Saad, Kandala Sridhar, Li Xiufeng, Mair Ross W, Mangia Silvia, Marcus Daniel, Mascali Daniele, Moeller Steen, Nichols Thomas E, Robinson Emma C, Salat David H, Smith Stephen M, Sotiropoulos Stamatios N, Terpstra Melissa, Thomas Kathleen M, Tisdall M Dylan, Ugurbil Kamil, van der Kouwe Andre, Woods Roger P, Zöllei Lilla, Van Essen David C, and Yacoub Essa. Extending the human connectome project across ages: Imaging protocols for the lifespan development and aging projects. Neuroimage, 183:972–984, December 2018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- He Kaiming, Chen Xinlei, Xie Saining, Li Yanghao, Dollár Piotr, and Girshick Ross. Masked autoencoders are scalable vision learners, 2021. URL https://arxiv.org/abs/2111.06377.
- Henschel Leonie, Conjeti Sailesh, Estrada Santiago, Diers Kersten, Fischl Bruce, and Reuter Martin. Fastsurfer - a fast and accurate deep learning based neuroimaging pipeline. NeuroImage, 219:117012, 2020. ISSN 1053–8119. doi: 10.1016/j.neuroimage.2020.117012. URL https://www.sciencedirect.com/science/article/pii/S1053811920304985. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Henschel Leonie, Kügler David, and Reuter Martin. Fastsurfervinn: Building resolutionin-dependence into deep learning segmentation methods—a solution for highres brain mri. NeuroImage, 251:118933, 2022. ISSN 1053–8119. doi: 10.1016/j.neuroimage.2022.118933. URL https://www.sciencedirect.com/science/article/pii/S1053811922000623. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Holmes Colin J, Hoge Rick, Collins Louis, Roger, Toga Arthur W, and Evans Alan C. Enhancement of MR images using registration for signal averaging. Journal of Computer Assisted Tomography, 22(2):324–333, 1998. [DOI] [PubMed] [Google Scholar]
- Howell Brittany R, Styner Martin A, Gao Wei, Yap Pew-Thian, Wang Li, Baluyot Kristine, Yacoub Essa, Chen Geng, Potts Taylor, Salzwedel Andrew, Li Gang, Gilmore John H, Piven Joseph, Smith J Keith, Shen Dinggang, Ugurbil Kamil, Zhu Hongtu, Lin Weili, and Elison Jed T. The UNC/UMN baby connectome project (BCP): An overview of the study design and protocol development. Neuroimage, 185:891–905, January 2019. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Isensee Fabian, Jaeger Paul F, Kohl Simon A A, Petersen Jens, and Maier-Hein Klaus H. nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation. Nat. Methods, 18(2):203–211, February 2021. [DOI] [PubMed] [Google Scholar]
- Jack Clifford R Jr, Bernstein Matt A, Fox Nick C, Thompson Paul, Alexander Gene, Harvey Danielle, Borowski Bret, Britson Paula J, Whitwell Jennifer L, Ward Chadwick, Dale Anders M, Felmlee Joel P, Gunter Jeffrey L, Hill Derek L G, Killiany Ron, Schuff Norbert, Fox-Bosetti Sabrina, Lin Chen, Studholme Colin, DeCarli Charles S, Krueger Gunnar, Ward Heidi A, Metzger Gregory J, Scott Katherine T, Mallozzi Richard, Blezek Daniel, Levy Joshua, Debbins Josef P, Fleisher Adam S, Albert Marilyn, Green Robert, Bartzokis George, Glover Gary, Mugler John, and Weiner Michael W. The alzheimer’s disease neuroimaging initiative (ADNI): MRI methods. J. Magn. Reson. Imaging, 27(4): 685–691, April 2008. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Jenkinson Mark, Beckmann Christian F, Behrens Timothy E J, Woolrich Mark W, and Smith Stephen M. FSL. Neuroimage, 62(2):782–790, August 2012. [DOI] [PubMed] [Google Scholar]
- Jernigan Terry L, Brown Timothy T, Hagler Donald J Jr, Akshoomoff Natacha, Bartsch Hauke, Newman Erik, Thompson Wesley K, Bloss Cinnamon S, Murray Sarah S, Schork Nicholas, Kennedy David N, Kuperman Joshua M, McCabe Connor, Chung Yoonho, Libiger Ondrej, Maddox Melanie, Casey BJ, Chang Linda, Ernst Thomas M, Frazier Jean A, Gruen Jeffrey R, Sowell Elizabeth R, Kenet Tal, Kaufmann Walter E, Mostofsky Stewart, Amaral David G, Dale Anders M, and Pediatric Imaging, Neurocognition and Genetics Study. The pediatric imaging, neurocognition, and genetics (PING) data repository. Neuroimage, 124(Pt B):1149–1154, January 2016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kennedy David N, Haselgrove Christian, Hodge Steven M, Rane Pallavi S, Makris Nikos, and Frazier Jean A. CANDIShare: a resource for pediatric neuroimaging data. Neuroinformatics, 10(3):319–322, July 2012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Kim Yeun, Joshi Anand A, Choi Soyoung, Joshi Shantanu H, Bhushan Chitresh, Varadarajan Divya, Haldar Justin P, Leahy Richard M, and Shattuck David W. BrainSuite BIDS app: Containerized workflows for MRI analysis. September 2024.
- Landman Bennett and Warfield Simon. Miccai 2012 workshop on multi-atlas labeling. In MICCAI Grand Challenge and Workshop on Multi-Atlas Labeling, Nice, France, 2012. CreateSpace Independent Publishing Platform. [Google Scholar]
- Lerch Jason P, van der Kouwe André J W, Raznahan Armin, Paus Tomáš, Johansen-Berg Heidi, Miller Karla L, Smith Stephen M, Fischl Bruce, and Sotiropoulos Stamatios N. Studying neuroanatomy using MRI. Nat. Neurosci, 20(3):314–326, February 2017. [DOI] [PubMed] [Google Scholar]
- Li Hongwei Bran, Conte Gian Marco, Anwar Syed Muhammad, Kofler Florian, Ezhov Ivan, van Leemput Koen, Piraud Marie, Diaz Maria, Cole Byrone, Calabrese Evan, Rudie Jeff, Meissen Felix, Adewole Maruf, Janas Anastasia, Kazerooni Anahita Fathi, La-Bella Dominic, Moawad Ahmed W, Farahani Keyvan, Eddy James, Bergquist Timothy, Chung Verena, Shinohara Russell Takeshi, Dako Farouk, Wiggins Walter, Reitman Zachary, Wang Chunhao, Liu Xinyang, Jiang Zhifan, Familiar Ariana, Johanson Elaine, Meier Zeke, Davatzikos Christos, Freymann John, Kirby Justin, Bilello Michel, Fathallah-Shaykh Hassan M, Wiest Roland, Kirschke Jan, Colen Rivka R, Kotrotsou Aikaterini, Lamontagne Pamela, Marcus Daniel, Milchenko Mikhail, Nazeri Arash, Weber Marc-André, Mahajan Abhishek, Mohan Suyash, Mongan John, Hess Christopher, Cha Soonmee, Javier Villanueva-Meyer Errol Colak, Crivellaro Priscila, Jakab Andras, Albrecht Jake, Anazodo Udunna, Aboian Mariam, Yu Thomas, Chung Verena, Bergquist Timothy, Eddy James, Albrecht Jake, Baid Ujjwal, Bakas Spyridon, Linguraru Marius George, Menze Bjoern, Iglesias Juan Eugenio, and Wiestler Benedikt. The brain tumor segmentation (BraTS) challenge 2023: Brain MR image synthesis for tumor segmentation (BraSyn). ArXiv, June 2023. [Google Scholar]
- Ma Jun, He Yuting, Li Feifei, Han Lin, You Chenyu, and Wang Bo. Segment anything in medical images. Nat. Commun, 15(1):654, January 2024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Marcus Daniel S, Fotenos Anthony F, Csernansky John G, Morris John C, and Buckner Randy L. Open access series of imaging studies: longitudinal MRI data in nondemented and demented older adults. J. Cogn. Neurosci, 22(12):2677–2684, December 2010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Morey Rajendra A, Petty Christopher M, Xu Yuan, Hayes Jasmeet Pannu, Wagner H Ryan 2nd, Lewis Darrell V, LaBar Kevin S, Styner Martin, and McCarthy Gregory. A comparison of automated segmentation and manual tracing for quantifying hippocampal and amygdala volumes. Neuroimage, 45(3):855–866, April 2009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nikolov Stanislav, Blackwell Sam, Zverovitch Alexei, Mendes Ruheena, Livne Michelle, De Fauw Jeffrey, Patel Yojan, Meyer Clemens, Askham Harry, Romera-Paredes Bernadino, Kelly Christopher, Karthikesalingam Alan, Chu Carlton, Carnell Dawn, Boon Cheng, D’Souza Derek, Moinuddin Syed Ali, Garie Bethany, McQuinlan Yasmin, Ireland Sarah, Hampton Kiarna, Fuller Krystle, Montgomery Hugh, Rees Geraint, Suleyman Mustafa, Back Trevor, Hughes Cían Owen, Ledsam Joseph R, and Ronneberger Olaf. Clinically applicable segmentation of head and neck anatomy for radiotherapy: Deep learning algorithm development and validation study. J. Med. Internet Res, 23(7):e26151, July 2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rutherford Saige, Fraza Charlotte, Dinga Richard, Kia Seyed Mostafa, Wolfers Thomas, Zabihi Mariam, Berthet Pierre, Worker Amanda, Verdi Serena, Andrews Derek, Han Laura KM, Bayer Johanna MM, Dazzan Paola, McGuire Phillip, Mocking Roel T, Schene Aart, Sripada Chandra, Tso Ivy F, Duval Elizabeth R, Chang Soo-Eun, Penninx Brenda WJH, Heitzeg Mary M, Burt S Alexandra, Hyde Luke W, Amaral David, Nordahl Christine Wu, Andreasssen Ole A, Westlye Lars T, Zahn Roland, Ruhe Henricus G, Beckmann Christian, and Marquand Andre F. Charting brain growth and aging at high spatial precision. eLife, 11:e72904, feb 2022. ISSN 2050–084X. doi: 10.7554/eLife.72904. URL https://doi.org/10.7554/eLife.72904. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schoemaker Dorothee, Buss Claudia, Head Kevin, Sandman Curt A., Davis Elysia P., Chakravarty M. Mallar, Gauthier Serge, and Pruessner Jens C.. Hippocampus and amygdala volumes from magnetic resonance images in children: Assessing accuracy of freesurfer and fsl against manual segmentation. NeuroImage, 129:1–14, 2016. ISSN 1053–8119. doi: 10.1016/j.neuroimage.2016.01.038. URL https://www.sciencedirect.com/science/article/pii/S1053811916000537. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Somerville Leah H, Bookheimer Susan Y, Buckner Randy L, Burgess Gregory C, Curtiss Sandra W, Dapretto Mirella, Elam Jennifer Stine, Gaffrey Michael S, Harms Michael P, Hodge Cynthia, Kandala Sridhar, Kastman Erik K, Nichols Thomas E, Schlaggar Bradley L, Smith Stephen M, Thomas Kathleen M, Yacoub Essa, Van Essen David C, and Barch Deanna M. The lifespan human connectome project in development: A large-scale study of brain connectivity development in 5–21 year olds. Neuroimage, 183:456–468, December 2018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Su Dingjie, Liu Yihao, Zuo Lianrui, and Dawant Benoit. Mesh-prompted anatomy segmentation. In Medical Imaging with Deep Learning, 2025. URL https://openreview.net/forum?id=YuHh2JioCs. [Google Scholar]
- Tae Woo Suk, Kim Sam Soo, Lee Kang Uk, Nam Eui-Cheol, and Kim Keun Woo. Validation of hippocampal volumes measured using a manual method and two automated methods (FreeSurfer and IBASPM) in chronic major depressive disorder. Neuroradiology, 50(7): 569–581, July 2008. [DOI] [PubMed] [Google Scholar]
- Tarvainen Antti and Valpola Harri. Weight-averaged consistency targets improve semi-supervised deep learning results. CoRR, abs/1703.01780, 2017. URL http://arxiv.org/abs/1703.01780. [Google Scholar]
- Yushkevich Paul A, Amaral Robert S C, Augustinack Jean C, Bender Andrew R, Bernstein Jeffrey D, Boccardi Marina, Bocchetta Martina, Burggren Alison C, Carr Valerie A, Chakravarty M Mallar, Chételat Gaël, Daugherty Ana M, Davachi Lila, Ding Song-Lin, Ekstrom Arne, Geerlings Mirjam I, Hassan Abdul, Huang Yushan, Iglesias J Eugenio, Joie Renaud La, Kerchner Geoffrey A, LaRocque Karen F, Libby Laura A, Malykhin Nikolai, Mueller Susanne G, Olsen Rosanna K, Palombo Daniela J, Parekh Mansi B, Pluta John B, Preston Alison R, Pruessner Jens C, Ranganath Charan, Raz Naftali, Schlichting Margaret L, Schoemaker Dorothee, Singh Sachi, Stark Craig E L, Suthana Nanthia, Tompary Alexa, Turowski Marta M, Van Leemput Koen, Wagner Anthony D, Wang Lei, Winterburn Julie L, Wisse Laura E M, Yassa Michael A, Zeineh Michael M, and Hippocampal Subfields Group (HSG). Quantitative comparison of 21 protocols for labeling hippocampal subfields and parahippocampal subregions in in vivo MRI: towards a harmonized segmentation protocol. Neuroimage, 111:526–541, May 2015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhang Xuzhe, Wu Yuhao, Angelini Elsa, Li Ang, Guo Jia, Rasmussen Jerod M., O’Connor Thomas G., Wadhwa Pathik D., Jackowski Andrea Parolin, Li Hai, Posner Jonathan, Laine Andrew F., and Wang Yun. Mapseg: Unified unsupervised domain adaptation for heterogeneous medical image segmentation based on 3d masked autoencoding and pseudo-labeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5851–5862, June 2024. [DOI] [PMC free article] [PubMed] [Google Scholar]
- Zhao Ziheng, Zhang Yao, Wu Chaoyi, Zhang Xiaoman, Zhou Xiao, Zhang Ya, Wang Yanfeng, and Xie Weidi. Large-vocabulary segmentation for medical images with text prompts. NPJ Digital Medicine, 8(1):566, 2025. [DOI] [PMC free article] [PubMed] [Google Scholar]
