Skip to main content
Alzheimer's & Dementia : Diagnosis, Assessment & Disease Monitoring logoLink to Alzheimer's & Dementia : Diagnosis, Assessment & Disease Monitoring
. 2026 May 25;18(2):e70345. doi: 10.1002/dad2.70345

Refining cognitive dispersion metrics for brain–behavior prediction in aging and mild cognitive impairment

Truc Tran Thanh Nguyen 1, Yu‐Ling Chang 2,3,4,5,6,✉
PMCID: PMC13239138  PMID: 42255962

Abstract

INTRODUCTION

This study addresses three key issues in the standardization of cognitive dispersion: its operationalizations though intra‐individual standard deviation (ISD) versus coefficient of variation (CoV), its reliability, and its dependence on the size of the neuropsychological battery.

METHODS

Cognitive dispersion was calculated in 318 older adults (Mage = 70.7, 61% female). Linear regression tested whether ISD, in the context of mean cognitive performance, provided greater explanatory power than CoV for brain morphometry. We further evaluated 2‐year reliability of dispersion and compared psychometric properties across four battery sizes.

RESULTS

We found that ISD, modeled jointly with mean cognitive performance, was a stronger predictor of entorhinal cortex thickness than CoV, which obscured critical mean–dispersion interactions.

DISCUSSION

These findings suggest that ISD, rather than CoV, offers a more valid quantification of cognitive dispersion, and that dispersion measures are most informative when derived from adequately comprehensive neuropsychological batteries.

Keywords: cognitive aging, cognitive dispersion, intra‐individual variability, mild cognitive impairment, neuropsychological tests

Highlights

  • Intra‐individual standard deviation (ISD) outperformed coefficient of variation (CoV) in capturing brain–behavior associations.

  • CoV obscured critical interactions between mean ability and variability.

  • Dispersion showed only moderate 2‐year reliability in older adults.

  • Larger test batteries yielded more robust and valid dispersion measures.

  • Findings inform standardization of dispersion for clinical applications.

1. BACKGROUND

Cognitive aging is a natural and dynamic process that can confer both beneficial and detrimental effects on the cognitive function of older adults. 1 To better understand the mechanisms and risk factors underlying age‐related cognitive impairment and late‐life dementia, research has increasingly focused on the variability between and within individuals. In this context, cognitive dispersion, which refers to the degree of intra‐individual variability (IIV) in performance across multiple cognitive domains or tasks at a single occasion, has emerged as a sensitive marker of early cognitive decline. While it is psychometrically normal for individuals to display strengths and weaknesses, with a few low test scores not necessarily indicating impairment, 2 pronounced difficulties in maintaining consistent performance across tasks have been associated with pathological aging. 3 , 4 , 5 , 6 , 7 , 8 , 9 , 10 , 11 , 12 , 13 Supporting this view, a large‐scale meta‐analysis found significantly greater cognitive dispersion in individuals with mild cognitive impairment (MCI) and dementia compared to cognitively normal controls, 14 and longitudinal studies have further shown that elevated dispersion predicts subsequent cognitive decline and conversion to MCI or dementia. 6 , 13 , 15 , 16 , 17 , 18 , 19

Despite these promising findings, cognitive dispersion has not yet been widely integrated into clinical practice. To address this gap, the present study focuses on three critical issues that may limit its clinical utility. 14 The first issue pertains to how dispersion is best operationalized. At present, dispersion is primarily quantified using either the intra‐individual standard deviation (ISD) or the coefficient of variation (CoV). 10 , 14 ISD represents the standard deviation of standardized scores across a neuropsychological battery, while the CoV is obtained by dividing the ISD by the individual's mean cognitive performance across all test scores. Accordingly, CoV is often considered a measure of cognitive variability adjusted for overall cognition function. While the rationale for calibrating variability for mean performance is well grounded, CoV may be problematic because it conflates ISD, mean performance, and their interaction, making it difficult to determine the source of any significant demographic, behavioral, or biomarkers associations. 20 Previous studies have handled this issue inconsistently: some included mean performance as a covariate, 3 , 5 , 7 , 8 , 13 others did not, 12 , 15 , 17 and still others relied primarily on CoV. 21 , 22 Very few, however, have explicitly tested the interaction between mean performance and ISD. Here, we argue that interpreting ISD in relation to mean performance is critical, and that relying solely on CoV risks obscuring meaningful Mean × ISD interactions, given that mean level and variability represent distinct constructs of cognition. 14

The second issue concerns whether dispersion is stable over time. Although higher dispersion at baseline has been associated with increased risk of dementia, 17 the predictive value of this finding remains uncertain because evidence regarding the stability or reliability of cognitive dispersion is limited. 23 , 24 , 25 Establishing whether dispersion exhibits a stable or variable trajectory across time is essential to determine its robustness as a predictor of future cognitive changes.

The third issue pertains to the influence of neuropsychological test composition on dispersion estimates. Different studies have employed batteries ranging from three measures across three tests 8 to as many as 30 measures from 19 tests, 26 introducing substantial heterogeneity (Figure S1). While larger test batteries may capture cognitive profiles more comprehensively, they may also narrow the range of dispersion scores 23 and increase the likelihood of obtaining at least one abnormal score, thereby inflating dispersion indices. 2 Understanding the extent to which battery composition influences dispersion is therefore critical for advancing standardization and ensuring comparability across studies and clinical settings.

Understanding these three issues is crucial for determining whether cognitive dispersion is a robust and clinically useful marker of early cognitive decline. Accordingly, in the present study we examined dispersion in a well‐characterized cohort of older adults with normal cognition (NC) and MCI, using a comprehensive neuropsychological battery to address questions of operationalization, temporal stability, and test composition.

2. METHODS

2.1. Participants

A total of 323 older adults (155 with NC and 168 with MCI) with available neuropsychological data were included in this study. Participants in the NC group were recruited from local communities, whereas those in the MCI group were recruited from the memory clinics of local hospitals. The diagnosis of MCI was based on the following criteria: (1) objective cognitive impairment, defined as performance at least one standard deviation (SD) below normative means on two or more measures within a single cognitive domain; 27 and (2) preserved activities of daily living. Within the MCI group, 61 individuals met criteria for single‐domain amnestic MCI, 82 for multiple‐domain amnestic MCI, and 25 for nonamnestic MCI. Participants were excluded if they had a history of dementia or other neuropsychiatric disorders, learning disability, head injury with loss of consciousness, substance abuse, or significant visual or auditory deficits that interfered with neuropsychological testing. The study protocol was approved by the Institutional Review Board of National Taiwan University Hospital, and written informed consent was obtained from all participants.

2.2. Neuropsychological assessment

Neuropsychological assessments were administered by licensed neuropsychologists or trained technicians in a quiet office setting. Cognitive performance was evaluated across five domains: memory, attention, executive function, visuospatial function, and language. A total of 21 measures derived from 16 standardized neuropsychological tests were included in the present analyses (Table 1). Detailed descriptions of each test, including task procedures, reliability, validity, and references, are available in our previous work. 28

TABLE 1.

Neuropsychological measures included in the present study

Cognitive domain Test Measure(s)
Memory 1. WMS‐III logical memory 1. Immediate recall (LM I)
2. Delayed recall (LM II)
2. California Verbal Learning Test, 2nd edition 3. Total learning (Trials 1–5)
4. Delayed recall
Attention & executive function 3. Color trails test 5. Part 1
6. Part 2
4. WAIS‐III digit span 7. Forward task
8. Backward task
5. WMS‐III spatial span 9. Forward task
10. Backward task
6. WAIS‐III digit symbol substitution test 11. Total correct symbols/digits
7. WAIS‐III arithmetic 12. Total scores
8. Modified card sorting test 13. Categories completed

9. D‐KEFS design fluency

14. Total correct unique designs (“switch” condition)
10. WAIS‐III similarities 15. Total correct answers
11. WAIS‐III matrix reasoning 16. Total correct answers
Visuospatial 12. WAIS‐III block design 17. Accuracy score
Language 13. Boston naming test 18. Total scores
14. WAIS‐III vocabulary 19. Total scores
15. Phonetic fluency 20. Total scores
16. Semantic fluency 21. Total animal & fruit items

Abbreviations: D‐KEFS, Delis‐Kaplan Executive Function System; WAIS‐III, Wechsler Adult Intelligence Scale, 3rd edition; WMS‐III, Wechsler Memory Scale, 3rd edition.

2.3. Calculation of cognitive dispersion

Raw neuropsychological scores were transformed into age‐ and education‐adjusted z‐scores (mean = 0, SD = 1) using regression coefficients derived from the NC group. Notably, sex‐correction was not performed as sex explained only a minimal amount of variance on all raw scores (less than 9%) and the dispersion indices (less than 0.7%; Table S1). For measures in which higher numbers indicate poorer performance (i.e., the Color Trails Test), scores were inverted so that higher values consistently reflected better performance across all measures. The z‐scores were then converted into T‐scores (mean = 50, SD = 10). Three indices were derived: (1) Overall Test Battery Mean (OTBM), the average of T‐scores across all neuropsychological measures; (2) ISD, the standard deviation of T‐scores across all measures, with higher values reflecting greater cognitive dispersion; and (3) CoV, the ISD divided by OTBM, representing dispersion relative to mean performance. Notably, five participants (all with MCI) with extreme values (ISD ≥ 3 SDs above the cohort mean) were excluded from analysis, consistent with previous studies, 5 , 29 resulting in a final analytic sample of 318 participants. The neuropsychological profile of these excluded participants is considered in the Supplementary Results (Figure S2).

RESEARCH IN CONTEXT

  1. Systematic review: The authors reviewed the literature using traditional (e.g., PubMed) sources to examine the clinical applicability of cognitive dispersion. The review revealed no consensus regarding its standardization, particularly with respect to operational definitions and neuropsychological test composition.

  2. Interpretation: Cognitive dispersion is best operationalized as intra‐individual standard deviation (ISD) rather than coefficient of variation (CoV). ISD provided a more sensitive marker of brain–behavior associations, while CoV obscured important interactions between dispersion and overall cognitive ability. The moderate 2‐year reliability of dispersion may complicate its use as a longitudinal marker of cognitive decline. Additionally, dispersion measures derived from larger neuropsychological batteries demonstrated superior psychometric properties and stronger associations with brain morphometry compared with those from smaller batteries.

  3. Future directions: Future studies should examine the temporal stability of dispersion and evaluate its utility as a marker of adverse cognitive trajectories and clinical outcomes.

2.4. Magnetic resonance imaging acquisition and processing

Participants underwent magnetic resonance imaging (MRI) scanning on a 3.0 T Siemens Magnetom Trio system (Siemens Medical Solutions, Erlangen, Germany). High‐resolution T1‐weighted images were acquired, visually inspected for quality, and processed with FreeSurfer version 6.0 (Martinos Center for Biomedical Imaging, Charlestown, MA, USA) to obtain cortical gray matter thickness and volumetric maps according to the Destrieux atlas. 30 , 31 Our analyses focused on brain regions previously identified as vulnerable to aging, including the hippocampus (volume) and cortical gray matter regions of the lateral prefrontal cortex (PFC), orbitofrontal cortex, entorhinal cortex, and parahippocampal cortex (thickness). 28 , 32 , 33 To control for potential confounding effects, age and sex were regressed out of all volumetric and thickness measures; hippocampal volume was additionally adjusted for estimated total intracranial volume (eTIV). Standardized residuals (z‐scores) were then used in all region‐of‐interest (ROI) analyses.

2.5. Statistical analyses

To address our study aims, three sets of analyses were performed. First, we examined whether ISD or CoV better explained variance in brain morphometric measures. Linear regression was used to examine associations between dispersion indices and a priori ROI measures. For each ROI, four models were constructed with the following independent variables: (1) OTBM; (2) OTBM and ISD as main effects (referred to as OTBM + ISD or O + I); (3) OTBM and ISD with an interaction term (OTBM × ISD or O × I); and (4) CoV. These models tested whether dispersion accounted for variance in ROIs beyond overall cognitive ability (O + I vs. OTBM); whether mean ability and dispersion exerted interactive rather than additive effects (O × I vs. O + I); and whether any synergistic effects outperformed mean‐adjusted dispersion (O × I vs. CoV). Critically, the O × I interaction term was included to test whether the association between dispersion and brain ROIs varies as a function of overall mean cognitive performance (i.e., moderation), such that dispersion may carry different implications at different levels of OTBM. Comparisons among nested models (OTBM vs. O + I vs. O × I) were based on adjusted R 2 and analysis of variance (ANOVA) F‐tests, whereas comparisons between non‐nested models (O × I vs. CoV) were based on the Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC). The model providing the best fit (highest adjusted R 2 or lowest AIC/BIC) was considered optimal. 34 , 35 , 36 To address multiple comparisons, we applied false discovery rate (FDR < 0.05) correction to the ROI‐based regression results. 37 We also evaluated effect sizes of independent variables in the O × I and CoV models using η 2, with thresholds of > 1% (small), > 6% (moderate), and > 14% (large). 38

Second, to examine the stability of dispersion, we analyzed data from a subset of 102 participants (50 NC, 52 MCI) who completed a follow‐up visit approximately 2 years after baseline. Dispersion was calculated as described in Section 2.3. Associations between dispersion indices at the two time points were examined using partial correlations, adjusting for the interval length.

Third, to evaluate the impact of test battery composition on dispersion, we recalculated dispersion indices using three subsets of neuropsychological measures 7 , 8 , 39 : small (three measures from three tests), medium (six measures from four tests), and large (nine measures from seven tests; Table S2). For each ISD index, we (1) tested distributional properties (Gaussian vs. non‐normal) with the Shapiro‐Wilk test, (2) examined intercorrelations, and (3) assessed associations with brain ROIs using the regression models described above.

All analyses were performed in R version 4.4.2, with statistical significance defined as P < 0.05 unless otherwise specified. Further methodological details, including MRI preprocessing and assessment for multicollinearity, are provided in the Supplementary Methods.

3. RESULTS

3.1. Participant characteristics

Table 2 presents the demographic and neuropsychological characteristics of the sample. Participants had a mean age of 70.7 ± 7.4 years (range = 50–94) and an average of 13.3 ± 3.3 years of education (range = 6–22). More than half of the participants were female (195/318, 61.3%). Compared with NC participants, those with MCI were significantly older, had fewer years of education, and demonstrated a less favorable cognitive profile, as reflected by lower OTBM and higher ISD and CoV (all ps ≤ 0.001).

TABLE 2.

Demographic, clinical, and neuropsychological characteristics of participants

Parameter

Total

(N = 318)

NC

(n = 155)

MCI

(n = 163)

p‐Value Effect size
Demographic
Age 70.70 (7.39) 69.01 (6.73) 72.31 (7.64) <0.001 −0.45
Education 13.28 (3.34) 14.15 (2.72) 12.45 (3.66) <0.001 0.52
Female, n (%) 195 (61) 99 (64) 96 (59) 0.42 0.05
Clinical
MMSE* 28.06 (1.67) 28.57 (1.36) 27.54 (1.81) <0.001 0.64
CDR‐SOB* 0.75 (1.25) 0.32 (0.57) 1.15 (1.55) <0.001 −0.70
Framingham risk score 12.93 (12.3) 10.96 (12.7) 14.81 (11.7) <0.001 −0.31
Neuropsychological
Memory
WMS‐III logical memory I (0–75) 33.69 (13.02) 39.96 (10.73) 27.72 (12.2) <0.001 0.14
CVLT‐II total learning (0–80) 42.05 (12.95) 49.72 (9.01) 34.76 (11.89) <0.001 0.27
WMS‐III logical memory II (0–50) 20.30 (10.22) 25.54 (8.25) 15.31 (9.4) <0.001 0.17
CVLT‐II delayed recall (0–16) 8.68 (4.47) 11.43 (2.62) 6.07 (4.3) <0.001 0.30
Attention
Color trails test 1 (seconds) 62.06 (31.95) 50.92 (18.09) 72.65 (38.13) <0.001 0.05
WAIS‐III digit span forward (0–9) 7.63 (1.27) 7.92 (1.13) 7.36 (1.34) 0.10 0.01
WMS‐III spatial span forward (0–9) 5.51 (1.08) 5.78 (0.99) 5.25 (1.1) 0.004 0.02
WAIS‐III DSS (0–133) 56.66 (17.62) 65.26 (14.79) 48.48 (16.18) <0.001 0.15
Executive function
Color trails test 2 (seconds) 123.31 (48.26) 101.66 (31.26) 143.90 (52.49) <0.001 0.11
WAIS‐III digit span backward (0–8) 4.92 (1.45) 5.49 (1.43) 4.38 (1.25) <0.001 0.08
WMS‐III spatial span backward (0–9) 5.10 (1.15) 5.41 (1.09) 4.80 (1.13) <0.001 0.03
WAIS‐III arithmetic (0–22) 12.36 (4.03) 14.23 (3.47) 10.59 (3.73) <0.001 0.11
MCST categories achieved (0–8) 4.64 (1.99) 5.64 (1.45) 3.70 (1.97) <0.001 0.17
D‐KEFS design fluency (0–12) 5.95 (2.42) 6.86 (1.97) 5.07 (2.49) <0.001 0.07
WAIS‐III similarities (0–33) 19.79 (5.74) 21.87 (4.97) 17.81 (5.73) <0.001 0.05
WAIS‐III matrix reasoning (0–26) 13.00 (5.46) 15.75 (4.76) 10.38 (4.76) <0.001 0.16
Visuospatial function
WAIS‐III block design (0–68) 31.45 (10.21) 35.16 (9.26) 27.93 (9.84) <0.001 0.06
Language
Boston naming test (0–30) 27.67 (3.03) 28.75 (1.53) 26.64 (3.67) <0.001 0.05
WAIS‐III vocabulary (0–66) 43.06 (10.97) 46.99 (8.61) 39.32 (11.67) <0.001 0.05
Phonetic fluency 23.51 (6.82) 25.69 (6.44) 21.43 (6.54) <0.001 0.04
Semantic fluency 29.98 (7.36) 32.86 (6.44) 27.25 (7.16) <0.001 0.07
Dispersion
Overall test battery mean 50.13 (5.47) 53.50 (3.29) 46.91 (5.19) <0.001 1.50
Intra‐individual standard deviation 8.23 (1.63) 7.78 (1.41) 8.65 (1.71) <0.001 −0.56
Coefficient of variation 0.17 (0.05) 0.15 (0.03) 0.19 (0.05) <0.001 −0.99

Abbreviations: CDR‐SOB, Clinical Dementia Rating–Sum of Boxes; CVLT‐II, California Verbal Learning Test–second edition; D‐KEFS, Delis–Kaplan Executive Function System; DSS, digit symbol substitution; MCI, mild cognitive impairment; MCST, Modified Card Sorting Test; NC, normal cognition; WAIS‐III, Wechsler Adult Intelligence Scale, 3rd Edition; WMS‐III, Wechsler Memory Scale, 3rd edition.

Data are presented as mean (standard deviation) unless otherwise noted. Higher scores indicate better function or performance, except for the Color Trails Test. Possible minimum and maximum values of neuropsychological measures are provided in parentheses where applicable. For demographic and clinical variables, P values were computed using t‐test, Wilcoxon signed‐rank test, or χ 2 test; effect sizes were calculated as Cohen's d or Cramer's V. For neuropsychological variables, P values were computed using multiple linear regression analyses adjusted for age and education; effect sizes were expressed as eta‐squared (η²). *Data missing for 111/318 (35%) participants on the MMSE and 82/318 (26%) participants on the CDR‐SOB.

3.2. Evaluating ISD and CoV in relation to brain measures

Figure 1 and Table S3 present the linear regression analyses examining associations between dispersion indices and brain measures. Several important findings emerged. First, F‐tests comparing nested models demonstrated that the O × I model provided superior explanatory power relative to the O + I model for approximately half of the ROIs assessed, including left hippocampal volume, bilateral entorhinal, left parahippocampal, and left lateral PFC regions. The O × I model also outperformed the OTBM‐alone model for these ROIs, as well as for the right lateral PFC and right orbitofrontal regions (Table S3 and Figure 1A). After FDR correction, significant associations were retained for left hippocampal volume and bilateral entorhinal regions in the O × I vs. O + I models, and for these regions as well as bilateral PFC and right orbitofrontal regions in the O × I vs. OTBM‐alone models.

FIGURE 1.

FIGURE 1

Evaluation of regression models for association between cognitive dispersion and brain ROIs. (A) Adjusted R 2 across models with four sets of independent variables: OTBM, OTBM + ISD (O + I), OTBM × ISD (O × I), and CoV. (B) Differences of AIC (red) and BIC (blue) between the CoV and OTBM × ISD models, calculated as AIC/BICCoV–AIC/BICOTBM × ISD. Negative values indicate smaller AIC/BIC for the CoV model (i.e., better fit), and vice versa. Dashed vertical lines at ± 2 demonstrate the threshold for substantial evidence in favor of a candidate model. 35 , 36 (C) Effect sizes of the independent variables in the OTBM × ISD model (first three columns, black box; OTBM:ISD indicates the interaction term) and the CoV model (last column, blue box), expressed as eta‐squared (η 2). According to Cohen, 38 η 2 > 1% is small, > 6% moderate, and > 14% large. FDR‐corrected significant P values are shown in bold, while those with P > 0.05 are shaded in light gray. (D) Interaction effect between OTBM and ISD on left entorhinal cortex thickness. Low and high ISD correspond to 1 SD below and above the mean ISD score, respectively; medium ISD corresponds to the mean. AIC, Akaike information criterion; BIC, Bayesian information criterion; CoV, coefficient of variation; FDR, false discovery rate; ISD, intra‐individual standard deviation; L, left; OTBM, overall test battery mean; PFC, prefrontal cortex; R, right; ROI, region of interest.

Second, comparisons between the O × I and CoV models, which are non‐nested, were evaluated using AIC and BIC (Figure 1B–C). ΔAIC revealed strong support for the O × I model in explaining variance in the left hippocampal and bilateral entorhinal cortices. In contrast, ΔBIC favored the CoV model for nearly all ROIs, with the exception of the left entorhinal cortex (Figure 1B). Third, the relationship between left entorhinal cortex thickness and the OTBM × ISD interaction is illustrated in Figure 1D. Conceptually, the interaction term tests whether the association between dispersion and brain morphometry depends on overall mean cognitive performance (i.e., moderation). In our data, among participants with low OTBM, lower ISD was associated with greater entorhinal thickness relative to higher ISD, suggesting that dispersion explains additional variance in entorhinal thickness beyond mean performance. By contrast, at higher OTBM levels, the dispersion‐entorhinal association was attenuated, suggesting that dispersion may be more informative closer to the lower end of the mean‐performance spectrum. Critically, this interaction effect would have been overlooked if CoV had been relied upon as the sole dispersion measure.

Finally, results for the left orbitofrontal and right hippocampal ROIs showed significant associations in the OTBM‐alone and CoV models, but not for the ISD in either the O + I or O × I models (Table S3). This suggests that CoV is primarily driven by OTBM rather than ISD, highlighting that CoV is more strongly influenced by mean performance than by variability. These findings underscore the limitations of CoV as a dispersion measure, given that it is intended to index variability rather than central tendency.

We did not find evidence of high multicollinearity in our regression models (see Supplementary Results and Figure S3). We also conducted subgroup analyses, which showed that significant associations between dispersion indices and brain ROIs were primarily observed in the MCI group, but not in the NC group (see Supplementary Results and Table S4).

Taken together, these findings demonstrate that CoV does not effectively isolate variability from central tendency, limiting its interpretability as a dispersion measure. In contrast, incorporating ISD in interaction with mean performance provides a more nuanced account of cognitive–brain associations and reveals effects that would otherwise remain undetected.

3.3. Temporal stability of cognitive dispersion indices

To examine the stability of cognitive dispersion, we analyzed data from a subset of 102 participants (50 NC, 52 MCI), who completed at least one follow‐up visit (Table S5). On average, the two visits were 16.3 ± 2.8 months apart (range = 10.5–23.9 months). Participants with and without follow‐up data had similar baseline characteristics, except for scores of the Color Trails Test, which were better among those with follow‐up data (Table S6). Figure 2 illustrates the relationships of OTBM, ISD, and CoV between baseline and follow‐up. OTBM demonstrated high reliability (R = 0.88), whereas ISD and CoV showed moderate stability over time (Rs = 0.43 and 0.56, respectively; all ps < 0.001). In addition, Fisher z‐tests revealed that participants with NC and MCI differed significantly in the test‐retest correlation for OTBM (P = 0.025) but not for ISD (p = 0.98) or CoV (p = 0.59). These results indicate that test‐retest reliability for dispersion is poorer than that of mean cognitive performance among both NC and MCI groups.

FIGURE 2.

FIGURE 2

Correlations of OTBM, ISD, and CoV between baseline and follow‐up visits. R values represent partial Pearson correlation coefficients, adjusted for the interval between visits. CoV, coefficient of variation; ISD, intra‐individual standard deviation; MCI, mild cognitive impairment; NC, normal cognition; OTBM, overall test battery mean.

3.4. Influence of neuropsychological battery size on dispersion

Among the ISD indices derived from four subsets of neuropsychological measures, only the one calculated from the largest battery approximated a Gaussian distribution (Figure 3A). OTBM values were highly consistent across batteries (Rs = 0.72–0.91, all ps < 0.001). By comparison, ISD values from batteries with fewer measures correlated only weakly with those from the other three indices (Rs = 0.15–0.71, all ps ≤ 0.005, Figure 3B). Furthermore, left entorhinal cortex thickness was significantly associated with OTBM × ISD when dispersion was estimated from the medium, large, and full batteries, but not from the small battery (Table S7). These findings suggest that a sufficient number of measures are necessary to ensure reliable estimation of dispersion effects.

FIGURE 3.

FIGURE 3

Influence of neuropsychological battery size on dispersion indices. (A) Distributions of ISD derived from small, medium, large, and full batteries, with corresponding Shapiro‐Wilk test statistics. Only the ISD from the full battery displayed an approximately Gaussian distribution. (B) Zero‐order correlations of OTBM (left) and ISD (right) calculated from different battery sizes. OTBM values showed strong correlations across all batteries (all ps < 0.001), whereas ISD values were only weakly to moderately correlated depending on battery size. All p values were < 0.001, except for the correlation between ISD–Small and ISD–Large (P = 0.005). ISD, intra‐individual standard deviation; OTBM, overall test battery mean.

4. DISCUSSION

Our study advances the characterization of cognitive dispersion by demonstrating its added value when modeled in interaction with mean cognitive performance. Rather than treating variability and central tendency as independent constructs, we show that considering ISD in the context of OTBM provides a more sensitive marker of brain–behavior relationships. In particular, the interaction of ISD and OTBM was a stronger predictor of bilateral entorhinal cortex thickness than CoV alone. Moreover, while mean cognitive performance exhibited a markedly stable test‐retest reliability profile, dispersion showed poorer reliability across both NC and MCI groups. Finally, dispersion indices derived from larger neuropsychological batteries offered clear psychometric advantages and showed robust associations with brain ROIs—associations that were not observed when smaller batteries were used.

Our work contributes to the ongoing debate on the heterogeneity of dispersion quantifications. 10 , 14 , 20 , 23 , 40 Several advantages of CoV relative to ISD are worth noting. Because ISD and mean‐level performance are expressed in the same unit, their ratio (i.e., CoV) is unitless and can be easily expressed as a percentage, which often makes it more intuitively interpretable. 23 CoV also provides a convenient method for calibrating dispersion by overall cognitive ability, which may explain why BIC statistics tended to favor CoV when compared with OTBM and ISD in modeling brain morphometry. However, our findings demonstrate that relying solely on CoV risks misinterpretation, as apparent associations may be driven primarily by mean performance rather than variability. By modeling ISD and OTBM separately, we can test whether dispersion carries incremental information beyond mean performance, and critically, whether its neurobiological relevance depends on the level of mean cognition (i.e., a moderation effect). This distinction is important because dispersion may not be uniformly informative across the full range of cognitive ability; rather, it may be most informative nearer the lower end of the mean‐performance spectrum, where variability could index greater vulnerability to dysregulated cognitive systems. In this context, a substantial Mean × ISD interaction can reveal a graded profile across the mean‐dispersion spectrum (meanhigh–dispersionlow > meanhigh–dispersionhigh > meanlow–dispersionlow > meanlow–dispersionhigh), with potential clinical utility for stratifying individuals who have similar mean performance but different dispersion‐related neurobiological profiles (Figure S4).

To our knowledge, only one prior study has systematically benchmarked different quantifications of IIV, concluding that CoV should be avoided due to its nonlinear association with mean performance in an age‐heterogenous older adults sample, which undermines its transparency as an operationalization of IIV. 20 Although that study focused on inconsistency (trial‐to‐trial variability within a task) rather than dispersion (variability across tasks at a single timepoint), its conclusion nevertheless resonates with our findings. Taken together, these results suggest that while CoV has pragmatic appeal, it may obscure important brain–behavior relationships that emerge when mean and dispersion are modeled as distinct but interacting constructs. Beyond the methodological implications, these distinctions may be clinically relevant, as more precise modeling of dispersion could sharpen our ability to stratify risk for structural brain changes and possibly dementia progression. These findings suggest that the combined assessment of mean performance and dispersion may serve as a more refined index of subtle neural compromise, particularly in medial temporal regions critically implicated in Alzheimer's disease.

We further observed that the associations between dispersion and brain ROIs were present primarily among individuals with MCI rather than those with NC. This pattern is consistent with previous reports suggesting that dispersion is more sensitive to abnormal brain morphometric or functional connectivity in higher‐risk group, such as MCI, 3 positive amyloid‐beta, 7 , 41 or cognitively unimpaired individuals with elevated subjective cognitive complaints, 12 than in healthy controls. Thus, in our sample, dispersion appeared more strongly linked to brain integrity once clinical risk was present, rather than within the NC range. Another noteworthy finding was that dispersion was more consistently associated with brain morphometry in the left hemisphere. This lateralized pattern aligns with prior studies demonstrating asymmetric cortical thinning and relative left hemisphere vulnerability in aging, Alzheimer's disease, and other neurodegenerative disorders. 28 , 42 , 43 , 44 Because regional morphometric measures are commonly averaged across hemispheres, 3 , 7 future research that model left and right hemisphere ROIs separately may help clarify whether cognitive dispersion preferentially tracks left‐hemisphere vulnerability and whether such asymmetry has implications for subsequent neurodegenerative processes.

The differential reliability observed between mean performance and dispersion highlights the potentially diminished predictive utility of dispersion when modeled alongside mean performance. 20 , 45 Additionally, this result adds nuance to prior longitudinal findings in which baseline dispersion was associated with subsequent cognitive decline or conversion to MCI or dementia. 6 , 13 , 15 , 16 , 17 , 18 , 19 Placing our findings alongside recent work provides further perspective. 24 , 25 For example, one large non‐clinical lifespan study (N = 2229, ages 33–83) reported 9‐year reliabilities of 0.75 for OTBM and 0.34 for ISD. 24 Another study in 238 individuals with MCI (ages 55–91) estimated a reliability of 0.69 for ISD across an average 1.5‐year interval (range 0.6–5.6 years). 25 Notably, in our data, dispersion showed comparable test‐retest reliability in NC and MCI participants, which complicated the common speculation that dispersion reflects primarily “random measurement error” in NC but “pathological variability” in MCI subgroups. 46 More broadly, the marked heterogeneity across studies, in terms of sample size, clinical characteristics, dispersion quantification, and evaluation intervals, suggests that reliability estimates may be highly cohort‐ and battery‐specific. Clinically, this variability cautions against using dispersion as a stand‐alone marker. Establishing a clearer stability profile, ideally with multiwave longitudinal designs, will be critical for determining whether dispersion indices can be reliably integrated into risk models for cognitive decline, or whether they should remain primarily exploratory measures.

One critical step toward the standardization of dispersion metric procedures is to determine how the psychometric attributes of individual neuropsychological tests within a battery, as well as the interplay among tests, influence dispersion measures. 14 , 23 To date, little attention has been given to the empirical significance of the number of (sub)tests in a cognitive battery. Some researchers have stressed that this issue must be addressed before more definitive clinical applications—e.g., generating norms—can be proposed, 23 while others have argued that the number of tests is less critical. 14 Our results help reconcile this contradiction: we found that larger batteries tended to yield dispersion measures that followed Gaussian distributions and showed stronger associations with brain ROIs. The former finding aligns with prior claims that “more tests administered will decrease dispersion”, 23 since extreme values from a small number of tests disproportionately inflated ISD. Similar effects have been reported elsewhere: Kiselica observed highly non‐normal ISD distributions when 12 measures (from seven tests) were used in over 4000 participants, 10 while Buchholz found normal ISD distributions when 30 scores (from 19 tests) were used in 327 individuals. 26 While our findings do not mandate that dispersion must always be computed from comprehensive batteries, they indicate that dispersion derived from very small batteries may not provide psychometrically meaningful indices. From a psychometric perspective, Gaussian remains the ideal distribution. 47 Clinically, this implies that dispersion measures based on limited test sets—such as those frequently used in screening settings—should be interpreted with caution, as they may reflect statistical artifacts rather than genuine cognitive variability.

The strengths of our study include the use of a comprehensive neuropsychological battery and rigorous analytic strategies to examine dispersion quantification. Nonetheless, several limitations warrant mentioning. First, our longitudinal analysis was based on only one follow‐up time point and baseline NC/MCI classification, wherein we did not account for diagnostic progression between visits (e.g., NC to MCI, or reversion from MCI to NC). Second, when evaluating how battery composition influenced dispersion, we focused primarily on the number of tests. Future work should consider domain composition (e.g., relative representation of memory vs. executive function tests) to assess whether certain domains contribute disproportionately to dispersion indices. Finally, while we showed that smaller batteries may yield psychometrically unstable dispersion measures, larger batteries pose challenges of their own, including increased likelihood of obtaining one or more abnormal scores, greater vulnerability to measurement error from less reliable tests, and potential fatigue effects in participants. 2 , 48 Clinically, this suggests that dispersion indices may be most informative when derived from moderately comprehensive, well‐balanced test batteries that minimize both statistical distortion and patient burden.

In conclusion, we demonstrate that ISD and CoV are not interchangeable metrics of cognitive dispersion and highlight the considerable heterogeneity introduced by different operationalizations and test battery compositions. Clarifying these distinctions is essential for building a stronger theoretical framework of dispersion and guiding its use in research and clinical contexts. Future efforts aimed at standardizing dispersion metrics will be critical to refine their interpretability and to determine their potential value as markers of neural integrity and clinical outcomes.

CONFLICT OF INTEREST STATEMENT

The authors declare no conflicts of interest. Author disclosures are available in the supporting information.

CONSENT STATEMENT

The study was approved by the Institutional Review Board of National Taiwan University Hospital. All participants provided written informed consent.

Supporting information

Supporting Information: dad270345‐sup‐0001‐SuppMat.pdf

Supporting Information: dad270345‐sup‐0001‐SuppMat.pdf

DAD2-18-e70345-s001.pdf (321.7KB, pdf)

ACKNOWLEDGMENTS

The authors acknowledge the Imaging Center for Integrated Body, Mind, and Culture Research, National Taiwan University, for MRI facility support. Part of the results from this study were presented at the Alzheimer's Association International Conference in Toronto, Canada, on July 28, 2025.

Nguyen TTT, Chang Y‐L. Refining cognitive dispersion metrics for brain–behavior prediction in aging and mild cognitive impairment. Alzheimer's Dement. 2026;18:e70345. 10.1002/dad2.70345

Funding information

This work was supported by the National Science and Technology Council, Taiwan (grant numbers 112‐2410‐H‐002‐201‐MY3, 114‐2423‐H‐002‐009, 114‐2223‐E‐002‐002, 114‐2424‐H‐002‐020‐DR). This research was also supported by the NTU‐UT Joint Research Center for LIFE_Leveraging Artificial Intelligence for Future Healthcare Excellence, National Taiwan University (grant number NTU‐ICRP‐114L7507).

REFERENCES

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supporting Information: dad270345‐sup‐0001‐SuppMat.pdf

Supporting Information: dad270345‐sup‐0001‐SuppMat.pdf

DAD2-18-e70345-s001.pdf (321.7KB, pdf)

Articles from Alzheimer's & Dementia : Diagnosis, Assessment & Disease Monitoring are provided here courtesy of Wiley

RESOURCES