Skip to main content
Nature Portfolio logoLink to Nature Portfolio
. 2026 Sep 4;2(1):85. doi: 10.1038/s44400-026-00139-y

The feasibility and factor structure of a multi-domain cognitive test battery in older Kenyan adults

Rachel W Maina 1,2,#, Kevin Kipkoech 1,#, Wambui Karanja 1,2, Cynthia Smith 1, Jasmit Shah 1, Lucy W Kamau 1, Matenyeh Kaba 3,4, Shireen Javandel 3,4, Elena Tsoy 3,4, Chinedu T Udeh-Momoh 1,4,5, Karen Blackmon 1,6,✉
PMCID: PMC13545010  PMID: 42701381

Abstract

This study evaluated the feasibility and internal factor structure of an adapted multi-domain and multi-modal (paper/digital) cognitive test battery for cognitively unimpaired multilingual Kenyan adults. The battery was translated into Swahili, back-translated, culturally adapted with community input, and refined through expert review. It was administered to 148 adults (37.26% male; mean age 57.68 years, range 45–79) using paper-based measures of learning, memory, attention, fluency, and visuoconstruction, along with tablet-based tasks assessing attention, set-shifting, symbol-matching, visuoperception, and visual associative learning. Participants found the tests relevant and acceptable, recommending language simplification, short breaks, and flexibility in language use. High completion rates and positive engagement supported feasibility. Exploratory factor analyses identified a four-factor structure explaining 61.4% of variance, reflecting executive control, rote memory, contextual memory, and visuospatial functioning with preliminary evidence for acceptable model fit, pending cross-validation. Findings indicate feasibility and provide initial evidence for construct validity based on the internal structure of the test battery. These results advance multidimensional cognitive phenotyping for brain health assessment in Sub-Saharan Africa.

Subject terms: Health care, Neuroscience, Psychology, Psychology

Introduction

Multi-domain cognitive assessments are needed for diagnosis, treatment planning, and outcome tracking of people with dementia1. Although brief screening may be the first step in identifying whether cognitive impairment is present, a comprehensive multi-domain assessment is essential for characterizing cognitive phenotypes linked to different dementia syndromes. When paired with Alzheimer’s disease biomarkers, cognitive phenotyping can determine pretest probability of amyloid positivity and indicate whether multimorbidity should be considered2. Moreover, multi-domain assessments can monitor disease progression3 and track response to interventions4. Validation of multi-domain assessments in global settings is essential for ensuring that populations around the world can benefit from available interventions, a core component of the World Health Organization (WHO) dementia research blueprint5.

Early and accurate diagnosis of cognitive impairment is critical for achieving maximum treatment benefit6. However, valid cognitive diagnostics are elusive in many low- and middle-income countries (LMICs) across Sub-Saharan Africa7. This is often attributed to lower levels of educational attainment, multilingualism, and cultural differences relative to high income countries (HICs), where most cognitive tests were developed8. When digital tests are deployed, limited familiarity with technology can also compromise their validity9, particularly in older adults10,11. Therefore, it is essential to enlist stakeholders from target populations to determine the face validity of tools and to assess the feasibility of administration procedures, prior to psychometric validation12,13.

Guidelines provided by the International Test Commission (ITC) outline a process for validating psychological assessments in novel populations that involves translation, back-translation, adaptation, and robust psychometric validation14. This process is further specified for neuropsychological tools15. Accordingly, there have been efforts to translate, adapt, validate, and harmonize cognitive assessments across Sub-Saharan Africa. One such initiative is the Harmonized Cognitive Assessment Protocol (HCAP), which has coordinated efforts to standardize the assessment of cognitive aging to facilitate cross-national comparisons and tracking of dementia trends globally16. Kenyan adaptation of the HCAP revealed unique cultural features that need to be considered when administering cognitive tests in low-resource communities, such as flexibility in language use to accommodate multilingualism, timing of testing to accommodate local socio-cultural rhythms, and modifications to items that require literacy to accommodate older adults with no formal education13. These important modifications need to be prospectively applied prior to formal data collection to ensure valid results and fair interpretation.

Modifications to administration instructions, practice items, scoring methods, and stimuli are necessary to ensure cultural appropriateness but may change the nature of the constructs being assessed17. For this reason, the resulting factor structure needs to be evaluated and compared against globally recognized neuropsychological constructs. In other words, we need to be sure we are still measuring what we think we are measuring across different cultures. For example, multilingualism is not just a concern for translation but also has direct psychometric implications. Language switching is common in multilingual settings and can introduce inhibitory control and set-shifting demands in ways that may change the nature of attention and fluency tasks18,19. Thus, it is critical that the factor structure of neuropsychological test batteries is scrutinized early in the validation process.

In addition, deployment of digital tools requires scrutiny in novel settings, particularly in older adults with varying degrees of digital literacy11. Digital tests offer promise for capturing executive constructs and have shown comparable diagnostic performances with the paper-and-pencil tests in the identification of mild cognitive impairment (MCI) and dementia10; however, they may also introduce modality variance that could influence latent factor structure in test-naïve populations. Therefore, it is important to evaluate whether factors cluster by test modality (digital versus paper–pencil) or by neuropsychological domains.

In this study, we evaluate the feasibility and internal factor structure of an adapted multi-modal and multi-domain cognitive test battery among cognitively unimpaired (CU) multilingual older adults in Nairobi, Kenya. The battery includes HCAP measures, tests from the National Alzheimer’s Coordinating Center (NACC) Uniform Data Set (UDS), as well as digital cognitive assessments designed with input from global stakeholders to capture key executive functions, such as inhibitory attentional control and cognitive set shifting20–22. We aimed to determine whether classic paper-based and novel digital tools were considered fair and feasible by local stakeholders and if the resulting test battery factor structure conforms to globally recognized neuropsychological domains. Given the multilingual nature of the Kenyan population, we hypothesized that measures of simple attention and verbal fluency would share variance with measures of cognitive set-shifting and inhibitory control.

Results

Qualitative adaptation

A focus group discussion (FGD) was carried out with eight community members, leaders, and local experts, and elicited the following feedback on the relevance, clarity, linguistic, and cultural appropriateness of the test battery. Primary themes are summarized below, with exemplary quotes.

The first extracted theme was relevance and engagement. Participants found the tests stimulating and reflective of everyday cognitive demands, noting that they “awakened the mind” and effectively measured attention, memory, and concentration. The assessments were viewed as enjoyable and informative, enhancing awareness of brain health.

  • “It was cool, and it has awakened my mind, knowing things are appropriate, seeing if I remember or others I don’t remember.” (Participant 1)

  • “Very interesting. It’s like I’ve tested my mind, that is, my brain, as it continues. I have seen that I forget quickly. But after some time when I focused on what I was doing, I saw that now my mind has returned to it.” (Participant 2)

  • “If the tool was meant to gauge the level of brains, I qualify it, because it worked.” (Participant 6)

The second extracted theme was task comprehension and tolerance. Instructions were generally clear, and participants valued demonstrations and practice trials. Repetition by examiners improved confidence and task understanding. However, reaction-time tasks and long sessions were mentally taxing, prompting adjustments in pacing and inclusion of rest breaks.

  • “The tool shows you first what you will do, before you go to the test itself and that is what was efficient for me. I could understand.” (Participant 5)

  • “I used to understand but when the timing comes, yes, timing and remembering, that’s what confuses you.” (Participant 6)

  • “I saw that the test was very long with many questions. So, if you are going to use it for a setting, you will have to summarize it. It will have to be shorter. Because it can be time consuming.” (Participant 3)

The third extracted theme was language and cultural fit. Participants emphasized that overly formal Kiswahili reduced accessibility and that rigid monolingual testing risked assessing language proficiency rather than cognition. Flexibility to respond in English or Kiswahili was strongly preferred. These insights informed simplified translations, contextually grounded phrasing, and allowance for natural code-switching, which was later quantified as a study variable.

  • “I prefer Kiswahili and, don’t bring the deepest Kiswahili, the one I can hear. But if you put the formal one, I will not be able to say anything. Let people answer in the language they know best.” (Participant 8)

  • “And the culture, it was playful…there could be a leeway, an integration of other languages which can help people be able to, rather not be tested in their knowledge or education in a particular language but rather, their brain function.” (Participant 3)

  • “Still on that brave man story, the Kiswahili is okay but now there is a place where they have used words I did not understand.” (Participant 7)

The fourth extracted theme was testing experience. Initial apprehension was common, but supportive examiner interaction promoted engagement and comfort. Rapport building was highlighted as essential for successful testing with older adults.

  • “I’m coming back as a nursery child. I didn’t know anything and again I was completely afraid, but the examiner was great, s/he helped me understand, s/he knows how to handle old people like me.” (Participant 8)

  • “I’m actually thinking it should be more and even more time, so that we do less, we go and drink tea a little, we go back and do it.” (Participant 1)

Overall, feedback from the FGD led to simpler wording, allowance for flexibility in language use and for repetition of instructions, and revised wording and pacing, to enhance the comprehensibility, acceptability, and ecological validity of the adapted battery.

Sample characteristics for quantitative analyses

To evaluate the feasibility and factor structure of the resulting neuropsychological test battery, we enrolled 155 participants; seven were excluded based on the following criteria: age < 45 years (n = 1), presence of an exclusionary medical condition (n = 1), and IDEA cognitive screen failue screen failure (n = 5), resulting in a sample of 148 older adults (Table 1). All participants endorsed fluency in the language of testing (Swahili), which was the first language in 13%, the second language in 82%, and the third language in 5%. Participants endorsed fluency in one to seven languages (median = 3), which illustrates the multilingual nature of the sample.

Table 1.

Sociodemographic information of the study sample (N = 148)

Variables
Gender, n (%) 92 females (62%)/56 males (38%)
Age in years, mean (SD, range) 56.34 (SD = 9.00, range = 45–80)
Education in years, mean (SD, range) 10.99 (SD = 3.53, range = 3–23)

Feasibility assessments of the adapted battery provided insights into test usability and implementation challenges and included evaluation of error types, language switching, practice trial success, and learning curves.

Errors were more common for digital (11.9%) than paper-based (5.4%) tests. For the digital tests, 9.2% of errors were attributable to technical factors such as application crashes, which represent modifiable challenges rather than inherent test limitations. Other errors were related to examinee (2.0%) or environmental (0.7%) factors and were comparatively infrequent. By contrast, paper-based errors were primarily examinee-related (4.7%), including poor effort, inadequate visual acuity, or difficulties with understanding tasks.

Language switching occurred frequently in tasks requiring sequencing of numbers (Digit Span Forward 97.4%; Backward 98.0%) during both instructions and responses, and in temporal sequencing tasks (Months Forward and Backward 89.5%) primarily during responses. The CERAD Word List Delayed Recall also showed high language switching rates (90.8%) during instructions only. This reiterates that allowance for language switching during instructions and responses may be necessary to preserve ecological validity in multilingual populations.

Practice trials were successful across digital tasks, with high pass rates for Flanker (100%), Match (100%), and Set Shifting (96.7%). This provides evidence of test instruction comprehensibility. Pass rates were lower for Line Orientation (85.6%), suggesting challenges in conveying geometric concepts (e.g., parallel lines).

Learning curves were apparent. CERAD Word List Immediate Recall improved by 95.4% across three trials, while TabCAT Birdwatch performance improved by 63.4% over seven trials. These positive learning curves suggest sufficient engagement with both paper-based and digital learning trials.

Exploratory factor analysis

After outlier removal (4 outliers), the Little missing completely at random (MCAR) test indicated a p-value of 0.01, showing the data were not missing completely at random. Set Shifting had the highest missingness (n = 15, 10.14%), followed by Flanker (n = 4, 2.70%) and Birdwatch (n = 3, 2.03%). For all other variables, missingness was 0%. Logistic regression results (Table 2) indicated that higher missingness on Set Shifting was significantly associated with lower education (β = −0.26, p = 0.007). We imputed missing data using K-nearest neighbors (KNN) since deletion would result with 14.86% data loss and only 126 participants would be retained for analysis. In addition, deletion of cases with data not missing at random risks selection bias towards participants with higher education. Sensitivity analysis of complete cases is provided in Supplementary Materials.

Table 2.

Demographic predictors of outcome data missingness

Variable Predictor β (SE) p-value
Set Shifting Age 0.03 (0.03) 0.339
Gender −1.08 (0.61) 0.077
Education −0.26 (0.10) 0.007

Multicollinearity check showed that reaction time variables for the Grooved Pegboard in dominant and non-dominant hands were above the inclusion threshold (r > 0.9). The remaining variables were correlated within an acceptable range for exploratory factor analysis (EFA), per Fig. 1. Normality checks indicated unacceptably high skewness in Vegetable Naming (10.04), Pegboard Dominant Hand (9.71), Pegboard Non-dominant Hand (9.28), Trails A (7.48), Months Forward (4.00), and Months Backward (3.17) while the rest had moderate to low skewness (−3.0> & <3.0). Variables with high multicollinearity or skewness were excluded from the EFA model.

Fig. 1. Correlation matrix of neuropsychological test variables.

Fig. 1

The heatmap displays pairwise Pearson correlations between all neuropsychological test variables considered for inclusion in the exploratory factor analysis (EFA). Correlation coefficients are represented on a color gradient ranging from dark blue (strong positive correlation) to dark red (strong negative correlation), with white indicating a correlation near zero. TT timed test.

Examination of the scree plot suggested an elbow after the fourth factor (Fig. 2A). Moreover, Kaiser’s rule, which supports retention of components with eigenvalues > 1, supported a four-factor solution. Parallel analysis (Fig. 2B) suggested a three-factor solution, with the fourth eigenvalue falling marginally below the threshold derived from simulated random data. We retained the four-factor solution on the basis of Kaiser’s rule, theoretical considerations, and the interpretive coherence of a distinct visuospatial factor, pending cross-validation in an independent sample.

Fig. 2. Scree plots to determine the optimal number of factors for extraction.

Fig. 2

Scree plots, eigenvalues (Kaiser’s rule), and parallel analysis were used to determine the optimal number of factors for extraction. A The scree plot displays eigenvalues (y-axis) plotted against factor number (x-axis) for all factors extracted using principal axis factoring (PAF). A clear inflection point (elbow) is observed after the fourth factor, indicating that four factors account for a meaningful and non-redundant proportion of variance in the data. This visual criterion was corroborated by Kaiser’s rule, which supports retention of factors with eigenvalues exceeding 1.0, represented by the dashed horizontal reference line. B Parallel analysis of the scree plot shows factor-analysis (FA) eigenvalues from our actual dataset (blue line) and a simulated random/resampled dataset (orange line). Although the parallel analysis supports a three-factor solution, a rule of thumb is to retain factors that are larger than those expected by chance, and the orange line shows marginal support for a fourth factor. The four-factor model was carried forward for oblimin rotation and interpretation on the basis of the scree elbow, theoretical considerations, and interpretive coherence, pending cross-validation in an independent sample.

The Kaiser–Meyer–Olkin (KMO) measure of sampling adequacy was 0.88, indicating appropriate sampling for factor analysis. The number of variables included in the EFA was 16 with a participant-to-variable ratio of 9.2:1; whereas, the participant-to-factor ratio for the four-factor solution was 37:1. The participant-to-variable ratio was marginally below the 10:1 guideline, but the mean communality was 0.791 (range 0.578–0.970), which is above the ≥0.60 threshold for stable factor solutions23,24. The overall MSA values for individual items ranged from 0.77 to 0.93, all above the acceptable threshold of 0.5. Bartlett’s test was highly significant (p < 0.001), confirming that the correlation matrix is suitable for factor analysis.

A 4-Factor structure explained 61.4% of the total variance (PA1: 21.2% of variance; PA3: 18.1% of variance; PA2: 13.8% of variance; and PA4: 8.3%). Figure 3 shows factor loadings of this 4-factor model. Factor 1 loadings included measures of attention, executive function, visual associative memory, and language from TabCAT and paper-based tests. Factor 3 loadings included visual and verbal rote memory tests. Factor 2 loadings included story-based contextual memory tests. Factor 4 loadings included measures of visuoconstruction and visual recognition memory. Model fit was acceptable (χ² = 75.037, df = 80, p = 0.64; RMSEA = 0; CFI = 1.000; TLI = 1.018; SRMR = 0.054). Non-significant chi-square (p = 0.636) indicated that the model adequately reproduced observed correlations. Factor correlations indicated related but distinct cognitive domains per Table 3. The model showed acceptable internal consistency (McDonald’s omega) across all 4 factors (Table 3). A higher number of languages spoken was associated with higher PA3 (r = 0.19; p = 0.02) but no other factor scores. There was no difference between those who were tested in their first, second, or third language on PA1 [F(2,144) = 1.34, p = 0.27], PA2 [F(2,144) = 0.78, p = 0.46], PA3 [F(2,144) = 2.37, p = 0.10], or PA4 [F(2,144) = 0.98, p = 0.38] factor scores, although small subgroup sizes limit comparison.

Fig. 3. Factor loading diagram from the four-factor exploratory factor analysis.

Fig. 3

The figure displays standardized factor loadings for each neuropsychological test variable on the four retained factors following principal axis factoring with oblimin rotation. Factor 1 (PA1; executive control) includes measures of attention, executive function, visual associative learning, and semantic fluency from both paper-based and tablet-based modalities. Factor 2 (PA2; contextual memory) includes story-based episodic memory tests. Factor 3 (PA3; rote memory) includes verbal and visual item-based memory tests. Factor 4 (PA4; visuospatial functioning) includes measures of visuoconstruction and visual recognition memory. The four-factor model explained 61.4% of total variance (PA1: 21.2%; PA3: 18.1%; PA2: 13.8%; PA4: 8.3%). A loading threshold of 0.30 was used to define meaningful factor membership, indicated by the arrows. PA principal axis factor.

Table 3.

Factor correlations and internal consistency

PA1 PA3 PA2 PA4
PA1 1.00 0.61 0.51 0.53
PA3 0.61 1.00 0.55 0.53
PA2 0.51 0.55 1.00 0.48
PA4 0.53 0.53 0.48 1.00
McDonald’s omega 0.952 0.978 0.944 0.721

Discussion

We found that a multi-modal (paper-based and tablet-based) and multi-domain cognitive test battery was adaptable, feasible, and aligned with expected neuropsychological domains (memory, visuospatial, and executive functions) in older Kenyan adults. By integrating qualitative and quantitative methods using a staged approach, we derived several practical and scientific insights to inform cognitive testing in a multilingual Sub-Saharan African context.

One practical insight was that code-switching must be considered when implementing cognitive testing in a multilingual setting like Nairobi. In an FGD, participants emphasized that tasks requiring the sequencing of numbers prompt code-switching between English and Kiswahili. This raises concerns for reduced ecological validity if language switching is constrained. In addition, participants raised concerns about the complexity of instructions and the long duration of test sessions. By simplifying instructions, enhancing practice tests, and allowing for multiple breaks and flexible language use, we achieved high rates of test engagement and completion, exceeding 90% across nearly all measures. Lower completion rates on digital tests were mostly accounted for by technical errors, representing modifiable barriers that can be mitigated by offline delivery options and quick remediation of software bugs. Errors on paper-based tasks were mostly examinee-related, with sensory limitations (e.g., vision/hearing loss) as a primary source of error. This suggests the need to provide vision/hearing support (e.g., reading glasses or hearing amplification) in settings where access to regular eye/ear exams may be limited.

Another practical insight was the importance of considering both formal education and digital literacy when digital tasks are included in a test battery. Although overall test incompletion was rare and largely attributable to technical errors, incompletion of the digital Set Shifting task was significantly predicted by lower educational attainment (Table 2). This pattern constitutes preliminary evidence of education-related measurement bias. When task non-completion is systematically associated with a participant characteristic, scores on that measure may not adequately represent the true range of functioning in that subgroup. This warrants careful consideration in future validation work, particularly in samples with greater educational heterogeneity. A step-down option, such as a paper-based executive function measure that does not require prior digital familiarity, could be considered for participants with limited technology exposure or lower educational attainment. In addition, a measure of digital literacy should be implemented in future studies to evaluate the impact on performance across digital measures.

Practice trial success was high across digital tasks, indicating that most participants readily understood task requirements. This reinforces the feasibility of deploying digital assessments at scale with minimal onboarding burden, even among older adults. Positive learning curves on a word list learning task and a tablet-based visual associative memory task suggest good participant engagement across both paper-based and digital measures. This is encouraging, given that the tablet-based measures we utilized incorporate alternate versions to minimize practice effects across repeat testing, which allows for longitudinal tracking of cognitive trajectories in older Kenyan adults.

A key scientific insight that emerged from exploratory factor analysis is that the internal factor structure of the test battery is broadly consistent with globally recognized neuropsychological domains of executive functions, episodic memory, and visuospatial functions, although further confirmatory factor analysis is needed. Both paper-based and digital tools shared latent factors with high loadings, which aligns with prior studies10 and supports the use of mixed modalities to measure a single cognitive construct. Despite alterations made during the adaptation phase, the tools showed interpretable loadings on the hypothesized domains.

As expected in a multilingual setting, we found that language, attention, and executive functioning tests did not differentially load on distinct latent factors. This may be due to the inherent “executive control” demands of code-switching, which were common on digit span and verbal fluency tasks. Allowance for language switching is recommended, as it more closely approximates everyday language use in Nairobi and reduces inhibitory control demands18,19. However, the act of switching between languages can introduce switch costs, similar to what is observed when shifting cognitive set25,26. This could have contributed to presumed language tasks (e.g., Animal Fluency) and simple attention tasks (e.g., Digit Span Forward) sharing latent factor loadings with executive control measures (e.g., Digit Span Backward and Set Shifting).

An additional finding that emerged from factor loadings was of two distinct latent factors underlying different measures of episodic memory, which is consistent with research in animals and humans showing the non-unitary and dissociable nature of episodic memory27. Contextualized (relational, narrative) memory and item-based (rote) memory rely on overlapping but partially dissociable medial temporal cortical systems28,29. Narrative encoding is linked to default mode network (DMN) activation30,31, likely because the emotional and relational content of incoming narrative material is processed relative to stored autobiographical memories. Narrative recall tasks place heavy demands on relational binding, recruiting posterior hippocampal, parahippocampal, retrosplenial, posterior cingulate/precuneus, and medial prefrontal cortical circuitry that encodes, organizes, and retrieves contextualized events29.

In contrast, rote learning and recall rely more on item-centric representations within the anterior temporal, perirhinal, amygdala, temporal pole, lateral orbitofrontal, and dorsofrontal system32. Inclusion of the complex figure recall task on this rote memory factor suggests that the absence of narrative scaffolding for the abstract design configuration likely recruits an overlapping item-centric network, extending into right parietal regions33. Thus, our EFA solution is aligned with evidence for two partially distinct large-scale memory networks: hippocampal–posterior medial DMN regions for contextual/narrative recall and anterior temporal–prefrontal regions for item/rote recall that differentially support the episodic memory clusters we observed in our CU older adult cohort.

Finally, results from EFA suggest that the Benson complex figure test copy trial and recognition memory trial load on a single latent factor, which most likely reflects visuoperceptual and visuoconstruction abilities. Reduced performance on these measures has been demonstrated in people with frontotemporal and Alzheimer’s type dementia with distinct, syndrome-specific neuroanatomical substrates of performance34. Moreover, complex figure test performance distinguishes people with atypical Alzheimer’s disease variant (i.e., posterior cortical atrophy syndrome) from people with typical amnestic Alzheimer’s variant, even a few years after symptom onset35. The presence of a latent visuospatial factor in our multi-domain test battery may strengthen its potential for differential cognitive phenotyping of common and uncommon dementia syndromes, pending cross-validation in independent samples.

Findings from EFA should be considered preliminary. Analyses were conducted in cognitively unimpaired healthy controls. There is a possibility that the same factor structure may not be replicated at the population level or in clinical samples where variability is introduced by differing degrees of health and cognitive functioning. However, the confounding effects of disease are minimized in our approach, which allows for comparison to factor structures derived in cognitively unimpaired control cohorts across global settings. Future studies will expand the current results to testing measurement invariance across people with and without clinically diagnosed dementia syndromes.

The near-perfect fit indices obtained in the primary EFA should be interpreted cautiously, as they may partly reflect the influence of data pre-processing steps on the inter-variable correlation matrix. Although we included a sensitivity analysis that demonstrates robustness of the factor structure to imputation (Supplementary Materials), cross-validation in an independent sample is needed to further validate the stability of the factor solution. Relatedly, parallel analysis supported a three-factor solution, whereas the scree elbow, Kaiser’s rule, theoretical considerations, and interpretive coherence of the visuospatial factor supported retention of four factors. Cross-validation in an independent sample should formally compare three- and four-factor solutions to determine the most parsimonious and generalizable structure. Additional work is also needed to evaluate the convergent and divergent validity of factors against external measures of structural and functional brain health.

The study’s strengths include its mixed-methods and phased approach, which leverages insights gained from qualitative and quantitative analyses to inform interpretation of the latent factor structure. This design ensured that tools were not only psychometrically sound but also socially and linguistically grounded. Inclusion of an objective feasibility assessment is another strength, with results informing practical approaches to future dementia research implementation in limited-resource and multilingual settings for improved participant engagement and measurement validity. Future implementation should evaluate whether similar outcomes are achievable with delivery by community health workers, as our reliance on bachelor’s-level examiners introduces a resource constraint to scalability.

In conclusion, neuropsychological test adaptation in limited-resource and multilingual settings benefits from utilizing a comprehensive, mixed-method approach and inclusion of relevant community stakeholders in decision-making. With appropriate adaptation, hybrid paper-based and digital-based cognitive test batteries are feasible in older adults and may be a scalable solution for multi-domain cognitive phenotyping in sub-Saharan Africa. Similarities between the factors we derived and those observed across neuropsychological test implementation in global settings lay the foundation for harmonization with international datasets. Our results provide a methodological and analytical framework for addressing the WHO dementia research blueprint Strategic Goal 6, which calls for advancing research and development of clinical assessments that are relevant across diverse populations.

Methods

Study design and setting

This mixed-methods study was conducted at the Aga Khan University, Brain and Mind Institute (BMI) in Nairobi, Kenya. The design incorporated a community-based participatory research (CBPR) approach to qualitatively adapt neurocognitive assessment tools, followed by quantitative feasibility assessment, and preliminary evaluation of structural validity using exploratory factor analysis in a community-based sample.

Ethics approval and consent

Ethical approval was obtained from the Aga Khan University Institutional Scientific and Ethics Review Committee (2023-ISERC-0419-2) and the National Commission for Science, Technology and Innovation (NACOSTI; License No. NACOSTI/P/24/32915), in accordance with the Declaration of Helsinki. Written informed consent was obtained from all participants, with additional consent for audio-recording and anonymous note-taking of focus groups.

Participants

In collaboration with community partner organizations and awareness campaigns on social media, we recruited community-dwelling older adults across Nairobi County, representing diverse socioeconomic strata, literacy levels, and linguistic backgrounds.

For initial qualitative evaluation to inform tool adaptation, focus group participants were recruited who were ≥45 years, fluent in Kiswahili or English, and included community leaders from local dementia foundations, older adults, and local experts (e.g., psychologists, nurses, social scientists). Exclusion criteria included inability to consent or cognitive impairment severe enough to preclude meaningful participation.

For quantitative analyses, CU participants were also community-dwelling adults aged ≥45 years. Inclusion criteria were fluency in Kiswahili or English, stable health for ≥3 months, preserved functional ability (walking, bilateral hand use), and capacity to provide consent. Exclusion criteria included diagnosed dementia, major neurological or psychiatric disorders, significant head injury, substance abuse, self-reported infections (HIV, hepatitis C, tuberculosis, syphilis, COVID-19), neoplastic disease with chemotherapy/radiation, or sensory/motor impairments that would interfere with testing.

Transcultural adaptation

Adaptation of the cognitive measures followed ITC guidelines14 through a staged process that is described in more detail in Supplementary Materials. In brief, forward–backward translation into Kiswahili was performed by bilingual social science experts, and discrepancies were resolved by expert consensus. Test instructions, stimuli, and procedures were evaluated by community stakeholders using van Ommeren et al.36 criteria for Comprehensibility, Acceptability, Relevance, and Technical Equivalence.

Feedback from a community-based focus group discussion (FGD) guided iterative refinements in wording, instructions, and test flow. Issues were systematically logged by facilitators and used to optimize clarity and enhance the cultural relevance of the test battery. An examiner manual was developed with standardized prompts, explanations of the test purpose, and strategies for managing fatigue or participant discomfort. A summary table of major linguistic, cultural, and procedural adaptations is provided in Supplementary Materials (S1).

Measures

Cognitive screening was conducted using the Identification of Dementia in Elderly Africans (IDEA) tool to establish participants as CU, using the established cut-off score of <1037. The adapted test battery included both paper-based and digital (iPad-delivered) assessments spanning key neuropsychological domains (Table 4).

Table 4.

Multi-modal multi-domain neuropsychological test battery

Test Subtests Description
Consortium to Establish A Registry for Alzheimer’s Disease (CERAD) 10-word List Learning Test List learning immediate This test measures rote verbal learning/encoding and episodic memory. We used a version adapted for use in Kenya as part of the HCAP/LOSHAK study13, with the exception of adding a 20-item recognition trial (10 targets, 10 foils) after the delayed word recall trial to facilitate differentiation of consolidation-based from retrieval-based memory processes.
List learning delayed
List learning recognition
Story recall/brave man Braveman immediate This test measures contextual, schema-supported, verbal episodic memory. We utilized a version that aligns with the HCAP/LOSHAK procedures and stimuli, which adapted a brief narrative from the East Boston Memory Test to include more contextually relevant content, while preserving general length and complexity13.
Braveman delayed
Categorical verbal fluency test Animal fluency This test measures language and executive functions, specifically lexical access, semantic memory, and executive control. The Animal Fluency test has been adapted for use in Kenya13, and we applied the same administration procedures, including the allowance of different languages at the response phase, for both trials of the test.
Vegetable fluency
Digit span test Digit span forward This test measures simple and complex attention. We utilized a version that is aligned with the NACC UDS version 3.045.
Digit span backward
Benson complex figure test Benson immediate This test measures visuoconstruction and visual memory. We utilized procedures and stimuli that are aligned with the UDS version 3.045.
Benson delayed
Benson recognition
Grooved pegboard Pegboard non-dominant hand TTa This test measures fine motor speed, manual dexterity, and hand–eye coordination. We utilized the original version of this test that involves placing 25 pegs into small holes on a board46.
Pegboard dominant hand TTa
Months forward and backward test Months forward TTa This test measures concentration and attention. It involves the participant rapidly verbalizing the 12 months of the year chronologically, first from January to December, then from December to January.
Months backward TTa
Trail making test Trails A TTa This test measures attention and processing speed. Trial A of this test has 25 circles numbered 1–25. It involves the participant drawing connecting lines between these circles in ascending order.
Tablet-based cognitive assessment tool [TabCAT] Birdwatch This is an iPad-based cognitive assessment platform developed at the University of California, San Francisco (UCSF) (https://tabcathealth.com/platform). For the purpose of our study, we utilized 5 tests assessing visual associative memory (Birdwatch), executive (Set Shifting, Flanker, Match), and visuospatial (Line Orientation) functions44,47.
Set Shifting
Flanker
Match
Line Test

a TT time taken (s)

Assessments were conducted in quiet, distraction-free testing rooms in a university clinic setting. All examiners had at least a bachelor’s level of education and received standardized training, supported by a detailed manual outlining administration procedures, standardized prompts, and strategies for managing participant fatigue or confusion. At the conclusion of training, examiners needed to pass two supervised practice sessions and a quality assurance checklist. Ongoing quality control procedures included periodic supervision and review of administration practices to ensure adherence to protocol.

Each TabCAT task contained embedded onboarding instructions and practice trials. Examiners guided participants through practice trials to ensure that instructions and response formats were clearly understood. If misunderstandings occurred during the practice phase, examiners provided clarification using standardized prompts outlined in the examiner manual and repeated the practice trials as many times as needed. Once participants demonstrated understanding of the task requirements, the scored trials proceeded without further assistance. These procedures were implemented to minimize the influence of unfamiliarity with touchscreen devices on task performance.

Study visits lasted approximately 150 min: 30 min for consent and interview, 90 min for cognitive testing, and 30 min for questionnaires. Core demographic, medical, and functional variables were recorded alongside cognitive test scores and validity indicators.

Data analyses

Qualitative data from FGDs were transcribed, translated as needed, anonymized, and coded by independent researchers. Deductive thematic analysis was guided by the van Ommeren et al.36 domains and informed iterative adaptations to the test battery.

Quantitative data were analyzed in RStudio (version 2024.12.0)38. Descriptive statistics summarized participant characteristics and feasibility outcomes. To assess feasibility, we analyzed error types (technical, examiner, environmental, examinee-related), task completion rate (including timeouts), occurrence of language switching (during instructions, responses or both), practice trial success rates, and learning curves (performance change across repeated trials).

Data cleaning for exploratory factor analyses (EFA) involved evaluation of outliers, missing data, multicollinearity, and skewness. For outlier detection, cognitive variables were regressed on age and evaluation proceeded in three steps: (1) visual inspection of scatter plots to identify points whose removal substantially altered effect sizes, (2) examination of standardized residuals to detect extreme values, and (3) identification of data points with z-scores exceeding <−4 and >4. Case-by-case review by research assistants flagged values likely reflecting administration errors. Data points deemed impossible or exhibiting poor model fit according to these criteria were excluded.

Missing data patterns were evaluated using Little’s missing completely at random (MCAR) test and logistic regression identified variables associated with missingness. Missing values were then handled using k-nearest neighbor (KNN) imputation, which preserves multivariate relationships by estimating values based on participants with similar characteristics. This approach minimized bias and ensured complete data for EFA. Sensitivity analysis of complete cases was conducted to evaluate the robustness of the model structure to imputation.

Variables falling beyond multicollinearity and skewness thresholds (inter-item r > 90 and skewness beyond <−3 and >3, respectively) were excluded from the EFA model as high skewness and multicollinearity can obscure relationships between variables. High skewness on timed metrics was considered in the context of cultural validity concerns around the speed-accuracy trade-off in Sub-Saharan African settings39–43. Qualitative research in South Africa shows prioritization of accuracy over speed; adults describe difficulty working both quickly and accurately on tests they are performing for the first time44. In quantitative normative modeling of neuropsychological data from healthy South African controls, slower performance on speed tasks was interpreted as an indicator of bias and reduced construct validity40, which is consistent with broader critiques of speed-based metrics in cross-cultural neuropsychological assessment41–43. When included in an EFA model that is predominantly comprised of accuracy-based metrics, processing speed measures may form artifactual factors that reflect the speed-accuracy distinction rather than meaningful cognitive dimensions, obscuring the latent structure of the accuracy-based constructs and reducing the interpretability of the solution.

To determine the optimal factor solution, we examined scree plots, eigenvalues (Kaiser’s rule), and parallel analysis. EFA was conducted using principal axis factoring (PAF) with oblimin rotation. We then checked for internal consistency of the factor structure using McDonald’s omega.

Supplementary information

44400_2026_139_MOESM1_ESM.pdf (48.2KB, pdf)

MainaR_Feasibility_FactorStructure_NPSY_Kenya_R1_Supplement.

Acknowledgements

All research participants and organizations that facilitated recruitment, including Futbol Mas and Nivishe Institute. Also recognize the invaluable support of Monica Kinyanjui, Zul Merali, Sam Musembi, Yasmin Aboyo, Agnes Kariuki, Prof. Anthony Ngugi, Dr. Roselytar Rianga, Catherine Bikeri, Sarah McDonagh, and Lucy Thaithi. Research reported in this publication was supported by the National Institute on Aging of the National Institutes of Health under Award Number UG3AG090679 and the Global Brain Health Institute at the University of California, San Francisco (UFRA-424|CA-0241758). The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health or Global Brain Health Institute.

Author contributions

Conceptualization and study design: C.U. and K.B. Funding acquisition: C.U. and K.B. Project management: W.K. Data collection: K.K., C.S., and W.K. Study implementation: R.M., W.K., K.K., C.S., C.U., J.S., L.K., E.T., and K.B. Manuscript writing: R.M., W.K., K.K., L.K., S.J., and K.B. Data cleaning and analysis: R.M., J.S., L.K., E.T., and K.B. Manuscript review: R.M., W.K., K.K., C.S., C.U., J.S., L.K., S.J., E.T., M.K., and K.B. Approval of final manuscript and consent for publication: R.M., W.K., K.K., C.S., C.U., J.S., L.K., S.J., E.T, M.K., and K.B.

Data availability

The datasets generated during the current study are not publicly available due to restrictions imposed by Kenyan data protection legislation, specifically the Kenya Data Protection Act (2019), and the terms of the data transfer agreements between Aga Khan University and the Global Brain Health Institute, which require that all data sharing, including deidentified data, comply with agreed protocols and authorization procedures, but are available from the corresponding author on reasonable request and with the appropriate agreements in place.

Code availability

The code that supports the findings of this study is available upon reasonable request.

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

These authors contributed equally: Rachel W. Maina, Kevin Kipkoech.

Supplementary information

The online version contains supplementary material available at https://doi.org/10.1038/s44400-026-00139-y.

References

  • 1.Alzola, P. et al. Neuropsychological assessment for early detection and diagnosis of dementia: current knowledge and new insights. J. Clin. Med.13, 3442 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 2.Bouteloup, V. et al. Cognitive phenotyping and interpretation of Alzheimer blood biomarkers. JAMA Neurol.82, 506–515 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Jutten, R. J. et al. A composite measure of cognitive and functional progression in Alzheimer’s disease: design of the Capturing Changes in Cognition study. Alzheimers Dement. (NY)3, 130–138 (2017). [DOI] [PMC free article] [PubMed]
  • 4.Ngandu, T. et al. A 2 year multidomain intervention of diet, exercise, cognitive training, and vascular risk monitoring versus control to prevent cognitive decline in at-risk elderly people (FINGER): a randomised controlled trial. Lancet385, 2255–2263 (2015). [DOI] [PubMed] [Google Scholar]
  • 5.World Health Organization. A Blueprint for Dementia Research (World Health Organization, 2022).
  • 6.Rasmussen, J. & Langerman, H. Alzheimer’s disease – why we need early diagnosis. Degener. Neurol. Neuromuscul. Dis.9, 123–130 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Nightingale, S. et al. Cognitive impairment in people living with HIV: consensus recommendations for a new approach. Nat. Rev. Neurol.19, 424–433 (2023). [DOI] [PubMed] [Google Scholar]
  • 8.Nyamayaro, P., Chibanda, D., Robbins, R. N., Hakim, J. & Gouse, H. Assessment of neurocognitive deficits in people living with HIV in sub-Saharan Africa: a systematic review. Clin. Neuropsychol.33, 1–26 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 9.Maina, E. N. Language and cultural heritage, a Kenyan perspective. Aedon 97–108 10.7390/106343 (2022). [DOI]
  • 10.Chan, J. Y. C., Yau, S. T. Y., Kwok, T. C. Y. & Tsoi, K. K. F. Diagnostic performance of digital cognitive tests for the identification of MCI and dementia: a systematic review. Ageing Res. Rev.72, 101506 (2021). [DOI] [PubMed] [Google Scholar]
  • 11.Gwala, N. & Mawela, T. Investigating the internet skills of older adults in South Africa. Artha J. Soc. Sci.23, 49–78 (2024). [Google Scholar]
  • 12.Daga, M., Mohanty, P., Krishna, R. & Swarna Priya, R. M. Statistical validation in cultural adaptations of cognitive tests: a multi-regional systematic review. In Challenges and Opportunities in Artificial Intelligence (ed. Mishra, S., Singh, A. K. & Prajapati, P.) Ch. 3 (CRC Press, 2024). 10.1201/9781003781196 (2025). [DOI]
  • 13.Riang’a, R. M. et al. Contextualization of Harmonized Cognitive Assessment Protocol (HCAP) in an aging population in rural low-resource settings in Africa: experiences and strategies adopted to optimize effective adaption of cognitive tests in Kenya. Alzheimers Dement.21, e70552 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 14.ITC Guidelines for Translating and Adapting Tests, 2nd edn. Int. J. Test.18, 101–134 (2018).
  • 15.Nguyen, C. M. et al. Neuropsychological application of the International Test Commission guidelines for translation and adapting of tests. J. Int. Neuropsychol. Soc.30, 621–634 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 16.Gross, A. L. et al. Harmonisation of later-life cognitive function across national contexts: results from the Harmonized Cognitive Assessment Protocols. Lancet Healthy Longev.4, e573–e583 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Maina, R. et al. Psychometric evaluation of the computerized battery for neuropsychological evaluation of children (BENCI) among school aged children in the context of HIV in an urban Kenyan setting. BMC Psychiatry23, 373 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 18.Prior, A. & Gollan, T. H. Good language-switchers are good task-switchers: evidence from Spanish-English and Mandarin-English bilinguals. J. Int. Neuropsychol. Soc.17, 682–691 (2011). [DOI] [PubMed] [Google Scholar]
  • 19.Declerck, M., Grainger, J., Koch, I. & Philipp, A. M. Is language control just a form of executive control? Evidence for overlapping processes in language switching and task switching. J. Mem. Lang.95, 138–145 (2017). [Google Scholar]
  • 20.Tsoy, E. et al. Global perspectives on brief cognitive assessments for dementia diagnosis. J. Alzheimers Dis.82, 1001–1013 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Muñoz-Najar, A. et al. Peruvian validation and standardization of the TabCAT-brain health assessment. Front. Public Health13, 1600131 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Ogbuagu, C. et al. Cultural adaptation of the brain health assessment for early detection of cognitive impairment in Southeast Nigeria. Front. Dement.3, 1423957 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 23.MacCallum, R. C., Widaman, K. F., Zhang, S. & Hong, S. Sample size in factor analysis. Psychol. Methods4, 84–99 (1999). [Google Scholar]
  • 24.Costello, A. B. & Osborne, J. Best practices in exploratory factor analysis: four recommendations for getting the most from your analysis. Pract. Assess. Res. Eval.10, 7 (2005). [Google Scholar]
  • 25.Gollan, T. H., Montoya, R. I. & Werner, G. A. Semantic and letter fluency in Spanish-English bilinguals. Neuropsychology16, 562–576 (2002). [PubMed] [Google Scholar]
  • 26.Rosselli, M. et al. A cross-linguistic comparison of verbal fluency tests. Int. J. Neurosci.112, 759–776 (2002). [DOI] [PubMed] [Google Scholar]
  • 27.Ranganath, C. & Ritchey, M. Two cortical systems for memory-guided behaviour. Nat. Rev. Neurosci.13, 713–726 (2012). [DOI] [PubMed] [Google Scholar]
  • 28.Eichenbaum, H., Yonelinas, A. P. & Ranganath, C. The medial temporal lobe and recognition memory. Annu. Rev. Neurosci.30, 123–152 (2007). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Davachi, L. Item, context and relational episodic encoding in humans. Curr. Opin. Neurobiol.16, 693–700 (2006). [DOI] [PubMed] [Google Scholar]
  • 30.Song, H., Park, B.-Y., Park, H. & Shim, W. M. Cognitive and neural state dynamics of narrative comprehension. J. Neurosci.41, 8972–8990 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 31.Azarias, F. R., Almeida, G. H. D. R., de Melo, L. F., Rici, R. E. G. & Maria, D. A. The journey of the default mode network: development, function, and impact on mental health. Biology (Basel)14, 395 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 32.Montaldi, D. & Mayes, A. R. The role of recollection and familiarity in the functional differentiation of the medial temporal lobes. Hippocampus20, 1291–1314 (2010). [DOI] [PubMed] [Google Scholar]
  • 33.Mock, N. et al. Nonverbal memory tests revisited: neuroanatomical correlates and differential influence of biasing cognitive functions. Cortex164, 63–76 (2023). [DOI] [PubMed] [Google Scholar]
  • 34.Possin, K. L., Laluz, V. R., Alcantar, O. Z., Miller, B. L. & Kramer, J. H. Distinct neuroanatomical substrates and cognitive mechanisms of figure copy performance in Alzheimer’s disease and behavioral variant frontotemporal dementia. Neuropsychologia49, 43–48 (2011). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 35.Charles, R. F. & Hillis, A. E. Posterior cortical atrophy: clinical presentation and cognitive deficits compared to Alzheimer’s disease. Behav. Neurol.16, 15–23 (2005). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 36.van Ommeren, M. et al. Preparing instruments for transcultural research: use of the Translation Monitoring Form with Nepali-speaking Bhutanese refugees. Transcult. Psychiatry36, 285–301 (1999). [Google Scholar]
  • 37.Gray, W. K. et al. Population normative data for three cognitive screening tools for older adults in sub-Saharan Africa. Dement. Neuropsychol.15, 339–349 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 38.R. Core Team. R: A Language and Environment for Statistical Computing (R Foundation for Statistical Computing, 2024).
  • 39.Aghvinian, M. et al. Taking the test: a qualitative analysis of cultural and contextual factors impacting neuropsychological assessment of Xhosa-speaking South Africans. Arch. Clin. Neuropsychol.36, 976–980 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 40.Ferrett, H. L. et al. The cross-cultural utility of foreign- and locally-derived normative data for three WHO-endorsed neuropsychological tests for South African adolescents. Metab. Brain Dis.29, 395–408 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 41.Nell, V. Cross-Cultural Neuropsychological Assessment: Theory and Practice (Erlbaum, 2000).
  • 42.Ardila, A. Cultural values underlying psychometric cognitive testing. Neuropsychol. Rev.15, 185–195 (2005). [DOI] [PubMed] [Google Scholar]
  • 43.Grieve, K. W. & van Eeden, R. A preliminary investigation of the suitability of the WAIS-III for Afrikaans-speaking South Africans. SAJP40, 262–271 (2010). [Google Scholar]
  • 44.Sanderson-Cimino, M. et al. Development and validation of the TabCAT-EXAMINER: a tablet-based executive functioning battery for research and clinical trials. J. Int. Neuropsychol. Soc.31, 242–253 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 45.Weintraub, S. et al. Version 3 of the Alzheimer Disease Centers’ neuropsychological test battery in the Uniform Data Set (UDS). Alzheimer Dis. Assoc. Disord.32, 10–17 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 46.Russell, E. W. & Starkey, R. I. Halstead Russell Neuropsychological Evaluation System (HRNES): Manual (Western Psychological Services, 1993).
  • 47.Tsoy, E. et al. BHA-CS: a novel cognitive composite for Alzheimer’s disease and related disorders. Alzheimers Dement. (Amst.)12, e12042 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

44400_2026_139_MOESM1_ESM.pdf (48.2KB, pdf)

MainaR_Feasibility_FactorStructure_NPSY_Kenya_R1_Supplement.

Data Availability Statement

The datasets generated during the current study are not publicly available due to restrictions imposed by Kenyan data protection legislation, specifically the Kenya Data Protection Act (2019), and the terms of the data transfer agreements between Aga Khan University and the Global Brain Health Institute, which require that all data sharing, including deidentified data, comply with agreed protocols and authorization procedures, but are available from the corresponding author on reasonable request and with the appropriate agreements in place.

The code that supports the findings of this study is available upon reasonable request.


Articles from Npj Dementia are provided here courtesy of Nature Publishing Group

RESOURCES