Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Sep 12.
Published before final editing as: Psychol Assess. 2026 Sep 10:10.1037/pas0001501. doi: 10.1037/pas0001501

Computerized Adaptive Testing of Neurocognitive Deficits in Psychosis

Michael L Thomas 1, Emily T Sturm 1, Maiele Mignard 1, Anastasia G Sares 1, John R Duffy 1,2, Jazmin M Diaz 1, Don Rojas 1, Kristina T Legget 3,4, Jason R Tregellas 3,4, Jared W Young 5,6
PMCID: PMC13565237  NIHMSID: NIHMS2198862  PMID: 42721333

Abstract

Computerized adaptive testing (CAT) is a technology that has been used to improve the precision and efficiency of individual difference estimates for cognitive, achievement, personality, and clinical symptom assessments. Motivated by the psychometric limitations of experimental cognitive tasks, we determined whether experimental cognitive tasks, such as those being developed for the National Institute of Mental Health’s Research Domain Criteria (RDoC) initiative and similar projects, could also benefit from CAT. Participants were 64 patients diagnosed with psychosis spectrum disorders and 50 healthy controls who completed at least one of six experimental cognitive tasks: a probabilistic learning task (PLT), a continuous performance task (CPT), a Flanker task, a Sternberg task, a self-ordered pointing task (SOPT), and an N-back task. Results indicated strong correlations between ability estimates for fixed and CAT versions of most tasks, except the PLT. CAT significantly reduced standard error for the CPT, Sternberg, and SOPT, but not for the N-back, Flanker, or PLT. The tasks benefiting most from CAT were those with more fine-grained levels of difficulty. These findings suggest that adaptive testing can improve the psychometric properties of experimental cognitive tasks relevant to the RDoC framework. This improvement depends on the granularity of conditions and difficulty levels, however, highlighting an important consideration for integrating CAT into cognitive psychopathology research.

Keywords: adaptive testing, item response theory, cognitive-psychometrics, psychosis, research domain criteria (RDoC)


Standardized cognitive outcome measures are essential in clinical trials (Clarke, 2007). They have established psychometric properties, yield scores that can be compared to normative data, and facilitate harmonization and meta-analyses. In contrast, ad hoc and non-standardized outcome measures are interpreted without quantitative reference to prior samples and cannot be easily combined for aggregate analyses. Psychometric limitations of non-standardized measures also heighten threats of measurement confounds that can slow the pace of discovery in mental health research (Chapman & Chapman, 1973). Despite this, neuropsychological and cognitive neuroscience studies frequently create and employ measures that lack standardization and psychometric refinement. These experimental measures—while often tailored to specific research questions—can compromise findings due to poor reliability. Even conceptually well-established cognitive tasks can produce highly unreliable measures of individual and group differences (Hedge et al., 2018; Thomas et al., 2017). This confound is especially concerning in psychiatric research, where precise and consistent tools are crucial for advancing prevention and treatment.

The National Institute of Mental Health (NIMH) aims to use experimental cognitive tasks as outcomes measures as part of the Research Domain Criteria (RDoC) initiative (National Advisory Mental Health Council Workgroup on Tasks and Measures for Research Domain Criteria (RDoC), 2016). This approach includes developing and refining translational and experimental tasks that are designed to measure specific cognitive processes irrespective of disease. Experimental cognitive tasks have strong construct representation (e.g., Bhakta & Young, 2017; Whitely, 1983) and can be used to identify deficits in cognitive processes linked to well-defined brain systems (Chudasama & Robbins, 2006; D’Esposito, 2007; Goodglass & Kaplan, 1979). However, many of these tasks are altered for use in each study and therefore have unknown psychometric properties due to lack of consistency. A report by the National Advisory Mental Health Council Workgroup on Tasks and Measures noted that several tasks being considered for the RDoC either had poor or insufficient evidence of their psychometric properties. This is especially true in neuroimaging studies, as the tasks must be adapted to scanner limitations.

A promising approach to standardize and psychometrically evaluate experimental cognitive tasks is to use item response theory (IRT) to develop computerized adaptive tests (CATs) (Moore et al., 2023; Thomas et al., 2021). CATs can be administered on any compatible device (e.g., computer, tablet, or smartphone) that supports the software. CATs use participants’ item responses in real time to estimate their latent ability and the standard error of those estimates, and then use this information to select the next items (from a pre-calibrated item bank) that are most likely to yield the most precise estimates (Segall, 2009), essentially narrowing the task to the most sensitive difficulty window. CATs not only improve assessment efficiency, but they can also ensure that measurement error stays below a predefined threshold for each examinee. This threshold typically corresponds to a minimum acceptable reliability. In other words, CATs ensure that disparate administrations of test stimuli (e.g., different item sets, conditions, etc.) are optimally tailored to each participant, thereby maximizing psychometric precision. Moreover, because stimuli prosperities are scaled from a common measurement model, scores obtained from different versions or administrations of a task can be directly compared (i.e., they are already equated and harmonized on the same measurement scale).

CATs are used commonly in educational achievement assessment (Van Wijk et al., 2024; Yang et al., 2022), but have seen only limited applications in neurocognitive assessment (Buckley et al., 2017). Interestingly, while adaptive testing has been commonly used in psychophysical research (Leek, 2001), the methods have historically been distinct from classic IRT approaches and are not commonly incorporated into neurocognitive assessments. One reason CATs are not more widely used in neuropsychology and cognitive neuroscience research is that they typically require extensive psychometric work with large samples to estimate item parameters and calibrate the instruments; in fact, samples of 1,000 or more participants are common (Martin & Lazendic, 2018). While the time and cost requirements for CATs may be feasible for for-profit test developers or for tests intended for mass administration, it is unlikely that sufficient time and funding will be commonly available for the development of CATs designed for experimental cognitive neuroscience research. To address this limitation, Thomas et al. (2021) outlined an approach to scoring cognitive tasks and developing CATs that does not rely on large calibration samples. The key idea is to adapt experimental task conditions rather than items. Thus, instead of calibrating dozens or even hundreds of test items, the approach calibrates manipulable characteristics of cognitive tasks that increase or decrease difficulty. Examples of manipulable characteristics include memory load, stimulus presentation time, and difficulty of probabilistic feedback. Thomas et al. (2021) previously demonstrated the approach using archival data from a commonly used working-memory measure, the N-back task, collected from a diverse sample. Compared to nonadaptive testing, the simulation demonstrated the CAT version made the task 36% more efficient or produced 23% more precise score estimates.

We propose uniting CAT technology with RDoC-recommended tasks, converting experimental cognitive tasks into psychometrically rigorous outcome measures; improving the precision and reliability while preserving the strong construct representation of the original experimental paradigms. Until now, the proposed modeling and CAT approach has only been evaluated in simulation, not in a live study. Therefore, the present study evaluated whether adaptive manipulation of task parameters improves measurement of cognitive ability in clinically relevant samples, including individuals with psychosis spectrum disorders (PSDs) and healthy controls (HCs). Specifically, we aimed to determine whether adaptive testing reduces the standard error (SE) of cognitive score estimates, yielding more accurate and reliable measurements. We also aimed to determine whether fixed and adaptive versions of tasks produce comparable estimates and show similar correlations with criterion variables (e.g., diagnosis, external measures of ability and achievement, and symptom ratings).

Methods

Participants

Data were analyzed from 64 PSDs and 50 HCs. We recruited male and female adults aged 18-70 years with adequate hearing and eyesight who were fluent in English and able to perform the tasks required. Exclusion criteria were: inability to understand or give consent, positive drug toxicology screen (except marijuana) or signs of intoxication at time of study visit, refusal to provide drug toxicology screen, pregnancy, previous significant head injury, significant extrapyramidal symptoms or tardive dyskinesia, and significant medical or neurological diagnoses (except self-reported diabetes). Controls were additionally excluded if they were enrolled in special education courses during school or met diagnostic criteria for any psychosis spectrum disorder or bipolar disorder. Clinical diagnoses were verified using a structured clinical interview. Additional information about diagnostic inclusions and exclusions is provided in the Supplement. Written consent was obtained from all participants. Research procedures were reviewed and approved by the Colorado Multiple Institutional Review Board.

Transparency and Openness

This study’s design and its analysis were not preregistered. Data used in this analysis are publicly available through the NIMH Data Archive (Thomas, 2020). Code used for analyses is available from the corresponding author upon request.

Design

Study personnel administered a clinical interview and all behavioral assessments in a private testing office. At the first visit, participants provided informed consent, including completion of the University of California San Diego Brief Assessment of Capacity to Consent (UBACC; Jeste et al., 2007), and administered a urine drug screen. Membership in the PSD group was confirmed using the DSM-5 Mini-International Neuropsychiatric Interview for Psychotic Disorders (MINI; Sheehan et al., 1998), version 7.0.2. The first visit also included a cognitive assessment using the MATRICS Consensus Cognitive Battery (MCCB; Nuechterlein et al., 2008), followed by approximately one hour of computerized cognitive tasks. The second visit required participants to complete whichever version of the task they had not yet completed. The second visit occurred, on average, 44 days after the first visit. Task administration was pseudorandomized, creating multiple administration schedules to balance task order and whether the fixed or adaptive version was given on visit 1 or visit 2 for each task. Participants were assigned sequentially to one of these schedules once they confirmed attendance for visit 1.

Experimental Tasks

The six tasks included in this battery are outlined in detail in a previous publication (Duffy et al., 2024). Below, we provide a concise description focusing on task characteristics that are relevant to the current study (i.e., fixed vs. adaptive testing). For the fixed versions of each task, difficulty levels were selected to span the same approximate range of difficulty as the adaptive version while also providing approximately even coverage across that range and not being too hard or too easy for most participants. Thus, the fixed tasks were designed to represent realistic, broadly applicable versions of each task rather than versions optimized for the specific ability distribution of the present sample. This approach allowed us to compare CAT against fixed-format tasks that sampled the same approximate lower and upper difficulty bounds, while acknowledging that alternative fixed versions with different difficulty levels could perform differently depending on the population being assessed. Different measurement models were used across tasks to reflect differences in task structure and response demands. All tasks were scored in a manner that is common in the literature. For tasks in which correct responses could occur through guessing, guessing was modeled as an item-level parameter. For tasks that required participants to distinguish target from non-target or signal from noise trials, signal detection parameters were used to separate discrimination from response bias. The one learning task (see below) was scored using a separate learning process model because performance depends on learning from feedback rather than on guessing or signal detection. A summary of all task design choices is repotted in Supplemental Table 1.

Probabilistic Learning Task (PLT).

The PLT was selected as a measure of the probabilistic and reinforcement learning subconstruct within the RDoC reward learning construct and has a rodent counterpart (Cavanagh et al., 2021). The task began with a learning block in which participants saw a pair of white geometric shapes (e.g., a square and a circle) on a black background and had to choose the shape that was “Correct” (i.e., rewarded) more often. In each trial, left/right positions were randomized. Participants were instructed to use the left and right arrows on the keyboard to choose a shape and to select the one that was “Correct” more often. Shapes were displayed for 3 seconds, during which a decision had to be made (or else the response was scored as incorrect). Participants received feedback after each trial indicating whether the shape they selected was “Correct” or “Incorrect.” The test segment was similar to the learning block, but consisted of 20 trials for each of 4 blocks, each block with different pairs of shapes. The test segment lasted 6 minutes and 56 seconds. Each shape had a certain probability of producing “Correct,” and the probabilities for a pair of shapes were complementary (e.g., in the annulus/triangle pair, annulus was correct on 70% of trials and triangle was correct 30% of trials). The probabilities were 60/40, 70/30, 80/20, 90/10 for the fixed version of the task. The CAT version adaptively selected from 55/45, 60/40, 65/35, 70/30, 75/25, 80/20, 85/15, 90/10, and 95/05 probabilities between blocks in a manner that optimized measurement precision for each participant. Scoring was based on a modified IRT two-parameter logistic model (2PL) where the discrimination parameters were fixed and coded following to win-stay, lose-shift (WSLS) scoring strategy. A win-stay response refers to a trial in which the participant was rewarded on the previous trial (i.e., they received the feedback “Correct”) and then selects the same shape in the current trial. A lose-shift response refers to a trial wherein the participant was not rewarded on the previous trial and then selects the opposite shape in the current trial. These are both considered as indicators of learning. In the model, discrimination parameters are defined dynamically based on the subject’s previous response; item responses that involved win-stay or lose-shift responses were positively weighted and item responses that involved win-shift or lose-stay responses were negatively weighted. Thus, the ability parameter in the model indicated an examinees preference towards win-stay, lose-shift responding.

Continuous Performance Test (CPT).

The 4-Choice Continuous Performance Test (4C-CPT), a variant of the 5-Choice Continuous Performance Test (5C-CPT; Young et al., 2009), that has predictive validity (MacQueen et al., 2018) and was found to have the most translatable cross-species paradigm with good utility (Salmon et al., 2024), was selected as a measure of cognitive control. Within the RDoC framework, CPTs are defined as measures of goal selection, updating, representation, and active maintenance subconstructs under the cognitive control construct. Our 4C-CPT variant provides compatibility with common MRI scanner response box configurations (i.e., 4 buttons). During the task, participants saw a white circle appear above one of four white lines on a black background, and were instructed to press a corresponding button quickly, or withhold their response if four circles appeared at once. Participants were first shown how to respond to the circles appearing above the lines, which were organized from left to right in an arch. Participants rested their hands on four adjacent keys (left-handed: a, s, d, f or right-handed: j, k, l, ;) which were paired with each of the lines on the semi-circle (first, second, third, or fourth). When a circle appeared above a line, the participants pressed the corresponding key. For example, if the circle appeared above the first line, a right-handed participant would press the j key. If all four circles appeared simultaneously, they were instructed to not press any key. A backward mask, consisting of four white squares, was presented after the circle(s) to ensure that participants could not determine the circle location from afterimages—shown to reveal deficits in people living with HIV (Brody et al., 2025) while deficits were not observed in standard stimuli (Pocuca et al., 2020); this challenge enabled manipulation of task difficulty. The testing segment lasted 5 minutes and 1 second and included 16 test blocks of 8 trials each. The backward-mask onset times were 50 ms, 150 ms, 250 ms, and 350 ms for the fixed version of the task. The CAT version adaptively selected backward-mask onset times ranging from 50 ms to 400 ms, in 10-ms increments between blocks, to optimize measurement precision for each participant. Each set of stimuli was presented for a total of 1 second. The original version of the task used a narrower range of difficulty levels: 50 ms, 100 ms, 150 ms, and 200 ms. However, early feedback suggested that even the 250-ms onset was too difficult for some participants, especially patients. We therefore extended the difficulty range after the first 19 participants had completed testing. These first 19 participants were omitted from the analyses. Scoring was based on a signal-detection weighted IRT 2PL model (Thomas et al., 2018) where the discrimination parameters were fixed to produce measures of discriminability (d′) and bias (C).

Flanker Task.

The Flanker task was selected as a measure of the effect of noise (presence of flanker stimuli) on target identification when no visual search is necessary (Eriksson et al., 2015). Within the RDoC framework, the Flanker is defined as a measure of the cognitive control construct further categorized into the response selection subconstruct, and also has a rodent version (Robble et al., 2021). Participants used the arrow keys on a keyboard to indicate the orientation of an arrow in the center of the screen. For each trial, participants were presented with one white target image (center arrow) and 6 white distracter images (either white dots or off-center arrows) in a horizontal line on a black background. The test phase lasted approximately 5 minutes and included 20 blocks containing 6 trials each. Our task employed a parametric manipulation of congruence consisting of the following levels: congruent (the center arrow is shown and all distractors match the center arrow), neutral (only the center arrow was shown [along with 6 distractor dots]), incongruence (the center arrow is pointed in a different direction than the distractor arrows), and partial congruence (the center arrow is pointed in one direction and distractor arrows are a mix of the same and the other direction). The fixed version of the task included all four congruency levels (with an equal number of items per level). The CAT version adaptively selected from each of these conditions between blocks (with multiple possible runs of each congruency type to choose from). The flanker was scored using an IRT three-parameter logistic (3PL) model where the discrimination parameters were all fixed to one. Guessing was fixed to 0.5 to account for the forced choice left vs. right response option.

Sternberg Task.

The Sternberg Task (also known as the Sternberg Item Recognition Paradigm) is a well-established paradigm that was selected as a measure of working memory (Sternberg, 1975). Although many variations of the original task exist, a key element is the manipulation of the number of items to be remembered (i.e., the set size or load). Within the RDoC framework, the Sternberg Task is defined as a measure of the working memory construct and is associated with two subconstructs: active maintenance and interference control. Participants viewed a learning screen where various white upper-case letters were displayed on a 4x4 grid on a black background. After a maintenance period, the next screen presented a single letter (probe), and participants indicated if the probe was a letter they had seen within the previous array. Participants were instructed to memorize the presented array (encoding phase), hold the letters in memory (maintenance phase) and then, upon seeing the probe, press the right arrow key if it was a letter they had seen or the left arrow key if it was a letter they had not seen (retrieval phase). The test phase lasted approximately 5 minutes and 41 seconds and included 24 trials. In the fixed version of the task, memory load ranged between even values of 2 and 12 letters with an equal number of trials (4) per load level. In the CAT version, memory load was adaptively chosen from these load levels between items to optimize measurement precision. Scoring was based on a modified signal-detection weighted IRT 2PL model (Thomas et al., 2018) where the discrimination parameters were fixed to produce measures of discriminability (d′) and bias (C).

Self-Ordered Pointing Task (SOPT).

The SOPT was selected as a measure of visuospatial working memory (Petrides & Milner, 1982). Within the RDoC framework, the SOPT is defined as a measure of the working memory construct and the active maintenance and interference control subconstructs. During each block, participants used a computer mouse or trackpad to select a shape from an array of white abstract images on a black background. After each selection, the same images were rearranged on the screen, and the participant was asked to choose a new image. Participants were instructed to select each image only once, and attempted as many selections as the number of images in the array. The test phase lasted approximately three to eight minutes (since image selection was self-paced) and included eight randomized blocks. The fixed version included memory loads of three, six, nine, and twelve, with two runs of each load level. The CAT version adaptively selected load between items to optimize measurement precision. The SOPT was scored using a 3PL model where the discrimination parameters were all fixed to one. In this model, a single latent ability drives performance. A unique aspect of this model is that the SOPT guessing probability is based on the current load of the task (e.g., with three items guessing is fixed to 1/3, with six items guessing is fixed to 1/6, etc.).

N-Back Task.

The N-back task was developed with the goal of measuring working memory in older adult populations (Kirchner, 1958). Within the RDoC framework, the N-Back task is defined as a measure of the working memory construct and is associated with the subconstruct of interference control. Participants watched 3-letter pseudowords appear in sequence on the screen and used the spacebar to indicate if they had seen the current word a certain number of words ago (either 1, 2, 3, or 4 words back, depending on the block). Participants began with a demonstration in which 3-letter pseudowords (e.g., CAC, NOX, GUX, VIV) were presented on the screen one at a time. The words appeared as white text centered on a black background. Participants were instructed on the definition of 1-back, 2-back, 3-back, and 4-back conditions. For each condition, participants had to press the spacebar when a pseudoword was repeated after 1, 2, 3 or 4 words, and they practiced each condition with feedback on their performance. For example, in the 1-back condition, if ‘GUX’ was followed immediately by another ‘GUX’, participants would press the spacebar. In the 2-back condition, if ‘GUX’ was followed by ‘VIV’ and then ‘GUX’ reappeared, participants would only press the spacebar when ‘GUX’ reappeared. The test phase lasted approximately 5 minutes and 41 seconds and included 4 blocks of 25 items. The fixed condition included one block of each of the four n-back conditions (1-, 2-, 3-, and 4-back) in a randomized order. In the CAT version, memory load was adaptively chosen from these load levels between blocks to optimize measurement precision. Scoring was based on a modified signal-detection weighted IRT model (Thomas et al., 2018) where the discrimination parameters were fixed to produce measures of discriminability (d′) and bias (C).

Adaptive Testing

Tasks were programmed in PsychoPy (Peirce, 2009). Adaptive administration of task conditions was implemented using the cogirt package for R (Thomas, 2024) in combination with custom Python code. Specifically, the cog_cat function was used to select task conditions (not items; see above) in a manner that would optimize measurement precision for the abilities of interest. The adaptive function applied a D-optimal approach (Segall, 2009). To interface between PsychoPy and R, a custom Python code snippet was added to PsychoPy. This code launched R from within Python using the os module and the Rscript terminal command. Item responses were passed to R via CSV files, which were then processed by the cog_cat function to select the appropriate next task conditions. The number of trials for adaptive testing was restricted to match the number of trials used in the fixed version of each task to facilitate an equivalent comparison between versions. The exact number of conditions and items administered (see description above) was chosen primarily based on feasibility, with the goal of keeping each test version to approximately 6 minutes in duration.

Data Analysis

Data were analyzed using the cogirt package for R (Thomas, 2024). The cogirt package provides IRT-based scoring for cognitive measures, such as those used in the present study. The logic and statistical strategy for this approach were described in detail and validated using simulation by Thomas et al. (2021).

Briefly, the package is designed to leverage the structure and scoring approaches common in neurocognitive testing. In these tests, both within and between conditions, items are expected to be highly similar (e.g., three-letter pseudowords in the N-back task); item difficulties are therefore treated as random effects. However, experimental manipulations are applied across different item sets (or even the same sets). In the present application, the CPT manipulates backward mask onset time; the Flanker manipulates congruency level; the N-Back, Sternberg, and SOPT manipulate memory load; and the PLT manipulates reward probability. In such tasks, the levels of these experimental factors, rather than item characteristics, are the primary determinants of test properties (e.g., difficulty). Therefore, adaptive testing is optimized by selecting levels or conditions of the experimental manipulation that maximize measurement precision, rather than by selecting items based on psychometric characteristics alone.

This methodology is operationalized in the modeling approach through a structural model for the task manipulation. Specifically, contrast codes from R’s native contr.code function are used to quantify the effects of the experimental manipulation. The contrast coding strategy and number of contrasts for each task were based on a preliminary study using fixed online versions of the tasks (Duffy et al., 2024). In that study, manipulations of task difficulty generally produced non-linear changes in performance, and therefore polynomial contrasts (contr.poly), including linear and quadratic terms, were used to capture changes in performance as a function of task difficulty.

Parameter estimation in cogirt uses the Metropolis-Hastings Robbins-Monro (MHRM) algorithm, a partially Bayesian estimation strategy that allows the user to provide mean and variance priors for model parameters to identify the measurement scale and ensure stable convergence. Priors were based on the preliminary study (Duffy et al., 2024). All other package defaults were used for estimation. The cogirt package was used to estimate latent ability/trait parameters for each participant on each task. The package also produces standard error (SE) values for each estimated person parameter.

Model fit was evaluated using Yen’s Q3 and posterior predictive checking. Yen’s Q3 reflects the residual correlation between item pairs after model estimation, with values close to zero indicating minimal local dependence. Historically, absolute Q3 values greater than .2 or .3 have been used as heuristic cutoffs to flag potentially problematic residual correlations (Christensen et al., 2017). However, residual correlations can be unstable in sparse data matrices such as those observed in the current study, leading to artificially extreme values. Accordingly, we used posterior predictive checking to approximate the sampling distribution of Q3 under the fitted model and to compute posterior predictive p-values (PPPs) for the observed residual correlations, using a nominal decision threshold of .05. It should be noted that the goal of this study, unlike many applications of IRT, was not to remove bad items or identify an optimal measurement model for each task, but rather to evaluate whether measurement of parameters commonly used to score these tasks could be improved through adaptive administration. Once the study began, scoring models were fixed and were not modified to ensure that all participants’ data were directly comparable. Scoring models were also not altered after data collection, as post hoc modification would not provide a valid test of the effectiveness of real-world CAT implementation.

Fixed testing and CAT were compared across several metrics. Pearson correlations were used to compare estimates from the fixed and CAT versions of each task to assess whether the two methodologies produced comparable results. Paired-samples t-tests were used to compare the standard error (SE) of the individual score estimates from each version, testing the hypothesis that CAT would yield more precise estimates. Paired-samples t-tests were also used to compare participants’ pleasantness ratings (on a scale of 1 - very unpleasant to 7 - very pleasant) for each task to determine whether CAT versions were perceived less favorably. Pleasantness questions were only administered beginning with participant 26; consequently, these data are based on a smaller subsample. Linear mixed-effects models fitted using the R lme4 package (Bates et al., 2015) were used to determine whether group moderated the effect of task type on SE. Finally, Pearson correlations were used to examine group differences in criterion-related validity groups, including group (diagnosis), symptoms, and cognitive functioning. Because these analyses were exploratory and intended to describe patterns of association rather than test a prespecified family of confirmatory hypotheses, we did not apply familywise-error corrections.

Results

Demographic and clinical characteristics of the sample are reported in Table 1. Data were collected from 118 participants but due to positive urine screening after some measures had been collected (n = 2) and review of MINI determining diagnosis fit exclusion parameters (n= 2), 114 participants were used for analysis (64 PSDs and 50 HCs). Of the 114 participants used for analysis, 70 (61%) participants completed both versions of every task. In the cases where all tasks were not administered, the most common cause was technical problems with at least one version of any task (n = 34) and in other cases was caused by failure to complete the second day of testing (n = 10). Thus, the vast majority of missing data were missing completely at random (i.e., technical problems). In terms of specific tasks, 83% completed both versions of the N-Back task, 83% completed both versions of the Sternberg task, 81% completed both versions of the SOPT, 84% completed both versions of the CPT, 82% completed both versions of the Flanker task, and 86% completed both versions of the PLT.

Table 1.

Demographic and Clinical Characteristics

Healthy Control Psychosis Spectrum Disorder Effect Size
N 50 64
Age (M) 43.04 41.22 −0.13
Sex (% Male) 58% 75% 0.18
Years of Education (M) 16.36 13.75 −1.35†††
Mother’s Years of Education (M) 14.39 14.02 −0.13
WRAT Reading Standard Score (M) 106.51 99.63 −0.56††
MCCB Average T-Score (M) 51.60 43.45 −1.24†††
SAPS (M) 0.14 1.75 1.91†††
SANS (M) 0.42 1.79 1.76†††
Race 0.12†
  White 72% 72%
  More than One Race 12% 13%
  Black 6% 8%
  Not Reported 4% 3%
  American Indian/Alaskan Native 2% 2%
  Asian 4% 2%
  Native Hawaiian/Pacific Islander 0% 2%
Hispanic/Latino 18% 25% 0.08

Note: Effect-size differences between patients and controls are reported as Cohen’s d for continuous variables and Cramer’s V for discrete variables (sex, Race, Hispanic/Latino). Magnitudes are classified as negligible, small, medium, or large using conventional cutoffs: d (0.20†, 0.50††, 0.80†††) and V (0.10†, 0.30††, 0.50†††). Positive d indicates patients > controls. WRAT = Wide Range of Achievement Test; MCCB = MATRICS Consensus Cognitive Battery; SAPS = Scale for the Assessment of Positive Symptoms; SANS = Scale for the Assessment of Negative Symptoms

In Table 1, we report demographic and clinical characteristics of both groups. PSDs were similar to HCs in age and mother’s (and father’s) years of education. However, PSDs were generally less highly educated and more likely to be male (both small effects). As expected, PSDs were more likely to be prescribed antipsychotic medications, had higher levels of both positive and negative symptoms, and had worse average MCCB T-scores by approximately a full standard deviation (all large effects). Discrepancies in WRAT reading scores were only about half of a standard deviation on the scaled score (medium effect). Supplementary Table 2 reports additional information including demographic and clinical characteristics of the sample by group and task, the percentage of valid administrations by group and task, and mean performance by group and task.

In terms of model fit, the median absolute Q3 value was below .10 for all tasks except the Sternberg (.105) and CPT (.143), suggesting that, for the majority of items, violations of local independence were minor. Posterior predictive model checking PPPs, however, suggested a mixed story. PPPs for the N-Back (0.080) and Flanker (0.097) tasks were close to the nominal .05 criterion, showing only minor evidence is misfit. PPPs for the PLT (0.101), Sternberg (0.141), and SOPT (0.164) fell in a more moderate range of misfit, whereas the CPT exhibited the poorest fit (.236). The elevated PPP for the CPT and other tasks likely reflect a combination of high task performance and sparse contingency tables—known challenges for residual-based model fit assessment under near-ceiling item performance—rather than substantive model misspecification. In fact, average PPPs for the CPT were correlated with task performance (r = .44), consistent with this interpretation. Additional information and discussion of model fit is provided in the Discussion section and in Supplemental Material.

Although adaptive testing is tailored to individual response profiles, we generally expect PSDs to receive less difficult conditions than HCs due to poorer task performance. For the adaptive PLT, PSDs showed descriptively but not significantly lower ability estimates (t(94.59) = 1.37, p = .17, d = 0.28) and were, on average, administered only slightly less difficult probabilistic reward levels (HCs = 70.94% vs. PSDs = 70.40%). For the adaptive CPT, PSDs showed significantly lower ability estimates (t(55.76) = 2.49, p = .016, d = 0.59) and were administered a less difficult (slower) average backward mask onset time (HCs = 105 ms vs. PSDs = 124 ms). For the adaptive Flanker task, PSDs showed significantly lower ability estimates (t(80.32) = 3.15, p = .002, d = 0.62). The distribution of administered Flanker conditions was broadly similar across groups. Compared to the PSD group, the HC group received a slightly higher proportion of partially incongruent trials (36 percent vs. 33 percent) and a slightly lower proportion of fully incongruent trials (28 percent vs. 31 percent). The proportion of congruent trials was nearly identical across groups (16 percent in HC vs. 15 percent in PSD), and neutral trials accounted for a slightly higher proportion of trials in the PSD group (21 percent vs. 19 percent in HC). For the Sternberg task, PSDs showed significantly lower ability estimates (t(92.47) = 2.32, p = .022, d = 0.47) and were also administered a lower average task load (HCs = 8.90 vs. PSDs = 8.24). For the adaptive SOPT, PSDs showed a non-significant trend toward lower ability estimates (t(75.14) = 1.66, p = .10, d = 0.35). However, both groups received the same average task load (HCs = 10.93 vs. PSDs = 10.93). Closer inspection of the results indicated that participants in both groups were consistently pushed toward the more difficult SOPT conditions, likely because the lower difficulty conditions permitted excessive guessing, limiting the precision of ability estimation. In other words, the SOPT appears to require higher difficulty levels for all participants relative to the fixed version in order to function effectively under adaptive administration, and our version of the task was effectively maxed out on the most difficult conditions available. This is a design limitation that can be corrected in future studies. For the N Back task, PSDs showed lower ability estimates on the CAT version of the task relative to the fixed version (t(88.25) = 3.91, p = .0002, d = 0.82) and were also administered a lower average task load (HCs = 2.81 vs. PSDs = 2.51).

Table 2 reports correlations, mean score estimates, score variability, SEs, and pleasantness ratings for the fixed and adaptive versions of each task. Ability estimates from fixed and adaptive versions were moderately to strongly correlated for the CPT, Flanker, N-Back, Sternberg, and SOPT, with all correlations above .70. In contrast, the fixed and adaptive versions of the PLT were weakly correlated. Mean score estimates were generally similar across formats for the Flanker, Sternberg, SOPT, and N-Back, but somewhat larger differences were observed for the CPT and PLT. Score variability also differed across formats to some degree, with SD ratios ranging from 0.742 for the CPT to 1.142 for the N-Back. In answer to one of the key aims of this study, adaptive testing significantly reduced SE for the CPT, Sternberg, and SOPT, but not for the PLT, Flanker, or N-Back. Notably, the CPT, Sternberg, and SOPT all have the widest range of difficulty and finest gradation of difficulty levels. Pleasantness ratings were significantly lower for the adaptive version of the CPT compared with the fixed version, although the effect was small. No other tasks showed significant differences in pleasantness ratings.

Table 2.

Comparison of Fixed and Adaptive Versions of Each Task

Score Estimate
Standard Error of Estimate
Pleasantness Ratings
rFixed,CAT Fixed M CAT M M Diff. Fixed SD CAT SD SD Ratio Fixed M CAT M t p d Fixed M CAT M t p d
PLT 0.186 −0.965 −0.79 0.175 0.673 0.704 1.096 0.180 0.187 −0.772 0.442 −0.078 3.849 4.119 −1.003 0.317 −0.154
CPT 0.807 1.305 1.16 −0.145 1.044 0.899 0.742 0.357 0.267 5.052 < .001 0.516 4.447 3.795 2.447 0.015 0.378
Flanker 0.751 1.781 1.837 0.056 1.069 1.127 1.112 0.380 0.390 −0.423 0.673 −0.044 5.036 5.059 −0.098 0.922 −0.015
Sternberg 0.794 1.751 1.717 −0.034 0.506 0.449 0.788 0.606 0.563 4.967 < .001 0.510 4.482 4.464 0.084 0.933 0.013
SOPT 0.708 0.083 0.102 0.02 0.665 0.6 0.814 0.399 0.318 9.815 < .001 1.023 5.149 4.952 0.873 0.384 0.133
N-back 0.798 1.414 1.544 0.13 1.02 1.09 1.142 0.338 0.348 −1.506 0.136 −0.157 3.105 3.414 −1.241 0.216 −0.189

Note: rFixed,CAT = correlation between scores for fixed and CAT versions of tasks; PLT = probabilistic learning task; CPT = continuous performance task; SOPT = self-ordered pointing task.

SE plots by ability estimates for each task are shown in Figure 1. Each figure should be interpreted as displaying the observed relationship between estimated ability and SE for the participants in this sample, rather than the full theoretical SE function for the task. In principle, standard errors are expected to show a characteristic “U”-shaped pattern, with SE lowest near the region of greatest measurement precision and increasing toward the tails of the ability distribution. However, because the plotted values reflect only the participants observed at specific estimated ability levels, the full lower and upper tails of the SE function may not be represented. Thus, an apparently monotonic increase in SE across the observed range should not be interpreted as evidence that the underlying SE function lacks a lower tail; rather, it indicates that the observed sample primarily occupies one side of the function. For most tasks, the patterns in Figure 1 are indicative of ceiling effects, with SE highest among participants with higher ability levels. In other words, from a psychometric standpoint, the tasks were generally too easy. The results imply that—for CPT, Sternberg, and SOPT—the CAT versions of these tasks successfully pushed examinees toward more difficult conditions (e.g., higher memory load and quicker backward mask onset time). As a result, the increase in standard error near the ceilings was less pronounced (i.e., less bowed).

Figure 1.

Figure 1

Standard Error by Ability Estimates for Each Task

Note. CAT = computerized adaptive test; PLT = probabilistic learning task; CPT = continuous performance task; SOPT = self-ordered pointing task. “Ability” refers to the intentional parameter for each model.

The interaction between group and task version was non-significant for the PLT (b = 0.007, SE = 0.018, df = 95, t = 0.389, p = 0.698), non-significant for the CPT (b = −0.007, SE = 0.036, df = 94, t = −0.203, p = 0.840), non-significant for the Flanker (b = −0.022, SE = 0.045, df = 91, t = −0.492, p = 0.624), non-significant for the Sternberg (b = − 0.020, SE = 0.017, df = 93, t = −1.208, p = 0.230), non-significant for the SOPT (b = −0.009, SE = 0.017, df = 90, t = −0.558, p = 0.578), and non-significant for the N-Back (b = 0.004, SE = 0.013, df = 90, t = 0.322, p = 0.748). Thus, the effectiveness of adaptive testing, or lack thereof, was not group dependent—a crucial characteristic for clinical usefulness.

Criterion-related validity for fixed and adaptive versions is summarized in Table 3. Group differences generally favored HCs (i.e., lower ability in PSDs), with significant effects for most tasks; exceptions were PLT Fixed, CPT Fixed, and CPT CAT, which were nonsignificant but still trending in the direction of lower ability estimates for PSDs. Higher symptom severity was associated with lower ability across tasks: SANS scores showed consistent, significant negative correlations for all measures, and SAPS scores did as well except for PLT CAT. Associations with MCCB were positive and significant for all measures except the CPT CAT (but was still marginally significant in the positive direction). For the WRAT, significant positive correlations were observed for PLT (both versions), Sternberg (both versions), SOPT CAT, and N-Back (both versions), with Flanker Fixed approaching significance. Overall, the pattern of validity correlations suggests that the fixed and adaptive versions of each task generally have similar relationships with important outcome variables.

Table 3.

Criterion-Related Validity of Fixed and Adaptive Tasks

Psychosis Diagnosis SAPS SANS MCCB WRAT

r p r p r p r p r p

PLT Fixed −0.125 0.207 −0.223 0.023 −0.245 0.014 0.241 0.015 0.223 0.045
PLT CAT −0.191 0.053 −0.078 0.43 −0.221 0.027 0.277 0.005 0.24 0.031
CPT Fixed −0.142 0.147 −0.326 0.001 −0.309 0.002 0.333 0.001 0.089 0.435
CPT CAT −0.137 0.167 −0.254 0.009 −0.289 0.003 0.176 0.079 0.045 0.686
Flanker Fixed −0.259 0.008 −0.307 0.002 −0.424 < .001 0.344 < .001 0.197 0.079
Flanker CAT −0.292 0.003 −0.385 < .001 −0.393 < .001 0.422 < .001 −0.037 0.744
Sternberg Fixed −0.384 < .001 −0.53 < .001 −0.554 < .001 0.413 < .001 0.341 0.001
Sternberg CAT −0.243 0.014 −0.326 0.001 −0.361 < .001 0.307 0.002 0.233 0.04
SOPT Fixed −0.244 0.013 −0.326 0.001 −0.348 < .001 0.359 < .001 0.129 0.248
SOPT CAT −0.226 0.023 −0.326 0.001 −0.414 < .001 0.357 < .001 0.264 0.019
N-Back Fixed −0.292 0.003 −0.411 < .001 −0.488 < .001 0.5 < .001 0.257 0.021
N-Back CAT −0.399 < .001 −0.446 < .001 −0.486 < .001 0.425 < .001 0.304 0.006

Note: CAT = computerized adaptive test; PLT = probabilistic learning task; CPT = continuous performance task; SOPT = self-ordered pointing task; SAPS = Scale for the Assessment of Positive Symptoms; SAPS = Scale for the Assessment of Negative Symptoms; MCCB = MATRICS Consensus Cognitive Battery; WRAT = Wide Range of Achievement Test.

Discussion

This study aimed to determine whether adaptive testing could improve the measurement precision of cognitive outcome measures used in the RDoC project and in studies of cognitive psychopathology more broadly. Overall, the results were mixed. Adaptive testing improved measurement precision for three of the six tasks examined—those with the widest range of difficulty and the finest gradation of difficulty levels. Ratings of task pleasantness and criterion-related validity correlations between adaptive and fixed versions of each task were generally similar, with only a few exceptions.

The CPT, Sternberg, and SOPT all benefited from adaptive testing. For the CPT, difficulty increased as the backward mask was presented more quickly, making the task harder relative to a slower backward mask onset time (longer presentation time). For the Sternberg and SOPT tasks, difficulty was manipulated by increasing memory load. For all three, the CAT versions were able to effectively increase difficulty levels for high-performing subjects, thereby improving measurement precision at the upper end of the distribution (Figure 1). An important caveat is that the CPT scoring model exhibited the poorest model fit (i.e., the highest degree of local dependence) among the tasks examined. This result appears to be driven primarily by data sparsity and high performance rather than by fundamental model misspecification (see Supplemental Material), but caution is warranted in future applications of signal detection-based scoring models to CPT data. Additionally, SOPT CAT administrations were near the ceiling of difficulty, suggesting that the adaptive version is still not optimally efficient.

Interestingly, while the Sternberg, SOPT, and N-Back are all measures of working memory that manipulate memory load, the N-Back was the only one of the three that did not benefit from adaptive testing. This outcome is particularly surprising given that a previous simulation of N-Back data suggested otherwise (Thomas et al., 2021). In that previous simulation study, the fixed version of the task included conditions ranging from 1-back to 3-back, whereas the adaptive version drew from a wider range of 1-back to 5-back. In the current study, however, both the fixed and adaptive versions included 1-back through 4-back conditions. It is possible that when the fixed version already spans an optimal range of difficulty (i.e., 1-back to 4-back), adaptively selecting difficulty conditions in real time adds little additional benefit. While IRT-based analyses and simulations can still help identify optimal ranges of difficulty conditions, our results suggest that some tasks—such as the N-Back—may not require “live” adaptation if the fixed version is appropriately specified for the population of interest. The key challenge remains knowing a priori which conditions would best capture the full ability range of the target population. It is also possible that with more time constraints (i.e. shorter task), the CAT version would have performed better.

The Flanker task also did not benefit from adaptive testing. We manipulated difficulty by varying congruence. Although our previous report, from a larger online sample, suggest that Flanker stimulus congruence does impact performance (Duffy et al., 2024), the effects were relatively minor, with the vast majority of participants still scoring very high on the measure. In the current study, performance was again quite high for the Flanker (Table 1), suggesting that this version of the task remains too easy to make effective use of adaptive testing. Notably, the Flanker had the most pronounced ceiling effect in comparison to the other measures (Figure 1). Alternative versions of the Flanker that manipulate other aspects of the stimuli may perform better in an adaptive testing context.

The adaptive version of the PLT also did not improve precision relative to the fixed version. Notably, unlike the other five tasks, the ability estimates produced by the fixed and adaptive PLT were only weakly correlated. We suspect this is due to the nature of the PLT, which requires subjects to learn as the task progresses. There is also likely a “learning to learn” process, in which subjects become increasingly aware of task requirements and optimize their learning strategies over time (Meeter et al., 2006). This learning process may unfold differently depending on whether the examinee is administered a fixed versus an adaptive version of the task. Therefore, unlike the other five measures—where the CAT versions appeared equivalent in terms of construct validity—the CAT version of the PLT does not appear to be equivalent to the fixed version. These findings raise concerns about applying conventional CAT methods to tasks in which performance depends on learning across trials. Because adaptive administration changes the sequence of difficulty and reinforcement experiences, it may alter the learning process itself rather than simply improving measurement efficiency. Thus, fixed and adaptive PLT scores may reflect partially different constructs. Future adaptive versions of reinforcement-learning tasks may need to preserve key features of the learning sequence, explicitly model learning trajectories, or use alternative adaptive designs that improve precision without disrupting the core learning process.

A clear recommendation from this study is that while adaptive testing can effectively improve the measurement precision of experimental cognitive tasks used in neuropsychological and cognitive neuroscience research, some measures are expected to benefit more than others; a finding that supports previous reports (e.g., Di Sandro et al., 2024). Perhaps more importantly, our findings also suggest that some versions of tasks may even measure different constructs depending on whether they are administered in a fixed versus adaptive format. Thus, adaptive versions of tasks should be thoroughly investigated and psychometrically evaluated before being widely used.

Importantly, this research has been largely based on tasks that are also available in animals (except the N-back task), enabling the possibility of translational studies determining neural mechanisms that contribute to performance and/or differences related to psychiatric conditions (e.g., Roberts & Young, 2022). While not routinely conducted, IRT-based adaptive testing could be utilized in touchscreen-based versions of these tasks (Dexter et al., 2025; Young, 2023), where the same principles of manipulating task difficulty to optimize measurement precision should similarly apply. Indeed, a central goal of the RDoC initiative is to bridge the gap between preclinical studies and research on human psychopathology. Of course, this possibility would require validation before being implemented in practice, but offers additional future translational research into these cognitive domains impacted in psychiatric conditions.

An important caveat to the findings reported above is that our use of CAT in the present study restricted the number of test conditions (and therefore items) that could be administered in both the fixed and adaptive versions; this was intended to provide a fair comparison between testing conditions. In many real-world applications of CAT, however, the number of administered items/conditions is allowed to vary, with testing continuing until a predetermined minimum level of estimation error is reached. In other words, CAT may still have added value in such settings, even if it is not more efficient than fixed testing. One possible variation of the CAT approach is to always administer the fixed version of a task first and then engage adaptive testing only if further information is needed to produce acceptable score precision.

Results and recommendations from this study should be considered in light of several limitations. First, administration of adaptive tests is technically challenging and expensive. For this reason, very few small, non-commercial projects have investigated this technology, especially outside of its traditional settings. Typically, such efforts involve multiple study waves of test refinement before adaptive testing is employed. It is possible that calibrating item parameters on data from a larger, in-person cohort more closely matched to the current study’s participants, along with multiple waves of refinement, would have improved CAT performance for some tasks. Second, we focused on just six cognitive tasks. While several of these tasks are among the most commonly used in neurocognitive research—and were specifically selected and recommended by the RDoC group—they represent only a small subsample of the many tasks that could potentially be optimized using IRT and CAT technology. The results of this study may help guide future efforts, but they should not be taken as definitive for measures not examined here. Third, the results should be interpreted in light of the specific populations and samples recruited, which may be limited in terms of diversity; accordingly, the findings may not generalize to the broader population. Related to this, calibration was based on an online sample, whereas task performance was evaluated in an in-person mixed clinical and non-clinical sample. Although the expected relationships between difficulty/load and performance replicated in the current sample (Supplemental Figure 1), the calibration parameters may not have generalized across samples. Differences in demographic characteristics, clinical status, testing context, motivation, or device/input characteristics could influence item difficulty. As a result, adaptive item selection and ability estimation may have been somewhat less efficient than would be expected if the tasks had been calibrated directly in the target population. Motivated by this concern, we conducted a sensitivity analysis to determine how robust the findings were to calibration inaccuracy. The results, reported in the Supplement, generally suggested that large deviations from the true calibration parameters would be required to alter our substantive conclusions. Still, future work should evaluate whether recalibration or population-specific calibration improves measurement precision and adaptive testing performance in clinical samples. Fourth, when CAT outperformed fixed testing, or vice versa, these findings should be interpreted as comparisons with the specific fixed versions evaluated in this study. Other fixed versions of each task could perform better or worse depending on the difficulty levels selected and how well those levels align with the ability distribution of the population being assessed. Fifth, the present findings should not be interpreted as evidence that fixed and adaptive versions of these tasks are clinically interchangeable. Although fixed and adaptive scores were often strongly correlated and showed broadly similar criterion-related validity patterns, correlations alone do not establish classification equivalence. Future studies should evaluate whether impairment classifications, clinical cutoff scores, and decision rules are preserved across fixed and adaptive administrations before these scores are used interchangeably in applied clinical contexts. Finally, fit indices across tasks ranged from evidence of mild misfit (e.g., N-Back) to more moderate misfit (e.g., CPT). Although prior work suggests that the practical consequences of model misfit can be quite limited (Sinharay & Haberman, 2014), including evidence that only severe misfit meaningfully affects adaptive testing performance (Reese, 1999), such effects are highly case-specific and warrant further investigation for the models and tasks examined in the current study.

In conclusion, IRT- and CAT-based technologies hold promise for improving the measurement precision of experimental cognitive tasks that currently suffer from poor psychometric properties. Caution is warranted however, in deploying these methods that are likely domain-specific, and definitive evidence is needed before claims of their superiority can be confirmed.

Supplementary Material

Supplemental Material

Public Significance Statement.

Cognitive functioning is an important dimension of mental health, but commonly used tests can be inefficient or imprecise for some individuals. This study shows how item response theory and computerized adaptive testing can improve measurement of cognitive performance across a broad range of ability. More precise and efficient cognitive assessment may strengthen research on mental illness and support better evaluation of interventions designed to improve cognition.

Acknowledgments

Research reported in this publication was supported by the National Institute of Mental Health of the National Institutes of Health under award number R01MH121546.

The manuscript was edited with the support of AI tools including ChatGPT and Google Gemini.

References

  1. Bates D, Mächler M, Bolker B, & Walker S. (2015). Fitting linear mixed-effects models using lme4. Journal of Statistical Software, 67(1), 1–48. 10.18637/jss.v067.i01 [DOI] [Google Scholar]
  2. Bhakta SG, & Young JW. (2017). The 5 choice continuous performance test (5C-CPT): A novel tool to assess cognitive control across species. Journal of Neuroscience Methods, 292, 53–60. 10.1016/j.jneumeth.2017.07.011 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Brody AL, Mischel AK, Sanavi AY, Wong A, Bahn JH, Minassian A, Morgan EE, Rana B, Hoh CK, Vera DR, Kotta KK, Miranda AH, Pocuca N, Walter TJ, Guggino N, Beverly-Aylwin R, Meyer JH, Vasdev N, & Young JW. (2025). Cigarette smoking is associated with reduced neuroinflammation and better cognitive control in people living with HIV. Neuropsychopharmacology: Official Publication of the American College of Neuropsychopharmacology, 50(4), 695–704. 10.1038/s41386-024-02035-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Buckley RF, Sparks KP, Papp KV, Dekhtyar M, Martin C, Burnham S, Sperling RA, & Rentz DM. (2017). Computerized cognitive testing for use in clinical trials: A comparison of the NIH Toolbox and Cogstate C3 batteries. The Journal of Prevention of Alzheimer’s Disease, 1–9. 10.14283/jpad.2017.1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  5. Cabranes JA, Ancín I, Santos JL, Sánchez-Morla E, García-Jiménez MA, Rodríguez-Moya L, Fernández C, & Barabash A. (2013). P50 sensory gating is a trait marker of the bipolar spectrum. European Neuropsychopharmacology, 23(7), 721–727. 10.1016/j.euroneuro.2012.06.008 [DOI] [PubMed] [Google Scholar]
  6. Cardno AG, & Owen MJ. (2014). Genetic relationships between schizophrenia, bipolar disorder, and schizoaffective disorder. Schizophrenia Bulletin, 40(3), 504–515. 10.1093/schbul/sbu016 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Cavanagh JF, Gregg D, Light GA, Olguin SL, Sharp RF, Bismark AW, Bhakta SG, Swerdlow NR, Brigman JL, & Young JW. (2021). Electrophysiological biomarkers of behavioral dimensions from cross-species paradigms. Translational Psychiatry, 11(1), 482. 10.1038/s41398-021-01562-w [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Chapman LJ, & Chapman JP. (1973). Problems in the measurement of cognitive deficit. Psychological Bulletin, 79(6), 380–385. [DOI] [PubMed] [Google Scholar]
  9. Christensen KB, Makransky G, & Horton M. (2017). Critical Values for Yen’s Q3: Identification of local dependence in the Rasch model using residual correlations. Applied Psychological Measurement, 41(3), 178–194. 10.1177/0146621616677520 [DOI] [PMC free article] [PubMed] [Google Scholar]
  10. Chudasama Y, & Robbins TW. (2006). Functions of frontostriatal systems in cognition: Comparative neuropsychopharmacological studies in rats, monkeys and humans. Biological Psychology, 73(1), 19–38. 10.1016/j.biopsycho.2006.01.005 [DOI] [PubMed] [Google Scholar]
  11. Clarke M. (2007). Standardising outcomes for clinical trials and systematic reviews. Trials, 8(1), 39. 10.1186/1745-6215-8-39 [DOI] [PMC free article] [PubMed] [Google Scholar]
  12. Clementz BA, Parker DA, Trotti RL, McDowell JE, Keedy SK, Keshavan MS, Pearlson GD, Gershon ES, Ivleva EI, Huang L-Y, Hill SK, Sweeney JA, Thomas O, Hudgens-Haney M, Gibbons RD, & Tamminga CA. (2022). Psychosis Biotypes: Replication and Validation from the B-SNIP Consortium. Schizophrenia Bulletin, 48(1), 56–68. 10.1093/schbul/sbab090 [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. D’Esposito M. (2007). From cognitive to neural models of working memory. Philosophical Transactions of the Royal Society B: Biological Sciences, 362(1481), 761–772. 10.1098/rstb.2007.2086 [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Dexter TD, Roberts BZ, Ayoub SM, Noback M, Barnes SA, & Young JW. (2025). Cross-species translational paradigms for assessing positive valence system as defined by the RDoC matrix. Journal of Neurochemistry, 169(1), e16243. 10.1111/jnc.16243 [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Di Sandro A, Moore TM, Zoupou E, Kennedy KP, Lopez KC, Ruparel K, Njokweni LJ, Rush S, Daryoush T, Franco O, Gorgone A, Savino A, Didier P, Wolf DH, Calkins ME, Cobb Scott J, Gur RE, & Gur RC. (2024). Validation of the cognitive section of the Penn computerized adaptive test for neurocognitive and clinical psychopathology assessment (CAT-CCNB). Brain and Cognition, 174, 106117. 10.1016/j.bandc.2023.106117 [DOI] [PMC free article] [PubMed] [Google Scholar]
  16. Duffy JR, Sturm ET, Sares AG, Sarabia L, Delao EM, M Becker K, Colmenares AM, Manavi RM, Rojas DC, Tregellas JR, Young JW, & Thomas ML. (2024). Psychometric evaluation of cognitive and positive valence tasks chosen for the NIMH Research Domain Criteria Project. Assessment, 10731911241280770. 10.1177/10731911241280770 [DOI] [PubMed] [Google Scholar]
  17. Eriksson J, Vogel EK, Lansner A, Bergström F, & Nyberg L. (2015). Neurocognitive Architecture of Working Memory. Neuron, 88(1), 33–46. 10.1016/j.neuron.2015.09.020 [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Goodglass H, & Kaplan E. (1979). Assessment of Cognitive Deficit in the Brain-Injured Patient. In Gazzaniga MS. (Ed.), Neuropsychology (Vol. 2, pp. 3–22). Springer; US. 10.1007/978-1-4613-3944-1_1 [DOI] [Google Scholar]
  19. Hedge C, Powell G, & Sumner P. (2018). The reliability paradox: Why robust cognitive tasks do not produce reliable individual differences. Behavior Research Methods, 50(3), 1166–1186. 10.3758/s13428-017-0935-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Jeste DV, Palmer BW, Appelbaum PS, Golshan S, Glorioso D, Dunn LB, Kim K, Meeks T, & Kraemer HC. (2007). A new brief instrument for assessing decisional capacity for clinical research. Archives of General Psychiatry, 64(8), 966. 10.1001/archpsyc.64.8.966 [DOI] [PubMed] [Google Scholar]
  21. Kirchner WK. (1958). Age differences in short-term retention of rapidly changing information. Journal of Experimental Psychology, 55(4), 352–358. 10.1037/h0043688 [DOI] [PubMed] [Google Scholar]
  22. Leek MR. (2001). Adaptive procedures in psychophysical research. Perception & Psychophysics, 63(8), 1279–1292. 10.3758/BF03194543 [DOI] [PubMed] [Google Scholar]
  23. Li T, Xie C, & Jiao H. (2017). Assessing fit of alternative unidimensional polytomous IRT models using posterior predictive model checking. Psychological Methods, 22(2), 397–408. 10.1037/met0000082 [DOI] [PubMed] [Google Scholar]
  24. MacQueen DA, Minassian A, Kenton JA, Geyer MA, Perry W, Brigman JL, & Young JW. (2018). Amphetamine improves mouse and human attention in the 5-choice continuous performance test. Neuropharmacology, 138, 87–96. 10.1016/j.neuropharm.2018.05.034 [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Malaspina D, Owen MJ, Heckers S, Tandon R, Bustillo J, Schultz S, Barch DM, Gaebel W, Gur RE, Tsuang M, Van Os J, & Carpenter W. (2013). Schizoaffective Disorder in the DSM-5. Schizophrenia Research, 150(1), 21–25. 10.1016/j.schres.2013.04.026 [DOI] [PubMed] [Google Scholar]
  26. Martin AJ, & Lazendic G. (2018). Computer-adaptive testing: Implications for students’ achievement, motivation, engagement, and subjective test experience. Journal of Educational Psychology, 110(1), 27–45. 10.1037/edu0000205 [DOI] [Google Scholar]
  27. Meeter M, Myers CE, Shohamy D, Hopkins RO, & Gluck MA. (2006). Strategies in probabilistic categorization: Results from a new way of analyzing performance. Learning & Memory, 13(2), 230–239. 10.1101/lm.43006 [DOI] [PMC free article] [PubMed] [Google Scholar]
  28. Moore TM, Di Sandro A, Scott JC, Lopez KC, Ruparel K, Njokweni LJ, Santra S, Conway DS, Port AM, D’Errico L, Rush S, Wolf DH, Calkins ME, Gur RE, & Gur RC. (2023). Construction of a computerized adaptive test (CAT-CCNB) for efficient neurocognitive and clinical psychopathology assessment. Journal of Neuroscience Methods, 386, 109795. 10.1016/j.jneumeth.2023.109795 [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. National Advisory Mental Health Council Workgroup on Tasks and Measures for Research Domain Criteria (RDoC). (2016). Behavioral assessment methods for rdoc constructs. https://www.nimh.nih.gov/sites/default/files/documents/about/advisory-boards-and-groups/namhc/reports/rdoc_council_workgroup_report.pdf [Google Scholar]
  30. Nuechterlein KH, Green MF, Kern RS, Baade LE, Barch DM, Cohen JD, Essock S, Fenton WS, Frese FJ, Gold JM, Goldberg T, Heaton RK, Keefe RSE, Kraemer H, Mesholam-Gately R, Seidman LJ, Stover E, Weinberger DR, Young AS, … Marder SR. (2008). The MATRICS Consensus Cognitive Battery, Part 1: Test Selection, Reliability, and Validity. American Journal of Psychiatry, 165(2), 203–213. 10.1176/appi.ajp.2007.07010042 [DOI] [PubMed] [Google Scholar]
  31. Peirce JW. (2009). Generating stimuli for neuroscience using PsychoPy. Frontiers in Neuroinformatics, 2. 10.3389/neuro.11.010.2008 [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Petrides M, & Milner B. (1982). Deficits on subject-ordered tasks after frontal- and temporal-lobe lesions in man. Neuropsychologia, 20(3), 249–262. 10.1016/0028-3932(82)90100-2 [DOI] [PubMed] [Google Scholar]
  33. Pocuca N, Young JW, MacQueen DA, Letendre S, Heaton RK, Geyer MA, Perry W, Grant I, Minassian A, & Translational Methamphetamine AIDS Research Center (TMARC). (2020). Sustained attention and vigilance deficits associated with HIV and a history of methamphetamine dependence. Drug and Alcohol Dependence, 215, 108245. 10.1016/j.drugalcdep.2020.108245 [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Reese LM. (1999). Impact of local item dependence on item response theory scoring in computerized adaptive testing (LSAC Computerized Testing Report 98-08). Law School Admission Council. https://files.eric.ed.gov/fulltext/ED467817.pdf [Google Scholar]
  35. Robble MA, Schroder HS, Kangas BD, Nickels S, Breiger M, Iturra-Mena AM, Perlo S, Cardenas E, Der-Avakian A, Barnes SA, Leutgeb S, Risbrough VB, Vitaliano G, Bergman J, Carlezon WA, & Pizzagalli DA. (2021). Concordant neurophysiological signatures of cognitive control in humans and rats. Neuropsychopharmacology: Official Publication of the American College of Neuropsychopharmacology, 46(7), 1252–1262. 10.1038/s41386-021-00998-4 [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Roberts B, & Young JW. (2022). Translational cognitive systems: Focus on attention. Emerging Topics in Life Sciences, 6(5), 529–539. 10.1042/ETLS20220009 [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Salmon C, Li S, Burrows EL, & Johnson KA. (2024). Translational validity of neuropsychological tasks of sustained attention between rodents and humans: A systematic review of three rodent tasks. Journal of Neurochemistry, 168(9), 2170–2189. 10.1111/jnc.16117 [DOI] [PubMed] [Google Scholar]
  38. Santelmann H, Franklin J, Bußhoff J, & Baethge C. (2015). Test–retest reliability of schizoaffective disorder compared with schizophrenia, bipolar disorder, and unipolar depression—A systematic review and meta-analysis. Bipolar Disorders, 17(7), 753–768. 10.1111/bdi.12340 [DOI] [PubMed] [Google Scholar]
  39. Segall DO. (2009). Principles of Multidimensional Adaptive Testing. In van der Linden WJ & Glas CAW. (Eds.), Elements of Adaptive Testing (pp. 57–75). Springer; New York. 10.1007/978-0-387-85461-8_3 [DOI] [Google Scholar]
  40. Sheehan DV, Lecrubier Y, Sheehan KH, Amorim P, Janavs J, Weiller E, Hergueta T, Baker R, & Dunbar GC. (1998). The Mini-International Neuropsychiatric Interview (M.I.N.I.): The development and validation of a structured diagnostic psychiatric interview for DSM-IV and ICD-10. The Journal of Clinical Psychiatry, 59 Suppl 20, 22–33;quiz 34-57. [PubMed] [Google Scholar]
  41. Sinharay S, & Haberman SJ. (2014). How often is the misfit of item response theory models practically significant? Educational Measurement: Issues and Practice, 33(1), 23–35. 10.1111/emip.12024 [DOI] [Google Scholar]
  42. Smith MJ, Barch DM, & Csernansky JG. (2009). Bridging the gap between schizophrenia and psychotic mood disorders: Relating neurocognitive deficits to psychopathology. Schizophrenia Research, 107(1), 69–75. 10.1016/j.schres.2008.07.014 [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Sternberg S. (1975). Memory scanning: New findings and current controversies. Quart.J.Exp.Psychol, 27(1), 1–32. 10.1080/14640747508400459 [DOI] [Google Scholar]
  44. Thomas ML. (2020). An Adaptive Testing Platform for Optimizing RDoC Experimental Cognitive Measures ( 10.15154/pyp0-7t31) [Dataset]. NIMH Data Archive. [DOI] [Google Scholar]
  45. Thomas ML. (2024). cogirt: Item Response Theory for Cognitive Testing [Computer software]. https://CRAN.R-project.org/package=cogirt [Google Scholar]
  46. Thomas ML, Brown GG, Gur RC, Moore TM, Patt VM, Risbrough VB, & Baker DG. (2018). A signal detection–item response theory model for evaluating neuropsychological measures. Journal of Clinical and Experimental Neuropsychology, 40(8), 745–760. 10.1080/13803395.2018.1427699 [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Thomas ML, Brown GG, Patt VM, & Duffy JR. (2021). Latent variable modeling and adaptive testing for experimental cognitive psychopathology research. Educational and Psychological Measurement, 81(1), 155–181. 10.1177/0013164420919898 [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. Thomas ML, Patt VM, Bismark A, Sprock J, Tarasenko M, Light GA, & Brown GG. (2017). Evidence of systematic attenuation in the measurement of cognitive deficits in schizophrenia. Journal of Abnormal Psychology, 126(3), 312–324. 10.1037/abn0000256 [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Van Wijk EV, Donkers J, De Laat PCJ, Meiboom AA, Jacobs B, Ravesloot JH, Tio RA, Van Der Vleuten CPM, Langers AMJ, & Bremers AJA. (2024). Computer Adaptive vs. Non-adaptive Medical Progress Testing: Feasibility, Test Performance, and Student Experiences. Perspectives on Medical Education, 13(1). 10.5334/pme.1345 [DOI] [PMC free article] [PubMed] [Google Scholar]
  50. Whitely SE. (1983). Construct validity: Construct representation versus nomothetic span. Psychological Bulletin, 93(1), 179–197. 10.1037/0033-2909.93.1.179 [DOI] [Google Scholar]
  51. Wood AJ, Carroll AR, Shinn AK, Ongur D, & Lewandowski KE. (2021). Diagnostic Stability of Primary Psychotic Disorders in a Research Sample. Frontiers in Psychiatry, 12, 734272. 10.3389/fpsyt.2021.734272 [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Yang ACM, Flanagan B, & Ogata H. (2022). Adaptive formative assessment system based on computerized adaptive testing and the learning memory cycle for personalized learning. Computers and Education: Artificial Intelligence, 3, 100104. 10.1016/j.caeai.2022.100104 [DOI] [Google Scholar]
  53. Yen WM. (1984). Effects of local item dependence on the fit and equating performance of the three-parameter logistic model. Applied Psychological Measurement, 8(2), 125–145. 10.1177/014662168400800201 [DOI] [Google Scholar]
  54. Young JW. (2023). Development of cross-species translational paradigms for psychiatric research in the Research Domain Criteria era. Neuroscience and Biobehavioral Reviews, 148, 105119. 10.1016/j.neubiorev.2023.105119 [DOI] [PubMed] [Google Scholar]
  55. Young JW, Light GA, Marston HM, Sharp R, & Geyer MA. (2009). The 5-Choice Continuous Performance Test: Evidence for a Translational Test of Vigilance for Mice. PLoS ONE, 4(1), e4227. 10.1371/journal.pone.0004227 [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplemental Material

RESOURCES