Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2025 Jan 30.
Published in final edited form as: Child Neuropsychol. 2024 Mar 21;31(1):1–30. doi: 10.1080/09297049.2024.2329435

A short executive functioning questionnaire in the context of early childhood screening: psychometric properties

Alyssa R Palmer a,*, Amanda W Kalstabakken a,b, Rebecca Distefano c,d,#, Stephanie M Carlson a, Samuel P Putnam e, Ann S Masten a
PMCID: PMC11451573  NIHMSID: NIHMS2021885  PMID: 38511396

Abstract

Early childhood executive functioning (EF) predicts later adjustment and academic achievement. However, measuring EF consistently and efficiently across settings in early childhood can be challenging. Most researchers use task-based measures of EF, but these methods present practical challenges that impede implementation in some settings. The current study of 380 3–5-year-old children in the United States evaluated the psychometric properties of a new 14-item parent-reported measure of EF in a diverse urban school district. This questionnaire aimed to capture a normative range of EF skills in ecologically valid contexts. There was evidence for two specific subscales – one that measures children’s EF challenges and another that measures children’s EF skills. Results suggested that several items demonstrated differential item functioning by age and race. After adjusting for measurement differences across demographic groups and controlling for age at screening, the EF challenges subscale was more strongly related to task-based measures of EF than was the EF skills subscale. EF challenges predicted third-grade math achievement, controlling for demographic variables and a performance-based measure of children’s early cognitive and academic skills. Results suggest that this parent report of EF could be a useful and effective early childhood screening tool.

Keywords: Early childhood screening, executive functioning, parent report, academic achievement, school readiness


Executive functioning (EF) has been increasingly recognized as an important domain of developmental competence for preschool-aged children. EF is defined as a collection of heterogeneous top-down neurocognitive processes involved in goal-directed behaviors. Common subcomponents of EF include inhibition, working memory, attention focusing, and attention shifting. Inhibition is the ability to stop a prepotent response or behavior. Working memory is the ability to hold limited amounts of information in mind and manipulate that information. Attention focusing is the ability to sustain attention on a specific stimulus, while attention shifting is the ability to flexibly shift focus given task demands (Diamond, 2013; Miyake et al., 2000). The current study aimed to introduce and evaluate the psychometric properties of a short parent-reported scale of EF for preschool-aged children that we will refer to as the Short Executive Functioning Questionnaire (SEFQ).

EF skills rapidly develop in the preschool period (3–5 years old; Carlson, 2005; Garon et al., 2008). Throughout this developmental period, children have increasing abilities to hold and manipulate information in mind, refrain from engaging in prohibited activities, and more readily control their attention and emotions. Deficits in EF skills have been related to general learning problems, social-emotional difficulties, and psychiatric disorders (Johnson et al., 2010; Snyder et al., 2015; Willcutt et al., 2005; Zelazo, 2020). EF skills during the preschool period are foundational for learning and later academic achievement (Pérez-Edgar et al., 2020; Zelazo et al., 2016). EF skills during the preschool period have been associated with academic skills in kindergarten (Kalstabakken et al., 2021; Willoughby et al., 2017) and in later school years (Best et al., 2011; Jacob & Parkinson, 2015; McClelland et al., 2006; Wolf & McCoy, 2019).

Theoretical orientation

The theoretical approach to EF in the present study was guided by the Cognitive Complexity and Control revised theory of EF development (Carlson et al., 2013; Zelazo et al., 2003). This perspective suggests that the development of EF in preschool-aged children is accounted for by increases in their ability to understand and apply complex sets of rules. This rapid increase in skills is due to the development of neural regions capable of increasingly complex reflection via the iterative re-processing of information. The current study was also influenced by the salience and role of EF in the literature pertinent to adaptive functioning in the context of stress, motivation, and high emotional arousal. Tasks involving motivation and emotional arousal likely demand different EF processes compared to tasks void of those components. A parent’s report of their child’s EF skills is grounded in their observations of their children during daily tasks. These daily tasks often involve children experiencing frustration or sadness due to blocked goals such as waiting to receive a reward or being required to be quiet or still. It is important to incorporate measures that consider young children’s EF skills in such contexts (Zelazo & Carlson, 2012).

From a practical perspective, the SEFQ items were intended to capture EF skills in a brief parent-reported format with items indexing readily observable child behavior. the SEFQ items were intended to capture specific observable skills in real-world settings reflecting both cool and hot EF skills. EF skills defined as “cool” assess “top-down” neurocognitive skills relatively devoid of motivational and emotional content, such as “good at memory games.” Items assessing “hot” EF skills refer to top-down regulation of behavior in emotionally arousing contexts, such as “can save candy/treats for later.” There is some evidence that hot and cool EF skills tap into overlapping but distinct neural and physiological processes (Zelazo & Carlson, 2012). Some SEFQ questions were written to mirror widely used task-based measures of hot EF, for example, asking if a child can wait to open a present (e.g., Gift Delay; Kochanska et al., 2000) or save a treat for later (e.g., the Marshmallow Test; Shoda et al., 1990).

Measuring EF

Universal patterns of and individual differences in EF development have primarily been identified via studies using direct assessments. However, it can be challenging to incorporate these tasks into practice-based settings. This may be particularly true for young children and those from marginalized backgrounds, who often show floor effects on existing task-based measures of EF (Akshoomoff et al., 2014; Carlson, 2005). Over the past decade, computer-based measures that incorporate adaptive testing methods have been created to address floor effects. These include the Minnesota Executive Function Scale (Carlson & Zelazo, 2014), and beta versions of downward extensions of the National Institutes of Health Toolbox measures of EF (Distefano et al., 2021, 2023; Kalstabakken et al., 2021). These tasks effectively lowered the floor of EF tasks by scaffolding children’s ability to understand the rules.

However, computer-based measures of EF are often given in controlled environments that may not reflect a child’s EF skills in the affectively laden context of daily life (McCoy, 2019). Further, direct assessment of young children presents a number of challenges, including fatigue, resistance to completing tasks, and inconsistent performance. For example, 3-year-old children tend to have lower test–retest reliability on EF assessments (Willoughby & Blair, 2011) compared to older children. For measures to have utility for educators or practitioners, they need to yield helpful information while minimizing logistical barriers including administration time and staff training. A potential low-burden alternative to using task-based measures is to use questionnaire-based reports from a parent, other caregivers, or teachers.

Questionnaire methods are often more readily implemented in settings like pediatrician offices or early childhood screening protocols because they have fewer logistical and administrative barriers. Further, questionnaires administered electronically can be provided to families remotely, reducing the need for transportation to assessment locations. This method can increase access to screening for families who may have not been able to access those services otherwise. Parent report questionnaires could also increase ecological validity by asking about EF skills in everyday contexts (Mahone & Hoffman, 2007; McCoy, 2019). Further, questionnaires that capture the normative range of EF may be helpful in studies that want to include low-burden measures while following groups of children longitudinally. The current study aimed to increase the usability of EF assessments in practice-based settings by evaluating the psychometric properties of a 14-item EF parent-reported questionnaire. We also aimed to evaluate the measure’s predictive utility for later academic achievement as compared to previous studies evaluating the predictive utility of task-based measures of EF.

There are a number of measures of EF aimed at school-aged children including: Childhood Executive Functioning Inventory (CHEXI; Thorell & Catale, 2014); Comprehensive Executive Function Inventory (CEFI; Naglieri & Goldstein, 2014); and Barkley Deficits in Executive Functioning Scale – Children and Adolescents (BDEFS-CA; Barkley, 2012). However, few of these measures were designed for preschool-aged children or to function well with diverse, less-educated, or low-income parents. It is worth noting that the CHEXI has been assessed for use with preschool-aged children (Camerota et al., 2018), and demonstrated acceptable factor structure and strong measurement invariance by child sex and income. The scale was weakly correlated with task-based measures of EF and there are no tests of longitudinal predictive utility of the measure from this age range. The current study aims to add a very short questionnaire screener of EF specifically designed for the preschool age range that can be used with diverse families.

Among the lengthy measures of EF in the preschool age range is the widely used, 63-item Behavior Rating Inventory of Executive Functioning-Preschool Version (BRIEF-P; Gioia et al., 2000). The BRIEF-P measures EF via five subscales: inhibitory control, working memory, shifting, planning, and emotional control. In validation studies, the BRIEF-P demonstrated good internal consistency and temporal stability (Gioia et al., 2000). The BRIEF-P was designed for clinical purposes and thus (understandably) deficit-focused on problem behaviors. Instructions ask parents how often a particular behavior “has been a problem in the past 6 months.” The BRIEF-P has substantial overlap with diagnostic criteria for Attention Deficit Hyperactivity Disorder and was not designed to capture developmentally normative EF-related behaviors (Nilsen et al., 2017; Thorell & Catale, 2014). Perhaps as a result of these features of the questionnaire, the BRIEF-P often correlates more strongly with concurrent measures of behavioral disruptions and impairments than it is correlated with EF skills measured via task-based performance (McAuley et al., 2010). This suggests that the BRIEF-P may be capturing more variance associated with metrics of behavioral disruption and impairment rather than variance associated with EF skills.

Another lengthy measure of EF in young children is the Rating of Everyday Executive Functioning (REEF; Nilsen et al., 2017). This 71-item scale captures the normative range of preschool children’s EF. Items focus on observable skills in specific daily contexts and the REEF has demonstrated strong convergent and divergent validity. REEF scores are moderately related to task-based measures of EF (r = 0.37), and more strongly related to those task-based measures compared to its relationship with questionnaire measures about child fear expression (r = −0.11), sadness expression (r = −0.09), smiling and laughing (r = 0.24), and peer problems (r = −0.28; Nilsen et al., 2017). Results also suggested a strong relationship between the REEF and other questionnaire-based measures of EF with correlation strength ranging from 0.46 to 0.66. This suggests that the REEF may be a better measure and capture more variance associated with EF skills rather than behavioral or emotional disruption. The current study assessed whether the considerably shorter SEFQ was also more related to task-based measures of EF compared to its association with questionnaire measures of social-emotional problems.

Divergent validity

As a test of divergent validity, the current study also assessed how the SEFQ is related to effortful control (EC) and broader indicators of parent-reported social-emotional functioning. Some researchers posit that EC and EF are measuring the same construct due to both including aspects of attention-focusing and attention-shifting skills (Zhou et al., 2012). However, the developers of the leading EC measure, the Children’s Behavior Questionnaire-Very Short Form (CBQ-VSF; Putnam & Rothbart, 2006), have argued persuasively that EC is distinct from EF. Most fundamentally, EC stems from the literature on temperament (Rothbart, 2011; Zentner & Bates, 2008), and captures aspects of emotional expression, arousal, and general regulation traits, which EF typically does not (Gartstein & Rothbart, 2003). EC also does not include the EF components of working memory in the midst of changing task demands. So, while we anticipated that these two constructs would be significantly correlated, our goal was to produce a parent-reported measure that more distinctly captures the unique aspects of EF.

Further, several previous studies have found that parent reports of EF are more strongly related to parent reports of social-emotional functioning compared to task-based assessments of EF (Camerota et al., 2018; Nilsen et al., 2017). To assess for divergent validity of the SEFQ, we also wanted to assess the ability of the SEFQ to uniquely measure EF and not more broadly capture social and emotional skills. In the current study, we used a widely used measure of social-emotional development – the Ages and Stages Questionnaire – Social Emotional (ASQ-SE; Squires et al., 2002).

Measurement invariance

One drawback of parent-reported measures is the potential for biased responses related to parental functioning and parents’ beliefs about child development (e.g., Lohaus et al., 2020). Further, parent-reported measures could function differentially across demographically diverse groups of individuals, such that the measure may mean different things to different participants. It is essential to ensure that a measure evaluates the same construct across groups of children and developmental periods (i.e., measurement invariance, Zedeck, 2014, p. 211) for it to have utility in practice settings to assess normative to abnormal ranges of functioning. Without assuring measurement invariance, any substantive group differences are inextricably confounded with measurement differences and thus one measure may be measuring slightly different things across various groups of participants.

For those reasons, the current study included measurement invariance analyses using Moderated Non-linear Factor Analysis (MNLFA; Bauer, 2017; Gottfredson et al., 2019). This method allowed us to test for and then adjust any race, sex, income, or age-related differences in how families respond to items. This approach aims to reduce biases in assessments that may historically have been interpreted as real differences in the skill rather than cultural differences in the interpretation of items (Reynolds et al., 2021). Measurement invariance may be especially important to consider for questionnaire-based measures, given that task-based measures are thought to be a key method for obtaining less-biased indicators of EF and predictors of later performance in racially and economically diverse communities (Blair & Razza, 2007; Masten et al., 2012; Obradović, 2010; Obradović & Willoughby, 2019).

Measurement invariance proceeds across a series of increasingly stringent assessments including configural, metric, and scalar invariance (Millsap & Olivera-Aguilar, 2012). Configural invariance assures that at a base level the questionnaire indicators all have the same factor structure across age and group. Metric invariance then is assuring the item loadings (e.g., the magnitude of relations between items and the latent construct) are equivalent across group and child age. Finally, scalar invariance assesses whether the intercept (e.g., mean) of each indicator is identical across groups or age. All three types of invariance are assessed in the current study.

Predictive validity

In addition to evaluating concurrent validity, divergent validity, and measurement invariance, we were also interested in determining how well the measure predicts children’s long-term academic skills. There is significant evidence that EF skills are related to school readiness, successful transitions to formal schooling, and trajectories of academic success (Espy et al., 2004; Kalstabakken et al., 2021; McClelland et al.,2006, 2013; Morrison et al., 2010; St Clair-Thompson & Gathercole, 2006; Willoughby et al., 2017; Zelazo et al., 2016). EF skills support self-regulation and learning that in turn allows children to be flexible in solving problems in academic spaces. EF has been positively linked to both reading and math outcomes across childhood (e.g., Jacob & Parkinson, 2015; Kalstabakken et al., 2021; Zelazo & Carlson, 2020). In this study, we were able to evaluate how well the SEFQ predicted third-grade math and reading assessments. Third-grade benchmark tests of proficiency predict long-term future academic success, including school attendance, engagement, graduation, and college enrollment (Goldhaber et al., 2020; Hernandez, 2011; Lesnick et al., 2010). Thus, preschool assessments that predict third-grade math and reading skills may be a harbinger of important later-life outcomes.

Current study

The goal of the current study was to evaluate the initial psychometric properties of the SEFQ, with four aims. Aim one was to evaluate the factor structure of the rating scale. Based on an exploratory factor analysis on a subsample of the current data (Wenzel et al., 2013; see supplemental material), we expected to find two distinct factors for this scale: one with positively worded items denoting EF skills and a second comprised of reverse-scored items denoting EF challenges. We used confirmatory factor analysis to evaluate how the two expected EF subscales functioned in relation to the CBQ-VSF EC scale. We expected factors to be related but that three separate factors (CBQ-VSF EC, SEFQ-skills, SEFQ-challenges) would demonstrate good item loadings and acceptable model fit.

Aim two was to evaluate whether the new parent-reported measure of EF demonstrated measurement invariance across sex, child age, race, and estimated household income. Given the novelty of this questionnaire, it is possible that items may function differently across diverse participants. Meeting measurement invariance criteria is rare among the majority of questionnaires administered to diverse populations (Bauer, 2017; Willoughby et al., 2012). However, we did not have specific expectations about which items would vary. We also expected there to be mean-level differences at the latent level across potential moderators of interest, consistent with past evidence that parent- and self-reported EF scores are associated with age (older children score higher; Willoughby et al., 2012), race (minoritized racial groups score lower; Nesbitt et al., 2013), income (low-income children score lower; Nesbitt et al., 2013), and in the case of parent reports of EF, the sex of the child (male children score lower; Gaillard et al., 2021; Garone et al., 2016; Gioia et al., 2002).

Aim three was to evaluate the convergent and divergent validity of these scales with concurrent task-based measures of EF and parent-reported questionnaires. We predicted that the EF skills subscale would be more strongly related to task-based EF measures, compared to the EF challenges subscale. This prediction is based on previous findings in the REEF that questionnaire scales involving skill-based languages that focus on normative EF skills are more strongly related to task-based measures of EF compared to more challenge-focused assessments like the BRIEF-P (Nilsen et al., 2017). We also predicted that the EF challenges subscale would be more strongly related to the Ages and Stages Questionnaire – Social Emotional (ASQ-SE) compared to the EF skills subscale due to the aforementioned finding that negatively worded items were more strongly related to social-emotional challenges (McAuley et al., 2010; Nilsen et al., 2017).

Aim four of this study was to evaluate how well the parent-reported indicators of EF predicted third-grade math and reading achievement test scores. We expected that the SEFQ would predict third-grade math and reading achievement. We then assessed if the subscales had unique predictive value above the typically used developmental readiness assessment. These analyses were exploratory. There is often significant overlap in constructs measured by task-based measures of developmental readiness and EF skills, and it was unclear how well a parent-reported questionnaire would function while modeled with a well-validated task-based measure of developmental readiness.

Methods

Participants

This study draws on data from a sample of 606 families in Minneapolis, Minnesota, in the United States, who agreed to participate in a study during their early childhood screening with the Minneapolis Public School (MPS) district. At the time of screening, children who participated were 36–75 months old (M = 54.90, SD = 9.72). Children were screened at multiple community-based sites to ensure the sample was representative of children in MPS. Participants for the current study were restricted to 380 children who received all measures in English, self-identified their race as Black (n = 169; 43%) or White (n = 211; 57%), and were 3–5 years old. There were not enough individuals who spoke other languages (Spanish n = 49; Hmong n = 36; Somali n = 23), identified as another race (Asian n = 59; American Indian/Alaska Native n = 37; Multiracial n = 55), or were 6 years old (n = 4) to meaningfully test for variation in the assessment instruments across those dimensions. Children in the final sample were 54.12 months old on average (SD = 9.71), and 53.94% were female (n = 205). The mean income based on census tract data was $53,249 (SD = $34,834) annually. According to the United States Census Bureau (n.d.), the median household income in the United States during 2013 was comparable at $52,320. Additionally, the current study aimed to over-sample children experiencing homelessness, and 10% (n = 39) of the sample were screened while staying in emergency homeless shelters.

Procedures

Data were collected in July and August in 2012 and 2013. Families were asked to participate in the study when they arrived for their early childhood screening, and 91% of the eligible families consented. Parents also consented for researchers to follow children administratively through third grade. Families completed early childhood screening as usual, and then assessors for the district administered the Peg Tapping task (i.e., a brief table-top measure of EF) (Diamond & Taylor, 1996). During this time, parents completed a series of questionnaires about their child. Children then completed additional computerized EF measures administered by research staff. Following screening, children’s unique education ID was used to merge their screening data with state-wide school records through third grade with the youngest participants reaching third grade during the 2018/2019 school year, pre-pandemic.

Measures

Demographics

Parents reported on their family’s demographic information. Sex was coded as 1 for male and 0 for female. Race was noted as 1 for Black children and 0 for White children. An income variable was created based on participant addresses using Geocoding software (Social Explorer Professional, 2014) because the school district does not request information about income or parent education during screening. Based on census block data, the average household income was estimated for each housed family. We did not want to exclude children who were experiencing homelessness, so we coded family income for all analyses as one standard deviation above the mean or more (n = 100), within the mean range (n = 159), or one standard deviation below the mean or less (n = 100). Families experiencing homelessness were included in the one standard deviation below the mean group. In models predicting third-grade outcomes, concurrent free or reduced lunch status (coded as 1) was included as a control variable.

Short executive functioning scale (SEFQ)

The SEFQ was developed in a formative project focused on EF measures for young children that was part of preliminary work for the later discontinued National Children’s Study (Masten et al., 2011). The goal was to develop a brief parent-reported measure of the EF that could function as an add-on to the Very Short Form of the Children’s Behavior Questionnaire (CBQ-VSF; Putnam & Rothbart, 2006) or as a stand-alone measure of EF. The SEFQ items were developed by a team with expertise in EF development and assessment of young children. The item pool was intended to include a normative range of every day, easy-to-understand behaviors reflecting both “cool” and “hot” EF skills that would be user-friendly for parents from diverse backgrounds. This scale ultimately included 14 items designed to address the following domains of development: attention focusing, attention and behavioral shifting when new rules are presented, inhibitory control, delay of gratification, emotional control, and working memory. During scale development, prior to the present study, items were piloted for usability with a small sample of nine parents of children attending preschools for disadvantaged families.

In the present study, administration of the SEFQ items directly followed a subset of six items from the effortful control scale of the CBQ-VSF (described below). Parents rated their children from 1 (extremely untrue of your child) to 7 (extremely true of your child). The SEFQ (14 items) had an alpha of 0.79, suggesting good internal reliability. Items are displayed in Table 1. Raw composites of the full scale as well as two subscales were also created by averaging items together. The full questionnaire with instructions is available in Supplemental Materials.

Table 1.

Short executive functioning questionnaire items.

Items Subscale Dropped Items

1 Can save candy or treats for later EF Skills Dropped
2 Has difficulty waiting in line for something EF Challenges
3 Can change from one activity to another when it is time to do so EF Skills
4 Gives up or quits a task when it becomes boring or difficult EF Challenges
5 Can take turns in a game even when excited EF Skills
6 Cannot wait to open presents EF Challenges Dropped
7 Is good at memory games EF Skills
8 When excited, s/he gets out of control EF Challenges
9 Can switch roles during games or pretend play. EF Skills
10 Listens without interrupting EF Skills
11 Will participate in an activity s/he does not like EF Skills
12 Needs reminding of the rules when playing a game. EF Challenges
13 When angry or frustrated, can keep emotions under control EF Skills
14 Has trouble sitting still when s/he is told to (at movies, church, etc) EF Challenges

Child behavior questionnaire – very short form (CBQ-VSF)

In the current study, we were only able to include six items from the CBQ-VSF EC factor (Putnam & Rothbart, 2006) because our community partners requested that we limit the number of questions added to the screening protocol. We chose three items in the attention-focusing subscale and three items in the inhibitory control subscale. Because we wished to include inversely worded items, one item was changed slightly from the CBQ-VSF. Item six is the reverse of “Is good at following instructions,” and was changed to “Has a hard time following instructions.” These items were selected because they were expected to have the most overlap with EF questions. These six items had an alpha of 0.71, suggesting adequate internal consistency.

Direct task-based executive functioning tasks

Three EF tasks were administered to children, including one tabletop task and two computerized tasks from the NIH Toolbox with developmental extensions.

Peg tapping.

Peg Tapping (Diamond & Taylor, 1996) is a table-top measure of EF. The task assesses a child’s ability to inhibit the natural tendency to directly copy an administrator’s actions and instead follow the opposite rule. An administrator uses a wooden dowel to tap one or two times on a table. Across 16 counterbalanced trials, children were told to tap the dowel once when the administrator tapped the dowel twice and to tap the dowel twice when the administrator tapped the dowel once. This task works well with children experiencing economic disadvantage (e.g., Blair & Razza, 2007) and homelessness (e.g., Masten et al., 2012; Obradović, 2010). This measure had an alpha of 0.93, suggesting good internal consistency.

Flanker inhibitory control and attention task.

The National Institutes of Health (NIH) Toolbox Flanker with developmental extension (Flanker-Dext; Anderson et al., 2021) provides a computer-based measure of attentional control and behavioral inhibition (Rueda et al., 2004). Children were presented with five fish and told to tap an arrow key that matched the way the middle fish was pointing. Flanking fish either pointed in the same or opposite direction as the middle fish. The developmental extension lowers the measurement floor of the traditional measure by providing an easier version of the task if the child cannot pass practice trials. Easier versions explain the concept of the middle, enlarge the center fish, space out the fish, or change the flanker fish colors. Scores could range from −5 to 10, where scores below 0 represent the range added by the developmental extension, and scores above 5 reflect reaction time for those who were highly accurate. Flanker-Dext has good test–retest reliability (Distefano et al., 2021), concurrent validity with other measures of EF (Kalstabakken et al., 2021), and predictive validity for third-grade academic benchmark assessments (Distefano et al., 2023).

Dimensional change card sort (DCCS).

The NIH Toolbox DCCS with developmental extension is a computer-based measure of cognitive flexibility and control (Carlson et al., 2021) that was based on an earlier version of the Minnesota Executive Function Scale (MEFS; Beck et al., 2011; Carlson & Zelazo, 2014). Children were presented with stimuli that varied by shape and color. Children began with the original version of the NIH Toolbox DCCS, where they sorted pictures by shape and then were asked to sort the same pictures by color. If they passed this round, they went on to sort pictures by more frequently alternating rules. If children failed the initial rounds of the task, they were presented with easier levels of sorting according to rule complexity. Scores could range from −5 to 10. The DCCS-Dext has good test–retest reliability (Distefano et al., 2021), concurrent validity (Kalstabakken et al., 2021), and predictive validity for third-grade academic benchmarks (Distefano et al., 2023).

Ages and stages questionnaire – social emotional (ASQ-SE)

As part of this district’s routine screening battery, primary caregivers reported on their children’s social and emotional skills via the ASQ-SE (Squires et al., 2002). This is a widely utilized and validated tool aimed at identifying children with social-emotional challenges. It covers many skills, including self-regulation (i.e., the ability to adjust to new environmental conditions), compliance with directives, ability to initiate communication, adaptive functioning (e.g., basic skills like sleeping and eating), ability to respond to situations without guidance, affect expression, and general social skills. It demonstrated high internal consistency in our sample (α = 0.86). Parents completed either the 36-month, 48-month, or 60-month version of the questionnaire based on their child’s age; these questionnaires varied in length from 34 to 36 items. Following Squires et al. (2002), each item was scored a 0, 5, or 10 based on the frequency of the behavior (more frequent behaviors scored higher), and 5 points were added if parents indicated that this was an area of concern. Therefore, higher scores index worse behavior ratings. To best align with the use of the tool in clinical and educational settings, we used cut scores that denoted clinically concerning social-emotional issues for each form. A score of 1 indicated that children were exhibiting heightened social-emotional challenges.

Minneapolis preschool screening instrument – revised (MPSI-R)

The Minneapolis Public School district uses a developmental screener developed by the district and subsequently revised. The MPSI-R was designed to measure multiple aspects of early childhood development for the purpose of screening prior to kindergarten, assessing motor skills, cognitive skills, language, and literacy. The MPSI-R has demonstrated good reliability, construct validity, and predictive validity (Minneapolis Public Schools, 2007, 2016), and is widely used by other districts across the state.

Minnesota comprehensive assessments (MCA)

The MCAs are Minnesota’s adaptive tests in reading, mathematics, and science that are used to meet federal and state legislative requirements for education. Children take these assessments beginning in third grade. The current study used third-grade assessments of reading and mathematics. Children in this sample were tested in third grade from the 2015–2016 academic year through the 2018–2019 academic year (before the COVID-19 pandemic). This span of school years used the same version of the MCAs. Scores used in the analyses were t-scores.

Missing data

The proportions of missing data from screening sessions were minimal, ranging from 0% (i.e., child age) to 14% (i.e., Flanker-Dext). Third-grade data were available for 79% of the sample. Children with third-grade data did not differ from children without third-grade data on measures of EF or academic skills, permitting the assumption that the data were missing at random. In order to address all missing data, we used multiple imputation with the R program MICE to generate parameter estimates across 20 data sets using fully conditional specifications (Van Buuren & Groothuis-Oudshoorn, 2011). We pooled the data using Rubin’s rules to account for within and between imputation variability (Rubin, 2004; Schafer & Graham, 2002). Results using listwise deletion were similar. Results using imputed data are reported.

Planned analyses

Confirmatory factor analysis and configural invariance

We first applied confirmatory factor analysis on the SEFQ using the package lavaan (Rosseel, 2012) in R version 3.6.1 (R Core Team, 2021) independently and then with the CBQ-VSF EC questions loaded on their own EC factor. We then tested for configural invariance of the SEFQ across child age at screening, sex, income, and race. This means we were evaluating if all indicators loaded in the same direction and if the model fit was similar across all sub-groups of children. We constrained all latent factor means to 0 and variances to 1. All other parameters were freely estimated. We consulted fit indices, parameter estimates, and modification indices for evidence of heterogeneous factor structure. If modification indices and theoretical analysis indicated that two items should covary, we averaged those two items together in future analyses. This is due to limitations of MNLFA in being able to accommodate item covariances.

Automated moderated nonlinear factor analysis

To test for and subsequently adjust for any measurement non-invariance across demographic groups, we conducted two MNLFA models (Figures 1 and 2) using the automated MNLFA R package version 1.1.0 (Cole et al., 2021; Gottfredson et al., 2019). We tested for differences in factor mean, item intercepts, and item factor loadings as a function of race, sex, family income, and age at screening. We regressed latent factors, variances, and item-level indicators onto the covariates simultaneously. This allows us to account for true variation at the latent level that is not a result of DIF directly. First, in an initial pass, we tested for differential item functioning (DIF) for indicator loadings and intercepts individually for each item, with the remaining indicators constrained to invariance (i.e., no DIF). We then tested for DIF simultaneously across all indicators. Given multiple testing, we used a Benjamini–Hochberg correction to adjust for inflated type I error rates. All models were fitted to the data using a robust maximum likelihood estimator and Monte Carlo integration. A final model was generated by retaining all significant covariate effects on the factors (mean and variance) and items (threshold and loadings). Factor scores were then extracted from these final models and used in correlation, partial correlation, and regression analyses.

Figure 1.

Figure 1.

Structural model of the final results of the aMNLFA analyses for the Short Executive Functioning Questionnaire (SEFQ) items on the EF skills subscale. Items 7 and 11 were averaged due to significant co-variation.

Figure 2.

Figure 2.

Structural model of the final results of the aMNLFA analyses for included Short Executive Functioning Questionnaire (SEFQ) items on the EF challenges subscale. R on items denotes reverse scored items. Higher scores indicate fewer challenges.

Concurrent and longitudinal validity

We then evaluated concurrent associations between extracted factor EF questionnaire scores and task-based measures via correlations, and partial correlations accounting for child age at screening and any other moderating demographic variables that were present in the aMNLFA analyses (e.g., race). We also analyzed correlational associations with unadjusted total EF questionnaire scores. For longitudinal predictive utility, we used a series of hierarchical linear regression models to evaluate if the questionnaire-based EF factor scores predicted third-grade academic skills, and if they had unique predictive utility in addition to the typically used developmental readiness assessment (i.e., the MPSI-R) and direct measures of EF. We also explored how the MNLFA-adjusted EF questionnaire scores functioned in comparison to unadjusted scores.

Results

Descriptive statistics for all measures are presented in Table 2. An examination of means and standard deviations indicates that the adjusted SEFQ skills and challenges subscales and the unadjusted total scores are normally distributed.

Table 2.

Correlations with concurrent indicators of Executive functioning, developmental readiness, and third-grade academic achievement.

Bivariate Correlations
Bivariate Correlations with Age Partialed Out
Mean (SD) [n] Range [%] 1. Skills Adj 2. Chal. Adj. 3. Total Unadj. 4. Skills Unadj. 5. Chal Unadj. 1. Skills Adj. 2. Chal Adj. 3. Total Unadj. 4. Skills Unadj. 5. Chal. Unadj

1. SEFQ Skills Adjusted 0.31 (0.05) −3.51–2.39
2. SEFQ Challenges Adjusted 0.20 (0.05) −2.23–2.20 0.43*** 0.47***
3. SEFQ Unadjusted Total 4.59 (0.49) 1.00–7.00 0.78*** 0.84*** 0.80*** 0.85***
4. SEFQ Skills Unadjusted 5.14 (0.84) 1.00–7.00 0.96*** 0.42*** 0.37*** 0.94*** 0.49*** 0.85***
5. SEFQ Challenges Unadjusted 4.04 (1.12) 1.00–7.00 0.49*** 0.95*** 0.20*** 0.46*** 0.43*** 0.95*** 0.86*** 0.47***
5. Peg Tapping 9.06 (0.30) 0.00–16.00 0.17** 0.20** 0.21*** 0.21*** 0.19*** 0.10 0.20** 0.19** 0.14* 0.19***
6. DCCS-DEXT 0.77 (0.19) −5.00–9.00 0.13 0.27*** 0.23*** 0.23*** 0.23*** 0.09 0.29*** 0.24*** 0.16** 0.25***
7. Flanker-DEXT 1.77 (0.17) −4.00–8.00 0.17* 0.22*** 0.22*** 0.20*** 0.18** 0.05 0.21*** 0.17** 0.10 0.19**
8. CBQ Effortful Control 5.62 (0.87) 2.00–7.00 0.62*** 0.44*** 0.54*** 0.55*** 0.17** 0.56*** 0.42*** 0.57*** 0.58*** 0.39***
9. ASQ-SE Cut Score [53] [14.1%] −0.49*** −0.50*** −0.58*** −0.52*** −0.47** −0.39*** −0.47*** −0.50*** −0.41*** −0.45***
10. MPSI-R 45.87 (0.81) 2.00–64.00 0.18** 0.22*** 0.22*** 0.37*** 0.20*** 0.10 0.22*** 0.20*** 0.14* 0.20***
11. MCA Math 51.92 (18.22) 15–92 0.04 0.17** 0.16* 0.18** 0.26*** 0.14* 0.26*** 0.26*** 0.19** 0.26***
12. MCA Reading 47.73 (25.98) 1–99 0.11 0.23*** 0.24*** 0.13* 0.21*** 0.09 0.19** 0.19*** 0.14* 0.19***
*

p<.05

**

p<.01

***

p<.001

SEFQ = Shore Executive Functioning Questionnaire; DCCS = Dimensional Change Card Sort; CBQ = Child Behavior Questionnaire; ASQ-SE: Ages and Stages Questionnaire – Social Emotional; MPSI-R: Minneapolis Preschool Screening Instrument – Revised; MCA: Minnesota Comprehensive Assessment; Chal.: SEFQ Challenges Scale; Skills: SEFQ Skills Scale; Adj.: Adjusted; Unadj: Unadjusted; SEFQ skills and challenges are the extracted factor scores from MNLFA analyses.

Confirmatory factor analysis and configural invariance

A previous study (Wenzel et al., 2013) conducted an exploratory factor analysis with approximately half of the data, finding that two factors best described the data. One factor was made up of positively worded items indicative of a child’s EF skills and the other was made up of negatively worded items indicative of a child’s EF challenges (Table 1). Negatively worded items were reverse-scored so that higher scores indicated better EF (i.e., fewer EF challenges). Given these exploratory factor analysis findings, first, we conducted a preliminary CFA with the two predicted subscales on the SEFQ as well as the six items included from the CBQ-VSF EC. This in part was done to assure separable latent constructs of both SEFQ subscales from EC, given that items were administered concurrently. Results suggested modest model fit (CFI = 0.84; TLI = 0.82; RMSEA = 0.07 [0.06–0.08], p < .001; SRMR = 0.06). EC was moderately correlated with EF skills (r = 0.36, p < .001) and with EF challenges (r = 0.28, p < .001). EF skills and challenges were moderately correlated with each other (r = 0.39, p < .001). Given that the model fit was not ideal, we evaluated modification indices, which suggested that a reverse-scored item (“Has a hard time following instructions”) on the EC scale belonged to items on the EF challenges scale. When this item was allowed to load on that scale, the model fit was acceptable (CFI = 0.89; TLI = 0.88; RMSEA = 0.06 [0.05–0.07], p = .13; SRMR = 0.05). Given the closeness of model fit and the long-standing psychometric work on the CBQ-VSF (Backer-Grøndahl et al., 2016; de la Osa et al., 2014), we dropped the added reverse scored item for EC and proceeded with the 5-item version of the EC scale in all subsequent analyses. The alpha for the 5-item version of the EC is 0.73. Largely, the results suggested separate but related constructs across EC, EF skills, and EF challenges.

Building on those findings and measurement structure, we next conducted separate CFAs for EF skills and challenges. Modification indices and poor model fit in the initial CFA of EF skills (positively worded items) suggested that two items had additional covariance beyond what was commonly explained by a latent factor. One was an item noting a child’s ability to perform well at memory games, and the other was an item indicating that a child will participate in games they do not like. Given the limitations of including covariation in aMNLFA, we averaged the two items’ scores. Although not ideal, it provides the best option to allow both variables to be included in latent factor estimation. Additionally, we evaluated configural invariance by child age at screening, sex, income, and race. An item asking if children can save treats for later had a standardized loading lower than 0.30 for children with income one standard deviation below the mean. This item was functioning differently for that group and so it was dropped from future analyses. The final model for positively worded items included seven items, demonstrated acceptable model fit (CFI = 0.97; TLI = 0.95; RMSEA = 0.07 [0.04–0.10], p = .16; SRMR = 0.03) and showed configural invariance across demographic groups. The alpha of the seven unadjusted items was 0.74.

In the model for EF challenges, items were reverse-scored and entered into a CFA. The items initially showed adequate model fit (CFI = 0.99; TLI = 0.99; RMSEA = 0.03 [0.00–0.07], p = .73; SRMR = 0.02). However, one item referring to a child’s ability to wait to open presents had a low standardized loading (<0.30) for 4-year-old children and male children, so it was dropped from the scale. The final model for negatively worded items included five items and demonstrated acceptable model fit (CFI = 0.99; TLI = 0.98; RMSEA = 0.57 [0.00–0.10], p = .46; SRMR = 0.02), with configural invariance across demographic groups. The alpha for the five unadjusted items was 0.75.

Measurement invariance – aMNLFA

Executive functioning skills subscale

Significant findings of non-invariant items and the impact of demographic variables on factor means for items on the EF skills subscale are shown in Figure 1. Observed DIF effects were on the predicted value of the item when the latent variable is zero (i.e., intercept) and on the predicted change in items associated with a one-unit shift in the latent variable (i.e., loadings). Therefore, even after controlling for the level of the latent variable, four-item means and loadings still differed by race and child age at screening.

When the latent variable was at zero (i.e., the mean), older children compared to younger children had a higher average score on an item about the child’s ability to take turns in a game when they are excited, and they had a lower average score on an item about listening without interrupting. This means that older children were systematically rated as having the ability to take turns on games, and systematically rated as having less developed listening skills. Black children compared to White children had higher average scores on three items: taking turns in a game when they are excited, being good at memory games, and doing activities the child does not like.

There were also significant differences in how much certain items contributed to the overall latent factor of EF skills across child age and race. Older children showed less change in an item about taking turns when there was a one-unit shift in the latent EF skills score. This means that this item provided less information to the latent factor for older children. When there was a one-unit shift in the latent variable, on average, Black children had smaller changes on items about being good at memory games and participating in activities they do not like. There were no mean level differences in the latent EF skills factor for Black compared to White children. However, older children were more likely to have higher EF skills scores (B = 0.27, SE = 0.08, p < .01).

Executive functioning challenges subscale

Significant findings of non-invariant items and the impact of demographic variables on the EF challenges factor mean are depicted in Figure 2. Items on the EF challenges subscale were reverse-scored so that higher scores denoted fewer challenges. Observed DIF effects were observed on both item intercepts and loadings for two items differing by child race. Black children were rated as having more challenges on an item about being able to stick with a task when it was hard compared to White children. This item and an item about waiting in line also contributed less information to the latent factor for Black children compared to White children. At the mean latent level, Black children had higher scores – indicating fewer EF challenges – compared to White children (B = 0.35, SE = 0.13, p < .01).

Concurrent associations

We evaluated how MNLFA-adjusted and non-adjusted parent-reported scores on the SEFQ were related to task-based measures of EF, developmental readiness screening, and a parent report of social-emotional functioning. Bivariate correlations and partial correlations when accounting for age are located in Table 2. Correlations among all study variables are provided in Supplemental Materials. We also evaluated partial correlations between concurrent measures when controlling for race, but they did not significantly differ from initial bivariate correlations. Results suggest that the SEFQ was more strongly related to other parent-reported measures of social-emotional challenges (ASQ-SE) and CBQ-VSF EC than to direct measures of EF, with correlations ranging from 0.44 to 0.62 in bivariate correlations and a 0.39–0.57 in correlations adjusted for age. This was counter to initial hypotheses, although follows a similar pattern to the BRIEF-P measure of EF. Further, the ASQ-SE and CBQ-VSF EC were moderately correlated with task-based measures of EF. This strength of the relationship is comparable to the strength of the relationship between the SEFQ scales and task-based measures of EF.

When adjusting for age, results suggested that the EF challenges subscale was related to performance on task-based assessments of EF. After controlling for age, EF skills were unrelated to task-based measures of EF. We evaluated whether there were any differences between the correlations between task-based measures and the two parent-reported EF subscales using the Meng et al. (1992) transformation of a Fisher’s z test. After adjusting for child age, results suggested significantly stronger relations between the EF challenges subscale and DCCS-Dext (z = −3.89, p < .001), Flanker-Dext (z = −3.06, p < .01), the MPSI-R developmental readiness (z = −2.30, p < .05), and EC (z = 3.18, p < .01), compared to the parallel relations with the EF skills subscale. The correlation did not differ for the relation between EF challenges and EF skills subscales with Peg tapping nor with the ASQ-SE. Results also indicated that the unadjusted EF total score as well as unadjusted EF challenges scales had similar relations as the EF challenges subscale with all concurrent variables.

Longitudinal predictive utility

We next compared adjusted and unadjusted EF scores as predictors of third-grade math and reading achievement. We found that both the aMNFLA-adjusted EF skills and EF challenges subscales (Table 3) as well as unadjusted EF skills, unadjusted EF challenges, and unadjusted EF total score (Table 5) of the SEFQ predicted third-grade math achievement, while controlling for child race, sex, free or reduced lunch status, and age at screening. When we included both the EF skills and challenges subscales in the same model, neither was significant; however, there was a positive significant increase in the amount of variance explained. This suggests that it was the shared variance between the SEFQ EF skills and challenges that predicted third-grade math achievement. Neither subscale of the SEFQ predicted third-grade reading achievement.

Table 3.

Longitudinal predictive utility of the SEFQ-adjusted factor scores to third-grade math.

MCA 3rd Grade Math
B (SE) t p R 2 ΔR 2

Block 1 .36
 Intercept 57.08 (4.69) 12.17 <.001
 Age at Screening 0.10 (0.08) 1.15 .25
 Black −9.78 (2.30) −4.25 <.001
 Male 0.87 (1.66) 0.52 .60
 Free or Reduced Lunch −14.43 (2.43) −5.95 <.001
Block 2a .39 .03***
 Intercept 60.14 (4.71) 12.77 <.001
 Age at Screening 0.01 (0.09) 0.14 .89
 Black −10.15 (2.27) −4.48 <.001
 Male 1.50 (1.67) 0.90 .37
 Free or Reduced Lunch −13.58 (2.43) −5.58 <.001
 SEFQ Skills Adjusted 3.38 (0.91) 3.71 <.01
Block 2b .38 .02***
 Intercept 57.56 (4.62) 12.46 <.001
 Age at Screening 0.06 (0.08) 0.75 .45
 Black −9.38 (2.26) −4.14 <.001
 Male 1.08 (1.64) 0.66 .51
 Free or Reduced Lunch −13.42 (2.43) −5.53 <.001
 SEFQ Challenges Adjusted 3.16 (0.92) 3.43 <.001
*

p < .05

**

p < .01

***

p < .001

SEFQ = Short Executive Functioning Questionnaire; SEFQ adjusted scores are the extracted factor scores from MNLFA analyses.

Table 5.

Longitudinal predictive utility of the SEFQ-unadjusted average scores to third-grade math.

MCA 3rd Grade Math
B (SE) t p R2 ΔR2

Block 1 .36
 Intercept 57.08 (4.69) 12.17 <.001
 Age at Screening 0.10 (0.08) 1.15 .25
 Black −9.78 (2.30) −4.25 <.001
 Male 0.87 (1.66) 0.52 .60
 Free or Reduced Lunch −14.43 (2.43) −5.95 <.001
Block 2c .38 .02***
 Intercept 45.27 (5.74) 7.88 <.001
 Age at Screening 0.02 (0.09) 0.29 .77
 Black −9.46 (2.28) −4.15 <.001
 Male 1.49 (1.68) 0.89 .38
 Free or Reduced Lunch −13.79 (2.44) −5.66 <.001
 SEFQ Skills Unadjusted 2.96 (0.86) 3.42 <.01
Block 2d .38 .02***
 Intercept 47.28 (5.40) 8.75 <.001
 Age at Screening 0.07 (0.08) 0.83 .41
 Black −9.56 (2.27) −4.21 <.001
 Male 0.96 (1.64) 0.58 .56
 Free or Reduced Lunch −13.36 (2.42) −5.51 <.001
 SEFQ Challenges Unadjusted 2.67 (0.74) 3.59 <.001
Block 2e .39 .03***
 Intercept 43.33 (5.73) 7.56 <.001
 Age at Screening 0.03 (0.08) 0.31 .75
 Black −9.40 (2.26) −4.16 <.001
 Male 1.37 (1.65) 0.83 .41
 Free or Reduced Lunch −13.31 (2.43) −5.48 <.001
 SEFQ Total Unadjusted 3.61 (0.89) 4.07 <.001
*

p < .05

**

p < .01

***

p < .001

SEFQ = Short Executive Functioning Questionnaire.

Finally, we evaluated if parent-reported EF scores explained any additional variance in third-grade math achievement, above typically used developmental readiness assessments (Tables 4 and 6). Results suggested that the adjusted EF challenges subscale, adjusted EF skills subscale, unadjusted EF challenges subscale and unadjusted EF total score had additional predictive utility above the MPSI-R. Each explained approximately 1% of the variance of math scores in third grade when controlling for the MSPI-R.

Table 4.

Longitudinal predictive utility of the SEFQ-adjusted factor scores to third-grade math, controlling for developmental readiness scores.

MCA 3rd Grade Math
B (SE) t p R 2 ΔR 2

Block 1 .45
 Intercept 61.18 (4.47) 13.70 <.001
 Age at Screening −0.48 (0.12) −3.82 <.001
 Black −6.35 (2.18) −2.91 <.01
 Male 1.62 (1.61) 1.01 .32
 Free or Reduced Lunch −11.55 (2.38) −4.85 <.001
 MSPI-R 0.52 (0.08) 6.20 <.001
Block 2a .46 .01*
 Intercept 62.67 (4.50) 13.92 <.001
 Age at Screening −0.48 (0.13) −3.84 <.001
 Black −6.85 (2.17) −3.16 <.05
 Male 1.93 (1.63) 1.18 .24
 Free or Reduced Lunch −11.28 (2.39) −4.71 <.001
 MSPI-R 0.48 (0.09) 5.64 <.001
 SEFQ Skills Adjusted 2.01 (0.88) 2.27 <.05
Block 2b
 Intercept 61.28 (4.43) 13.83 <.001 .46 .01*
 Age at Screening −0.47 (0.13) −3.75 <.001
 Black −6.26 (2.16) −2.89 <.01
 Male 1.72 (1.59) 1.08 .28
 Free or Reduced Lunch −10.99 (2.36) −4.66 <.001
 MPSI-R 0.49 (0.09) 5.67 <.001
 SEFQ Challenges Adjusted 2.26 (0.91) 2.48 <.05
*

p < .05

SEFQ = Short Executive Functioning Questionnaire; MPSI-R = Minneapolis Preschool Screening Instrument – Revised; SEFQ adjusted scores are the extracted factor scores from MNLFA analyses.

Table 6.

Longitudinal predictive utility of the SEFQ-unadjusted average scores to third-grade math, controlling for developmental readiness scores.

MCA 3rd Grade Math
B (SE) t p R 2 ΔR 2

Block 1 .45
 Intercept 61.18 (4.47) 13.70 <.001
 Age at Screening −0.48 (0.12) −3.82 <.001
 Black −6.35 (2.18) −2.91 <.01
 Male 1.62 (1.61) 1.01 .32
 Free or Reduced Lunch −11.55 (2.38) −4.85 <.001
 MSPI-R 0.52 (0.08) 6.20 <.001
Block 2c .46 .01
 Intercept 54.55 (5.62) 9.70 <.001
 Age at Screening −0.48 (0.13) −3.80 <.001
 Black −6.42 (2.17) −2.96 <.01
 Male 1.90 (1.64) 1.16 .25
 Free or Reduced Lunch −11.41 (2.39) −4.77 <.001
 MPSI-R 0.49 (0.08) 5.72 <.001
 SEFQ Skills Unadjusted 1.59 (0.83) 1.91 .06
Block 2d .46 .01**
 Intercept 53.63 (5.38) 9.97 <.001
 Age at Screening −0.47 (0.13) −3.73 <.001
 Black −6.37 (2.17) −2.94 <.01
 Male 1.64 (1.59) 1.03 .30
 Free or Reduced Lunch −10.91 (2.35) −4.63 <.001
 MSPI-R 0.49 (0.08) 5.72 <.001
 SEFQ Challenges Unadjusted 1.99 (0.73) 2.71 <.01
Block 2e .46 .01*
 Intercept 52.32 (5.71) 9.16 <.001
 Age at Screening −0.47 (0.13) −3.76 <.001
 Black −6.42 (2.17) −2.96 <.01
 Male 1.86 (1.61) 1.16 .25
 Free or Reduced Lunch −11.12 (2.38) −4.67 <.001
 MSPI-R 0.48 (0.09) 5.49 <.001
 SEFQ Total Unadjusted 2.23 (0.87) 2.57 <.05
*

p < .05

**

p < .01

SEFQ = Short Executive Functioning Questionnaire; MPSI-R = Minneapolis Preschool Screening Instrument – Revised.

Discussion

This study provides initial evidence of validity for the SEFQ scale in measuring children’s EF skills in early childhood. Although the original intention of the 14-item SEFQ was to broadly index EF, a two-factor solution best fit the data, with one scale noting EF skills (7 items) and the other EF challenges (5 items). There was evidence for construct validity: both the EF challenges subscale and the full unadjusted SEFQ scores were significantly associated with task-based measures of EF, even after accounting for children’s age. In contrast, the EF skills subscale was only correlated with Peg Tapping and Flanker-Dext at the bivariate level, and these relations were non-significant after controlling for children’s age. A similar pattern emerged for the long-term predictive validity of the SEFQ. The EF skills subscale and the EF challenges subscale, as well as the unadjusted SEFQ total mean and subscale scores (Tables 5 and 6), predicted third-grade math achievement. Only the unadjusted EF skills subscale did not predict additional variance in third-grade achievement when also controlling for a measure of developmental readiness typically used in screening.

Contrary to hypotheses, the EF challenges subscale (reverse-scored) – not the EF skills – was significantly associated with all three of the task-based EF measures, as well as math achievement in third grade. Parents may be more attuned and better able to report on child behaviors that are disruptive or challenging to manage compared to more neutral behaviors or skills, such as being good at memory games. Indeed, items on the EF skills subscale demonstrated slightly higher means and lower variability compared to the EF challenges subscale. This means that parents more uniformly rated their child as having these EF skills. A commonly used parent-reported EF measure, the BRIEF-P, also consists of only negatively worded items (e.g., “resists changes in routine”; Gioia et al., 2000). The BRIEF-P also adjusted their written instructions from the BRIEF format for older children to denote how often specific behaviors have been a problem. For parent-reported measures in this age range, using negatively worded items may be more effective in eliciting nuanced and variable ratings because parents may be better attuned to challenge-based questions rather than strengths.

In addition, the SEFQ included items that were intended to assess children’s hot EF skills. Two of these items (i.e., saving treats for later, waiting to open presents) did not have adequate factor loadings, so they were dropped from subsequent analyses. These hot EF items were based on commonly used delay of gratification tasks in the laboratory, but they might not translate well to a questionnaire. In laboratory delay of gratification tasks, for example, an examiner explicitly instructs children to wait, often with the promise of a reward (e.g., Mischel et al., 1989). Thus, it would perhaps be unusual for a preschool-aged child to delay gratification of their own volition. These items could be reworded to specify that children can delay gratification when asked or when there is an additional potential reward. Further, it is possible that these items are not valid for children facing higher levels of deprivation – like many of those in our urban sample – because waiting for a reward is not adaptive in the context of their lives (Duran & Grissmer, 2020). It is also a limitation of the current study that we were unable to include a task-based measure of hot EF like those mentioned above.

Measurement invariance results suggested that some items functioned differently depending on child age and race. Given the rapid development of EF skills during the preschool period (Carlson, 2005), it can be particularly challenging to create a measure that meaningfully captures individual differences in EF across such a wide age range. Overall, we found only a few differences in item functioning in younger versus older children in the EF skill subscale, and no relationship between age and the EF challenges subscale, which offers encouraging support for the use of the SEFQ with preschoolers. These results are promising given how rapidly EF skills are developing and changing in the age range of 3–5 years. It suggests that the EF challenges scale can be used to capture the EF construct consistently across the 3–5-year-old age period.

There were more marked differences in item functioning across child race compared to child age. Overall, parents of Black children systematically rated their children higher on being good at memory games, switching roles during games, and participating in activities they do not like compared to White children. Items about playing games also contributed less information to the total EF skills score for Black children compared to White children. Additionally, parents systematically rated Black children as having more trouble sticking with boring tasks compared to White children. That item, as well as an item about waiting in line for something, also contributed less information to the total EF challenges score for Black children compared to White children. It would be beneficial to gather qualitative data on these specific items with parents from diverse backgrounds to better understand the reasons for this variation. However, in general, the results suggest that the current version has adequate construct validity across the groups assessed. In particular, the EF challenges subscale performed better than the EF skills subscale because it demonstrated fewer item and loading biases based on child demographic differences.

For aim three, we found that the bivariate correlations between the EF challenges subscale and the task-based EF measures ranged from r = .20 to .29. The magnitude of these associations is comparable to the meta-analytic effect found by Toplak et al. (2013) in their comparison of report-based and task-based measures of EF among both children and adults. Divergence of report-based and task-based measures of EF may be due to differences in the conditions under which EF is assessed. Task-based measures tend to be given under optimal conditions, whereas report-based measures reflect typical EF across a host of situations (Gioia et al., 2008). Thus, the negatively worded SEFQ subscale seems to be functioning at least as well as other report-based measures of EF, with the added benefit of being much shorter and working reasonably well in a diverse population. Future work would benefit from assessing how the SEFQ also relates to hot EF task-based measures.

We found limited divergent validity of the SEFQ, and this is similar to many other parent-reported measures of EF including the CHEXI (Camerota et al., 2018). EF skills and challenges demonstrated largely separable factor structures to EC in aim one. However, the partial correlations for aMNFLA-adjusted EF scores (controlling for child age) with EC ranged from r = .39 to .58 and from r = −.39 to −.50 with the ASQ-SE scale (higher scores indicate more social-emotional problems). All of these measures are parent-reported, and parents appear to be reporting quite consistently across them. The SEFQ was also completed after the other two measures, so parents might have been primed to report a certain way after answering many other questions about their child. Further, the ASQ-SE and the EC measure had a similar strength relationship with EF tasks, suggesting that these two constructs are tapping something very similar to the EF parent-reported scales.

Additionally, there is still significant debate about the overlap between EF and EC. One recent study by Kälin and Roebers (2021) found significant, positive correlations between task-based measures of EF and EC and demonstrated that all tasks loaded onto one latent factor during the preschool period. However, other work has pointed to key differences in the constructs, namely the centrality of working memory to EF skills, while EC tended to be more focused on the emotional expressive components of self-regulation (e.g., Blair & Razza, 2007; Wolfe & Bell, 2004). The inclusion of the EC subscale was to evaluate whether the SEFQ was related to but also distinct from EC. The moderate relation between EC and both EF subscales could suggest the SEFQ is tapping constructs more aligned with the conceptual construct of hot EF, while also still capturing distinct variations and constructs. It is also important to note that given the practical limitations of working with a community partner to administer these assessments, we had external limitations on the length of the questionnaire we could administer. Given this limitation, we were not able to include the full CBQ-VSF EC scale and instead chose items that we thought conceptually may be most related to EF skills and items. This included items from the attention-focusing and inhibitory control subscales. As a result, it will be important to continue research evaluating the overlap and distinctiveness of EC and EF concepts and measures in preschool-aged children.

Consistent with the findings on construct validity, the EF challenges subscale also performed somewhat better than the EF skills subscale for long-term predictive validity. The EF challenges subscale explained significantly more variation in third-grade math achievement than the EF skills subscale, even after controlling for children’s demographic characteristics.

A recent study (Distefano et al., 2023) evaluated how the computer task-based EF measures included in the present study predicted academic achievement in an overlapping sample with the current study, and found that both the DCCS-Dext and Flanker-Dext predicted third-grade math and reading. The amount of variation in math scores explained by each alone was 10% for the DCCS-Dext and 12% for the Flanker-Dext. For reading, the DCCS-Dext accounted for 9% of the variation when entered alone, whereas the Flanker-Dext accounted for 8% of the variation. Although each study used a slightly different sample and covariates, these findings generally suggest better predictive utility of computer-based task-based measures of EF compared to parent-reported measures evaluated in the current manuscript, which individually added 2% to the explanation of variation in third-grade math scores when not controlling for MPSI-R scores.

Further, the EF challenges subscale continued to be a significant predictor of math achievement when included in a model with children’s performance on the MPSI-R. However, the EF challenges subscale only explained a limited amount of variation in math achievement when included in the model on its own (3%), whereas the MPSI-R independently explained 10% of the variation before adding EF. Thus, it would not be advisable to give the SEFQ in place of routine performance-based screening measures, but rather it could have value-added utility as a quick and easy questionnaire for parents in settings without well-established screening protocols. The SEFQ was not significantly associated with reading achievement. This result is consistent with previous research that has demonstrated stronger relations between task-based EF measures and math compared to reading (Allan et al., 2014; Ernst et al., 2022).

In general, the SEFQ challenges unadjusted average score appeared to be working well when compared to the adjusted score. Given the minimal evidence of non-invariance and its minimal effect on substantive conclusions drawn from adjusted or non-adjusted models, the SEFQ challenges averaged subscale (averaging only five items) can provide an appropriate even shorter (only 5-item) option for screening EF skills in the preschool age range. We reverse-scored items of the SEFQ challenges subscale to be consistent with the EF skills scale because we wanted the direction of all items to mean the same thing (e.g., higher is a more developed skill) in the initial factor analysis work, as well as for consistent interpretations of the subscales throughout the whole manuscript. Future users may choose not to reverse-score items, especially if using the challenges subscale on its own. If users choose not to reverse-score items, higher scores would reflect more EF challenges.

Additionally, given the initial promise of the SEFQ, there are a number of avenues for additional development and future research. We examined multiple psychometric properties of the SEFQ, but future studies may consider adjusting the scale in light of these findings and then investigating the test–retest reliability, as well as associations with other validated parent-reported measures of EF in preschoolers, such as the REEF (Nilsen et al., 2017). There is also a need to improve some of the items aimed at measuring hot EF skills. A revised version of the SEFQ could also remove or adjust items that indicated bias based on children’s age and race in consultation with parents. For example, parents might have different ideas about what a “memory game” is. However, it should be noted that the unadjusted EF questionnaire performed similarly to the adjusted subscales across concurrent, divergent, and longitudinal validity assessments. This may suggest that any differences in measurement across groups are small and likely do not impact substantive interpretations.

The current study had a number of strengths, including the diverse and relatively large sample of children participating from the general population for early childhood screening in a large midwestern city in the USA, in a state that requires early childhood screening before school entry. Given the high participation rates, the comparability of the sample’s census tract income to the national average income in that time period, and that the scales were normally distributed, we think that the results suggest that this measure will likely work well in both disadvantaged and advantaged families. Additionally, the current study included multiple task-based measures of EF to examine construct validity, and the inclusion of longitudinal administrative data to assess long-term predictive validity. However, there were also several limitations. Given the collaboration with a local school district to collect the data during early childhood screening, there were limits to the number of questionnaire items and tasks we could add to their protocol. We thus were unable to include another lengthy concurrent parent-reported measure of executive functioning such as the BRIEF-P. Future work is needed to compare this scale directly to other parent-reported scales in this age group. Additionally, our analyses of DIF were specific to English-speaking parents of 3–5-year-olds, as well as Black and White children, and thus are likely only generalizable to similar urban populations. There is a need to examine the extent to which the SEFQ works well among other diverse cultural groups. Researchers may also consider directly measuring cultural variation in play or beliefs about parenting and child-rearing rather than relying on nonspecific measures such as race or ethnicity as indicators of cultural variation. Further, because this school district does not ask parents about income for screening purposes, we only had a broad index of income that may have precluded the identification of differences associated with unmeasured disadvantages or adversity exposures. Studies with more direct reports by families of income, specific access to resources, or adverse life experiences might reveal differential effects of disadvantage on the SEFQ. As noted earlier, it also would have been informative to have a task-based measure of “hot” EF. The SEFQ was created to index both “cool” and “hot” EF, but it is still an open question as to whether it is significantly associated with children’s “hot” EF skills measured directly.

Despite these limitations, this study provides initial evidence that the SEFQ, particularly the 5-item SEFQ challenges subscale, offers a brief, valid parent-reported measure of EF for preschool-aged children. The associations of even the unadjusted SEFQ average scores with task-based measures are comparable in magnitude to other well-established EF questionnaires, and significantly predicted math achievement 3.5–6 years later. It is not always feasible for young children to complete large batteries of task-based EF assessments, nor for parents to complete lengthy surveys, so the SEFQ may be a useful alternative. Further, parents can provide unique insights into children’s EF skills during everyday situations, which could be more indicative of how they are likely to behave in a classroom setting filled with competing goals and distractions. With further development and refinement, the SEFQ shows promise in equitably screening for EF skills among diverse young children.

Supplementary Material

Supplementary Material

Acknowledgments

The authors would like to thank the families who participated and the staff at the Minneapolis Public Schools for their immense effort to ensure high-quality, routine, early childhood screening each year. We would also like to acknowledge Dr Mary K. Rothbart for her theoretical and intellectual contribution to the development of the Short Executive Functioning Scale. Dr Rothbart provided feedback on item development as well as study design. Additionally, this study’s design and hypotheses were preregistered; see https://osf.io/sr5cu/files/osfstorage/621800d5b702cd02d6fc4cc1

Funding

This study was supported by a National Science Foundation graduate fellowship (A.W.K.) and the University of Minnesota Interdisciplinary Doctoral Fellowship (A.R.P.). Additional research funding was provided by the Fesler-Lampert Chair in Urban and Regional Affairs (A.S.M), the Irving B. Harris Professorship (A.S.M), The Institute of Child Development Research Grant (A.R.P.), The Howard Diversity Scholars Fund (A.R.P.), and a Lorrain D Eyed Grant from The American Psychological Foundation (A.R.P.). Manuscript preparation was supported in part by the National Institute of Mental Health T32 Postdoctoral Training Grant [MH073517-16] in the science of child mental health treatment (A.R.P.).

Footnotes

Disclosure statement

No potential conflict of interest was reported by the author(s).

Supplemental data for this article can be accessed at https://doi.org/10.1080/09297049.2024.2329435.

References

  1. Akshoomoff N, Newman E, Thompson WK, McCabe C, Bloss CS, Chang L, Amaral DG, Casey BJ, Ernst TM, Frazier JA, Gruen JR, Kaufmann WE, Kenet T, Kennedy DN, Libiger O, Mostofsky S, Murray SS, Sowell ER, Schork N, … Jernigan TL (2014). The NIH toolbox cognition battery: Results from a large normative developmental sample (PING). Neuropsychology, 28(1), 1–10. 10.1037/neu0000001 [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Allan NP, Hume LE, Allan DM, Farrington AL, & Lonigan CJ (2014). Relations between inhibitory control and the development of academic skills in preschool and kindergarten: A meta-analysis. Developmental Psychology, 50(10), 2368–2379. 10.1037/a0037493 [DOI] [PubMed] [Google Scholar]
  3. Anderson JE, Kalstabakken AW, Zelazo PD, Carlson SM, Distefano R, & Masten AS (2021). Technical report on the developmental extension of the NIH toolbox flanker task. E-Prime version. [Google Scholar]
  4. Backer-Grøndahl A, Nærde A, Ulleberg P, & Janson H (2016). Measuring effortful control using the children’s behavior questionnaire–very short form: Modeling matters. Journal of Personality Assessment, 98(1), 100–109. 10.1080/00223891.2015.1056303 [DOI] [PubMed] [Google Scholar]
  5. Barkley RA (2012). Barkley deficits in executive functioning scale–children and adolescents (BDEFS-CA). Guilford Press. [Google Scholar]
  6. Bauer DJ (2017). A more general model for testing measurement invariance and differential item functioning. Psychological Methods, 22(3), 507–526. 10.1037/met0000077 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Beck DM, Schaefer C, Pang K, & Carlson SM (2011). Executive function in preschool children: Test–retest reliability. Journal of Cognition and Development, 12(2), 169–193. 10.1080/15248372.2011.563485 [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Best JR, Miller PH, & Naglieri JA (2011). Relations between executive function and academic achievement from ages 5 to 17 in a large, representative national sample. Learning and Individual Differences, 21(4), 327–336. 10.1016/j.lindif.2011.01.007 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Blair C, & Razza RP (2007). Relating effortful control, executive function, and false belief understanding to emerging math and literacy ability in kindergarten. Child Development, 78(2), 647–680. 10.1111/j.1467-8624.2007.01019.x [DOI] [PubMed] [Google Scholar]
  10. Camerota M, Willoughby MT, Kuhn LJ, & Blair CB (2018). The childhood executive functioning inventory (CHEXI): Factor structure, measurement invariance, and correlates in US preschoolers. Child Neuropsychology, 24(3), 322–337. 10.1080/09297049.2016.1247795 [DOI] [PubMed] [Google Scholar]
  11. Carlson SM (2005). Developmentally sensitive measures of executive function in preschool children. Developmental Neuropsychology, 28(2), 595–616. 10.1207/s15326942dn2802_3 [DOI] [PubMed] [Google Scholar]
  12. Carlson SM, & Zelazo PD (2014). Minnesota executive function scale: Test manual. Reflection Sciences. [Google Scholar]
  13. Carlson SM, Zelazo PD, Anderson JE, Kalstabakken AW, Distefano R, & Masten AS (2021). Technical report on the developmental extension of the NIH toolbox dimensional change card Sort task: E-Prime version. [Google Scholar]
  14. Carlson SM, Zelazo PD, & Faja S (2013). Executive function. In Zelazo PD (Ed.), The oxford handbook of developmental psychology (vol. 1): Body and mind (pp. 706–743). Oxford University Press. [Google Scholar]
  15. Cole VT, Gottfredson NC, Giordano M, & Janssen T (2021). Automated fitting of moderated nonlinear factor analysis (MNLFA) through the Mplus program. R package version 1.0.
  16. de la Osa N, Granero R, Penelo E, Domènech JM, & Ezpeleta L (2014). The short and very short forms of the children’s behavior questionnaire in a community sample of preschoolers. Assessment, 21(4), 463–476. 10.1177/1073191113508809 [DOI] [PubMed] [Google Scholar]
  17. Diamond A (2013). Executive functions. Annual Review of Psychology, 64(1), 135–168. 10.1146/annurev-psych-113011-143750 [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Diamond A, & Taylor C (1996). Development of an aspect of executive control: Development of the abilities to remember what I said and to “do as I say, not as I do.” Developmental Psychobiology, 29(4), 315–334. 10.1002/(SICI)1098-2302(199605)29:4&lt;315:AID-DEV2&gt;3.0.CO;2-T [DOI] [PubMed] [Google Scholar]
  19. Distefano R, Fiat AE, Merrick JS, Slotkin J, Zelazo PD, Carlson SM, & Masten AS (2021). NIH Toolbox executive function measures with developmental extensions: Reliability and validity with preschoolers in emergency housing. Child Neuropsychology, 27(6), 709–717. 10.1080/09297049.2021.1888905 [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Distefano R, Palmer AR, Kalstabakken AW, Hillyer CK, Seiwert MJ, Zelazo PD, Carlson SM, & Masten AM (2023). Predictive validity of NIH Toolbox executive function measures with developmental extensions: Pre-kindergarten screening to third grade benchmark tests of achievement. Developmental Neuropsychology, 48(8), 373–386. 10.1080/87565641.2023.2286353 [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Duran CAK, & Grissmer DW (2020). Choosing immediate over delayed gratification correlates with better school-related outcomes in a sample of children of color from low-income families. Developmental Psychology, 56(6), 1107–1120. 10.1037/dev0000920 [DOI] [PubMed] [Google Scholar]
  22. Ernst JR, Grenell A, & Carlson SM (2022). Associations between executive function and early math and literacy skills in preschool children. International Journal of Educational Research Open, 3, 100201. 10.1016/j.ijedro.2022.100201 [DOI] [PMC free article] [PubMed] [Google Scholar]
  23. Espy KA, McDiarmid MM, Cwik MF, Stalets MM, Hamby A, & Senn TE (2004). The contribution of executive functions to emergent mathematic skills in preschool children. Developmental Neuropsychology, 26(1), 465–486. 10.1207/s15326942dn2601_6 [DOI] [PubMed] [Google Scholar]
  24. Gaillard A, Fehring DJ, & Rossell SL (2021). A systematic review and meta-analysis of behavioural sex differences in executive control. European Journal of Neuroscience, 53(2), 519–542. 10.1111/ejn.14946 [DOI] [PubMed] [Google Scholar]
  25. Garon N, Bryson SE, & Smith IM (2008). Executive function in preschoolers: A review using an integrative framework. Psychological Bulletin, 134(1), 31–60. 10.1037/0033-2909.134.1.31 [DOI] [PubMed] [Google Scholar]
  26. Garon NM, Piccinin C, & Smith IM (2016). Does the BRIEF-P predict specific executive function components in preschoolers? Applied Neuropsychology: Child, 5(2), 110–118. 10.1080/21622965.2014.1002923 [DOI] [PubMed] [Google Scholar]
  27. Gartstein MA, & Rothbart MK (2003). Studying infant temperament via the revised infant behavior questionnaire. Infant Behavior and Development, 26(1), 64–86. 10.1016/S0163-6383(02)00169-8 [DOI] [Google Scholar]
  28. Gioia GA, Espy KA, & Isquith PK (2002). Behavior rating inventory of executive function – Preschool version. Psychological Assessment Resources. [Google Scholar]
  29. Gioia GA, Isquith PK, Guy SC, & Kenworthy L (2000). The behavior rating inventory of executive function. Psychological Assessment Resources. [Google Scholar]
  30. Gioia GA, Isquith PK, & Kenealy LE (2008). Assessment of behavioral aspects of executive function. In Anderson V, Jacobs R, & Anderson PJ (Eds.), Executive functions and the frontal lobes (pp. 179–202). Psychology Press. [Google Scholar]
  31. Goldhaber D, Wolff M, & Daly T (2020). Assessing the accuracy of elementary school test scores as predictors of students’ high school outcomes. Carnegie Foundation and the National Center for Analysis of Longitudinal Data in Education Research. https://caldercenter.org/sites/default/files/CALDER%20WP%20235-0520-2.pdf [Google Scholar]
  32. Gottfredson NC, Cole VT, Giordano ML, Bauer DJ, Hussong AM, & Ennett ST (2019). Simplifying the implementation of modern scale scoring methods with an automated R package: Automated moderated nonlinear factor analysis (aMNLFA). Addictive Behaviors, 94, 65–73. 10.1016/j.addbeh.2018.10.031 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. Hernandez DJ (2011). Double jeopardy: How third grade reading skills and poverty influence high school graduation. The Annie E. Casey Foundation. [Google Scholar]
  34. Jacob R, & Parkinson J (2015). The potential for school-based interventions that target executive function to improve academic achievement: A review. Review of Educational Research, 85(4), 512–552. 10.3102/0034654314561338 [DOI] [Google Scholar]
  35. Johnson ES, Humphrey M, Mellard DF, Woods K, & Swanson HL (2010). Cognitive processing deficits and students with specific learning disabilities: A selective meta-analysis of the literature. Learning Disability Quarterly, 33(1), 3–18. 10.1177/073194871003300101 [DOI] [Google Scholar]
  36. Kälin S, & Roebers CM (2021). Self-regulation in preschool children: Factor structure of different measures of effortful control and executive functions. Journal of Cognition and Development, 22(1), 48–67. 10.1080/15248372.2020.1862120 [DOI] [Google Scholar]
  37. Kalstabakken AW, Desjardins CD, Anderson JE, Berghuis KJ, Hillyer CK, Seiwert MJ, Carlson SM, Zelazo PD, & Masten AS (2021). Executive function measures in early childhood screening: Concurrent and predictive validity. Early Childhood Research Quarterly, 57, 144–155. 10.1016/j.ecresq.2021.05.009 [DOI] [Google Scholar]
  38. Kochanska G, Murray KT, & Harlan ET (2000). Effortful control in early childhood: Continuity and change, antecedents, and implications for social development. Developmental Psychology, 36(2), 220. 10.1037/0012-1649.36.2.220 [DOI] [PubMed] [Google Scholar]
  39. Lesnick J, Goerge RM, Smithgall C, & Gwynne J (2010). Reading on grade level in third grade: How is it related to high school performance and college enrollment. Chapin Hall at the University of Chicago. [Google Scholar]
  40. Lohaus A, Rueth JE, & Vierhaus M (2020). Cross-informant discrepancies and their association with maternal depression, maternal parenting stress, and mother-child relationship. Journal of Child and Family Studies, 29(3), 867–879. 10.1007/s10826-019-01625-z [DOI] [Google Scholar]
  41. Mahone EM, & Hoffman J (2007). Behavior ratings of executive function among preschoolers with ADHD. The Clinical Neuropsychologist, 21(4), 569–586. 10.1080/13854040600762724 [DOI] [PubMed] [Google Scholar]
  42. Masten AS, Carlson SM, Zelazo PD, Wenzel A, Anderson JE, Buckner M, & McGovern P (2011, August). Assessment of executive function for the national children’s study. Invited presentation for the National Children’s Study Research Day, https://videocast.nih.gov/watch=10467 [Google Scholar]
  43. Masten AS, Herbers JE, Desjardins CD, Cutuli JJ, McCormick CM, Sapienza JK Long JD, Zelazo PD (2012). Executive function skills and school success in young children experiencing homelessness. Educational Researcher, 41(9), 375–384. 10.3102/0013189×12459883 [DOI] [Google Scholar]
  44. McAuley T, Chen S, Goos L, Schachar R, & Crosbie J (2010). Is the behavior rating inventory of executive function more strongly associated with measures of impairment or executive function? Journal of the International Neuropsychological Society, 16(3), 495–505. 10.1017/S1355617710000093 [DOI] [PubMed] [Google Scholar]
  45. McClelland MM, Acock AC, & Morrison FJ (2006). The impact of kindergarten learning-related skills on academic trajectories at the end of elementary school. Early Childhood Research Quarterly, 21(4), 471–490. 10.1016/j.ecresq.2006.09.003 [DOI] [Google Scholar]
  46. McClelland MM, Acock AC, Piccinin A, Rhea SA, & Stallings MC (2013). Relations between preschool attention span-persistence and age 25 educational outcomes. Early Childhood Research Quarterly, 28(2), 314–324. 10.1016/j.ecresq.2012.07.008 [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. McCoy DC (2019). Measuring young children’s executive function and self-regulation in classrooms and other real-world settings. Clinical Child and Family Psychology Review, 22(1), 63–74. 10.1007/s10567-019-00285-1 [DOI] [PubMed] [Google Scholar]
  48. Meng XL, Rosenthal R, & Rubin DB (1992). Comparing correlated correlation coefficients. Psychological Bulletin, 111(1), 172–175. 10.1037/0033-2909.111.1.172 [DOI] [Google Scholar]
  49. Millsap RE, & Olivera-Aguilar M (2012). Investigating measurement invariance using confirmatory factor analysis. In Hoyle RH (Ed.), Handbook of structural equation modeling (pp. 380–391). Guilford Press. [Google Scholar]
  50. Minneapolis Public Schools. (2007). Technical manual for the Minneapolis preschool screening instrument—Revised.
  51. Minneapolis Public Schools. (2016). Minneapolis preschool screening instrument – Revised (MPSI-R) technical manual revised may 2016.
  52. Mischel W, Shoda Y, & Rodriguez MI (1989). Delay of gratification in children. Science, 244 (4907), 933–938. 10.1126/science.2658056 [DOI] [PubMed] [Google Scholar]
  53. Miyake A, Friedman NP, Emerson MJ, Witzki AH, Howerter A, & Wager TD (2000). The unity and diversity of executive functions and their contributions to complex “frontal lobe” tasks: A latent variable analysis. Cognitive Psychology, 41(1), 49–100. 10.1006/cogp.1999.0734 [DOI] [PubMed] [Google Scholar]
  54. Morrison FJ, Ponitz CC, & McClelland MM (2010). Self-regulation and academic achievement in the transition to school. In Posner M (Series Ed.), Calkins S, & Bell M (vol. Eds.) The developing human brain: Child development at the intersection of emotion and cognition (pp. 203–224). American Psychological Association. 10.1037/12059-011 [DOI] [Google Scholar]
  55. Naglieri JA, & Goldstein S (2014). Using the Comprehensive Executive Function Inventory (CEFI) to assess executive function: From theory to application. In Goldstein S & Naglieri J (Eds.), Handbook of executive functioning (pp. 223–244). Springer Science + Business Media. 10.1007/978-1-4614-8106-5_14 [DOI] [Google Scholar]
  56. Nesbitt KT, Baker-Ward L, & Willoughby MT (2013). Executive function mediates socio-economic and racial differences in early academic achievement. Early Childhood Research Quarterly, 28(4), 774–783. 10.1016/j.ecresq.2013.07.005 [DOI] [Google Scholar]
  57. Nilsen ES, Huyder V, McAuley T, & Liebermann D (2017). Ratings of Everyday Executive Functioning (REEF): A parent-report measure of preschoolers’ executive functioning skills. Psychological Assessment, 29(1), 50. 10.1037/pas0000308 [DOI] [PubMed] [Google Scholar]
  58. Obradović J (2010). Effortful control and adaptive functioning of homeless children: Variable-focused and person-focused analyses. Journal of Applied Developmental Psychology, 31(2), 109–117. 10.1016/j.appdev.2009.09.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Obradović J, & Willoughby MT (2019). Studying executive function skills in young children in low-and middle-income countries: Progress and directions. Child Development Perspectives, 13 (4), 227–234. 10.1111/cdep.12349 [DOI] [Google Scholar]
  60. Pérez-Edgar K, Vallorani A, Buss KA, & LoBue V (2020). Individual differences in infancy research: Letting the baby stand out from the crowd. Infancy, 25(4), 438–457. 10.1111/infa.12338 [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Putnam SP, & Rothbart MK (2006). Development of short and very short forms of the children’s behavior questionnaire. Journal of Personality Assessment, 87(1), 102–112. 10.1207/s15327752jpa8701_09 [DOI] [PubMed] [Google Scholar]
  62. R Core Team. (2021). R: A language and environment for statistical computing. R Foundation for Statistical Computing. https://www.R-project.org/ [Google Scholar]
  63. Reynolds CR, Altmann RA, & Allen DN (2021). The problem of bias in psychological assessment. In Mastering modern psychological testing (pp. 573–613). Springer. 10.1007/978-3-030-59455-8_15 [DOI] [Google Scholar]
  64. Rosseel Y (2012). Lavaan: An R package for structural equation modeling. Journal of Statistical Software, 48(2), 1–36. 10.18637/jss.v048.i02 [DOI] [Google Scholar]
  65. Rothbart MK (2011). Becoming who we are: Temperament and personality in development. GuilforPress. 10.5860/choice.49-2373 [DOI] [Google Scholar]
  66. Rubin DB (2004). Multiple imputation for nonresponse in surveys (Vol. 81). John Wiley & Sons. [Google Scholar]
  67. Rueda MR, Fan J, McCandliss BD, Halparin JD, Gruber DB, Lercari LP, & Posner MI (2004). Development of attentional networks in childhood. Neuropsychologia, 42 (8), 1029–1040. 10.1016/j.neuropsychologia.2003.12.012 [DOI] [PubMed] [Google Scholar]
  68. Schafer JL, & Graham JW (2002). Missing data: Our view of the state of the art. Psychological Methods, 7(2), 147. 10.1037/1082-989X.7.2.147 [DOI] [PubMed] [Google Scholar]
  69. Shoda Y, Mischel W, & Peake PK (1990). Predicting adolescent cognitive and self-regulatory competencies from preschool delay of gratification: Identifying diagnostic conditions. Developmental Psychology, 26(6), 978–986. 10.1037/0012-1649.26.6.978 [DOI] [Google Scholar]
  70. Snyder HR, Miyake A, & Hankin BL (2015). Advancing understanding of executive function impairments and psychopathology: Bridging the gap between clinical and cognitive approaches. Frontiers in Psychology, 6, 328. 10.3389/fpsyg.2015.00328 [DOI] [PMC free article] [PubMed] [Google Scholar]
  71. Social Explorer Professional. (2014). Social explorer. Retreived July 30, 2014, from https://www.socialexplorer.com/us-census-data
  72. Squires J, Bricker D, & Twombly E (2002). Ages & stages questionnaire: Social-emotional (ASQ: SE): A parent-completed, child-monitoring system for social-emotional behaviors. Paul H. Brookes Publishing Co. [Google Scholar]
  73. St Clair-Thompson HL, & Gathercole SE (2006). Executive functions and achievements in school: Shifting, updating, inhibition, and working memory. Quarterly Journal of Experimental Psychology, 59(4), 745–759. 10.1080/17470210500162854 [DOI] [PubMed] [Google Scholar]
  74. Thorell LB, & Catale C (2014). The assessment of executive functioning using the childhood Executive functioning inventory (CHEXI). In Goldstein S & Naglieri J (Eds.), Handbook of executive functioning (pp. 359–366). Springer. 10.1007/978-1-4614-8106-5_20 [DOI] [Google Scholar]
  75. Toplak ME, West RF, & Stanovich KE (2013). Practitioner review: Do performance-based measures and ratings of executive function assess the same construct? Journal of Child Psychology and Psychiatry, 54(2), 131–143. 10.1111/jcpp.12001 [DOI] [PubMed] [Google Scholar]
  76. U.S. Census Bureau. (n.d.). Historical income- household. Table H-3. Current population survey, 1968 to 2023 annual social and economic supplements. Retrieved December 19, 2023, from https://www.census.gov/data/tables/time-series/demo/income-poverty/historical-income-households.html
  77. Van Buuren S, & Groothuis-Oudshoorn K (2011). Mice: Multivariate imputation by chained equations in R. Journal of Statistical Software, 45(3), 1–55. 10.18637/jss.v045.i03 [DOI] [Google Scholar]
  78. Wenzel AJ, Sapienza JK, Carlson SM, Desjardins CD, Rothbart MK, & Masten AS (2013, March). A parent report scale of executive function in early childhood. Poster presented at the society for research in child development Biennial Meeting, Seattle, WA. [Google Scholar]
  79. Willcutt EG, Doyle AE, Nigg JT, Faraone SV, & Pennington BF (2005). Validity of the executive function theory of attention deficit/hyperactivity disorder: A meta-analytic review. Biological Psychiatry, 57(11), 1336–1346. 10.1016/j.biopsych.2005.02.006 [DOI] [PubMed] [Google Scholar]
  80. Willoughby M, & Blair C (2011). Test-retest reliability of a new executive function battery for use in early childhood. Child Neuropsychology, 17(6), 564–579. 10.1080/09297049.2011.554390 [DOI] [PubMed] [Google Scholar]
  81. Willoughby MT, Magnus B, Vernon-Feagans L, Blair CB, & The Family Life Project Investigators. (2017). Developmental delays in executive function from 3 to 5 years of age predict kindergarten academic readiness. Journal of Learning Disabilities, 50(4), 359–372. [DOI] [PMC free article] [PubMed] [Google Scholar]
  82. Willoughby MT, Wirth RJ, & Blair CB (2012). Executive function in early childhood: Longitudinal measurement invariance and developmental change. Psychological Assessment, 24 (2), 418. 10.1037/a0025779 [DOI] [PMC free article] [PubMed] [Google Scholar]
  83. Wolf S, & McCoy DC (2019). The role of executive function and social-emotional skills in the development of literacy and numeracy during preschool: A cross-lagged longitudinal study. Developmental Science, 22(4), e12800. 10.1111/desc.12800 [DOI] [PubMed] [Google Scholar]
  84. Wolfe CD, & Bell MA (2004). Working memory and inhibitory control in early childhood: Contributions from physiology, temperament, and language. Developmental Psychobiology, 44 (1), 68–83. 10.1002/dev.10152 [DOI] [PubMed] [Google Scholar]
  85. Zedeck S (Ed.). (2014). APA dictionary of statistics and research methods. In American Psychological Association. 10.1037/14336-000 [DOI] [Google Scholar]
  86. Zelazo PD (2020). Executive function and psychopathology: A neurodevelopmental perspective. Annual Review of Clinical Psychology, 16(1), 431–454. 10.1146/annurev-clinpsy-072319-024242 [DOI] [PubMed] [Google Scholar]
  87. Zelazo PD, Blair CB, & Willoughby MT (2016). Executive function: Implications for education (NCER 2017–2000). National Center for Education Research, Institute of Education Sciences, U.S. Department of Education. http://ies.ed.gov/ [Google Scholar]
  88. Zelazo PD, & Carlson SM (2012). Hot and cool executive function in childhood and adolescence: Development and plasticity. Child Development Perspectives, 6(4), 354–360. 10.1111/j.1750-8606.2012.00246.x [DOI] [Google Scholar]
  89. Zelazo PD, & Carlson SM (2020). The neurodevelopment of executive function skills: Implications for academic achievement gaps. Psychology & Neuroscience, 13(3), 273. 10.1037/pne0000208 [DOI] [Google Scholar]
  90. Zelazo PD, Müller U, Frye D, Marcovitch S, Argitis G, Boseovski J, Chiang JK, Hongwanishkul D, Schuster BV, & Sutherland A (2003). The development of executive function in early childhood. Monographs of the Society for Research in Child Development, 68(3), vii–137. 10.1111/j.0037-976x.2003.00260.x [DOI] [PubMed] [Google Scholar]
  91. Zentner M, & Bates JE (2008). Child temperament: An integrative review of concepts, research programs, and measures. International Journal of Developmental Science, 2(1–2), 7–37. 10.3233/DEV-2008-21203 [DOI] [Google Scholar]
  92. Zhou Q, Chen SH, & Main A (2012). Commonalities and differences in the research on children’s effortful control and executive function: A call for an integrated model of self-regulation. Child Development Perspectives, 6(2), 112–121. 10.1111/j.1750-8606.2011.00176.x [DOI] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Supplementary Materials

Supplementary Material

RESOURCES