Abstract
Definitions of dyslexia typically make reference to unexpected poor reading, although how best to operationalize unexpected remains an issue. When operationally defined as reading below expectations based on level of oral language, cases of unexpected poor reading make up fewer than half of cases of poor reading, and cases of unexpected poor reading occur throughout the range of reading proficiency. An implication is that what optimally predicts poor reading may not optimally predict unexpected poor reading. The goal of the three presented studies was to test this implication empirically. In Study 1, a model-based meta-analysis, phonological awareness accounted for 40% of the variance in decoding but only 1% of the variance in decoding that was unexpected based on level of vocabulary. Conversely, unexpected phonological awareness accounted for 34% of the variance in unexpected decoding but only 1% of the variance in decoding. An analogous pattern of results occurred for reading comprehension. In Study 2, a study of 766 children in kindergarten, first grade, and second grade, latent variables were used to represent oral vocabulary, phonological awareness, and decoding. As was seen in Study 1, unexpected decoding was better predicted by unexpected phonological awareness than by phonological awareness. In Study 3, a longitudinal study of 1,025 children followed from preschool through grade 2, the pattern of results mirrored those of Studies 1 and 2. An important implication of these studies is that typical assessments may be better at identifying poor reading than they are at identifying unexpected poor reading or dyslexia.
Keywords: early identification, dyslexia, reading-disability
In 1896, W.P. Morgan made what appears to be the first report of a case of developmental dyslexia (Fletcher et al., 2019). He described a 14-year-old boy who “has always been a bright and intelligent boy, quick at games, and in no way inferior to others of his age.” However, when it came to reading or spelling words, “He seems to have no power of preserving and storing up the visual impression produced by words—hence the words, though seen, have no significance for him. His visual memory for words is defective or absent, which is equivalent to saying he is “…word blind” (Morgan, 1896, p. 378). We now label this phenomenon developmental dyslexia, which refers to difficulty learning to read well in spite of adequate opportunity to learn and an absence of exclusionary factors (Wagner et al., 2022). In contrast, acquired dyslexia refers to the phenomenon in which formerly able readers lose the ability to read typically after head injury, stroke, or disease. Developmental dyslexia is not limited to alphabetic writing systems but appears to occur throughout the world regardless of the type of script used for writing (Snowling & Melby-Lervåg, 2016).
Definition of Developmental Dyslexia
A widely used definition of developmental dyslexia was proposed by a working group of the International Dyslexia Association (Lyon et al., 2003):
Dyslexia is a specific learning disability that is neurobiological in origin. It is characterized by difficulties with accurate and/or fluent word recognition and by poor spelling and decoding abilities. These difficulties typically result for a deficit in the phonological component of language that is often unexpected in relation to other cognitive abilities and the provision of effective classroom instruction. Secondary consequences may include problems in reading comprehension and reduced reading experience that can impede growth of vocabulary and background knowledge.
(p. 2)
Note that dyslexia is defined as a specific learning disability. Deficit areas are limited to specific processes (e.g., typically the phonological component of language) as opposed to general level of language functioning and specific areas of achievement (e.g., word recognition, spelling, and decoding). Consequently, dyslexia refers to performance that is often unexpected relative to performance in other areas (Fletcher et al., 2019; Grigorenko et al., 2020). Aspects of this definition may be incomplete or underspecified. Nevertheless, 30 influential international researchers in the field of dyslexia agreed on this consensus definition, and there has been little support for changes (Dickman, 2017). The British Dyslexia Association adopted a definition of dyslexia that was proposed in the Rose (2009) report. It states that dyslexia occurs across the range of intellectual abilities, and nothing in the definition adopted by the International Dyslexia Association restricts dyslexia to a particular range of intellectual abilities. In the US, the term specific reading disability was defined by legislation in 1968 as a disorder in one or more of the basic psychological processes involved in understanding or in using spoken or language, and this definition continues to the present day (see Fletcher et al., 2019, pp. 17-21 for a summary of relevant legislation in the US).
The definition of specific reading disability did not provide procedures to be used for identification. To be implemented in practice, a companion regulatory definition was provided by the U.S. Office of Education in 1977 that relied on a severe discrepancy between achievement and intellectual ability (U.S. Office of Education, 1977, p. G1082). Note that this regulatory definition required poor achievement relative to intellectual ability but not poor achievement compared to age or grade-levels standards. In addition to requiring a severe discrepancy between achievement and intellectual ability, the exclusionary clause for mental retardation was implemented as an IQ cut-off. If a student’s IQ was below the cut-off, it was not possible to identify the student as having a specific learning disability. Public Law 94-242 was most recently reauthorized in 2004 as the Individuals with Disabilities Education Improvement Act (IDEA), maintaining the original statutory definition provided earlier. One change however was that other methods besides finding a severe discrepancy were allowed for identification purposes. In addition to allowing but not mandating use of a severe discrepancy between achievement and intellectual ability, an individual could be considered to have a specific learning disability if the “child does not achieve adequately for the child’s age or meet state-approved grade-level standards when provided with learning experiences and instruction appropriate or insufficient progress is shown (e.g., poor response to instruction and intervention) and the child exhibits a pattern of strengths and weaknesses in performance, achievement, or both.” What changed from the original regulatory definition is inclusion of the criterion that performance must be poor relative to the child’s age or state-approved grade level standards.
Two other diagnostic systems that are used by clinicians outside the school environment are also relevant. The DSM-5 (American Psychiatric Association, 2013) includes specific learning disorder as a neurodevelopmental disorder that is diagnosed when “there are specific deficits in an individual’s ability to perceive or process information efficiently and accurately” (p. 32). Skills must be substantially below what is expected for the child’s age and cause problems in school, work, or everyday activities. ICD-11 (World Health Organization, 2019) also includes reading disorder as a diagnosis for individuals who have trouble with reading despite normal intelligence. The diagnosis requires that the deficit in reading exists relative to both chronological age and intellectual ability.
Complicating the process of determining an acceptable operational definition of dyslexia are the issues of how exclusionary criteria such as whether the poor performance is the result of environmental, cultural, or economic disadvantages are incorporated, and whether common assessments used are biased against individuals whose background or language status is different from that of the majority culture.
In summary, what is common to all proposed definitions of dyslexia, specific learning disability, or specific learning disorder is that deficits are specific rather than characteristic of performance more generally. For that reason, the impaired performance is unexpected. More recent regulatory definitions included a requirement that performance be poor relative to chronological age or grade-level expectations for unspecified reasons.
Beyond considering the fact that unexpected poor reading and spelling are the basis for nearly all definitions of dyslexia (Fletcher et al., 2019; Grigorenko et al., 2020), what additional evidence is there to support making the distinction between unexpected poor and simple poor performance? From the educational perspective of what and how to teach, it may well be the case that how to teach children who struggle to learn to read (i.e., more explicit instruction delivered with greater intensity) is pretty much the same regardless of the reason for their struggle (Fletcher et al., 2019). However, there are exceptions to this generalization. Connor et al. (2009) studied the effects of individualizing student instruction by adjusting time spent in code-focused versus meaning-focused reading instruction depending on the needs of individual students. The effects of providing instruction based on this child-by-instruction interaction were improved student outcomes.
When the focus expands beyond how to teach beginning reading, understanding what accounts for the observed poor reading can be of more value. For example, given existing interventions, an individual with very severe and long-lasting dyslexia is unlikely to attain normal levels of reading proficiency from any known form of instruction or intervention. Treatment resistors or nonresponders exist for every current intervention (Torgesen, 2000). Indeed, one proposed hallmark of dyslexia is inadequate response to effective instruction and intervention (Miciak & Fletcher, 2020). Fortunately, individuals whose struggle with reading the words on the page results in their reading comprehension being worse than their listening comprehension can use text-to-speech and other forms of assistive technology that have been shown to be effective for students with reading disability (Wood et al., 2018).
Implications for Early Identification of Risk for Dyslexia from a New Approach to Estimating its Prevalence
Proposed estimates of the prevalence of dyslexia have ranged from 3 to almost 20 percent of school-age children (Fletcher et al., 2019, Hoeft et al, 2015; Moll et al., 2014; Peterson & Pennington, 2012; Shaywitz et al., 1992; Snowling & Melby-Lervåg, 2016). One explanation for the wide range of prevalence estimates is differences in the severity of the reading problem that is required before an individual is considered to have dyslexia. There is consensus that prevalence estimates vary as a function of severity, but this consensus is not reflected in prevalence estimates. Rather, it is used as an explanation for the substantial differences in published prevalence.
Recently, Wagner et al. (2019, 2020) proposed that no single prevalence estimate is correct, but rather that there is a distribution of prevalence as a function of severity. Creating this distribution required picking an operational definition of dyslexia that would be sufficient for determining prevalence. To reflect the consensus that a core feature of dyslexia is that it is unexpected and thus different than mere poor reading (Grigorenko et al., 2020), poor reading comprehension relative to listening comprehension was used (Aaron, 1991; Badian, 1999; Bedford-Feuell et al., 1995; Erbeli et al., 2018; Spring & French, 1990; Stanovich, 1991b). Poor reading comprehension relative to listening comprehension was not proposed as sufficient for identification at the level of the individual, as it has become clear that no single criterion is sufficient for reliable identification of specific individuals with dyslexia. Rather, it was proposed as a proxy for dyslexia sufficient for estimating prevalence at the level of the population.
Wagner et al. (2020) used model-based meta-analysis to create composite correlation matrices and then created simulated datasets using the composite correlations as estimates of the population parameters needed to generate the simulated data. A scatterplot of listening comprehension and reading comprehension that resulted from this approach is presented in Figure 1. Points to the left of the vertical line represent cases of poor readers, operationally defined as scoring at or below the 20th percentile in reading comprehension. Points above the diagonal line represent cases of unexpected poor reading, operationally defined as scoring at or above 1.5 standard deviations above the mean in difference between listening and reading comprehension. One conclusion from this figure is that cases of unexpected poor reading (i.e., cases above the diagonal) make up less than half of the cases of poor reading (i.e., cases to the left of the vertical line). A second conclusion is that cases of unexpected poor reading occur across the reading spectrum as opposed to only existing at the lower tail of reading performance. A recent large-scale empirical study replicated both of these major findings (Odegard et al., 2022). An implication of the only partial overlap between poor and unexpected poor reading in Figure 1 is that what optimally predicts poor reading may not optimally predict unexpected poor reading. The goal of the three presented studies was to test this implication empirically.
Figure 1. Scatterplot of Listening (LC) and Reading Comprehension (RC) from Author(s) (2020).
Note. Points to the left of the vertical line represent scores of poor readers (i.e., the 20th %-ile in RC). Points above the diagonal line represent readers with listening comprehension better than reading comprehension (i.e., at or above 1.5 standard deviations above the mean in listening comprehension-reading comprehension discrepancy score).
STUDY 1
Study 1 was an analysis of a simulated dataset that was derived from a model-based meta-analysis of relations between reading comprehension, listening comprehension, decoding, vocabulary, and phonological awareness. We included both vocabulary and phonological awareness as well as decoding based on what has become known as the “triangle” model of word-level reading. This model has been applied to account for developmental dyslexia by Harm and Seidenberg (1999). According to the model, which is represented as a connectionist computer-simulation model, word-level reading depends on the relations among three kinds of representations. Orthographic representations are how knowledge about spellings of words is stored. Semantic representations are how knowledge about the meanings of words is stored. Finally, phonological representations are how knowledge about the pronunciations of words and word parts are stored. Using two ways of limiting the ability of the model to develop phonological representations resulted in impaired decoding but leaving the semantic part of the model (i.e., vocabulary knowledge) relatively intact.
Method
The results of the literature search of nationally normed standardized tests carried out by Wagner et al. (2020) served as the starting point for the present study. An initial search was carried out using the search terms “standardized measure(s)” and “norm referenced” and “reading” and English and intercorrelation, as well as search combinations with decoding, listening comprehension, reading comprehension and Phono*. The databases used were ProQuest, ERIC, Google Scholar and Pubmed. The search was not productive, and we revised our search by carrying out a Google search using the search string standardized measures of reading. This search yielded the Southwest Educational Development Laboratory (SEDL) reading assessment data base (www.sedl.org/reading/rad/list.html). We then found more assessments from an early reading assessment guiding tool on the Reading Rockets website (www.readingrockets.org) and from the Wrightslaw reading assessment list (www.wrightslaw.com/bks/aat/ch6.reading.pdf).
The search, which was completed in April of 2019, yielding 91 assessments. After applying the inclusionary criteria of (1) norm referenced; (2) nationally representative norming sample; (3) in English; (4) included subtests for measuring listening comprehension, reading comprehension, vocabulary, decoding, and phonological awareness; (5) correlation matrix of subtests and subtest reliability available; and (6) included data from multiple ages or grades, the following assessments remained: the Kaufman Test of Educational Achievement (KTEA III) (Kaufman et al., 2014), the Woodcock Johnson IV (WJ IV) (McGrew et al., 2011), the Iowa Test of Basic Skills (ITBS) (Hoover et al., 2003), The Wechsler Individual Achievement Test III (WIAT III) (Breaux, 2009), The Woodcock Reading Mastery Test III (WRMT III) (Woodcock, 2011), the Early Reading Diagnostic Test (ERDA) (The Psychological Corporation, 2003), The Oral and Written Language Scales II (OWLS II) (Carrow-Woolfolk, 2011), The Brigance Comprehensive Inventory of Basic Skills II (CIBS II) (French & Glascoe, 2010) and the Stanford Achievement Test 10 (SAT10) (Harcourt Educational Measurement, 2004). Additional details about the search can be found in Wagner et al. (2020).
A model-based meta-analysis (Becker & Aloe, 2019; Cheung, 2015) was carried out to calculate average weighted correlations between the measures of listening comprehension, reading comprehension, vocabulary, decoding, and phonological awareness. These were used subsequently as population parameters to generate the simulated dataset. Analyses were carried out using the metaSEM package in R.
Results
The composite correlation matrix generated from the first stage of the two-stage metaSEM analysis is presented in Table 1. The correlations were moderate in size. We then used an SPSS macro to create a simulated dataset of 5,000 cases drawn from a population with rho values equal to the composite correlations in Table 1. Once the simulated data were created, compute statements in SPSS were used to create six additional variables that represented unexpected performance: listening comprehension minus reading comprehension, listening comprehension minus decoding, listening comprehension minus phonological awareness, vocabulary minus reading comprehension, vocabulary minus decoding, and vocabulary minus phonological awareness. Correlations among these 11 variables are presented in Table 2.
Table 1.
Composite correlation matrix between key constructs from first stage metaSEM analysis
| RC | LC | DC | VC | PA | |
|---|---|---|---|---|---|
| RC | --- | ||||
| LC | 0.494 | --- | |||
| DC | 0.642 | 0.326 | --- | ||
| VC | 0.675 | 0.490 | 0.560 | --- | |
| PA | 0.617 | 0.439 | 0.633 | 0.552 | --- |
Note. RC = reading comprehension; LC = listening comprehension; DC = decoding; VC = vocabulary; PA = phonological awareness.
Table 2.
Correlations among original and derived variables from the simulated sample
| RC | LC | DC | VC | PA | LCMINRC | VCMINPA | VCMINDC | LCMINDC | LCMINPA | VCMINRC | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| RC | --- | ||||||||||
| LC | .49*** | --- | |||||||||
| DC | .64*** | .33*** | --- | ||||||||
| VC | .68*** | .49*** | .56*** | --- | |||||||
| PA | .61*** | .44*** | .63*** | .55*** | --- | ||||||
| LCMINRC | −.50*** | .50*** | −.31*** | −.18*** | −.17*** | --- | |||||
| VCMINPA | .06*** | .05*** | −.08*** | .47*** | −.47*** | −.01 | --- | ||||
| VCMINDC | .04* | .18*** | −.47*** | .47*** | −.09*** | .14*** | .59*** | --- | |||
| LCMINDC | −.13*** | .58*** | −.58*** | −.06*** | −.17*** | .70*** | .11*** | .56*** | --- | ||
| LCMINPA | −.11*** | .53*** | −.29*** | −.06*** | −.53*** | .64*** | .50*** | .25*** | .71*** | --- | |
| VCMINRC | −.40*** | −.01 | −.10*** | .40*** | −.08*** | .40*** | .51*** | .54*** | .08*** | .07*** | --- |
Notes. RC = reading comprehension; LC = listening comprehension; DC = decoding; VC = vocabulary; PA = phonological awareness; LCNIMRC = LC – RC; VCMINPA = VC – PA; VCMINDC = VC – DC; LCMINDC = LC – DC; LCMINPA = LC – PA; VCMINRC = VC – RC.
p < .05
p < .001
The pattern of results was that criterion measures of unexpected performance were more strongly related to predictors that also were measures of unexpected performance, and criterion measures of simple performance were more related to predictors that were also measures of simple performance. For example, beginning with decoding as the criterion to be predicted, phonological awareness was a significantly stronger predictor of decoding (r = .633) than of decoding that was unexpected based on vocabulary (r = −.086). Conversely, when the predictor was unexpected phonological awareness based on vocabulary, it was a significantly stronger predictor of unexpected decoding based on vocabulary (r = .587) than of decoding (r = −.077). Turning to reading comprehension as the criterion to be predicted, decoding (r = .642) and phonological awareness (r = .614) were significantly stronger predictors of reading comprehension than of unexpected reading comprehension based on listening comprehension with correlations of −.314, and −.174, respectively. When the criterion variable was unexpected reading comprehension, unexpected decoding (r = .704) and unexpected phonological awareness (r = .642), were significantly stronger predictors of unexpected reading comprehension than of simple reading comprehension with correlations of −.127, and −.113, respectively.
The results of Study 1 supported the hypothesis that what optimally predicts simple performance may not optimally predict unexpected performance. However, there are two limitations of Study 1 that deserve mention. First, there is a need to cross-validate these results using empirical datasets. Using model-based meta-analysis should produce more generalizable results than would be obtained from an individual empirical study. Nevertheless, assumptions are made when creating simulated datasets using population parameters derived from model-based meta-analysis that require cross-validation. For example, the variables in the simulated dataset were assumed to be normally distributed and that the relations between pairs of variables were linear. A second limitation is that the same observed variable, either vocabulary or listening comprehension, was used to create the relative predictors and corresponding relative criterion measures. Measure-specific variance associated with vocabulary or listening comprehension could have positively biased the strength of the predictive relations. These two limitations were addressed in Studies 2 and 3.
STUDY 2
The purpose of Study 2 was to extend the results of Study 1 to an evaluation of these questions using an empirical rather than a simulated dataset, and where constructs could be represented as latent variables, thereby providing evidence at the level of the construct instead of for a single measure of the construct. Recall the limitation of Study 1 that same observed variable, either vocabulary or listening comprehension, was used to create the relative predictors and corresponding relative criterion measures. For constructs represented by latent variables with multiple indicators, only common variance is represented by the latent variable. Indicator-specific variable is shunted to error variance terms in the model.
Method
Participants
Children in this study were part of a larger project involving assessments of reading, language, and cognitive abilities associated with decoding and reading comprehension in samples of preschool through fifth-grade children. Only the 766 children in kindergarten, first grade, and second grade were used in this study to better align with data from Study 3. Overall, children in the sample ranged in age from 61 to 124 months (M = 83.40, SD = 11.72). Girls represented 53% of the sample, and the sample was diverse in its racial and ethnic make-up; 77% of participants were White; 18% were Black/African American; 3% were Multiracial; 1% were Asian, 4% were classified as other/unknown, and 7% identified as Hispanic/Latino. There were 228 kindergarten children (M-age = 70.24 months, SD = 5.37), 302 first-grade children (M-age = 96.11 months, SD = 5.77), and 227 second-grade children (M-age = 96.11 months, SD = 6.80). Children were recruited across two years from 115 classrooms in 18 schools in north Florida, including schools in which greater than 50% of students were eligible for free-and-reduced-price lunch.
Measures
Decoding.
Children completed measures of word or nonword decoding: The Letter-Word Identification and Word Attack subtests of the Woodcock-Johnson Tests of Achievement-3rd ed. (WJ-III; Woodcock et al., 2001), the Sight-Word Efficiency subtest of the Test of Word Reading Efficiency (TOWRE; Torgesen et al., 1999), and the Test of Silent Contextual Reading Fluency (TOSCRF; Hammill et al., 2006). The Letter-Word Identification subtest of the WJ-III requires children to correctly name printed letters and pronounce printed words aloud. There are 76 items of increasing difficulty, and basal, based on child grade, and ceiling (i.e., six consecutive items are responded to incorrectly) rules are used. Internal consistency is high (α = .91) for the age range of children in this sample. The Word Attack subtest of the WJ-III requires children to read phonologically and orthographically regular nonwords aloud. It has 32 items. Basal, based on child grade, and ceiling (i.e., six consecutive items are responded to incorrectly) rules are used. Internal consistency is good (α = .87) for the age range of children in this sample. The Sight-Word Efficiency subtest of the TOWRE requires children to read aloud as many words as they can read from a list of 104 real words. Scores are the number of words read correctly within 45 seconds. Reliability of the measure is high (alternate-form rs = .93 - .97; test-retest rs = .84 - .97) for the age-range of children in this sample. On the TOSCRF, children are presented with 12 short passages in which words are printed in uppercase. with no spaces or punctuation (e.g., “READINGTHISSENTENCEISALITTLETOUGH”), and have three minutes to draw a line between as many words as possible. Because of the nature of the test, the manual only reports that inter-rater reliability for scoring is high.
Oral vocabulary.
For this study, four measures of vocabulary from a larger set of standardized assessments of oral language that children completed were used. The Expressive One-Word Picture Vocabulary Test (Brownell, 2000a) requires children to verbally label illustrations of objects, actions, or concepts. The test has 170 items, but a basal (based on child age) and a ceiling (i.e., six consecutive incorrect responses) rules are used. Internal consistency for the age-range of children in this sample is high (αs = .95 to .97). The Receptive One-Word Picture Vocabulary Test (Brownell, 2000b) requires children to select the one picture out of four that corresponds to a word spoken by the examiner. The test has 170 items, but a basal (based on child age) and a ceiling (i.e., six consecutive incorrect responses) rules are used. Internal consistency for the age-range of children in this sample is high (αs = .95 to .98). The Antonyms subtest of the Comprehensive Assessment of Spoken Language (Carrow-Woolfolk, 2008) assesses children’s word knowledge, retrieval, and production of antonyms. Children are asked by the examiner to “Tell me a word that means the opposite of ____.” The subtest has 55 items but a basal (based on child age) and a ceiling (i.e., five consecutive incorrect responses) rules are used. Reliability of the subtest is adequate to high (split-half rs = .82 - .90; test-retest rs = .80 - .95) for this age-range of children. The Expressive Vocabulary subtest of the Clinical Evaluation of Language Fundamentals (Semel et al., 2003) requires children to label illustrations of people, objects, and actions. The subtest has 27 items. Children begin at the item designated for their age group, and testing is discontinued seven consecutive incorrect responses. Reliability of the subtest is adequate to good (i.e., αs = .80 - .85; test-retest rs = .87 - .91) for this age-range of children.
Phonological awareness.
Children completed the Blending Words, Blending Nonwords, and Elision subtests of the Comprehensive Test of Phonological Processing (Wagner et al., 1999). The 20-item Blending Words subtest requires children to combine a series of phonological segments into a single word, starting with combining two syllables to make one word to combining 10 phonemes to make one word. The 18-item Blending Nonwords subtest requires children to combine a series of phonological segments into single nonwords, starting with combining two syllables to make one nonword to combining seven phonemes to make one nonword. The 20-item Elision subtest requires children to listen to a word, repeat that word, and then say the word with a specified sound removed, starting with removing a word from a compound word to removing initial, final, and medial phonemes from a word. For all three subtests, children start with the first item, and testing is discontinued after three consecutive errors. Internal consistency reliabilities for all subtests for this age range of children are moderate to high (i.e., αs = .79 - .92).
Procedure
Parents of all participating children provided written informed permission/consent for their children’s participation, and all classroom teachers consented to participate. Children in the larger study completed a variety of language, reading, and reading-related assessments. Assessments were administered to children in a quiet area of their schools by research assistants who were trained to criterion performance on the measures through didactic presentations, modeling, a performance assessment of test administration to an adult posing as a child, and live observations of assessment with feedback. Assessments were conducted in a non-fixed order over several 30- to 45-minute sessions, within a six- to eight-week period. Assessments were conducted based on children’s availability, the assessment staff available at a school on a particular day, taking account of children’s attentional and motivational capacities (e.g., breaks were taken when children appeared to become distracted). Due to the size of the full assessment battery in the larger study, a missing-by-design approach (Graham et al., 2006) was used to reduce the testing burden for individual children (see Lonigan et al., 2018, and supplemental materials for an expanded description of this approach).
Results
All analyses were conducted in Mplus 7.11 (Muthén & Muthén, 2013) using full information maximum likelihood and robust maximum likelihood estimation to account for missing data and non-normal variable distributions. The cluster-robust sandwich estimator option was used to correct for the nesting of children within their classrooms. Age-adjusted standardized scores on all measures were used in the analyses (i.e., standardized residuals created by regressing raw score for each task on child age in months within each grade) because in the larger project, some assessments were used above or below the age range for which standard scores were available, standard scores were not available for all tasks, in many cases, tabled standard scores represent relatively coarse age-standardization brackets (i.e., 6-month age brackets), and to provide a common metric across grades. The measurement model included three factors: Vocabulary, Phonological Awareness, and Decoding. This model provided adequate to good fit to the data for kindergarten (i.e., CFI = .94), first grade (i.e., CFI = .92), and second grade (i.e., CFI = .97) children, and it provided better fit than alternative measurement models (see Table S1-S3 in supplemental materials). The Decoding factor was substantially correlated with both the Vocabulary and the Phonological Awareness factors (see Table 3), and Vocabulary and Phonological Awareness factors were substantially correlated in kindergarten (r = .81), first grade (r = .62) and second grade (r = .63).
Table 3.
Concurrent correlations between latent variables or latent-variable residuals for phonological awareness and decoding for kindergarten, first grade, and second grade samples
| Grade Group | Predictor-Outcome Condition | |||
|---|---|---|---|---|
| Predictor | Raw - Raw | Raw - Residual | Residual – Residual |
Residual - Raw |
| Kindergarten (N = 228) | ||||
| Vocabulary | .57 | --- | --- | --- |
| Phonological Awareness | .71 | .30 | .51 | .42 |
| First Grade (N = 302) | ||||
| Vocabulary | .50 | --- | --- | --- |
| Phonological Awareness | .83 | .60 | .76 | .66 |
| Second Grade (N = 227) | ||||
| Vocabulary | .69 | --- | --- | --- |
| Phonological Awareness | .90 | .64 | .83 | .60 |
Note. All correlations are statistically significant at p < .001.
Structural models were examined to determine the unique contributions of phonological awareness to kindergarten, first-grade, and second-grade decoding. Models examining the unique contributions of the phonological awareness to decoding were constructed by creating latent variables representing the residuals of the Decoding or Phonological Awareness latent variables (i.e., variables residualized on the Vocabulary latent variable), and then using these residuals as dependent and independent variables in the analyses (Koch et al., 2018). Table 3 shows the zero-order and different unique associations of the Phonological Awareness factor with the Decoding factor for kindergarten, first-grade, and second-grade children. The association between the Phonological Awareness factor and the Decoding factor was reduced substantially after accounting for vocabulary (i.e., raw-residual column of Table 4). When accounting for the overlap of the Vocabulary factor with both the Phonological Awareness and Decoding factors, the association between the Phonological Awareness and Decoding factors was also reduced but to a lesser extent (i.e., the residual-residual column of Table 3).
Table 4.
Longitudinal correlations between latent variables or latent-variable residuals for preschool vocabulary, phonological awareness, and print knowledge with kindergarten, first-grade, and second-grade decoding outcomes
| Grade for Outcome | Predictor-Outcome Condition | |||
|---|---|---|---|---|
| Preschool Predictor | Raw - Raw | Raw – Residual | Residual - Residual |
Residual - Raw |
| Kindergarten | ||||
| Vocabulary | .42 | --- | --- | --- |
| Phonological Awareness | .56 | .21 | .43 | .39 |
| Print Knowledge | .71 | .48 | .64 | .58 |
| First Grade | ||||
| Vocabulary | .45 | --- | --- | --- |
| Phonological Awareness | .54 | .16 | .33 | .29 |
| Print Knowledge | .70 | .45 | .59 | .53 |
| Second Grade | ||||
| Vocabulary | .47 | --- | --- | --- |
| Phonological Awareness | .54 | .15 | .30 | .26 |
| Print Knowledge | .66 | .40 | .53 | .46 |
Notes. N = 1,025; all correlations are statistically significant at p < .001.
Discussion
Like the results of Study 1, the results of Study 2 demonstrated that the best predictor of unexpected reading was unexpected phonological awareness. Results from both studies reveal that there is substantial overlap between language skills, reading skills, and reading-related skills, like phonological awareness—even measured at the construct level (i.e., multiple indicators used to define the measured construct). Although these are unsurprising results, given the language basis of both phonological awareness and decoding, the results suggest that identification of unexpectedly low decoding skill may be impeded unless unexpectedly poor performance on a predictor, like phonological awareness is also taken into account.
STUDY 3
The results of Study 2 were consistent with the results of Study 1. The best predictors of unexpectedly poor decoding were unexpectedly poor performance on the predictor variables. The purposes of Study 3 were to determine if this pattern of predictive results at the construct level held true (a) when examining longitudinal predictors, as opposed to concurrent relations; (b) for other predictors of decoding such as early print knowledge, and (c) held when examining predictive relations between early childhood, prior to when most children are exposed to formal instruction in reading, and early elementary school.
Method
Participants
As part of a larger study of identifying and reducing risk for reading-related learning disabilities, the sample of children in this study was recruited from preschools in private centers and the Title I preschools of the local school district in north Florida. For this study, data from 1,025 children who were assessed at the end of preschool and at the end of kindergarten, first grade, and second grade were used. Girls represented 45% of the sample, and the sample was diverse in its racial and ethnic make-up; 49% of participants were White; 42% were Black/African American; 3% were Hispanic/Latino; 3% were Multiracial; 2% were Asian, and 1% were classified as other/unknown. Children ranged in age from 52 to 69 months (M = 61.18 months, SD = 3.65) at the end-of-preschool assessment, and averaged 71.89 (SD = 3.72), 83.91 (SD = 3.81), and 95.68 (SD = 3.84) months of age at the kindergarten, first-grade, and second-grade assessments, respectively.
Based on parent report from the 75% of the preschool sample who provided responses to a home-environment questionnaire, the median income level of $31,000 – $41,000 was below the median income for the local area at the time of the study (i.e., $51,000). Median maternal and paternal education was high school completion. Most children (73%) were from single-parent homes, and there was an average of 2.49 (SD = 1.17) children in the home. Fewer children completed kindergarten (n = 922), first-grade (n = 848), and first-grade (n = 755) than the preschool assessments; however, comparisons between those who did and did not complete the follow-up assessments on the preschool assessments used in this study did not reveal any statistically significant differences for kindergarten (ps > .06), first grade (ps > .23), or second grade (ps > .17).
Measures
At both the preschool and elementary-school assessments, multiple measures were used to assess the constructs of interest. In preschool, children’s early literacy skills were assessed, including oral vocabulary, phonological awareness, and print/alphabet knowledge. In kindergarten, first grade, and second grade, children’s word-level reading skills (i.e., decoding) were assessed.
Early Literacy Skills
Oral vocabulary measures.
Children completed the Receptive Vocabulary and the Definitional Vocabulary subtests from the Preschool Comprehensive Test of Phonological and Print Processing (PCTOPPP; Lonigan et al., 2002), which was the development version of the Test of Preschool Early Literacy (Lonigan et al., 2007) and the Basic Concepts subtest Clinical Evaluation of Language Fundamentals--Preschool (CELF-P; Wiig et al., 1992). On the Receptive Vocabulary subtest, children select one picture from four that corresponds to a word spoken by the examiner. On the Definitional Vocabulary subtest, children provide the name of a picture and are asked to give a brief definition of the word (e.g., if shown a picture of a shovel, the child must say “shovel” and answer the question, “what is it used for?”). Both subtests have high internal consistency reliability (αs > .91) and correlate with other measures of oral language (e.g., rs = .47-.62). The Basic Concepts subtest of the CELF-P assesses children’s knowledge of dimension and size, direction, location and position, number and quantity, and equality. The subtest uses basal and ceiling rules. Internal consistency reliability is moderate to high (i.e., αs = .81 to .96), and it has good evidence of concurrent validity with other norm-referenced measures (Wiig et al., 1992).
Phonological awareness measures.
Children completed the two phonological awareness subtests from the PCTOPPP. On the 18-item Elision subtest, which includes nine multiple-choice and nine free-response items that require elision at the word, syllable, or phoneme level, children indicate the word that results when a sound is removed from a word spoken by the examiner. On the 21-item Blending subtest, which includes nine multiple-choice items and 12 free-response items that require blending words, syllables, onset-rime, or phonemes, children indicate the word that results from combining sounds spoken by the examiner. Both subtests have good internal consistency (i.e., αs = .85-.87) and adequate concurrent validity correlations (i.e., rs > .53).
Print knowledge measures.
Children completed the Print Knowledge subtest of the PCTOPPP and the Alphabet subtest of the Test of Early Reading Ability, 3rd Edition (TERA-3; Reid et al., 2001). The Print Knowledge subtest of the PCTOPPP has 36 items, which include both multiple-choice and free-response items, that assess children’s understanding of what letters are (e.g., versus pictures or non-letter symbols), that print can be read, and both letter names and letter sounds. Internal consistency reliability of the subtest is high (αs = .94-.95 for this age group), and scores on the subtest are correlated concurrently with other measures of print knowledge and word decoding (rs = .40-.75). The Alphabet subtest of the TERA-3 assesses children’s knowledge of the alphabet, sound-letter correspondence, and rudimentary word-decoding skills. Internal consistency reliability for the TERA-3 is high (α ≥ .95).
Decoding Skills
In kindergarten, first grade, and second grade, children completed four measures of word-level reading, including the Word Identification (WID) and the Word Attack (WA) subtests of the Woodcock Reading Mastery-Revised (WRM-R; Woodcock, 1998) and the Sight Word Efficiency (SWE) and the Phonemic Decoding Efficiency (PDE) subtests of the TOWRE (Torgesen et al., 1999). Both the WID and WA subtests of the WRM-R are untimed measures of the ability to accurately read aloud single real words (WID) or pseudowords (WA) of increasing difficulty. WID has 106 items, but testing is discontinued after six consecutive errors. Split-half reliability coefficients are high (i.e., .91 - .98). WA has 45 items that are pseudowords or words of low frequency of occurrence in the English language ordered by increasing in difficulty; testing is discontinued after six consecutive errors. Split-half reliability coefficients are high (i.e., .89 - .94). The Sight-Word Efficiency subtest of the TOWRE requires children to read aloud as many words as they can read from a list of 104 real words. Scores are the number of words read correctly within 45 seconds. The Phonemic Decoding Efficiency subtest of the TOWRE requires children to read aloud as many nonwords as they can read from a list of 63 pseudowords. Scores are the number of nonwords read correctly within 45 seconds. Reliability of both subtests is high (alternate-form rs = .93 - .97; test-retest rs = .84 - .97) for the age-range of children in this sample.
Procedure
Parents of all participating children provided written informed permission/consent for their children’s participation, and all classroom teachers consented to participate. As part of the larger study, recruitment of children occurred in the fall of the preschool year, and children completed a variety of assessments at the beginning, middle, and end of their preschool year. Only the end-of-preschool assessments were used in this study. These children completed reading- and math-related assessments during the spring terms of their kindergarten, first-grade, and second-grade years. All assessments were administered to children in a quiet area of their preschools and schools by research assistants trained to criterion performance on the measures through didactic presentations, modeling, a performance assessment of test administration to an adult posing as a child, and live observations of assessment with feedback. Assessments were conducted over several 30- to 45-minute sessions, within a two- to three-week period.
Results
All analyses were conducted in Mplus 7.11 (Muthén & Muthén, 2013) using full information maximum likelihood and robust maximum likelihood estimation to account for missing data and non-normal variable distributions. The cluster-robust sandwich estimator option was used to correct for the nesting of children within their preschools. Raw scores on all measures were used in the analyses. Only raw scores were available for the PCTOPPP, and raw scores were used from other measures for consistency. The preschool measurement model included three factors: Vocabulary (Receptive and Definitional Vocabulary subtests from PCTOPPP and Basic Concepts subtest from CELF-P), Phonological Awareness (Elision and Blending subtests from PCTOPPP), and Print Knowledge (Print Knowledge subtest from PCTOPPP and Alphabet subtest from TERA). This model provided a very good fit to the data, χ2(11) = 18.47, p = .07, with comparative fit index (CFI) = 1.00, Tucker-Lewis index (TLI) = 1.00, and root mean square error of approximation (RMSEA) = .03 (90% confidence interval: .00 - .05), and it fit better than alternative measurement models (see Table S4 in supplemental materials). Correlations between factors were moderate to high (i.e., Vocabulary with Phonological Awareness = .87, Vocabulary with Print Knowledge = .66, Phonological Awareness with Print Knowledge = .72).
Longitudinal structural models were examined to determine the zero-order and unique contributions of Phonological Awareness and Print Knowledge factors to kindergarten, first-grade, and second-grade Decoding factors. Models examining the unique contributions of the preschool constructs to later decoding were constructed by creating latent variables representing the residuals of the Decoding, Phonological Awareness, and Print Knowledge latent variables (residualized on the Vocabulary latent variable) and then using these residuals as dependent and independent variables in the analyses (Koch et al., 2018). All structural models had adequate fit (i.e., CFI ≥ .96, TLI ≥ .94, and RMSEA ≤ .08). Table 4 shows the zero-order and different unique association of preschool predictors with kindergarten, first-grade, and second-grade Decoding. As seen in the second column (i.e., raw-raw column) of Table 4, all preschool predictors were substantial longitudinal correlates of Decoding across grades. The longitudinal predictive association between both phonological awareness and print knowledge was reduced when predicting decoding after accounting for vocabulary (i.e., raw-residual column of Table 3) and to a lesser extent using phonological awareness or print knowledge with the overlap with vocabulary removed when predicting decoding after accounting for vocabulary (i.e., the residual-residual column of Table 4).
Discussion
The results of Study 3 were highly similar to those of Study 2, despite the use of different measures, a sample whose initial assessments occurred prior to the onset of substantial formal instruction in reading, longitudinal prediction, and the inclusion of a measure of early understanding of print. The best predictors of unexpected decoding were unexpected phonological awareness and unexpected print knowledge. Conversely, the best predictors of decoding were phonological awareness and print knowledge. Somewhat surprisingly, this pattern extended to print knowledge, which should be less influenced by language skills than is phonological awareness; however, the decrease in predictive association depending on whether the predictor and criterion matched on unexpected versus straight performance appeared to be less for print knowledge than it was for phonological awareness.
General Discussion
Three studies were carried out that tested the hypothesis that what optimally predicts poor reading may not optimally predict unexpected poor reading. Study 1 was a model-based meta-analysis of key predictors of decoding and reading comprehension. The results were that measures of unexpected performance on the predictor of phonological awareness predicted criterion measures based on unexpected performance significantly better than they did criterion measures based on simple performance. Conversely, measures of simple performance predicted criterion measures based on simple performance significantly better than they did criterion measures based on unexpected performance. This was true for both decoding and reading comprehension as the criterion variables. Studies 2 and 3 were large-scale empirical studies with constructs represented by latent rather than observed variables, and unexpected performance was measured by residuals from regressing latent constructs on each other as opposed to simple differences between observed variables as was the case for Study 1. The pattern of results in both studies mirrored those in Study 1, although the differences were somewhat attenuated.
Referring back to the scatterplot between listening and reading comprehension presented in Figure 1, the distributions of listening comprehension, reading comprehension, and reading comprehension unexpected on the basis of listening comprehension are all multivariate normally distributed, and we treated them as continuous in our analyses. However, if one puts cut points on the distributions as was done in the figure, it is clear that the overlap between poor readers (i.e., cases represented by points to the left of the vertical line) and unexpected poor readers (cases above the diagonal line) is partial at best. It is therefore not surprising that optimal predictors of simple reading are unlikely to optimally predict unexpected reading.
The results of the three studies were remarkably consistent across a model-based meta-analysis, an empirical study using latent variables, and a second empirical study using latent variables with a longitudinal design. The decoding results are consistent with the triangle model of dyslexia in which learning to read at the word level involves connecting phonological, orthographic, and semantic features that represent individual words (Harm & Seidenberg, 1999). When parameters are adjusted to limit the ability of the simulation model to develop adequate phonological representations, the result is that performance in learning unfamiliar words (i.e., nonwords) is primarily affected. This mirrors the key deficit observed in children with phonologically based dyslexia (Castles & Coltheart, 1993), and explains why phonological awareness that is unexpected based on vocabulary knowledge is an important predictor of unexpected decoding.
A potential reservation about focusing on unexpected poor reading is that it seems to be an example of a definition based on discrepancy, and IQ-achievement discrepancy definitions have been shown to be problematic (Fletcher et al., 2019; Stanovich, 1991a, 1991b). There is a related argument that the cognitive and linguistic profiles of discrepant and non-discrepant poor readers are remarkably similar (Fletcher et al., 1994). Referring to Figure 1, the fact that phonological awareness does not differentiate poor readers (i.e., cases to the left of the vertical line) who are above the diagonal line representing unexpected impairment from poor readers who are below the diagonal line is consistent with the results of the current studies. What may well have differentiated the groups of poor and unexpected poor readers would be a measure of unexpected poor phonological awareness.
In addition, there are likely to be different developmental relations between phonological awareness and decoding for poor readers whose difficulty learning to read is caused by a deficit in phonological processing and poor readers whose deficit in phonological awareness may be a by-product of their poor reading. Developmental methodologies such as latent change score modeling can now distinguish leading and lagging indicators in studies of co-development, and there is support for different developmental relations for reading disabled and non-disabled controls. Ferrer et al. (2010) used this methodology to show that the kind of developmental coupling between cognition and reading that happens for typical readers is absent for individuals with dyslexic.
Turning to more practical problems associated with IQ-achievement discrepancy models, questions about the functional significance of IQ as it relates specifically to reading can be avoided by focusing on differences between listening and reading comprehension. Reading comprehension that is substantially below listening comprehension signals a potential benefit of using assistive technology.
Regarding the lessened reliability of difference scores relative to the scores that make up the difference score, this is a practical problem that can be addressed. What matters more than group-based measures of reliability that vary as a function of how variable the groups are is a measure of the precision of an individual’s score. If it were believed to be important to measure the difference between listening and reading comprehension, a computer adaptive sentence verification task could be used with sentences assigned randomly to reading or listening format. Presentation of sentences could continue until the precision of the difference met the desired criterion. At the word level, a similar approach could be used to measure differences between vocabulary and decoding or spelling by having individuals spell, read, and provide definitions for the same set of words. These approaches would also alleviate the problem that current measures of listening comprehension and reading comprehension differ in many respects and even trivial differences affect the reliability of difference scores (Macmann & Barnett, 1985). Having individuals do different tasks on the same set of words eliminates the problem of differential word knowledge affecting the correlation between task performance (Spencer et al., 2015).
Reconciling Science, Policy, and Practice
The fields of learning disabilities in general, and of dyslexia in particular, are remarkable examples of areas of science that have had profound effects on policy and practice regarding important problems in the everyday world. There are many areas of science that have little or no obvious application. When there are active interfaces between science, policy, and practice, it is important to understand their appropriate roles and relations. In general, we think science should inform policy and practice, recognizing that decisions about policy and practice necessarily involve other considerations such as resource limitations and fairness
For example, a consensus from the literature on dyslexia is that it occurs across the range of intellectual abilities, as expressed in the Rose (2009) report and the definition of dyslexia adopted by the British Dyslexia Association. We are aware of no research that indicates that having intellectual disabilities provides a protective factor that wards off learning disabilities including dyslexia. However, the definition of learning disabilities in the United States that was adopted in 1968 as part of Public Law 94-142 and continued with the Education of All Handicapped Children Act of 1975 included mental retardation as an exclusionary criterion by stating that the “The term does not include children who have learning disabilities, which are primarily the result of visual, hearing, or motor handicaps, or mental retardation, or emotional disturbance, or of environmental, cultural, or economic disadvantage.”
Science would suggest that some individuals with intellectual disability would also have dyslexia whereas others would not. An individual with both intellectual disability and dyslexia presumably would have more difficulty learning to decode words than would an individual with a comparable level of intellectual disability but not dyslexia. From a resource perspective, the argument in support of the exclusion might be that a child with intellectual disabilities will qualify for special education services on the basis of intellectual disability. This is true, but it is not obvious that a special education teacher who specializes in intellectual disability will receive training in identification and remediation of dyslexia. And by excluding individuals with intellectual disability from studies of dyslexia, there is a dearth of knowledge about effective word-level reading interventions for this population.
Students with above average language skills whose dyslexia pulls down their reading to somewhat below average may not qualify for services as a student with a specific learning disability if their deficit in reading does not meet the criterion of “the child does not achieve adequately for the child’s age or meet state-approved grade-level standards.” There are reasonable arguments to be made that limited resources should be reserved for students whose disability results in poor reading instead of only unexpected poor reading:
Imagine, for example, someone who is a near-genius in abilities overall but who is only above average in reading skills. If one views reading disability as defined in terms of the difference between IQ and reading skills, the near-genius actually might be classified as having a reading disability. In fact, the individual’s reading skills are way below the individual’s overall abilities, but so what? This is not an individual in whom one would want to invest any of society’s resources so as to bring his or her reading up to the level of his or her other abilities”
(Sternberg & Grigorenko, 2002, p. 74).
There is an obvious counterargument that it is indeed society’s role to help all children succeed to the maximum they can achieve, but regardless, we are aware of no evidence that average or above language skills or IQ provide a deterrent to also having dyslexia. At the very least, individuals with intellectual disabilities or above average language abilities should not be excluded from research on dyslexia.
Supplementary Material
Acknowledgments
The research described in this article was supported by Grant Number P50 HD52120 from the National Institute of Child Health and Human Development and Grant Number R305F100027 from the Institute of Education Sciences.
References
- Aaron PG (1991). Can reading disabilities be diagnosed without using intelligence tests? Journal of Learning Disabilities, 24(3), 178–186. 10.1177/002221949102400306 [DOI] [PubMed] [Google Scholar]
- American Psychiatric Association. (2013). Diagnostic and statistical manual of mental disorders (5th ed.). 10.1176/appi.books.9780890425596 [DOI] [Google Scholar]
- Badian NA (1999). Reading disability defined as a discrepancy between listening and reading comprehension: A longitudinal study of stability, gender differences, and prevalence. Journal of Learning Disabilities, 32(2), 138–148. 10.1177/002221949903200204 [DOI] [PubMed] [Google Scholar]
- Becker BJ, & Aloe AM (2019). Model-based meta-analysis and related approaches. In Cooper HM, Hedges LV, & Valentine JC (Eds.), The handbook of research synthesis and meta-analysis (pp. 339–363). Russell Sage. [Google Scholar]
- Bedford-Feuell C, Geiger S, Moyse S, & Turner M (1995). Use of listening comprehension in the identification and assessment of specific learning difficulties. Educational Psychology in Practice, 10(4), 207–214. 10.1080/0266736950100402 [DOI] [Google Scholar]
- Brownell R (2000a). Expressive One-Word Picture Vocabulary Test, third edition. Academic Therapy Publications, Inc. [Google Scholar]
- Brownell R (2000b). Receptive One-Word Picture Vocabulary Test. Academic Therapy Publications, Inc. [Google Scholar]
- Breaux KC (2009). Wechsler Individual Achievement Tests III (WIAT III) Technical Manual. Pearson. [Google Scholar]
- Carrow-Woolfolk E (2008). Comprehensive assessment of spoken language. Western Psychological Services. [Google Scholar]
- Carrow-Woolfolk E (2011). Oral and written language scales, (2nd ed.). Western Psychological Services. [Google Scholar]
- Castles A, & Coltheart M (1993). Varieties of developmental dyslexia. Cognition, 47(2), 149–180. 10.1016/0010-0277(93)90003-E [DOI] [PubMed] [Google Scholar]
- Cheung M (2015). Meta-analysis: a structural equation modeling approach. John Wiley & Sons. [Google Scholar]
- Connor CM, Piasta SB, Fishman B, Glasney S, Schatschneider C, Crowe E, Underwood P, & Morrison F (2009). Individualizing student instruction precisely: Effects of child x instruction interactions on first graders' literacy development. Child Development, 80(1):77–100. 10.1111/j.1467-8624.2008.01247.x [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dickman E (2017). Do we need a new definition of dyslexia? The Examiner (International Dyslexia Association), 6(1). https://dyslexiaida.org/do-we-need-a-new-definition-of-dyslexia/ [Google Scholar]
- Erbeli F, Hart SA, Wagner RK, & Taylor J (2018). Examining the etiology of reading disability as conceptualized by the hybrid model. Scientific Studies of Reading, 22(2), 167–180. 10.1080/10888438.2017.1407321 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Ferrer E, Shaywitz BA, Holahan JM, Marchione K, & Shaywitz SE. (2010). Uncoupling of reading and IQ over time: Empirical evidence for a definition of dyslexia. Psychological Science, 21(1):93–101. 10.1177/0956797609354084 [DOI] [PubMed] [Google Scholar]
- Fletcher JM, Lyon GR, Fuchs LS, & Barnes MA (2019). Learning disabilities: From identification to intervention (2nd ed.). Guilford. [Google Scholar]
- Fletcher JM, Shaywitz SE, Shankweiler DP, Katz L, Liberman IY, Stuebing KK, Francis DJ, Fowler AE, & Shaywitz BA (1994). Cognitive profiles of reading disability: Comparisons of discrepancy and low achievement definitions. Journal of Educational Psychology, 86(1), 6–23. 10.1037/0022-0663.86.1.6 [DOI] [Google Scholar]
- French BF, & Glascoe FP (2010). Brigance comprehensive inventory of basic skills II standardization and validation manual. Curriculum Associates. [Google Scholar]
- Graham JW, Taylor BJ, Olchowski AE, & Cumsille PE (2006). Planned missing data designs in psychological research. Psychological Methods, 11(4), 323–343. 10.1037/1082-989X.11.4.323 [DOI] [PubMed] [Google Scholar]
- Grigorenko EL, Compton D, Fuchs L, Wagner RK, Willcutt E, & Fletcher JM (2020). Understanding, educating, and supporting children with specific learning disabilities: 50 years of science and practice. American Psychologist, 75(1), 37–51. 10.1037/amp0000452 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hammill DD, Wiederholt JL, & Allen EA (2006). Test of Silent Contextual Reading Fluency. Pro-Ed. [Google Scholar]
- Harcourt Educational Measurement (2004). Stanford technical data report. Pearson. [Google Scholar]
- Harm MW & Seidenberg MS (1999). Phonology, reading acquisition, and dyslexia: Insights from connectionist models. Psychological Review, 106(3), 491–528. 10.1037/0033-295X.106.3.491 [DOI] [PubMed] [Google Scholar]
- Hoeft F, McCardle P, & Pugh K (2015). The Myths and Truths of Dyslexia in Different Writing Systems. International Dyslexia Association. The Examiner. https://dyslexiaida.org/the-myths-and-truths-of-dyslexia/ [Google Scholar]
- Hoover HD, Dunbar SB, & Frisbie DA (2003). The Iowa tests: Guide to research and development. Riverside Publishing. [Google Scholar]
- Kaufman AS, & Kaufman NL (with Breaux KV) (2014). Technical and Interpretive Manual. Kaufman Test of Educational Achievement–Third Edition (KTEA-3). NCS Pearson. [Google Scholar]
- Koch T, Holtmann J, Bohn J, & Eid M (2018). Explaining general and specific factors in longitudinal, multimethod, and bifactor models: Some Caveats and Recommendations. Psychological Methods, 23(3), 505–523. 10.1037/met0000146 [DOI] [PubMed] [Google Scholar]
- Lonigan CJ, Burgess SR, & Schatschneider C (2018). Examining the simple view of reading with elementary school children: Still simple after all these years. Remedial and Special Education, 39(5), 260–273. 10.1177/0741932518764833 [DOI] [Google Scholar]
- Lonigan CJ, Wagner RK, and Torgesen JK (2007). Test of Preschool Early Literacy. Pro-Ed. [Google Scholar]
- Lonigan CJ, Wagner RK, Torgesen JK, & Rashotte C, (2002). Preschool Comprehensive Test of Phonological and Print Processing. Florida State University. [Google Scholar]
- Lyon GR, Shaywitz SE, & Shaywitz BA (2003). A definition of dyslexia. Annals of Dyslexia, 53, 1–14. 10.1007/s11881-003-0001-9 [DOI] [Google Scholar]
- Macmann GM, & Barnett DW (1985). Discrepancy score analysis: A computer simulation of classification stability. Journal of Psychoeducational Assessment, 3(4), 363–375. 10.1177/073428298500300409 [DOI] [Google Scholar]
- McGrew KS, LaForte EM, & Schrank FA (2014). Technical Manual. Woodcock-Johnson IV. Riverside. [Google Scholar]
- Miciak J, & Fletcher JM (2020). The critical role of instructional response for identifying dyslexia and other learning disabilities. Journal of Learning Disabilities, 53(5), 343–353. 10.1177/0022219420906801 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Moll K, Kunze S, Neuhoff N, Bruder J, & Schulte-Korne (2014). Specific learning disorder: Prevalence and gender differences. PLoSOne, 9(7): e103537 10.1371/journal.pone.0103537 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Morgan WP (1896). A case of congenital word-blindness (inability to learn to read). British Medical Journal, 2, 1543–1544. doi: 10.1136/bmj.2.1871.1378 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Muthén LK, & Muthén BO (2013). Mplus (Version 7.11). Muthén & Muthén. [Google Scholar]
- Odegard TN, Farris EA, & Washington JA (2022). Exploring boundary conditions of the listening comprehension-reading comprehension discrepancy index. Annals of dyslexia, 72, 301–323. 10.1007/s11881-021-00250-0 [DOI] [PubMed] [Google Scholar]
- Petersen RL, & Pennington BF (2012). Developmental dyslexia. The Lancet, 379(9830), 1997–2007. 10.1016/S0140-6736(12)60198-6 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Reid D, Hresko W, & Hammill D (2001). Test of Early Reading Ability (third edition). Pro-Ed. [Google Scholar]
- Rose J, 2009. Identifying and teaching children and young people with dyslexia and literacy difficulties. An independent report. https://dera.ioe.ac.uk/14790/7/00659-2009DOM-EN_Redacted.pdf [Google Scholar]
- Semel E, Wiig EH, & Secord WA (2003). Clinical Evaluation of Language Fundamentals - fourth edition: CELF-4. Pearson. [Google Scholar]
- Shaywitz SE, Escobar MD, Shaywitz BA, Fletcher JM, & Makuch R (1992). Evidence that dyslexia may represent the lower tail of the normal distribution of reading ability. New England Journal of Medicine, 326, 145–150. 10.1056/NEJM199201163260301 [DOI] [PubMed] [Google Scholar]
- Snowling MJ, & Melby-Lervåg M (2016). Oral language deficits in familial dyslexia: A meta-analysis and review. Psychological Bulletin, 142(5), 498–545. 10.1037/bul0000037 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Spencer M, Muse A, Wagner RK, Foorman B, Petscher Y, Schatschneider C, Tighe EL, & Bishop D (2015). Examining the underlying dimensions of morphological awareness and vocabulary knowledge. Reading and Writing, 28, 959–988. 10.1007/s11145-015-9557-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Spring C, & French L (1990). Identifying children with specific reading disabilities from listening and reading discrepancy scores. Journal of Learning Disabilities, 23(1), 53–58. 10.1177/002221949002300112 [DOI] [PubMed] [Google Scholar]
- Stanovich KE (1991a). Discrepancy definitions of reading disability: Has intelligence led us astray? Reading Research Quarterly, 26(1), 7–29. 10.2307/747729 [DOI] [Google Scholar]
- Stanovich KE (1991b). Conceptual and empirical problems with discrepancy definitions of reading disability. Learning Disability Quarterly, 14(4), 269–280. 10.2307/1510663 [DOI] [Google Scholar]
- Sternberg RJ, & Grigorenko EL (2002). Difference scores in the identification of children with learning disabilities. it's time to use a different method. Journal of School Psychology, 40(1):65–83. 10.1016/S0022-4405(01)00094-2 [DOI] [Google Scholar]
- The Psychological Corporation. (2003). Early reading diagnostic assessment technical manual (ERDA) (2nd ed.). Pearson. [Google Scholar]
- Torgesen JK (2000). Individual differences in response to early interventions in reading: The lingering problem of treatment resisters. Learning Disabilities Research & Practice, 15(1), 55–64. 10.1207/SLDRP1501_6 [DOI] [Google Scholar]
- Torgesen JK, Wagner RK, & Rashotte CA (1999). Test of Word Reading Efficiency: TOWRE. Pro-Ed. [Google Scholar]
- U.S. Office of Education. (1968). First annual report of the National Advisory Committee on Handicapped Children, Washington, DC: U.S. Department of Health, Education and Welfare. [Google Scholar]
- U.S. Office of Education. (1977). Assistance to states for education for handicapped children: Procedures for evaluating specific learning disabilities. Federal Register, 42, G1082–G1085. [Google Scholar]
- Wagner RK, Edwards AA, Malkowski A, Schatschneider C, Joyner RE, Wood S, & Zirps FA (2019). Combining old and new for better understanding and predicting dyslexia. New Directions for Child and Adolescent Development, 165, 1–11. 10.1002/cad.20289 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wagner RK, Torgesen JK, & Rashotte CA (1999). Comprehensive Test of Phonological Processing (CTOPP). Pro-Ed. [Google Scholar]
- Wagner RK, Zirps FA, Edwards AA, Wood SG, Joyner RE, Becker BJ, Liu G, & Beal B (2020). The prevalence of dyslexia: A new approach to its estimation. Journal of Learning Disabilities, 53(5), 354–365. 10.1177/0022219420920377 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wagner RK, Zirps FA, & Wood SG (2022). Developmental dyslexia. In Snowling MJ, Hulme C, & Nation K (Eds.), The science of reading: a handbook (2nd. Ed., pp. 416–438). Wiley Blackwell. [Google Scholar]
- Wiig EH, Second W, & Semel E (1992). Clinical Evaluation of Language Fundamentals – Preschool. Psychological Corporation. [Google Scholar]
- Wood SG, Moxley JH, Tighe EL, Wagner RK (2018). Does use of text-to-speech and related read-aloud tools improve reading comprehension for students with reading disabilities? A meta-analysis. Journal of Learning Disabilities, 51(1), 73–84. 10.1177/0022219416688170 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Woodcock RW (1998). Woodcock Reading Mastery Tests – Revised (WRM-R). American Guidance. [Google Scholar]
- Woodcock RW (2011). Woodcock Reading Mastery Tests, third edition manual (WRMT III). NCS Pearson. [Google Scholar]
- Woodcock RW, McGrew KS, & Mather N (2001). Woodcock-Johnson III Tests of Achievement (WJIII). Riverside Publishing. [Google Scholar]
- World Health Organization (2019). International Statistical Classification of Diseases and Related Health Problems (11th ed.). https://icd.who.int/ [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.

