Abstract
Language impairment (LI) is highly heritable and aggregates in families. Genetic investigation of LI has revealed many chromosomal regions and genes of interest, though very few studies have focused on rare variant analysis in non-English speaking or non-European samples. We selected four candidate genes (TM4SF20, NFXL1, CNTNAP2 and ATP2C2) strongly suggested for specific language impairment (SLI), a subtype of LI, and investigated rare protein coding variants through Sanger sequencing of probands with LI ascertained from Pakistan. The probands and their family members completed a speech and language family history questionnaire and a vocabulary measure, the Peabody Picture Vocabulary Test-fourth edition (PPVT-4), translated to Urdu, the national language of Pakistan. Our study aimed to determine the significance of rare variants in these SLI candidate genes through segregation analysis in a novel population with a high rate of consanguinity. In total, we identified 16 rare variants (according to the rare MAF in the global population in gnomAD v2.1.1 database exomes), including eight variants with a MAF <0.5 % in the South Asian population. Most of the identified rare variants aggregated in proband’s families, one rare variant (c.*9T>C in CNTNAP2) co-segregated in a small family (PKSLI-64) and another (c.2465C>T in ATP2C2) co-segregated in the proband branch (PKSLI-27). The lack of complete co-segregation of most of the identified rare variants indicates that while these genes could be involved in overall risk for LI, other genes are likely involved in LI in this population. Future investigation of these consanguineous families has the potential to expand our understanding of gene function related to language acquisition and impairment.
Keywords: consanguinity, candidate gene investigation, language impairment, Pakistan
1. Introduction
The innateness of language has been the topic of a long-standing debate, rooted in language being a uniquely human trait (Chomsky, 1957; Lenneberg, 1967). With advances in molecular genetic methods, the innateness of language debate became part of biological inquiry, specifically through the investigation of individuals with language impairments (LI) and their families (Andres et al., 2020; Andres et al., 2019; Bartlett et al., 2004; Bartlett et al., 2002; Fisher et al., 1998; Rice et al., 1998; SLI Consortium, 2002, 2004; Tallal et al., 2001; Villanueva et al., 2011; Villanueva et al., 2008). Individuals with specific LI (SLI), a subtype of LI, struggle to acquire language despite average non-verbal intelligence (NV-IQ; National Institute on Deafness and Other Communication Disorders, 2019). Behavioral genetic studies supported the involvement of genetic factors in the expression of language traits (Bishop et al., 1995; Rice et al., 1998; Stromswold, 1998, 2001; Tallal et al., 2001). The estimated prevalence of SLI is 7-10 % in English-speaking populations (Norbury et al., 2016; Tomblin et al., 1997). Individuals with SLI show persistent delays in both receptive and expressive language compared to their age-matched peers and these delays remain throughout development (Rice, 2012, 2017; Rice & Hoffman, 2015). Once individuals with SLI begin to acquire language, their language development follows the same trajectory as individuals with typical language development, indicating that the key to understanding the delay could be in the genes that play a role in the earliest stages of development (Rice, 2012). However, our understanding of the genes involved in SLI remains limited.
Several candidate genes have been suggested for SLI, though only ATP2C2 (chromosome 16q24.1) and CNTNAP2 (chromosome 7q35-q36) have been replicated independently (Chen et al., 2017; Martinelli et al, 2021; Newbury et al., 2010; Newbury & Monaco, 2010; Newbury et al., 2011; Newbury et al., 2009; Reader et al., 2014; SLI Consortium, 2002, 2004; Smith, 2007; Vernes et al., 2008; Wang et al., 2015; Whitehouse et al., 2011). ATP2C2 encodes calcium-transporting ATPase, type 2c, member 2, which removes calcium and manganese from the cell (Missiaen et al., 2007; Newbury et al., 2009). Calcium regulates many neuronal processes including working memory and synaptic plasticity, while manganese dysregulation has been linked to other neurological disorders, such as epilepsy (Carl et al., 1986; Friedman et al., 2008; Mefford et al., 2010; Newbury & Monaco, 2010; Newbury et al., 2009; Strauss et al., 2006). CNTNAP2 (chromosome 7q35-q36) is a downstream target of FOXP2 (Fisher & Scharff, 2009), a gene implicated as causative for severe speech and language impairments in the KE family (Fisher et al., 1998). CNTNAP2 encodes neurexin, which plays an important role in connecting neurons and cortex development in human brains. In situ hybridization analysis shows elevated expression of FOXP2 and CNTNAP2 in fetal brain tissues (Rodenas-Cuadrado et al., 2014; Vernes et al., 2008). It is important to note that ATP2C2 and CNTNAP2 were identified and replicated in English-speaking samples (Newbury et al., 2010; Newbury & Monaco, 2010; Newbury et al., 2011; Newbury et al., 2009; Whitehouse et al., 2011). Language abilities are universal to humans. To gain a better understanding of biological markers involved in LI, behavioral and genetic studies of SLI need to expand to non-English speaking populations.
Some critical genetic findings of LI have emerged from non-Westernized and/or non-English speaking populations, setting the precedent for future genetic study of SLI to include behaviorally and genetically diverse populations. Two such investigations of LI resulted in the candidate genes, NFXL1 (chromosome 4p12) and TM4SF20 (chromosome 2q36.3). The former was identified through the study of an isolated founder population living on Robinson Crusoe Island off the coast of Chile, who speak Spanish and have a high rate of SLI (35%; Villanueva et al., 2008). Genome-wide linkage analysis and whole-exome analysis of the founder population identified rare variants in a novel candidate gene, NFXL1 (chromosome 4p12 (Villanueva et al., 2011; Villanueva et al., 2015). Functional studies reported elevated expression of NFXL1 in the cerebellum, which is predicted to be involved in higher cognitive function and in language development (Nudel, 2016). TM4SF20 was identified through genome-wide analysis in more than 15,000 children of numerous ethnicities with behavioral or motor developmental delays (Wiszniewski et al., 2013). A 4kb deletion on chromosome 2 was identified that also deleted a coding exon of TM4SF20 among individuals (probands) of Southeast Asian descent with language delay, many of whom were recent US immigrants and non-English speaking (Wiszniewski et al., 2013). The probands’ family members were sequenced for the TM4SF20 deletion, and the deletion was determined to be population specific to Southeastern subpopulations (Wiszniewski et al., 2013). Brain imaging of the probands and their family members showed that the presence of white matter hyperintensities (WMH) co-segregated with the 4kb deletion and language delay (Wiszniewski et al., 2013). These studies underscore the role of family-based investigation and the importance of expanding inclusion in the study of LI, both genetically and behaviorally.
Our knowledge of TM4SF20, NFXL1, CNTNAP2, and ATP2C2 and their association with LI in other populations is limited. In the current study, we performed Sanger sequencing of the protein coding exons of these genes in Pakistani probands with a history of LI. Probands included in this study are members of large consanguineous families. Consanguinity favors the expression of recessive alleles and increases the likelihood of genetic disorders. In fact, researchers have recently called for an increase in the study of consanguineous families, arguing that they may be the key to expanding our knowledge base and understanding of gene functions (Erzurumluoglu et al., 2016). Our preliminary study will identify the rate of rare variants in four candidate genes previously suggested in SLI using non-English speaking Pakistani probands and extended families with a history of LI.
2. Materials and Methods
2.1. Participants
We selected probands (N=38) from unrelated families with a history of LI and performed targeted Sanger sequencing of four candidate genes previously suggested for SLI. The probands ranged from 6 years 6 months to 18 years 2 months of age at the time of assessment. Additionally, Sanger sequencing was performed in available family members of the 16 unrelated probands, in which rare variants were identified in the candidate genes (Fig. 1 and Fig. 2). DNA was available from 28 parents, 38 siblings and 31 extended family members of the 16 unrelated probands (Fig. 1 and Fig. 2). Sequencing of variants in two individuals was unsuccessful (40_007 and 51_007; Fig. 2). The study was approved by the University of Kansas institutional review board (IRB #STUDY00143136).
Fig. 1.

Aggregation of identified candidate gene rare variants (MAF <0.5% in South Asians based on gnomAD v2.1.1 exomes) in probands’ families.
Note. Each pedigree is labelled on top with a unique ID. The cDNA and amino acid change are provided for each variant above the pedigree ID. PKSLI-10 is also shown on Fig. 2, as one variant observed has a MAF >0.5% in South Asians.
Fig. 2.

Aggregation of identified candidate gene rare variants (MAF <0.5% in Europeans and >0.5% in South Asians based on gnomAD v2.1.1 exomes) in probands’ families.
Note. Each pedigree is labelled on top with a unique ID. The cDNA and amino acid change are provided for each variant above the pedigree ID. PKSLI-04 is shown twice in this figure, as one variant observed on NFXL1 is shared with PKSLI-10. PKSLI-10 is also shown on Fig. 1, as one variant observed has a MAF <0.5% in South Asians.
2.1.1. Phenotype assessment
Examiners interviewed the parents using the Rice Family History Questionnaire and Interview Form (PhenX Toolkit protocol #200401; Hamilton et al., 2011; Rice et al., 1998). The family history questionnaire includes three questions about the target child/proband, followed by yes/no questions that account for any family member having language or academic related difficulties. The interview form provides a grid to detail if the probands and their family members have language or academic related difficulties, including whether the individual was slow to talk or had a learning disability (LD), intellectual disability (ID), hearing impairment or autism (only based on parent report, no clinical diagnoses).
Each available family member completed the Peabody Picture Vocabulary Test, fourth edition, (Dunn & Dunn, 2007) translated to Urdu (U-PPVT-4) and their standard scores were determined based on the US norms, with adjusted standard score cutoffs to define affected status as ≤ 80 in males and ≤ 75 in females, as used previously with this sample (Andres et al., 2019). The standard score cut-off for females was reduced to account for their reduced literacy rate compared to males in Pakistan (Chaudhry & Rahman, 2009). In the current study, all 38 probands completed the U-PPVT-4. In the 16 unrelated proband families used for genetic analysis in this study, U-PPVT-4 scores were available from 22 parents, 37 siblings, and 23 extended family members. In total, U-PPVT-4 scores from 17 individuals were unavailable and were shaded in gray (unknown language phenotype) on the pedigrees (Fig. 1). A population-based study of a control sample targeting age-stratified groups of individuals in Pakistan with typical language is in progress. Our preliminary data of the U-PPVT continues to indicate encouraging reliability and validity, as it did in our initial investigation of this clinical sample (on-going study; Andres et al., 2019).
2.1.2. Proband identification
Probands (N=38) were identified through schools and community centers in Punjab, Pakistan. We sent a brief introductory letter to schoolteachers and community leaders, which included information to identify children with limited language abilities compared to their same-age peers. The identification information was based on the family history questionnaire developed by Rice and colleagues (1998). We contacted families for behavioral data and DNA collection once participants’ information was received. In the current study, we assigned proband status to the participants based on the information provided by teachers or community leaders and parent report about the family history of speech and language abilities. The affected status was assigned based on the U-PPVT-4 performance (Fig. 1).
2.2. Genetic analyses
Saliva samples were collected from participants at the time of behavioral assessment, using the Oragene-Discover OGR-500 Kit from DNA Genotek (Oragene) and we extracted DNA following the standard manufacturer protocol. The DNA was not available from all the extended family members who completed behavioral assessment (Fig. 1).
Four candidate genes were selected for Sanger sequencing based on the genetics of SLI literature: TM4SF20 (NM_024795.4), NFXL1 (NM_152995.5), CNTNAP2 (NM_014141.5) and ATP2C2 (NM_001286527.2). Coding exons of the four selected candidate genes were sequenced in 38 probands through Sanger DNA sequencing using the ABI 3130xl Genetic Analyzer. The number of protein coding exons sequenced on each gene was: TM4SF20 (4), NFXL1 (22), CNTNAP2 (24) and ATP2C2 (28). Primers were designed for each exon using RefSeq from the UCSC Genome Browser and the online interface, Primer 3. Primers were designed to capture each end of the exon, which required sequencing the flanking intronic regions. Subsequently, Sanger sequencing of identified variants was performed in the available family members of probands in which rare variants were identified. The sequencing data were analyzed using SeqMan Pro, a program within the DNAStar suite. Reported variants were considered rare with a minor allele frequency (MAF) <0.5% in Europeans in the Genome Aggregation Database (gnomAD) v2.1.1 exomes (Table 1; Table 2; Karczewski et al., 2020). We highlighted a subset of rare variants with the MAF <0.5% in South Asians in gnomAD v2.1.1 exomes (Table 1). The variants’ MAFs from additional databases, including the 1000 Genomes Project, the NHLBI GO Exome Sequencing Project (ESP6500), Iranome and GenomeAsia 100K, was also reported when available (Table 1; Table 2; Clarke et al., 2017; Fattahi et al., 2019; GenomeAsia 100K Consortium, 2019). The Iranome and GenomeAsia 100K databases seek to increase the inclusion of Asians in genomic work; Asians represent > 40% of the world’s population but are less prevalent in the common genomic databases (Fattahi et al., 2019; GenomeAsia 100K Consortium, 2019). At the time of this publication, the Iranome database includes over 800 Iranians and the GenomeAsia 100K database has about 3500 samples, including more than 100 individuals from Pakistan (Fattahi et al., 2019; GenomeAsia 100K Consortium, 2019).
Table 1.
Rare variants identified in Pakistani probands with MAF < 0.5% in South Asians based on gnomAD v2.1.1 exomes.
| Chr & Gene |
Genomic position (hg38) |
c.DNA | AA change | SNP ID | Variant type | Proband | gnomAD
v2.1.1 exomes |
1000 Genomes (all) |
ESP6500 (all) |
Iranome | GenomeAsia | ||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| European | South Asian |
all | 100K South Asian |
||||||||||
| Chr4: NFXL1 | 47848167 | c.2732A>T | p.(N911I) | rs775580487 | Non-synonymous | 10_004 | 0 | 0.0001 | NA | NA | NA | NA | NA |
| Chr7: CNTNAP2 | 146116890 | c.14C>T | p.(P5L) | rs779710846 | Non-synonymous | 28_004 | 0 | 0.0004 | NA | NA | NA | NA | NA |
| 147121078 | c.854G>C | p.(G285A) | rs150918383 | Non-synonymous | 13_004 | 0.006^2 | 0.0009 | 0.0003 | 0.0033 | 0.0006 | NA | NA | |
| 148147553 | c.2617T>A | p.(S873T) | rs776896133 | Non-synonymous | 45_005 | 0 | 0.0002 | NA | NA | NA | 0.000287 | 0.00 | |
| 147869433 | c.2873G>C | p.(G958A) | NA | Non-synonymous | 15_003 | NA | NA | NA | NA | NA | NA | NA | |
| 148415625 | c.*9T>C | NA | rs539868299 | 3’UTR | 64_005 | 0 | 0.002^2 | 0.001 | NA | NA | 0.00115 | 0.0027 | |
| Chr16: ATP2C2 | 84448536 | c.1507C>T | p.(Q503X) | rs758765955 | Nonsense (Stop gain) | 28_004 | 0 | 0.00003 | NA | NA | NA | NA | NA |
| 84460698 | c.2465C>T | p.(P822L) | rs757310826 | Non-synonymous | 27_005 | 0 | 0.0001 | NA | NA | NA | NA | NA | |
Note. chr=chromosome, AA=amino acid, NA=Not available
homozygotes in gnomAD v2.1.1 database presented as exponents, 1000 Genomes, ESP6500, Iranome, GenomeAsia 100K – acquired on July 28, 2021
Table 2.
Rare variants identified in Pakistani probands with MAF < 0.5% in Europeans and > 0.5% in South Asians based on gnomAD v2.1.1 exomes.
| Chr & Gene |
Genomic position (hg38) |
c.DNA | AA change |
SNP ID | Variant type | Proband | gnomAD
v2.1.1 exomes |
1000 Genomes (all) |
ESP6500 (all) |
Iranome | GenomeAsia | ||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| European | South Asian |
all | 100K South Asian |
||||||||||
| Chr4: TM4SF20 | 227363746 | c.668A>C | p.(K223T) | rs137891000 | Non-synonymous | 03_005, 24_004 | 0.005^2 | 0.03^28 | 0.0054 | 0.0035 | 0.02375 | 0.007475 | 0.01174 |
| Chr4: NFXL1 | 47862904 | c.2258C>G | p.(T753R) | rs370816326 | Non-synonymous | 04_003, 10_004 | 0.00001 | 0.02^8 | 0.0056 | 0.000077 | 0.001884 | 0.009504 | 0.02011 |
| Chr7: CNTNAP2 | 146116828 | c.−49T>G | NA | rs549396215 | 5’UTR | 40_002, 51_004 | 0.001 | 0.02^3 | 0.003 | NA | 0.02642 | 0.007475 | 0.015883 |
| 147977962 | c.2356G>T | p.(V786L) | rs138517537 | Non-synonymous | 17_003, 44_002 | 0.00002 | 0.03^18 | 0.0086 | 0.000076 | 0.004375 | NA | NA | |
| Chr16: ATP2C2 | 84425752 | c.937G>A | p.(G313S) | rs531589939 | Non-synonymous | 04_003 | 0 | 0.006^4 | 0.000399 | NA | NA | 0.005817 | 0.007779 |
| 84425768 | c.953A>G | p.(K318R) | rs551531561 | Non-synonymous | 04_003 | 0 | 0.006^4 | 0.0004 | NA | NA | 0.003347 | 0.007779 | |
| 84439438 | c.1123G>A | p.(V375I) | rs202026876 | Non-synonymous | 43_001, 52_005 | 0.002 | 0.01^3 | 0.002 | 0.00109 | NA | 0.0046 | 0.009668 | |
| 84462127 | c.2807T>C | p.(L936P) | rs16973859 | Non-synonymous | 04_003 | 0.0002 | 0.01^5 | 0.017 | 0.01527 | 0.00501 | 0.006612 | 0.009668 | |
Note. chr=chromosome, amino acid, NA= Not available
homozygotes in gnomAD v2.1.1 database presented as exponents, 1000 Genomes, ESP6500, Iranome, GenomeAsia 100K – acquired on July 28, 2021
The functional effect of each exonic variant observed was estimated in five in silico programs (Table 3). The in silico programs included SIFT (Sorting Intolerant from Tolerant), PolyPhen-2 (Polymorphism Phenotyping v2), Mutation assessor, PROVEAN (Protein Variation Effect Analyzer) and MutationTaster2 (Adzhubei et al., 2010; Choi & Chan, 2015; Choi et al., 2012; Reva et al., 2011; Schwarz et al., 2014; Sim et al., 2012). Each of the five in silico programs mentioned above provide a prediction score to describe the functional effect of the variant. In addition, each exonic variant was automatically analyzed in the HOPE (Have (y)Our Protein Explained) web application, which provides context of the structural effects of the amino acid (AA) change by showing a 3D model and considering the conservation at the variant location (Venselaar et al., 2010).
Table 3.
| Variants with MAF <0.5% in South Asians based on gnomAD
v2.1.1 exomes |
Variants with
MAF <0.5% in Europeans and >0.5% in South
Asians based on gnomAD v2.1.1 exomes |
|||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Gene – chr | NFXL1 – chr4 |
CNTNAP2 – chr7 | ATP2C2 – chr16 | TM4SF20 – chr2 |
NFXL1 – chr4 |
CNTNAP2 – chr7 |
ATP2C2 – chr16 | |||||||||
| rsID/ chromosomal location |
rs775580487 | rs779710846 | rs150918383 | rs776896133 | chr7:147,869,433 | rs59868299 (3’UTR) |
rs758765955 | rs757310826 | rs137891000 | rs370816326 | rs549396215 (5’UTR) |
rs138517537 | rs531589939 | rs551531561 | rs202026876 | rs16973859 |
| SIFT (0 < s < 0.05) |
0.01 (D) |
0.32 (T) |
0.05 (D) |
0.04 (D) |
0.01 (D) |
NA | NA (D-STOP) |
0 (D) |
0.1 (T) |
0.06 (T) |
NA | 0.92 (T) |
0.01 (D) |
0.77 (T) |
0.04 (D) |
0.2 (T) |
| PolyPhen-2 (0.85 < s < 1) |
0.003 benign |
0 benign |
0.587 possibly damaging |
0.437 benign |
1.0 probably damaging |
NA | NA | 1.0 probably damaging |
0.954 possibly damaging |
0.82 possibly damaging |
NA | 0.288 benign |
0.988 probably damaging |
0.109 benign |
0.942 possibly damaging |
0.823 possibly damaging |
| Mutation
assessor (s > 3.5) |
0 (N) |
0 (N) |
2.165 (M) |
1.23 (L) |
2.74 (M) |
NA | NA | 4.45 (H) |
2.11 (M) |
1.545 (L) |
NA | 0.475 (N) |
2.435 (M) |
−0.365 (N) |
1.775 (L) |
1.59 (L) |
| PROVEAN (s ≤ −2.5) |
−0.33 (N) |
−0.88 (N) |
−2.67 (D) |
−2.30 (N) |
−5.23 (D) |
NA | −4.03 (D) |
−9.81 (D) |
−2.04 (N) |
−2.28 (N) |
NA | −0.63 (N) |
−5.51 (D) |
−0.49 (N) |
−0.84 (N) |
−2.75 (D) |
| Mutation Taster2 | 149 P (0.999) |
98 P (0.999) |
60 DC (0.999) |
58 DC (0.999) |
60 DC (0.999) |
NA | 6 DC (1.0) |
98 DC (0.999) |
78 P (0.999) |
71 P (0.606) |
NA | 32 DC (0.999) |
56 DC (0.999) |
26 DC (0.999) |
29 DC (0.999) |
98 P (0.994) |
Note. D = deleterious, T = tolerated, L = low [functional impact], M = medium [functional impact], N = neutral/[neutral functional impact], H = high [functional impact]; MutationTaster2: Grantham matrix score (0 – 215; amino acid comparison), P = polymorphism, DC = disease causing, probability (0.0-1.0) of the ‘security’ of the prediction (closer to 1 = more secure)
3. Results
3.1. Phenotype summary
Most parents reported that the probands were slow to talk and struggled with reading, spelling, language and/or learning. The parents of three probands (01_005, 19_002, 50_006) reported ID, while three others (18_003, 51_004, 57_002) reported that the proband stuttered, one of whom was also reported to have a hearing impairment, in addition to one proband who was only reported to have a hearing impairment (43_001). These notes are based on parent report, no clinical diagnoses were recorded. Rare variants were observed in two probands with other reported difficulties (43_001 and 50_006). The other five probands with parent report of other impairments were sequenced, but as no rare variants were observed, they are not reported in any of the results. Affected status for LI was determined based on performance on the U-PPVT-4. The average U-PPVT-4 standard score for the probands (N=38) was 53.9, about 20 or more standard score points lower than the average standard scores observed for parents (N=22, mean=72.1), siblings (N=35, mean=71.3) or extended family members (N=24, mean=80.1) with both DNA and U-PPVT-4 scores available.
3.2. Genetic analysis summary
Our analysis of sequencing data in the probands revealed 16 rare variants (Table 1, Table 2). In addition, we observed several common polymorphic coding variants and common intronic variants (Table 4). Seven variants had a MAF of < 0.5% in South Asians and one variant was novel, or not previously reported (Table 1, Fig. 1). The MAF of the novel variant was not found in other databases (Table 3). Of the seven variants with a MAF < 0.5% in South Asians, four were not reported in any of the other databases, while the other three were consistently < 0.5% when reported in other databases (1000 Genomes, ESP6500, Iranome and GenomeAsia 100K; Table 1). Sanger confirmation of the novel nonsynonymous variant observed in PKSLI-15 in CNTNAP2 (c.2873G>C) is shown in supplementary material (Fig. S1). The other eight variants had a MAF > 0.5% in South Asians, but < 0.5% in Europeans, according to the gnomAD v2.1.1 database and most of these eight variants also had a MAF > 0.5% in the Iranome and GenomeAsia 100K databases, consistent with the gnomAD findings for South Asians (Table 2, Fig. 2; Fattahi et al., 2019; GenomeAsia 100K Consortium, 2019). These 16 variants included 13 nonsynonymous variants, point mutations in the 3’UTR and in the 5’UTR, and one nonsense (stop gain) variant (Table 1, Table 2). In three probands (04_003, 10_004, 28_004), we observed more than one rare variant, while five variants were shared by more than one proband (Table 1, Table 2, Fig. 1, Fig. 2). The rare variants we identified in the four genes were all heterozygous in the probands. One family member (28_007) was homozygous for one variant (c.14C>T, p.P5L in CNTNAP2; Fig. 1); the change is predicted to be benign according to four of the five bioinformatic prediction tools reported, while the SIFT score predicted a deleterious effect (Table 3). No other family members were homozygous for any identified rare variants.
Table 4.
Number of additional variants by type in each candidate gene in all Pakistani probands (N=38).
| Gene | NM ID | Exonic variants | Intronic variants | Total | |||
|---|---|---|---|---|---|---|---|
| Non- synonymous |
Synonymous | UTR | Indel | Other | |||
| TM4SF20 | NM_024795.4 | 2 | 0 | 0 | 0 | 1 | 3 |
| NFXL1 | NM_152995.5 | 1 | 4 | 0 | 1 | 7 | 13 |
| CNTNAP2 | NM_014141.5 | 0 | 5 | 2 | 3 | 15 | 25 |
| ATP2C2 | NM_001286527.2 | 4 | 4 | 0 | 2 | 35 | 45 |
Note. Synonymous variant totals include those with a MAF < 0.5% in Europeans or South Asians, according to gnomAD v2.1.1 exomes.
The estimates from five in silico tools (SIFT, PolyPhen-2, Mutation assessor, PORVEAN and MutationTaster2) were inconsistent for most observed variants, which is not uncommon for these algorithms (Table 3; Dong et al., 2015). Deleterious effects were predicted by all five in silico algorithms for one rare nonsynonymous variant (c.2465C>T: p.P822L in ATP2C2; Table 1; Table 3), which co-segregated in the proband branch of PKSLI-27 (Fig. 1). No other non-synonymous variants showed patterns of co-segregation, but all variants were observed in additional family members (Fig. 1 and Fig. 2). The novel variant observed in CNTNAP2 (c.2873G>C), and the nonsense variant observed in ATP2C2 (c.1507C>T) were also consistently predicted to be damaging or deleterious (Table 3). The rare nonsense (stop gain) variant was only identified in the proband and a parent but not observed in the extended branches of PKSLI-28 (Table 1, Fig. 1). Two other variants (c.854G>C on CNTNAP2 and c.937G>A on ATP2C2) were predicted to be deleterious or damaging by four of the five in silico programs (all except Mutation assessor; Table 3). Four rare variants were observed in proband 04-003, including one with four deleterious scores (ATP2C2 c.937G>A), though the other three variants (c.2258C>G on NFXL1, c.953A>G and c.2807T>C on ATP2C2) were not predicted to be deleterious or damaging (Table 3). The rare variants observed in TM4SF20 (1 variant) and NFXL1 (2 variants) were each only predicted to be deleterious or damaging by one in silico algorithm, SIFT or PolyPhen-2 (Table 3).
Seven non-synonymous variants had a MAF > 0.5% in South Asians, four were observed in more than one unrelated proband and three were observed in ATP2C2 in a single proband (04_003; Table 2, Fig. 2).
The 3’UTR variant was observed in CNTNAP2 (c.*9T>C); the variant co-segregated with performance on the U-PPVT-4 in PKSLI-64 (Fig. 1) and the MAF of this variant was rare (< 0.5%) in South Asians (Table 1). The 5’UTR variant was also observed in CNTNAP2 (c.-49T>G), though it was common in South Asians, two probands and some of their family members showed this variant (Table 3). In silico prediction scores are not estimated for rare variants observed in the UTRs.
4. Discussion
This is the first investigation of SLI candidate genes in Pakistani probands with a history of LI. We identified multiple rare variants in the selected candidate genes that included point mutations in the 3’UTR and in the 5’UTR, a nonsense variant and several nonsynonymous variants. We also observed common single nucleotide polymorphisms (SNPs) in our study that were reported as risk alleles for SLI in previous association studies. However, our sample size was limited to measure the association of these SNPs with LI. Most of these variants are well conserved among other vertebrates and six are predicted to be damaging, according to at least three of the five in silico algorithms (Table 3). The HOPE web application model of the AA substitutions for the variants with multiple deleterious scores indicated that the size, structure and conservation changes were also likely to disturb the structure causing changes in expression, especially for those AA substitutions that were the least similar to any previously observed AA substitutions. However, the HOPE model summary for two (c.2873G>C on CNTNAP2 and c.2732A>T on NFXL1) of the 16 rare variants were inconsistent with the other in silico predictions.
We observed four total variants (rare, common, exonic and intronic) in TM4SF20 (Table 4). All four variants have MAF >0.5 % in South Asians, only rs137891000 has a MAF < 0.5 % in Europeans (gnomAD v2.1.1 exomes; Table 2 and Fig. 2). Other variants identified in TM4SF20 were heterozygous and observed in multiple probands (data not shown). A 4kb deletion overlapping with exon 3 in TM4SF20 in individuals with language delay was reported previously (Wiszniewski et al., 2013). Of note, one proband (half-Pakistani/half-Filipino) in this study inherited the 4kb deletion, despite the mother was not showing white matter abnormalities in brain imaging (Wiszniewski et al., 2013). Three probands in our study shared a common non-synonymous variant in exon 3 (data not shown), but no other variants were observed in exon 3 on TM4SF20. The candidate gene, TM4SF20 (2q36), is located adjacent to a linkage region (2q33.3-q35) we previously reported in a proband family under an autosomal recessive mode of inheritance (Andres et al., 2019; Wiszniewski et al., 2013). No co-segregating variants were identified in TM4SF20 in the proband families with strong linkage to chromosome 2q (PKSLI-05 and PKSLI-12) indicating other genes in this region might be a good target for future studies (Andres et al., 2019).
We identified two rare variants and 13 common polymorphic SNPs in NFXL1 (Table 1, Table 2). Both rare variants and several other common SNPs were observed in the same proband (10_004; Table 1, Fig. 1). The rare variant, rs370816326, was predicted to be damaging according to the PolyPhen-2 prediction, but it was prevalent in South Asians and shared by both 10_004 and 04_003 (Table 2, Fig. 2). The other variant of interest, rs775580487, was rare in South Asians, but only the SIFT score predicted a deleterious change (Table 1, Fig. 1). However, the HOPE 3D model suggested the AA substitution could be more damaging than the in silico scores predict. Specifically, the 3D model showed a smaller structure, which could limit the proteins interactions with other proteins. Additionally, the substituted AA (Isoleucine), or a similar AA, has not been previously observed in that location, indicating the change is more likely to be damaging, Sanger sequencing of the extended family members confirmed this variant in the proband’s father, but this variant was not observed in any other available family members. The U-PPVT-4 score for the proband, 10_004, was the lowest possible standard score, based on the US norms and the parents reported that the proband was slow to talk and had difficulties with language, reading, spelling, and learning. The proband’s father also performed below average on the U-PPVT-4. Three common SNPs (rs12651301, rs2053404 and rs6818556) that we identified in multiple Pakistani probands, including 10_004 (data not shown), were reported previously in the SLI Consortium (SLIC) cohort, indicating the utility of these SNPs for future population-based association studies of SLI (Villanueva et al., 2015).
We identified a total of 31 variants in CNTNAP2, the largest among those genes selected in this study (~2.3 Mb at chromosome 7q35-q36). We observed five rare/novel variants in this gene with MAF <0.5% in South Asians, each in a single proband (Table 1, Fig. 1). Four of the five in silico prediction scores predicted deleterious and damaging effects by the novel variant (c.2873G>C) and the HOPE model for this variant indicated a less flexibility for the protein, but the substituted AA (Alanine) has been observed at the position, suggesting it may not be as damaging as the scores and structure change predict. Two probands (17_003 and 44_002) shared a nonsynonymous variant (rs138517537), while two other probands (40_002, 51_004) shared a point mutation in the 5’UTR (rs549396215) with a MAF of >0.5% in the South Asian population (Table 2, Fig. 2). The probands (17_003, 44_002, 45_005, and 64_005) in which we observed rare variants, all performed below average on the U-PPVT-4 (Fig. 2) and their parents reported struggles with speech, reading, spelling and learning. Parents reported language difficulty in these probands as well, except in 44_002. Numerous studies have investigated the role of CNTNAP2 in studies of language disorders and related impairments, including reading disorder, stuttering and autism (Centanni et al., 2015; Newbury et al., 2011; Peter et al., 2011; Petrin et al., 2010; Scott et al., 2019; Vernes et al., 2008; Whalley et al., 2011; Whitehouse et al., 2011). We observed multiple common variants that were the target of CNTNAP2 investigations of Chinese children with reading disorder and Iraqi children with autism (Gu et al., 2018; Karmeet et al., 2015). Four common SNPs (rs2462603, rs10240503, rs3779031 and rs9648691) used in the targeted study of Chinese children with reading disorder, were also observed in multiple Pakistani probands (data not shown; Gu et al., 2018). We observed two common SNPs (rs3770931 and rs3779032) in multiple Pakistani probands that previously showed a strong association in Iraqi children with autism (Karmeet et al., 2015). Note that both studies targeted rs3779031 and showed variable risk associated with autism and reading impairment for this SNP (Gu et al., 2018; Karmeet et al., 2015). In the study of the Chinese sample, females with the minor allele ‘G’ at rs3779031 had a lower risk of reading disorder than those with the common ‘A’ allele (Gu et al., 2018). Again, these results highlight the possible value of these SNPs for population-based association studies of SLI.
In ATP2C2, we identified 45 total variants. Two rare variants were observed independently in two probands, 27_005 and 28_004 (Table 1, Fig. 1). Both probands performed below average on the U-PPVT-4, but parent report information was not available for 27_005. Neither of these variants was observed in family members of the other available branches, indicating this gene is less likely to be associated with LI in the Pakistani population (Fig. 1). A nonsynonymous variant (rs757310826) in PKSLI-27 was observed in most of the individuals who performed below average on the U-PPVT-4, though two affected individuals in other branches were wildtype (Fig. 1). All unaffected individuals and those with unknown phenotype status did not carry the variant. While increased penetrance of this variant was observed in one branch of PKSLI-27, the variant was not observed in other branches of the family indicating this gene could be a strong contributor to the risk of LI in this branch, but not the others. The AA substitution for the SNP observed in PKSLI-27 (p.P822L) produces a bigger structure, according to the HOPE web application. The wildtype AA (Proline) is a rigid structure and when disturbed, especially at this 100% conserved location, indicates the AA substitution is likely very damaging. A rare nonsense variant (p.Q503X) identified in 28_004 was only observed in the proband’s father, who was also affected on the U-PPVT-4, but no other available family members carried the nonsense variant (Fig. 1). Given the significance of a rare stop gain variant and high frequency of individuals with low U-PPVT-4 score in PKSLI-28, we plan to gather additional DNA from the extended family members. Four other variants (rs531589939, rs551531561, rs202026876, rs16973859) with a MAF > 0.5% in South Asians were observed in ATP2C2 (Table 2, Fig. 2). One variant (rs202026876) was shared by two probands (43_001 and 52_005), while the remaining three were identified in a single proband, 04_003 (Table 2, Fig. 2). The proband, 04 003, performed below average on the U-PPVT-4 and their parents reported that they were slow to talk. We identified these variants in the proband’s mother, who performed in the average range on the U-PPVT-4 (Fig. 2).
4.1. Limitations and Strengths
Limitations of the current study include targeting only four candidate genes, despite the numerous candidate genes previously reported, small sample size, lack of controls and limited phenotype information. Despite the limitations of this study, we identified rare variants in the proband families with a history of LI ascertained in Pakistan and built a foundation for future research of these proband families. We targeted the protein-coding regions of four candidate genes, chosen based on the most supportive evidence and replication of findings in multiple studies though several candidate genes have been proposed for SLI (Chen et al., 2017; Mountford et al., 2019; Newbury et al., 2010; Newbury & Monaco, 2010; Nudel, 2016; Villanueva et al., 2015; Wiszniewski et al., 2013). Previous studies have reported many intronic SNPs in these candidate genes as risk alleles, outside the area of focus in the current study, which was the protein-coding regions and flanking intronic DNA sequences. We reported several rare protein coding variants of high impact that aggregated in multiple proband families, but complete Mendelian segregation was not observed. While most probands performed below average on the U-PPVT-4, this phenotype information was limited by using standard scores determined from the US norms. The probands used in the current study were initially identified based on the information gathered from schoolteachers and parent report. Parent report of language difficulties and low performance on the U-PPVT-4 aligned with the information initially gathered from schoolteachers except in four probands, who performed well on the UPPVT-4 measure. We observed rare variants in two of these wellperforming probands (13_004 in Fig. 1, 24_004 in Fig. 2). Our future goal is to define the affected status in these families by calculating age-stratified means and standard deviations of population-based performance in the U-PPVT-4. We also aim to use additional parent report information and performance on the digit span to validate the U-PPVT-4 scores.
The current study was initiated when we first started enrolling probands in Pakistan. The rationale of using the candidate gene approach was that LI in these probands could be explained by rare genetic variants in genes suggested previously for a similar phenotype (Kwon & Goate, 2000; Mountford et al., 2019). As we extended enrollment to family members, we performed linkage and homozygosity mapping in some of the probands’ families included in the current study (Andres et al., 2019). The linkage and homozygosity regions observed in these families did not overlap with any of the candidate genes selected in the current study. All mapped regions were identified under the autosomal recessive mode of inheritance. Our results in the current study exclude these genes as candidates for future investigation in these families based on the absence of complete Mendelian segregation and rare homozygous variants. Future investigation using next-generation sequencing and analyzing variants in the linkage and homozygosity regions may identify novel SLI candidate genes, especially within regions shared by multiple families (2q13-q21.2, 8q21.11-q22.2 and 22q12.3-q13.31; Andres et al., 2019).
5. Conclusions
We identified several rare variants in the selected candidate genes in our proband families, indicating these genes may contribute to overall susceptibility of LI across populations. However, rare variants identified in the selected SLI candidate genes did not show co-segregation in the proband ascertained families suggesting additional genes are likely involved in the expression of LI. Future studies should continue to extend the genetic inquiry of language beyond English-speaking individuals of European descent. As researchers have called for increased study of consanguineous samples (Erzurumluoglu et al., 2016), further study of this sample could lead to novel SLI candidate genes and information about their function.
Supplementary Material
Acknowledgements
We are extremely thankful to the children and family members who participated in this research and to the students in Pakistan who collected the behavioral data and saliva samples from these participants. We would also like to thank Devan Crow and Michaela Russ for their work to organize the behavioral and genetic data from these families. Finally, we would like to thank Mabel L. Rice for her input on the initial letter to teachers for ascertainment of probands and the data collection of the family history information and Kathleen Kelsey Earnest for her input on the data collection of the U-PPVT-4 in these families.
Funding
This work was supported by the National Institutes on Deafness and Communication Disorders [T32DC000052; Mabel L. Rice and R21DC017830; MHR]; and start-up research funds provided by the University of Kansas, Lawrence to MHR.
Abbreviations
- LI
language impairment
- SLI
specific language impairment
- NV-IQ
non-verbal intelligence
- LD
learning disability
- ID
intellectual disability
- PPVT
Peabody Picture Vocabulary Test
- U-PPVT-4
Peabody Picture Vocabulary Test, fourth edition, translated to Urdu
- MAF
minor allele frequency
- UTR
untranslated region
- SNP
single nucleotide polymorphism
Footnotes
Declaration of Competing Interest
The authors do not have any competing interest to report.
Conflict of interest
None
Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final form. Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
References
- Adzhubei IA, Schmidt S, Peshkin L, Ramensky VE, Gerasimova A, Bork P, Kondrashov AS, Sunyaev SR 2010. A method and server for predicting damaging missense mutations. Nat. Methods, 74, 248–249. doi: 10.1038/nmeth0410-248 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Andres EM, Earnest KK, Smith SD, Rice ML, Raza MH 2020. Pedigree-based gene mapping supports previous loci and reveals novel suggestive loci in specific language impairment (SLI). J. Speech Lang. Hear. Res, 6312, 4046–4061. doi: 10.1044/2020_JSLHR-20-00102 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Andres EM, Hafeez H, Yousaf A, Riazuddin S, Rice ML, Basra MAR, Raza MH 2019. A genome-wide analysis in consanguineous families reveals new chromosomal loci in specific language impairment (SLI). Eur. J. Hum. Genet, 278, 1274–1285. doi: 10.1038/s41431-019-0398-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bartlett CW, Flax JF, Logue MW, Smith BJ, Vieland VJ, Tallal P, Brzustowicz LM 2004. Examination of potential overlap in autism and language loci on chromosomes 2, 7, and 13 in two independent samples ascertained for specific language impairment. Hum. Hered, 571, 10–20. doi: 10.1159/000077385 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bartlett CW, Flax JF, Logue MW, Vieland VJ, Bassett AS, Tallal P, Brzustowicz LM 2002. A major susceptibility locus for specific language impairment is located on 13q21. Am. J. Med. Genet, 711, 45–55. doi: 10.1086/341095 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Bishop DVM, North T, Donlan C 1995. Genetic basis of specific language impairment: Evidence from a twin study. Dev. Med. Child Neurol, 371, 56–71. doi: 10.1111/j.1469-8749.1995.tb11932.x [DOI] [PubMed] [Google Scholar]
- Carl GF, Keen CL, Gallagher BB, Clegg MS, Littleton WH, Flannery DB, Hurley LS 1986. Association of low blood manganese concentrations with epilepsy. Neurology, 3612, 1584–1587. doi: 10.1212/WNL.36.12.1584 [DOI] [PubMed] [Google Scholar]
- Centanni TM, Sanmann JN, Green JR, Iuzzini-Seigel J, Bartlett C, Sanger WG, Hogan TP 2015. The role of candidate-gene CNTNAP2 in childhood apraxia of speech and specific language impairment. Am. J. Med. Genet. B Neuropsychiatr. Genet, 1687, 536–543. doi: 10.1002/ajmg.b.32325 [DOI] [PubMed] [Google Scholar]
- Chaudhry IS, Rahman S 2009. The impact of gender inequality in education on rural poverty in Pakistan: an empirical analysis. Eur. J. Econ. Finance Administrative Sciences, 151, 174–188. [Google Scholar]
- Chen XS, Reader RH, Hoischen A, Veltman JA, Simpson NH, Francks C, Newbury DF, Fisher SE 2017. Next-generation DNA sequencing identifies novel gene variants and pathways involved in specific language impairment. Sci. Rep, 7, 1–17, Article 46105. doi: 10.1038/srep46105 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Choi Y, Chan AP 2015. PROVEAN web server: a tool to predict the functional effect of amino acid substitutions and indels. Bioinformatics, 3116, 2745–2747. doi: 10.1093/bioinformatics/btv195 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Choi Y, Sims GE, Murphy S, Miller JR, Chan AP 2012. Predicting the functional effect of amino acid substitutions and indels. PLoS One, 710, e46688. doi: 10.1371/journal.pone.0046688 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Chomsky N (1957). Syntactic structures. Mouton. [Google Scholar]
- Clarke L, Fairley S, Zheng-Bradley X, Streeter I, Perry E, Lowy E, Tasse AM, Flicek P 2017. The international Genome sample resource (IGSR): A worldwide collection of genome variation incorporating the 1000 Genomes Project data. Nucleic Acids Research, 45D1, D854–D859. doi: 10.1093/nar/gkw829 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dong C, Wei P, Jian X, Gibbs R, Boerwinkle E, Wang K, Liu X 2015. Comparison and integration of deleteriousness prediction methods for nonsynonymous SNVs in whole exome sequencing studies. Hum. Mol. Genet, 248, 2125–2137. doi: 10.1093/hmg/ddu733 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Dunn LM, Dunn DM (2007). PPVT-4: Peabody picture vocabulary test. Pearson Assessments [Google Scholar]
- Erzurumluoglu AM, Shihab HA, Rodriguez S, Gaunt TR, Day IN 2016. Importance of genetic studies in consanguineous populations for the characterization of novel human gene functions. Ann. Hum. Biol, 803, 187–196. doi: 10.1111/ahg.12150 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Exome Variant Server, NHLBI GO Exome Sequencing Project (ESP), Seattle, WA: (URL: http://evs.gs.washington.edu/EVS/) [July, 2021 accessed]. [Google Scholar]
- Fattahi Z, Beheshtian M, Mohseni M, Poustchi H, Sellars E, Nezhadi SH, Amini A, Arzhangi S, Jalalvand K, Jamali P, Mohammadi Z, Davarnia B, Nikuei P, Oladnabi M, Mohammadzadeh A, Zohrehvand E, Nejatizadeh A, Shekari M, Bagherzadeh M, Shamsi-Gooshki E, Borno S, Timmermann B, Haghdoost A, Najafipour R, Khorram Khorshid HR, Kahrizi K, Malekzadeh R, Akbari MR, Najmabadi H 2019. Iranome: A catalog of genomic variations in the Iranian population. Hum Mutat, 4011, 1968–1984. doi: 10.1002/humu.23880 [DOI] [PubMed] [Google Scholar]
- Fisher SE, Scharff C 2009. FOXP2 as a molecular window into speech and language. Trends Genet., 254, 166–177. doi: 10.1016/j.tig.2009.03.002 [DOI] [PubMed] [Google Scholar]
- Fisher SE, Vargha-Khadem F, Watkins KE, Monaco AP, Pembrey ME 1998. Localisation of a gene implicated in a severe speech and language disorder. Nat. Genet, 182, 168–170. doi: 10.1038/ng0298-168 [DOI] [PubMed] [Google Scholar]
- Friedman JI, Vrijenhoek T, Markx S, Janssen IM, van der Vliet WA, Faas BH, Knoers NV, Cahn W, Kahn RS, Edelmann L, Davis KL, Silverman JM, Brunner HG, van Kessel AG, Wijmenga C, Ophoff RA, Veltman JA 2008. CNTNAP2 gene dosage variation is associated with schizophrenia and epilepsy. Mol. Psychiatry, 133, 261–266. doi: 10.1038/sj.mp.4002049 [DOI] [PubMed] [Google Scholar]
- GenomeAsia 100K Consortium. 2019. The GenomeAsia 100K Project enables genetic discoveries across Asia. Nature, 5767785, 106–111. doi: 10.1038/s41586-019-1793-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- Gu H, Hou F, Liu L, Luo X, Nkomola PD, Xie X, Li X, Song R 2018. Genetic variants in the CNTNAP2 gene are associated with gender differences among dyslexic children in China. EBioMedicine, 34, 165–170. doi: 10.1016/j.ebiom.2018.07.007 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Hamilton CM, Strader LC, Pratt JG, Maiese D, Hendershot T, Kwok RK, Hammond JA, Huggins W, Jackman D, Pan H, Nettles DS, Beaty TH, Farrer LA, Kraft P, Marazita ML, Ordovas JM, Pato CN, Spitz MR, Wagener D, Williams M, Junkins HA, Harlan WR, Ramos EM, Haines J 2011. The PhenX Toolkit: get the most from your measures. Am. J. Epidemiol, 1743, 253–260. doi: 10.1093/aje/kwr193 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alfoldi J, Wang Q, Collins RL, Laricchia KM, Ganna A, Birnbaum DP, Gauthier LD, Brand H, Solomonson M, Watts NA, Rhodes D, Singer-Berk M, England EM, Seaby EG, Kosmicki JA, Walters RK, Tashman K, Farjoun Y, Banks E, Poterba T, Wang A, Seed C, Whiffin N, Chong JX, Samocha KE, Pierce-Hoffman E, Zappala Z, O'Donnell-Luria AH, Minikel EV, Weisburd B, Lek M, Ware JS, Vittal C, Armean IM, Bergelson L, Cibulskis K, Connolly KM, Covarrubias M, Donnelly S, Ferriera S, Gabriel S, Gentry J, Gupta N, Jeandet T, Kaplan D, Llanwarne C, Munshi R, Novod S, Petrillo N, Roazen D, Ruano-Rubio V, Saltzman A, Schleicher M, Soto J, Tibbetts K, Tolonen C, Wade G, Talkowski ME, Genome Aggregation Database C, Neale BM, Daly MJ, MacArthur DG 2020. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature, 5817809, 434–443. doi: 10.1038/s41586-020-2308-7 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Karmeet BK, Al-Kazaz AKA, Saber M 2015. Molecular genetics study on autistic patients in Iraq. Iraqi J. Sci, 561A, 119–124. [Google Scholar]
- Kwon JM, Goate AM 2000. The candidate gene approach. Alcohol Res. Health, 243, 164–168. https://www.ncbi.nlm.nih.gov/pubmed/11199286 [PMC free article] [PubMed] [Google Scholar]
- Lenneberg EH (1967). Biological foundations of language. John Wiley & Sons, Inc. [Google Scholar]
- Martinelli A, Rice ML, Talcott JB, Diaz R, Smith SD, Raza MH, Snowling MJ, Hulme C, Stein J, Hayiou-Thomas ME, Hawi Z, Kent L, Pitt SJ, Newbury DF, Paracchini S 2021. A rare missense variant in the ATP2C2 gene is associated with language impairment and related measures. Hum. Mol. Genet doi: 10.1093/hmg/ddab111 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Mefford HC, Muhle H, Ostertag P, von Spiczak S, Buysse K, Baker C, Franke A, Malafosse A, Genton P, Thomas P, Gurnett CA, Schreiber S, Bassuk AG, Guipponi M, Stephani U, Helbig I, Eichler EE 2010. Genome-wide copy number variation in epilepsy: novel susceptibility loci in idiopathic generalized and focal epilepsies. PLoS Genet., 65, e1000962. doi: 10.1371/journal.pgen.1000962 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Missiaen L, Dode L, Vanoevelen J, Raeymaekers L, Wuytack F 2007. Calcium in the Golgi apparatus. Cell Calcium, 415, 405–416. doi: 10.1016/j.ceca.2006.11.001 [DOI] [PubMed] [Google Scholar]
- Mountford HS, Villanueva P, Fernández MA, De Barbieri Z, Cazier JB, Newbury DF 2019. Candidate gene variant effects on language disorders in Robinson Crusoe Island. Ann. Hum. Biol, 462, 109–119. doi: 10.1080/03014460.2019.1622776 [DOI] [PubMed] [Google Scholar]
- National Institute on Deafness and Other Communication Disorders. (2019, October 21, 2019). Specific language impairment. Retrieved November 10, 2017 from [Google Scholar]
- Newbury DF, Fisher SE, Monaco AP 2010. Recent advances in the genetics of language impairment. Genome Med., 21, 6. doi: 10.1186/gm127 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Newbury DF, Monaco AP 2010. Genetic advances in the study of speech and language disorders. Neuron, 682, 309–320. doi: 10.1016/j.neuron.2010.10.001 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Newbury DF, Paracchini S, Scerri TS, Winchester L, Addis L, Richardson AJ, Walter J, Stein JF, Talcott JB, Monaco AP 2011. Investigation of dyslexia and SLI risk variants in reading- and language-impaired subjects. Behav. Genet, 411, 90–104. doi: 10.1007/s10519-010-9424-3 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Newbury DF, Winchester L, Addis L, Paracchini S, Buckingham LL, Clark A, Cohen W, Cowie H, Dworzynski K, Everitt A, Goodyer IM, Hennessy E, Kindley AD, Miller LL, Nasir J, O'Hare A, Shaw D, Simkin Z, Simonoff E, Slonims V, Watson J, Ragoussis J, Fisher SE, Seckl JR, Helms PJ, Bolton PF, Pickles A, Conti-Ramsden G, Baird G, Bishop DVM, Monaco AP 2009. CMIP and ATP2C2 modulate phonological short-term memory in language impairment. Am. J. Hum. Genet, 852, 264–272. doi: 10.1016/j.ajhg.2009.07.004 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Norbury CF, Gooch D, Wray C, Baird G, Charman T, Simonoff E, Vamvakas G, Pickles A 2016. The impact of nonverbal ability on prevalence and clinical presentation of language disorder: Evidence from a population study. J. Child Psychol. Psychiatry, 5711, 1247–1257. doi : 10.1111/jcpp.12573 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Nudel R 2016. An investigation of NFXL1, a gene implicated in a study of specific language impairment. J. Neurodev. Disord, 8, 13. doi: 10.1186/s11689-016-9146-9 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Peter B, Raskind WH, Matsushita M, Lisowski M, Vu T, Berninger VW, Wijsman EM, Brkanac Z 2011. Replication of CNTNAP2 association with nonword repetition and support for FOXP2 association with timed reading and motor activities in a dyslexia family sample. J. Neurodev. Disord, 31, 39–49. doi: 10.1007/s11689-010-9065-0 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Petrin AL, Giacheti CM, Maximino LP, Abramides DV, Zanchetta S, Rossi NF, Richieri-Costa A, Murray JC 2010. Identification of a microdeletion at the 7q33-q35 disrupting the CNTNAP2 gene in a Brazilian stuttering case. American Journal of Medical Genetics Part A, 152A12, 3164–3172. doi: 10.1002/ajmg.a.33749 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Reader RH, Covill LE, Nudel R, Newbury DF 2014. Genome-Wide Studies of Specific Language Impairment. Curr. Behav. Neurosci. Rep, 14, 242–250. doi: 10.1007/s40473-014-0024-z [DOI] [PMC free article] [PubMed] [Google Scholar]
- Reva B, Antipin Y, Sander C 2011. Predicting the functional impact of protein mutations: application to cancer genomics. Nucleic Acids Res., 3917, e118. doi: 10.1093/nar/gkr407 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rice ML 2012. Toward epigenetic and gene regulation models of specific language impairment: looking for links among growth, genes, and impairments. J. Neurodev. Disord, 41, 27. doi: 10.1186/1866-1955-4-27 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rice ML 2017. Overlooked by public health: Specific language impairment. Open Access Government. [Google Scholar]
- Rice ML, Haney KR, Wexler K 1998. Family histories of children with SLI who show extended optional infinitives. J. Speech Lang. Hear. Res, 412, 419–432. doi: 10.1044/jslhr.4102.419 [DOI] [PubMed] [Google Scholar]
- Rice ML, Hoffman L 2015. Predicting vocabulary growth in children with and without specific language impairment: a longitudinal study from 2;6 to 21 years of age. J. Speech Lang. Hear. Res, 582, 345–359. doi: 10.1044/2015_JSLHR-L-14-0150 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Rodenas-Cuadrado P, Ho J, Vernes S 2014. Shining a light on CNTNAP2: complex functions to complex disorders. Eur. J. Hum. Genet, 222, 171–178. doi: 10.1038/ejhg.2013.100 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Schwarz JM, Cooper DN, Schuelke M, Seelow D 2014. MutationTaster2: mutation prediction for the deep-sequencing age. Nat. Methods, 114, 361–362. doi: 10.1038/nmeth.2890 [DOI] [PubMed] [Google Scholar]
- Scott R, Sanchez-Aguilera A, van Elst K, Lim L, Dehorter N, Bae SE, Bartolini G, Peles E, Kas MJH, Bruining H, Marin O 2019. Loss of Cntnap2 Causes Axonal Excitability Deficits, Developmental Delay in Cortical Myelination, and Abnormal Stereotyped Motor Behavior. Cerebral Cortex, 292, 586–597. doi: 10.1093/cercor/bhx341 [DOI] [PubMed] [Google Scholar]
- Sim NL, Kumar P, Hu J, Henikoff S, Schneider G, Ng PC 2012. SIFT web server: predicting effects of amino acid substitutions on proteins. Nucleic Acids Res., 40Web Server issue, W452–457. doi: 10.1093/nar/gks539 [DOI] [PMC free article] [PubMed] [Google Scholar]
- SLI Consortium. 2002. A genomewide scan identifies two novel loci involved in specific language impairment. Am. J. Hum. Genet, 702, 384–398. doi: 10.1086/338649 [DOI] [PMC free article] [PubMed] [Google Scholar]
- SLI Consortium. 2004. Highly significant linkage to the SLI1 locus in an expanded sample of individuals affected by specific language impairment. Am. J. Hum. Genet, 746, 1225–1238. doi: 10.1086/421529 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Smith SD 2007. Genes, language development, and language disorders. Dev. Disabil. Res. Rev, 131, 96–105. doi: 10.1002/mrdd.20135 [DOI] [PubMed] [Google Scholar]
- Strauss KA, Puffenberger EG, Huentelman MJ, Gottlieb S, Dobrin SE, Parod JM, Stephan DA, Morton DH 2006. Recessive symptomatic focal epilepsy and mutant contactin-associated protein-like 2. N. Engl. J. Med, 35413, 1370–1377. doi: 10.1056/NEJMoa052773 [DOI] [PubMed] [Google Scholar]
- Stromswold K 1998. Genetics of Spoken Language Disorders. Hum. Biol, 702, 297–324. https://www.ncbi.nlm.nih.gov/pubmed/9549241 [PubMed] [Google Scholar]
- Stromswold K 2001. The Heritability of Language: A Review and Metaanalysis of Twin, Adoption, and Linkage Studies. Language, 774, 647–723. https://doi.org/DOI 10.1353/lan.2001.0247 [DOI] [Google Scholar]
- Tallal P, Hirsch LS, Realpe-Bonilla T, Miller S, Brzustowicz LM, Bartlett CW, Flax JF 2001. Familial aggregation in specific language impairment. J. Speech Lang. Hear. Res, 445, 1172–1182. doi: 10.1044/1092-4388(2001/091) [DOI] [PubMed] [Google Scholar]
- Tomblin JB, Records NL, Buckwalter P, Zhang X, Smith E, O’Brien M 1997. Prevalence of specific language impairment in kindergarten children. J. Speech Lang. Hear. Res, 406, 1245–1260. doi: 10.1044/jslhr.4006.1245 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Venselaar H, Te Beek TA, Kuipers RK, Hekkelman ML, Vriend G 2010. Protein structure analysis of mutations causing inheritable diseases. An e-Science approach with life scientist friendly interfaces. BMC Bioinformatics, 11, 548. doi: 10.1186/1471-2105-11-548 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Vernes SC, Newbury DF, Abrahams BS, Winchester L, Nicod J, Groszer M, Alarcon M, Oliver PL, Davies KE, Geschwind DH, Monaco AP, Fisher SE 2008. A functional genetic link between distinct developmental language disorders. N. Engl. J. Med, 35922, 2337–2345. doi: 10.1056/NEJMoa0802828 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Villanueva P, Newbury DF, Jara L, De Barbieri Z, Mirza G, Palomino HM, Fernández MA, Cazier JB, Monaco AP, Palomino H 2011. Genome-wide analysis of genetic susceptibility to language impairment in an isolated Chilean population. Eur. J. Hum. Genet, 196, 687–695. doi: 10.1038/ejhg.2010.251 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Villanueva P, Nudel R, Hoischen A, Fernández MA, Simpson NH, Gilissen C, Reader RH, Jara L, Echeverry MM, Francks C, Baird G, Conti-Ram sden G, O'Hare A, Bolton PF, Hennessy ER, SLI Consortium, Palomino H, Carvajal-Carmona L, Veltman JA, Cazier JB, De Barbieri Z, Fisher SE, Newbury DF 2015. Exome sequencing in an admixed isolated population indicates NFXL1 variants confer a risk for specific language impairment. PLoS Genet., 113, 1–24, Article e1004925. doi: 10.1371/journal.pgen.1004925 [DOI] [PMC free article] [PubMed] [Google Scholar]
- Villanueva P, Palomino HM, Palomino H 2008. High prevalence of specific language impairment in Robinson Crusoe Island. A possible founder effect. Rev. Med. Chile, 1362, 186–192. https://doi.org/s0034-98872008000200007 [PubMed] [Google Scholar]
- Wang G, Zhou Y, Gao Y, Chen H, Xia J, Xu J, Huen MSY, Siok WT, Jiang Y, Tan LH, Sun Y 2015. Association of specific language impairment candidate genes CMIP and ATP2C2 with developmental dyslexia in Chinese population. J. Neurolinguistics, 33, 163–171. doi: 10.1016/j.jneuroling.2014.06.005 [DOI] [Google Scholar]
- Whalley HC, O'Connell G, Sussmann JE, Peel A, Stanfield AC, Hayiou-Thomas ME, Johnstone EC, Lawrie SM, McIntosh AM, Hall J 2011. Genetic variation in CNTNAP2 alters brain function during linguistic processing in healthy individuals. Am. J. Med. Genet. B Neuropsychiatr. Genet, 156B8, 941–948. doi: 10.1002/ajmg.b.31241 [DOI] [PubMed] [Google Scholar]
- Whitehouse AJO, Bishop DVM, Ang QW, Pennell CE, Fisher SE 2011. CNTNAP2 variants affect early language development in the general population. Genes Brain Behav., 104, 451–456. doi: 10.1111/j.1601-183X.2011.00684.x [DOI] [PMC free article] [PubMed] [Google Scholar]
- Wiszniewski W, Hunter JV, Hanchard NA, Willer JR, Shaw C, Tian Q, Illner A, Wang X, Cheung SW, Patel A, Campbell IM, Gelowani V, Hixson P, Ester AR, Azamian MS, Potocki L, Zapata G, Hernandez PP, Ramocki MB, Santos-Cortez RL, Wang G, York MK, Justice MJ, Chu ZD, Bader PI, Omo-Griffith L, Madduri NS, Scharer G, Crawford HP, Yanatatsaneejit P, Eifert A, Kerr J, Bacino CA, Franklin AI, Goin-Kochel RP, Simpson G, Immken L, Haque ME, Stosic M, Williams MD, Morgan TM, Pruthi S, Omary R, Boyadjiev SA, Win KK, Thida A, Hurles M, Hibberd ML, Khor CC, Van Vinh Chau N, Gallagher TE, Mutirangura A, Stankiewicz P, Beaudet AL, Maletic-Savatic M, Rosenfeld JA, Shaffer LG, Davis EE, Belmont JW, Dunstan S, Simmons CP, Bonnen PE, Leal SM, Katsanis N, Lupski JR, Lalani SR 2013. TM4SF20 ancestral deletion and susceptibility to a pediatric disorder of early language delay and cerebral white matter hyperintensities. Am. J. Hum. Genet, 932, 197–210. doi: 10.1016/j.ajhg.2013.05.027 [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
