Skip to main content
PLOS One logoLink to PLOS One
. 2021 Mar 4;16(3):e0242754. doi: 10.1371/journal.pone.0242754

Neural representation of words within phrases: Temporal evolution of color-adjectives and object-nouns during simple composition

Maryam Honari-Jahromi 1,#, Brea Chouinard 2,#, Esti Blanco-Elorrieta 3,4, Liina Pylkkänen 3,4,5, Alona Fyshe 2,6,7,*
Editor: Nicola Molinaro8
PMCID: PMC7932185  PMID: 33661954

Abstract

In language, stored semantic representations of lexical items combine into an infinitude of complex expressions. While the neuroscience of composition has begun to mature, we do not yet understand how the stored representations evolve and morph during composition. New decoding techniques allow us to crack open this very hard question: we can train a model to recognize a representation in one context or time-point and assess its accuracy in another. We combined the decoding approach with magnetoencephalography recorded during a picture naming task to investigate the temporal evolution of noun and adjective representations during speech planning. We tracked semantic representations as they combined into simple two-word phrases, using single words and two-word lists as non-combinatory controls. We found that nouns were generally more decodable than adjectives, suggesting that noun representations were stronger and/or more consistent across trials than those of adjectives. When training and testing across contexts and times, the representations of isolated nouns were recoverable when those nouns were embedded in phrases, but not so if they were embedded in lists. Adjective representations did not show a similar consistency across isolated and phrasal contexts. Noun representations in phrases also sustained over time in a way that was not observed for any other pairing of word class and context. These findings offer a new window into the temporal evolution and context sensitivity of word representations during composition, revealing a clear asymmetry between adjectives and nouns. The impact of phrasal contexts on the decodability of nouns may be due to the nouns’ status as head of phrase—an intriguing hypothesis for future research.

Introduction

What is the relationship between the neural representation of a single word, occurring in isolation, and the representation of that same word in a combinatory context? In natural language, context can morph word meanings in many ways. One of the most obvious ways is via disambiguation: for example, in the phrase ‘term paper,’ ‘term’ means semester and ‘paper’ means ‘piece of writing’ even though both of these words have many other uses as well. Thus, the neural representations of ‘term’ and ‘paper’ may differ robustly depending on the context.

This study uses a decoding approach to address the general question of how combinatory contexts affect the neural representations of word meanings. We see value in starting with relatively straightforward cases, and thus did not investigate ambiguous cases such as those just mentioned. Instead, our combinatory contexts were all noun phrases comprised of a color-adjective and an object-describing noun, such as ‘blue cup.’ This type of composition has been studied extensively with time-sensitive magnetoencephalography (MEG), with results implicating the left anterior temporal lobe and the ventromedial prefrontal cortex as relatively stable correlates of composition [1].

Notably, unlike prior decoding work that investigated semantic representations during comprehension (e.g., [2]), the current MEG study involved language production to measure the planning of words millisecond by millisecond, during a picture naming task. Pictures of colored objects (e.g., white lamp) were named either with single nouns (“lamp”), single adjectives (“white”), or with combinations of those adjectives and nouns (“white lamp”). All pictures also contained a background color, allowing for an additional ‘list’ control, where participants named the background color plus the object shape (e.g., “green”, “lamp”). This created two-word utterances that do not form a conceptual combination of the sort that noun-adjective pairs create in natural language. This ‘list’ condition fully controlled for the number of uttered words and their lexical characteristics, allowing us to examine just the role of composition.

From prior MEG work on the planning of adjective-noun phrases in picture naming, we know that activity increases reflecting composition can be observed as early as 100-200ms after picture onset [35]. Estimates of the timing of lexical access in production fall into a similar time window: at around 150ms, both phonological and semantic properties of words are activated, as measured by MEG [6]. This suggests that lexical access and composition proceed in parallel. Against this background, context effects on lexical representations could start as early as 100-200ms in our study.

We used computational models of language (i.e. a word embedding model; [7]) and machine learning algorithms to detect semantic representations in the brain. We will use the term decodability to refer to the ability of our computational models to detect the semantic information related to the stimulus using a recording of brain activity. Decodability has been explored for people reading words in isolation [8, 9] phrases in isolation [2], sentences [10] and stories [11]. Decoding has also been successful using data collected while people listen to language [12, 13]. Our work is, to our knowledge, the first to use decoding techniques and MEG to detect the semantics of words before they are uttered.

Our design allowed us to quantify the relative decodability of noun and adjective neural representations in isolation, and when combined into adjective-noun lists or adjective-noun phrases. Classic syntactic theories hold that the features of the head of a phrase are inherited by the entire phrasal node [14], which in neural terms could mean that the representation of a head is stronger, and thus more decodable, than the representation of a modifier. In our study, this would result in the noun, as the head of the phrase, having greater decodability than the adjective. In contrast, theories in formal semantics posit that the intersective modification of nouns by adjectives proceeds via a fully symmetric predicate modification rule [15]. Thus, in our study, equal decodability of both the adjective and the noun within the phrases could be interpreted as a reflection of this type of symmetric semantic composition.

Additionally, decoding from a time-sensitive measurement like MEG allowed us to address neural representations across time. These analyses can be done in three ways. First we can train our model and then test using held-out data, where both train and test data come from the same time window (same-window-decoding). This typical decoding analysis allows us to measure the robustness of a neural signature within a time window. Second, we can also keep the time window constant, but choose held-out test data from another condition, testing how similar a representation is across conditions (across-condition). Third, we can choose a different time window from which to draw our held-out test data, testing the robustness of the pattern in time (resulting in a temporal generalization matrix, or TGM). This third analysis type allows us to test if a semantic representation is held constant in the brain over some time period, or if it re-emerges at different time points. Classic language production models propose a sequence of activated representations, proceeding from concept to lexeme to phonological representation and so forth [16] but say nothing about whether, for example, the conceptual representation stays active past the initial processing stages. Our data allowed us to ask whether representations detected by our models at a certain time after picture onset were also decodable at later times, or whether the processing stream only showed the characteristics of the classic models, where neural representations transform into new representations as we move towards articulation.

To summarize, we ask to what extent the neural representation of a word is the same when it is prepared for production as a single word compared to when it is prepared as part of a meaningful phrase. We do this using computational models of language meaning to compute decodability for adjectives and nouns when they are in isolation, and when they are part of a combinatory phrase vs. a non-combinatory list. We also evaluate the persistence and/or re-emergence of a word’s meaning, in isolation and in phrases and lists, using TGMs.

Methods

All experimental protocols were approved by New York University Institutional Review Board and conducted in accordance with the relevant guidelines and regulations, and all participants signed an informed consent form before taking part in the experiment. The data were originally collected to investigate the relationship between spoken and signed language with regards to the neural correlates of basic composition in language production [4]. In the current study, we used only the spoken language data. Our goal was to compare pre-utterance semantic representations in non-compositional versus compositional/phrasal contexts.

Participants

Nineteen right-handed monolingual native English speakers (9 female; ages: Mean: 25.6, 95 SD = 7.3), all neurologically intact with normal or corrected to normal vision, provided their written consent to participate in the original study [4].

Stimuli

Participants named pictures that depicted a colored object (e.g., white lamp) on a colored background (e.g., green) (Fig 1). All conditions used the same stimuli, and instructions at the beginning of the block differentiated the naming task to be performed in each condition. Non-compositional utterances were elicited by asking the participant to (i) say the name of the object color (i.e., “white”; adjective-only context); (ii) say the name of the object (e.g., “lamp”; noun-only context); or (iii) to say the background color, then pause, then the object name in a list-like fashion (e.g., “green”, “lamp”; list context). In the compositional context, participants were directed to describe the colored object on the screen by saying its color followed by its shape (e.g., “white lamp”; phrase context). Examples of images and utterances for each context appear in Fig 1A. There was also a control condition wherein the participant was instructed to say the background and object colors, but it was not analyzed here. The order of blocks was randomized across participants with the only constraint being that two blocks of the same condition never appeared consecutively. We controlled for frequency across word types (adjectives vs. nouns) using frequencies was extracted from Balota et al. [17]. Average noun frequency was 14396, average adjective frequency was 14976 (t = -0.47, p = .640742).

Fig 1. Experimental design and trials structure.

Fig 1

Participants named colored objects in three ways, depending on task instruction: as phrases (white lamp), as single nouns (lamp), as single adjectives (white) or as adjective-noun lists (green, lamp), naming the color of the background followed by the object name. Our analyses assessed the decodability of adjective and/or noun representations in these four contexts.

In total, the experiment consisted of 500 trials in which participants viewed one of 25 unique images created from a subset of five object shapes (bag, bell, cane, lamp, plane) and five colors (black, brown, green, red, white). Each image was also given a background color, which was counterbalanced so that background colors appeared an equal number of times, equally distributed across the different shapes. We also swapped the object and background colors to create a complementary set of 25 stimuli, thus creating 50 stimuli images in total. These 50 items were presented twice each for a total of 100 trials per condition. The items were presented in blocks of 25, for a total of 4 blocks per condition.

MEG procedure and preprocessing

Prior to the MEG recording, each participant’s head shape was digitized by a Polhemus dual source hand-held FastSCAN laser scanner (Polhemus, 112 VT, USA). The MEG data were recorded using a 208-channel axial gradiometer system (Kanazawa Institute of Technology, Kanazawa, Japan) at Neuroscience of Language Lab in NYU Abu Dhabi. MEG data were collected with a sampling frequency of 1000Hz (200 Hz low-passed filter). An MEG compatible microphone (Shure PG 81, Shure Europe GmbH) was used to record uttered speech of the participants. Each trial started with a fixation cross for 300ms, followed by the stimulus image, which was present until participant’s response or timeout (1500ms, see Fig 1B). Afterwards, a break of 1200ms was given until appearance of the fixation cross belonging to the next trial. Trials were epoched at 100ms before to 700ms after stimuli onset to avoid contamination via motion artifact coinciding with overt speech and noise was reduced using Continuously Adjusted Least-Squares Method [18]. Epochs were baseline corrected using the average of a 100ms interval prior to the stimulus onset. Unlike Blanco-Elorrieta et al. [4], we rejected only those trials with erroneous responses. MEG signals were band-passed using a Butterworth filter of order 20 between 0.1Hz and 40Hz. We used no ICA artefact rejection or blink / heart beat removal, as such artefacts are less problematic in decoding studies.

Our analysis operates on the within-participant average over trials of a particular word in a particular context. Based on the context and word-type of interest, we first chose a target word and then select trials with a target utterance containing that word. Within this set of selected trials, we averaged random groups of 5 epochs to minimize noise in the signal. This yielded 4 averaged epochs per noun and 20 averaged epochs in total. We followed the same procedure of averaging epochs for the adjectives. Trials were averaged within participant; we trained separate models for each participant, and report the average model performance. Further decoding analysis and statistical significance tests were conducted in Scikit-learn [19] and MNE-Python [20] and FieldTrip [21]. The code for all analysis is available at https://github.com/mahon94/compositionInBrain.

Decoding neural signatures using computational models of language

We calculated decoding accuracies to determine if the computer model could detect the semantic properties of the words to be uttered from the MEG data. Here, our computer model consists of word embeddings that represent the semantics of single words, and a regression model to map the MEG data to the dimensions of the word embeddings. We used Skip-gram word embeddings [22], which are derived from a neural network model trained on the Google News dataset (an internal Google dataset with one billion words) to predict context words. We were interested in how decoding accuracy varies over time, so we trained using 100ms windows of time, shifted in increments of 5ms across the full time window (i.e., -100 to +700ms relative to stimulus onset). An overview of the data organization, training procedure and 2 vs. 2 test appears in Fig 2.

Fig 2. Explanation of data, model, and testing procedures.

Fig 2

Values that are fixed appear as solid-colored rectangles, and values that are learned or predicted appear as dotted rectangles. A) The dimensions of the MEG data X (blue) and the word embedding matrix Y (green). B) The process for predicting one dimension (j) of the word embedding matrix Y. Note that this is corresponds to hj(X) in the in-text equations. C) Predicting all dimensions of a word embedding for MEG data sample xi. W is the concatenation of w vectors from B). D) The 2 vs 2 test. The 2 vs. 2 test measures how similar the predictions (y^a, y^b) are to their corresponding ground truth vectors (ya, yb) using a vector distance criterion d(v,u). If the correct matching of true to predicted vectors (blue lines) represents a smaller distance than the incorrect matching (red lines), the 2 vs 2 test passes.

Similar to the approach described in Fyshe et al. [2], we form the MEG dataset XRN×p where N = 20 is the total number of averaged epochs reshaped to vectors of size p = c×t, for c = 208 MEG gradiometer sensors with t = 100 time samples per window of analysis. We normalize each column of X to have mean 0 and standard deviation 1, and append a column of ones to account for the bias term. We then train d independent L2-regularized (ridge) regression models hj(X),j∈{1,2,…,d} to predict each column of the matrix of word embeddings YRN×d (d = 300 is dimension of the Skip-gram vector). The jth regression model is trained as follows:

hj(X)=Xwj,
wj=argminwXwyj22+λwTw
=(XTX+λI)1XTyj

Where wjRp,A22 indicates the squared two norm of A, and yj is the jth column of matrix Y. We determine the best performing regularization parameter λ separately for each column of the word embedding matrix using leave-one-out cross validation. Since Np, we speed up training using the kernel trick based on singular-value decomposition (SVD) demonstrated by Hastie & Tibshirani [23], and regularization helps to control overfitting. Note that every model was trained separately for each participant. Thus, the patterns underlying the decoding accuracies we observed may not be stable across people, and our analysis did not test for such stability. Rather, we tested for the presence of a pattern, and if the patterns generalize across time and condition within a participant’s data. Our methodology then tests if the average decoding accuracy (a function of the participant-specific patterns) shows stable patterns across participants.

As mentioned previously, the regression model can be trained and then tested within or across conditions. For simplicity, we will refer to each train/test regime using a pair of context names separated by a slash, with the word before the slash referring to the training context, and the word after the slash referring to the testing context. For example, if we were to both train and test in isolation, we would refer to it as isolation/isolation. If we were to train in the isolation context and test in the phrase context, we would refer to it as isolation/phrase.

Within-context accuracy (e.g., phrase/phrase) indicates the consistency of the representation across trials within a context. Across-context accuracy (e.g., isolation/phrase) indicates how consistent the neural representation is between the two contexts. Thus, high accuracy for nouns in a isolation/phrase analysis would indicate that the neural representation of the noun is similar in isolation and in list context. Accuracy was computed using the 2 versus 2 test.

Computing decoding accuracy using the 2 versus 2 test

On a dataset of N averaged epochs, we hold out 2 averaged epochs and the two corresponding target vectors (yi, yj). We train the model on the remaining N-2 averaged epochs and N-2 target vectors. Testing the model on the 2 held-out averaged epochs provides two predicted semantic vectors (y^i, y^j). The 2 vs. 2 test measures how similar the predictions (y^i, y^j) are to their corresponding ground truth vectors (yi, yj) using a vector distance criterion d(v,u). While any kind of distance metric can be used, we opt for cosine distance. In particular, the test passes if the following equation holds:

d(y^i,yi)+d(y^j,yj)<d(y^i,yj)+d(y^j,yi) (1)

where the distance of matching vectors is smaller than the distance of non-matching vectors. We award a score of 1 if the test passes, 0 if it fails, and 0.5 if the two summations are equal. The reported 2-vs-2 accuracy is the average score of the 2-vs-2 test on every possible pair in the dataset (20 choose 2, denoted (202)=190). Recall that our 20 data instances contain 4 samples for each word. For this reason, 30 of the 190 2-vs-2 pairs will pair two samples of the same word. Such a pairing renders the 2-vs-2 test degenerate because yi = yj and the two halves of Eq (1) are equal. Thus, there are a total of (202)30=160 valid 2-vs-2 pairs. An illustration of the 2 vs. 2 test appears in Fig 2D.

Statistical significance

To assess the statistical significance of the 2-vs-2 accuracy, we use permutation tests. In permutation tests we randomly shuffle the mapping of target utterances to MEG epochs. This simulates the scenario where there is no meaningful relation between the MEG recordings and the target utterances. We train and test our model on datasets built with 100 randomly shuffled utterance-to-epoch mappings, producing 100 decoding accuracies. As expected, the mean decoding accuracy on the shuffled datasets is at chance (50%). We then fit a normal kernel density function to the histogram of decoding accuracies to form a null distribution. From this null distribution, we calculate the p-value for the decoding accuracy of models trained with the original un-permuted labels. To determine if results are above chance, we correct the p-values for multiple comparisons over time using False Discovery Rate (FDR) with no dependency assumption (Benjamini-Hochberg-Yekutieli method; [24]).

When comparing the effect of context and word category on accuracy of the models, we find clusters of time where the 2-vs-2 accuracy differs significantly using a 2 x 3 ANOVA combined with the cluster permutation method [25]. We submit the 2-vs-2 accuracy of each time point to a 2 x 3 ANOVA (2 word categories by 3 conditions) to create p-values for main effects and interaction effects. For the cluster permutation method, we identify clusters of time where p < 0.05 for at least 3 adjacent time points (main and interaction effects considered separately). For each cluster, we assign a cluster-level statistic equal to the sum of F-values for all time points within the cluster. We report the largest time cluster in time window 0-400ms and 400-650ms to account for earlier and later effects. To correct for the final cluster-level p-value, we permute the accuracies by randomly assigning the word category and condition labels within each participant data for 10,000 times.

Temporal generalization matrices

Temporal generalization matrices (TGMs; [26]) were used to test if the patterns identified with our ridge regression models were stable across time and/or contexts [26]. For simplicity, we first describe how to use a TGM to evaluate across time but within context, and then generalize to evaluating across contexts.

To evaluate across time, instead of training and testing using data from the same time window, we form a matrix M where Mij contains the decoding accuracy of a model trained on a window centered at time i and tested on another window centered at time j. We leave out two averaged epochs from both time windows. We train the regression models using the N-2 remaining epochs from time window i, and test the regression models on the two left-out averaged epochs from time window j. If the neural representation of the word is consistent over time, then similar patterns will be leveraged by regression models trained on different time windows (thus yielding similar learned weights), resulting in above-chance decoding accuracy even when the train and test data are from differing time windows. If the representation of a word is stable over time, there will be high accuracy in blocks near-adjacent to the TGM diagonal, whereas if the representation re-emerges later in time there will be areas of high accuracy further from the diagonal (i.e., off-diagonal), separated from the diagonal by an area of lower accuracy.

We can also create cross-context TGMs, which test if representations of a certain word-type are consistent across contexts. Again, we form a matrix M where Mij contains accuracy of the prediction model trained on data from context A (e.g. phrase), time i and tested on data from context B (e.g. list), time j. In a cross-context TGM, if the representation of a word-type is similar in both contexts at the same time points (i.e., when i = j), there will be high accuracy along the diagonal. If the representations are similar across the contexts, but at different times (i.e., with some lag), we see high off-diagonal accuracy. (For an excellent tutorial on TGMs, see [26]. For a more language specific interpretation, see [27]).

Results

Trials containing behavioral errors were excluded from our analysis. Erroneous articulations included productions of wrong names and utterance repairs. Response accuracy was always above 97%. The average latency of speech onset for each context in increasing order was: adjective-only, 772ms; noun-only, 792ms; list, 897ms; and phrase, 917ms.

Effect of category and context on word decodability (within-condition, within-timepoint decoding)

The time courses of noun and adjective decodability are shown in Fig 3, broken down by context, with dots above the x-axes indicating windows of reliable, significantly above-chance accuracy. Here, training and testing data are from the same time window and the same condition. Zero indicates the onset of the picture stimuli.

Fig 3. Decodability across time for adjectives and nouns when presented within phrases, lists or in isolation as single words.

Fig 3

The grey shading indicates a significant main effect of category on decodability across all contexts, with nouns showing higher accuracy than adjectives in the mid-latency time-window of 240-395ms after picture onset. Dashed lines above the x-axes indicate when decoding accuracy was reliable for the nouns (red) and adjectives (blue). Blue shading indicates the intervals during which the main effect of context was observed, that is, higher decodability of both categories when occurring in two-word contexts (phrase or list). Though not shown, there is an interaction effect 100–190 ms.

When were nouns and adjectives in general decodable above chance? While nouns were reliably decodable in all three contexts for sustained periods of time (phrase: 105–365, 385, 420-460ms; list: 110–170, 195–215, 225–355, 380ms; isolation: 140–190, 200–215, 235–240, 325-650ms), adjective decodability was mostly limited to a late time-window close to articulation in all three tasks (phrase: 595–620, 630-650ms; list: 65–245, 520–530, 565-650ms; isolation: 430, 445-650ms). In addition, adjectives in lists showed an early peak of high accuracy (65-245ms), possibly due to the somewhat artificial attention that needed to be paid to the background colors in the list task. Though we did not explicitly test for the effort needed for each naming task, we find it plausible that this may increase neural demands in some way that could also increase decodability.

In the phrase condition, only nouns showed reliable decodability, and this lasted through much of the epoch. In lists, adjectives were robustly decodable in an early time-window of ~100-250ms, while the time course of noun decodability was similar to the phrase context. Finally, isolated single words showed the same contrast as phrases: more reliable and longer lasting decodability of nouns than adjectives, though at the end of the epoch, adjective decodability did reach significance.

The effect of category and context on decodability was evaluated with a 2 x 3 ANOVA with word-category (adjective, noun) and context (isolation, list, phrase) as factors (Fig 3). A main effect of word-category was significant at 240-395ms, with higher decoding accuracies for nouns than adjectives. A main effect of word-category was also observed at 580-650ms, with higher decoding accuracies for adjectives than nouns (Fig 3, gray shading). We found an interaction effect for this 2 X 3 ANOVA at 100-190ms (p < 0.00001; not illustrated).

As our main question pertained to the effect of context on single word representations, we conducted further, more targeted analyses contrasting the isolated word stimuli to the phrases and list trials in two separate 2 x 2 ANOVAs. A phrasal context indeed enhanced the decodability of both nouns and adjectives, as compared to an isolated word context, but so did a list context (isolation/isolation vs phrase/phrase: context effect at 245-295ms and 510-650ms. Similarly, isolation/isolation vs list/list: context effect at 60-175ms and 405-650ms.). Thus, we are not able to attribute this increase in decodability to composition specifically. All 2 x 2 Anova results appear in the S1 Appendix.

Generalizability of word representations across time in each context (within-condition decoding)

While the decoding analysis just described—with training and testing always using the same time point—did not reveal a compelling effect of composition on single word decodability, the analysis using TGMs did (Fig 4). In particular, during phrase planning, noun representations trained at early time points, starting at ~100ms, stayed active/decodable until about 400ms post picture onset (above chance accuracy regions outlined in black, Fig 4). This was not the case during the planning of lists, nor for adjective planning in any context.

Fig 4. Within-condition TGMs.

Fig 4

Within-condition TGMs showing the temporal generalizability of noun (A) and adjective (B) representations from training time X to testing time Y in the three contexts. When nouns occurred in phrases, their representations generalized between earlier and later time-points in a way that was not observed for nouns in non-phrasal contexts or for adjectives in any context. This is evidenced by the off-diagonal instances of reliable decoding in the Noun in Phrase results (A, left). The right-most column shows subtractions between phrasal and non-phrasal contexts, with black boxing indicating significant differences.

Generalizability of word representations across time and from isolation to two-word contexts (across-condition decoding)

Finally, our across-condition TGMs addressed the degree of similarity between word representations when produced as isolated words as opposed to when planned together with another word, in order to produce either a phrase or a list (Fig 5). Decoding accuracy of noun representations was reliable even when the training used isolated nouns and the testing used phrase data. The decoded representations also generalized across time, such that noun representations that were successfully decoded at 100-200ms disappeared and then re-emerged about a hundred milliseconds later, while representations at 300-400ms stayed active in a more sustained fashion until 500-600ms (above-chance regions outlined in black, Fig 5). The significantly above chance accuracy appears mostly below the diagonal, implying that the representation seen earlier in the isolation context matches the later representation in the phrase context. Isolated noun representations did not generalize well to list contexts, and did not show generalization across time (train isolation, test list in Fig 5A).

Fig 5. Across-condition TGMs showing the decodability and temporal generalizability of isolated word representations to phrase and list contexts.

Fig 5

Classifiers were trained on nouns and adjectives as they occurred in the isolation context and then tested when those same words occurred within phrases or lists. (A) Neural representations of nouns were sufficiently similar in isolation and in phrases such that decoding was reliable starting at 100ms and lasting till almost the end of the epoch. These representations also showed temporal generalizability starting at 100ms. Representations active at 100-200ms disappeared and then re-emerged about a hundred milliseconds later, while representations at 300-400ms stayed active in a more sustained fashion until the end of the epoch. (B) Adjective representations, in contrast, did not generalize from isolated contexts to phrasal contexts nearly as robustly. Mainly, shared representations across these two contexts were observed in a late time-window, close to articulation, at 500-600ms. This could reflect planning of the adjective articulation, which was the first word to be uttered in all three depicted contexts. In a similar late time-window, isolated adjective representations generalized to adjectives in the lists, though more weakly.

Adjective representations did not generalize from isolated contexts to phrasal contexts nearly as robustly as nouns. Mainly, evidence of shared representations across these two contexts were observed in a late time-window, close to articulation, at 500-600ms. This could reflect planning of the adjective articulation, which in all these contexts was the first word to be uttered. In a similar late time-window, isolated adjective representations generalized to adjectives in the lists, though more weakly.

Discussion

This work addressed the nature and time course of noun and adjective representations in phrasal, isolated word, and list contexts. How does a simple combinatory context affect the neural representation of a word? Are adjectives and nouns planned in symmetric fashion during language production, or do these word types elicit different activation time courses when measured with a decoding approach? Our study yielded three major findings. First, apart from a late time window shortly prior to articulation, nouns were generally more decodable than adjectives. Second, both adjectives and nouns were more decodable when the task required the production of two words, either as a phrase or a list. And finally, we used TGMs to evaluate the temporal evolution of specific semantic representations. As these representations are activated en route from picture onset to articulation, and nouns were planned as heads of phrases, the representations active soon after picture onset stayed active up to 400ms into the epoch. Such a profile was entirely absent when nouns were planned as single words or within lists, and for all cases of adjective planning. Our across-condition decoding also provided evidence of similar representations for isolated nouns and nouns in phrases, while such evidence was much weaker for the generalizability of isolated nouns to nouns in lists, or from isolated adjectives to adjectives in phrases or lists. In sum, our findings suggest that during production planning, the neural representations of nouns are more stable, and therefore more decodable, than those of adjectives, and that during phrase planning, noun representations generalize across time and contexts in ways that adjective representations do not.

Timeline of adjective and noun decodability

In general, across the full design, the first word to be uttered was decodable in a late time-window, shortly preceding speech onset. Given the late timing, the representations driving this result are likely motor related.

But earlier on 240-395ms during the language planning process, adjectives and nouns differed in their decodability. Within each context, nouns were consistently decodable, and were, for the most part, more decodable than adjectives. This was upheld by our 2 x 3 ANOVA, directly comparing adjective and noun decodability across all three contexts, which indicated greater noun than adjective decodability in early-to-mid-latency time windows. It appears, then, that noun representations stayed active for a protracted period of time, while adjective representations were only stable immediately prior to utterance onset. We also observed an effect of context, such that whenever either nouns or adjectives were planned as part of two-word utterances, whether they be lists or phrases, decoding accuracy was higher. Since we were not able to pinpoint this effect as directly relating to phrasal composition, it connects only loosely to our research question. It may stem from a higher level attention when planning two word expressions as opposed to single words. We leave this question for future work and focus our discussion on the higher decodability of nouns over adjectives.

Although color-adjectives and object-nouns differ in many ways, the restricted nature of our stimulus choices could have flattened out differences that one might observe in a more ecologically valid context. Nevertheless, a clear time-course difference was observed. With the current data, we cannot determine the cause of this, but multiple possibilities exist for future research to explore. Most interestingly, the difference in noun vs. adjective decodability could be driven by genuine semantic differences between the two word types. For example, objects are usually perceivable via multiple senses, we can feel them, see them, and perhaps hear them, but colors can be experienced only through vision. This could lead to less robust neural representations for colors. Relatedly, it has been shown that color dissociates from many other physical properties when comparing semantic representations in sighted and blind individuals [28]. Although the sighted and the blind appear to have similar representations for attributes such as shape and texture, this is not the case for color. It has been hypothesized that this may result from the lesser taxonomic value of color: color is a much weaker predictor of object kind than for example shape [28]. Thus the weaker associations between color and other object properties could also result in less detectable neural representations for color terms.

Color terms are also ambiguous in ways that we have not yet discussed [29]. For example, despite often occurring as textbook examples of a context-insensitive modifier, the interpretations of color terms are actually quite context-sensitive–compare red hair and red wine for example [30]. Although our experiment did not employ different hues, this underlying variability could nevertheless contribute to lesser decodability for color terms. There are also differences in ambiguity as regards syntactic category. Our study used two categories, nouns and adjectives. While our nouns (bag, bell, cane, lamp and plane) are very unlikely as adjectives in English, all our color-adjectives are actually also mass nouns (I like milk; I like blue). Given this, we cannot rule out the possibility that in the list context (red, bag), the participants were naming the background colors as nouns. If colors occurred both as nouns and as adjectives within the experiment, this also could have affected decodability.

Role of phrasal composition in the temporal and contextual generalizability of noun representations

In addition to addressing the general time course of word representations as they participate in combinatory phrasal planning, our method allowed us to examine the relationship between representations active at different times and between representations active in different contexts. Although our within-condition, within-timepoint analysis did not reveal compelling effects of phrasal composition on either noun or adjective representations; the generalizability of noun, but not adjective, representations was clearly enhanced by a phrasal context, in the following two ways.

First, the temporal generalizability of noun representations was enhanced by a phrasal context, both as compared to isolation and list contexts (Fig 4). This finding suggests a cascade of noun representations, many of which stay active for a while. In contrast, adjective representations showed almost no generalization across time, consistent with a model in which the representation of an adjective changes consistently across time, with new representations replacing the old ones, with no sustained activations (“chain” pattern in [26]). It is interesting that the compositional context produced the most temporally generalizable representations, as the hypothesis a priori may have been that composition would change the representation more over time. However, the stimuli adjectives are largely intersective, and so it is possible that such a compositional change is less apparent in this experiment. Nouns in phrases was also the only time we observed the amount of temporal generalization reported by Fyshe et al. [2], which showed large swaths of above-chance off-diagonal accuracy. This could be for several reasons, including that the Fyshe study used a reading paradigm that displayed one word at a time, and so the representations for adjectives and nouns were neatly and predictably separated in time.

Second, the contextual generalizability of isolated nouns to two-word contexts was higher when the two words formed a phrase (Fig 5). The same was not true of adjectives, which showed much weaker decodability even within condition (Fig 4). Particularly interesting in the across-condition decoding of nouns was evidence of reactivated representations: representations active at 100-200ms disappeared and then re-emerged about a hundred milliseconds later, providing some support to the 10 Hz oscillatory activity reported by Fyshe et al. [2]. In contrast, representations decoded at 300-400ms stayed active in a more sustained fashion until the end of the epoch. Though a theoretical understanding of this detailed pattern requires further experimentation, the general finding emerging from these results is that the head of the phrase, the noun, engages a much more stable set of neural representations than its modifier, the adjective. Given our highly controlled stimulus materials and symmetric nature of the paradigm, the asymmetry is striking, and further studies could search for contexts that eliminate that asymmetry. Assessing which aspects of the decoding results stem from word order would be straightforward with a language that uses a different word order. The syntactic relation of the two elements can also be altered using a language in which noun-adjective pairs can convey a predicative relation without an overt copula: boat (is) red (cf., [31]). In sum, our findings offer a description of the representational patterns of nouns and adjectives during English phrase planning, giving rise to a host of novel hypotheses for further investigation.

Conclusion

This work addresses the temporal evolution and context sensitivity of noun and adjective representations during phrase planning in production. We discovered a robust asymmetry between nouns and adjectives, with noun representations being generally more decodable, more consistent between isolated and phrasal contexts and more sustained over time in phrases than those of adjectives. While our findings are not yet highly theoretically constraining, they open up a rich space of testable hypotheses about the critical factors driving the observed contrasts, in terms of either the structural or semantic properties of the two word classes.

Supporting information

S1 Appendix

(DOCX)

Acknowledgments

We thank Chris Barker for useful discussion on the semantics of color terms and for related references.

Data Availability

The data underlying this study is available on OSF (https://osf.io/p7gc6/) and the code is available on GitHub (https://github.com/fyshelab/NeuralPhraseComposition).

Funding Statement

This research was supported by the NYUAD Research Institute (https://nyuad.nyu.edu/en/) under Grant G1001 (LP), by the Natural Sciences and Engineering Research Council of Canada (NSERC, https://www.nserc-crsng.gc.ca/index_eng.asp) through a Discovery Grant (AF), and the Canada CIFAR (Canadian Institute for Advanced Research, https://www.cifar.ca/) AI Chair program (AF). The computational work was supported in part by infrastructure made available by WestGrid (https://www.westgrid.ca/) and Compute Canada (https://www.computecanada.ca/) (AF). These funders played no role in the study design, data collection and analysis, decision to publish, nor preparation of the manuscript.

References

  • 1.Pylkkänen L. The neural basis of combinatory syntax and semantics. Science. 2019;366(6461): 62–66. 10.1126/science.aax0050 [DOI] [PubMed] [Google Scholar]
  • 2.Fyshe A, Sudre G, Wehb L, Rafidi N, Mitchell TM. The lexical semantics of adjective-noun phrases in the human brain. Hum brain mapp. 2019;40(15): 4457–4469. 10.1002/hbm.24714 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Blanco-Elorrieta E, Ferreira VS, Del Prato P, Pylkkänen L. The priming of basic combinatory responses in MEG. Cognition. 2018;170: 49–63. 10.1016/j.cognition.2017.09.010 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 4.Blanco-Elorrieta E, Kastner I, Emmorey K, Pylkkänen L. Shared neural correlates for building phrases in signed and spoken language. Scientific reports. 2018;8(1): 1–10. 10.1038/s41598-017-17765-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 5.Pylkkänen L, Bemis DK, Elorrieta EB. Building phrases in language production: An MEG study of simple composition. Cognition. 2014;133(2): 371–384. 10.1016/j.cognition.2014.07.001 [DOI] [PubMed] [Google Scholar]
  • 6.Miozzo M, Pulvermüller F, Hauk, O. Early parallel activation of semantics and phonology in picture naming: Evidence from a multiple linear regression MEG study. Cereb cortex. 2015;25(10): 3343–3355. 10.1093/cercor/bhu137 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 7.Erk K. Vector space models of word meaning and phrase meaning: A survey. Lang linguist compass. 2012;6(10):635–653. 10.1002/lnco.362 [DOI] [Google Scholar]
  • 8.Mitchell TM, Shinkareva SV, Carlson A, Chang K-M, Malave VL, Mason RA, et al. Predicting human brain activity associated with the meanings of nouns. Science. 2008;320(5880): 1191–1195. 10.1126/science.1152876 [DOI] [PubMed] [Google Scholar]
  • 9.Sudre G, Pomerleau D, Palatucci M, Wehbe L, Fyshe A, Salmelin R, et al. Tracking neural coding of perceptual and semantic features of concrete nouns. Neuroimage. 2012;62(1): 463–451. 10.1016/j.neuroimage.2012.04.048 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 10.Jat S, Tang H, Talukdar P, Mitchell T. Relating simple sentence representations in deep neural networks and the brain. arXiv:1906.11861v1. 2019. [cited year month day]. Available from: https://arxiv.org/abs/1906.11861 [Google Scholar]
  • 11.Wehbe L, Murphy B, Talukdar P, Fyshe A, Ramdas A, Mitchell T. Simultaneously uncovering the patterns of brain regions involved in different story reading subprocesses. PLoS ONE. 2014;9(11), 1–19. 10.1371/journal.pone.0112575 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 12.Huth AG, de Heer WA, Griffiths TL, Theunissen FE, Gallant JL. Natural speech reveals the semantic maps that tile human cerebral cortex. Nature. 2016;532(7600): 453–458. 10.1038/nature17637 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 13.Jain S, Huth A. Incorporating context into language encoding models for fMRI. Adv neural inf process syst 31 (NeurIPS 2018). 2018; 6629–6638. 10.1101/327601 [DOI] [Google Scholar]
  • 14.Chomsky N. Syntactic structures. Berlin: De Gruyter Mouton; 1957. [Google Scholar]
  • 15.Heim I, Kratzer A. Semantics in generative grammar (Vol. 1185). Oxford: Blackwell; 1998. [Google Scholar]
  • 16.Levelt WJM. Spoken word production: A theory of lexical access. Proceedings of the National Academy of Sciences. 2001;98(23); 13464–13471. 10.1073/pnas.231459498 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 17.Balota DA, Yap MJ, Hutchison KA, Cortese MJ, Kessler B, Loftis B, et al. The English lexicon project. Behavior res methods. 2007; 39(3): 445–459. 10.3758/bf03193014 [DOI] [PubMed] [Google Scholar]
  • 18.Adachi Y, Shimogawara M, Higuchi M, Haruta Y, Ochiai M. Reduction of non-periodic environmental magnetic noise in MEG measurement by continuously adjusted least squares method. IEEE trans appl supercond. 2001;11(1): 669–672. [Google Scholar]
  • 19.Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: Machine learning in Python. J mach learn res. 2011; 12: 2825–2830. [Google Scholar]
  • 20.Gramfort A, Luessi M, Larson E, Engemann D, Strohmeier D, Brodbeck C, et al. MEG and EEG data analysis with MNE-Python. Frontiers in neuroscience. 2013;7: 267. 10.3389/fnins.2013.00267 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 21.Oostenveld R, Fries P, Maris E, Schoffelen JM. FieldTrip: Open source software for advanced analysis of MEG, EEG, and invasive electrophysiological data. Comput intell and neurosci. 2011: 156869. 10.1155/2011/156869 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 22.Mikolov T, Corrado G, Chen K, Dean J. Efficient estimation of word representations in vector space. Proceedings of the International Conference on Learning Representations (ICLR 2013). 2013: 1–12.
  • 23.Hastie T, Tibshirani R. Efficient quadratic regularization for expression arrays. Biostatistics. 2004;5(3), 329–340. 10.1093/biostatistics/5.3.329 [DOI] [PubMed] [Google Scholar]
  • 24.Benjamini Y, Yekutieli D. The control of the false discovery rate in multiple testing under dependency. Ann stat. 2001;29(4): 1165–1188. 10.1214/aos/1013699998 [DOI] [Google Scholar]
  • 25.Maris E, Oostenveld R. Nonparametric statistical testing of EEG-and MEG-data. J neurosci methods. 2007;164(1): 177–190. 10.1016/j.jneumeth.2007.03.024 [DOI] [PubMed] [Google Scholar]
  • 26.King JR, Dehaene S. Characterizing the dynamics of mental representations: The temporal generalization method. Trends cogn sci. 2014;18: 203–210. 10.1016/j.tics.2014.01.002 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 27.Fyshe A. Studying language in context using the temporal generalization method. Philos trans r soc Lon B. 2019;375(1791). 10.1098/rstb.2018.0531 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 28.Kim JS, Elli GV, Bedny M. Knowledge of animal appearance among sighted and blind adults. Proceedings of the National Academy of Sciences. 2019;116(23): 11213–11222. 10.1073/pnas.1900952116 [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 29.Kennedy C, McNally L. Color, context, and compositionality. Synthese. 2010;174(1): 79–98. 10.1007/s11229-009-9685-7 [DOI] [Google Scholar]
  • 30.Cohen B, Murphy GL. Models of concepts. Cognitive science. 1984;8(1): 27–58. [Google Scholar]
  • 31.Matar S, Dirani J, Marantz A, Pylkkänen L. Dissociating syntactic processing and semantic composition in the left temporal lobe: MEG evidence from standard Arabic. Society for the Neurobiology of Language, Virtual Meeting. October 2020. Presentation.

Decision Letter 0

Nicola Molinaro

22 Dec 2020

PONE-D-20-34763

Neural representation of words within phrases: Temporal evolution of color-adjectives and object-nouns during simple composition

PLOS ONE

Dear Dr. Fyshe,

Thank you for submitting your manuscript to PLOS ONE. After careful consideration, we feel that it has merit but does not fully meet PLOS ONE’s publication criteria as it currently stands. Both Reviewers noticed some critical aspects and required clarifications that should be addressed to improve the overall quality of the Manuscript. Therefore, we invite you to submit a revised version of the manuscript that addresses the points raised during the review process.

Please submit your revised manuscript by Feb 05 2021 11:59PM. If you will need more time than this to complete your revisions, please reply to this message or contact the journal office at plosone@plos.org. When you're ready to submit your revision, log on to https://www.editorialmanager.com/pone/ and select the 'Submissions Needing Revision' folder to locate your manuscript file.

Please include the following items when submitting your revised manuscript:

  • A rebuttal letter that responds to each point raised by the academic editor and reviewer(s). You should upload this letter as a separate file labeled 'Response to Reviewers'.

  • A marked-up copy of your manuscript that highlights changes made to the original version. You should upload this as a separate file labeled 'Revised Manuscript with Track Changes'.

  • An unmarked version of your revised paper without tracked changes. You should upload this as a separate file labeled 'Manuscript'.

If you would like to make changes to your financial disclosure, please include your updated statement in your cover letter. Guidelines for resubmitting your figure files are available below the reviewer comments at the end of this letter.

If applicable, we recommend that you deposit your laboratory protocols in protocols.io to enhance the reproducibility of your results. Protocols.io assigns your protocol its own identifier (DOI) so that it can be cited independently in the future. For instructions see: http://journals.plos.org/plosone/s/submission-guidelines#loc-laboratory-protocols

We look forward to receiving your revised manuscript.

Kind regards,

Nicola Molinaro, Ph.D.

Academic Editor

PLOS ONE

Journal requirements:

When submitting your revision, we need you to address these additional requirements.

1. Please ensure that your manuscript meets PLOS ONE's style requirements, including those for file naming. The PLOS ONE style templates can be found at

https://journals.plos.org/plosone/s/file?id=wjVg/PLOSOne_formatting_sample_main_body.pdf and

https://journals.plos.org/plosone/s/file?id=ba62/PLOSOne_formatting_sample_title_authors_affiliations.pdf

2.In your Data Availability statement, you have not specified where the minimal data set underlying the results described in your manuscript can be found. PLOS defines a study's minimal data set as the underlying data used to reach the conclusions drawn in the manuscript and any additional data required to replicate the reported study findings in their entirety. All PLOS journals require that the minimal data set be made fully available. For more information about our data policy, please see http://journals.plos.org/plosone/s/data-availability.

Upon re-submitting your revised manuscript, please upload your study’s minimal underlying data set as either Supporting Information files or to a stable, public repository and include the relevant URLs, DOIs, or accession numbers within your revised cover letter. For a list of acceptable repositories, please see http://journals.plos.org/plosone/s/data-availability#loc-recommended-repositories. Any potentially identifying patient information must be fully anonymized.

Important: If there are ethical or legal restrictions to sharing your data publicly, please explain these restrictions in detail. Please see our guidelines for more information on what we consider unacceptable restrictions to publicly sharing data: http://journals.plos.org/plosone/s/data-availability#loc-unacceptable-data-access-restrictions. Note that it is not acceptable for the authors to be the sole named individuals responsible for ensuring data access.

We will update your Data Availability statement to reflect the information you provide in your cover letter.

3.Thank you for stating the following in your Competing Interests section: 

"No"

Please complete your Competing Interests on the online submission form to state any Competing Interests. If you have no competing interests, please state "The authors have declared that no competing interests exist.", as detailed online in our guide for authors at http://journals.plos.org/plosone/s/submit-now

 This information should be included in your cover letter; we will change the online submission form on your behalf.

Please know it is PLOS ONE policy for corresponding authors to declare, on behalf of all authors, all potential competing interests for the purposes of transparency. PLOS defines a competing interest as anything that interferes with, or could reasonably be perceived as interfering with, the full and objective presentation, peer review, editorial decision-making, or publication of research or non-research articles submitted to one of the journals. Competing interests can be financial or non-financial, professional, or personal. Competing interests can arise in relationship to an organization or another person. Please follow this link to our website for more details on competing interests: http://journals.plos.org/plosone/s/competing-interests

Reviewers' comments:

Reviewer's Responses to Questions

Comments to the Author

1. Is the manuscript technically sound, and do the data support the conclusions?

The manuscript must describe a technically sound piece of scientific research with data that supports the conclusions. Experiments must have been conducted rigorously, with appropriate controls, replication, and sample sizes. The conclusions must be drawn appropriately based on the data presented.

Reviewer #1: Yes

Reviewer #2: Yes

**********

2. Has the statistical analysis been performed appropriately and rigorously?

Reviewer #1: I Don't Know

Reviewer #2: Yes

**********

3. Have the authors made all data underlying the findings in their manuscript fully available?

The PLOS Data policy requires authors to make all data underlying the findings described in their manuscript fully available without restriction, with rare exception (please refer to the Data Availability Statement in the manuscript PDF file). The data should be provided as part of the manuscript or its supporting information, or deposited to a public repository. For example, in addition to summary statistics, the data points behind means, medians and variance measures should be available. If there are restrictions on publicly sharing data—e.g. participant privacy or use of data from a third party—those must be specified.

Reviewer #1: Yes

Reviewer #2: Yes

**********

4. Is the manuscript presented in an intelligible fashion and written in standard English?

PLOS ONE does not copyedit accepted manuscripts, so the language in submitted articles must be clear, correct, and unambiguous. Any typographical or grammatical errors should be corrected at revision, so please note any specific errors here.

Reviewer #1: Yes

Reviewer #2: Yes

**********

5. Review Comments to the Author

Please use the space provided to explain your answers to the questions above. You may also include additional comments for the author, including concerns about dual publication, research ethics, or publication ethics. (Please upload your review as an attachment if it exceeds 20,000 characters)

Reviewer #1: The present study proposes a decoding approach to study noun adjective brain representations from MEG recordings during a production task. The original data corresponds to a previously published study in which pictures had to be described only using nouns, adjectives, or noun adjectives in a compositional or list manner.

The authors track the neural representations of noun and adjectives over time, and explore the similarity between the representation of noun and adjectives before producing the words in isolation compared to the representation of the same words produced during compositional or list like contexts by training and testing the regression models across conditions.

They conclude that noun and adjective representations behave differently: nouns are more decodable and their representation is more consistent across time and context, whereas adjectives are less decodable across contexts and time, and hence their representation is more variable.

The manuscript is well written, the objectives and hypothesis are clearly stated, the methodological approach seems correctly conducted and is consistent with the author's questions. I specially value the collaboration between research groups and the repurpose of already collected data.

Overall the paper is good but I would suggest the authors to make more explicit some specifications of the statistical analyses (see below) as well as to enrich the discussion on the following points:

Variability of noun adjective representations in comprehension and production tasks.

Invariability of noun representation across time in this production task is somewhat unexpected. In the phrasal context the combination of noun and adjective would elicit a different representation of the noun (a lamp that is red, not any lamp), and the combined representation would have to be broken down into their constituent representations (red, lamp) to produce the correct articulation. Considering that the original article shows a composition activity, noun representation when modified by the adjective should correspond to a different representation than the noun in isolation. In this sense brain representation of noun and adjective in the list context would be expected to be more consistent across time than noun in phrases.

Moreover, how does this result on adjective variability and noun robustness of representation in the production task relate to the somewhat symmetrical result in the comprehension task studied in Fyshe et al., 2019.

Finally, did the authors explore how are the decoding accuracies for nouns and adjectives when training and testing the models across the context conditions? This could provide important information on the effect of composition on noun and adjective representation.

Accuracy before speech onset

The authors mention that the increased decoding accuracy for the adjective representation prior to word articulation is motor related. This would mean that the early representation of adjectives is also partly motor related? The hypothesis suggested by the authors would benefit from some discussion on word multimodal representation at early stages of word processing.

In relation to the methods sections some points should be made more clear and detailed so reproduction is possible:

M1. The preprocessing parameters seem to be different from the original paper (i.e: epochs length, filters). If this is the case, the authors should specify all of the information concerned with the preprocessing (i.e: trial rejection, all filter specifications). If the authors did not start their study from raw data I would suggest the authors to refer the readers to the original paper for the preprocessing details. Although note that in the original paper filter specifications that would allow replicating the preprocessing are missing (filter type, cutoff frequency, filter order, roll-off or transition bandwidth, direction of computation).

M2. Authors should be more explicit on the details of the permutation cluster analysis. The null distribution was constructed taking the cluster with the maximum statistic sum? or all clusters F sum were included?.

If the statistic F corresponds to the interaction obtained by a 3x2 ANOVA for the time points 100-190ms, how were the main effect cluster times determined?, this should be more thoroughly explained and justified.

M3. I would suggest the authors to incorporate the statistical results for the 2x2 ANOVAs contrasting the isolated word stimuli to the phrases and list trials

Minor issues are detailed below separated by sections

METHODS

M4. Please specify the order of presentation for the different condition blocks

M5. “with a sampling frequency of 1000Hz” This refers to the recording parameters not the filter specifications?

M6. Note 40HZ in all caps

M7. Misphrased: total of N choose 2 pairs in p.7

M8. Authors should cite software and statistical packages used to carry the analyses, providing information on versions, etc.

RESULTS

R1. This phrase is not clear: “A main effect of word-category was significant at 240-395ms, with higher decoding accuracies for nouns than adjectives and at 580-650, with higher decoding accuracies for adjectives than nouns (Figure 2, gray shading).”

R2. As separate models for each participant were trained authors should report the variability across subjects on top of the average model performance.

R3. Misphrased: “Isolated noun representations HAD DID not generalize well to list contexts, and did not show generalization across time”

DISCUSSION

D1. Replace an for a in “of an adjectives” in p.15

D2. Misphrased “which no sustained activations” in p.15

Reviewer #2: The paper by Honari-Jahromi and colleagues presents an important and under-researched question in cognitive neuroscience of language – how stable are the neural representations of the semantic properties of adjectives and nouns across both temporal, contextual and combinatorial dimensions. To answer this question they use a MEG data and a combination of decoding techniques. Results they present are interesting and compelling. Overall I am very happy with both the quality of the paper, novelty of the question and the analysis used to answer it, however there are several areas that would require improvement and clarifications.

Major

1. The methods section describing the analysis used (p. 7) requires a lot of clarification. The stages of the decoding analysis are not clear from the text and at present it would be difficult to replicate it. For instance, what is the dimensionality of the data and predictor matrixes? What distance d was used for vector comparison (Cosine? Mahalanobis?)? In each regression (n=300) the predicted single value (the nth dimension of the semantic vector) is derived/estimated from channel x time matrix? Was the data averaged over these 100ms? Was here some dimensionality reduction performed on the sensors? Or was the stimulus value estimated by integration over all 100 time points and channels (multiplied by the learned decoder matrix)? I think for the benefit of the reader an explanatory figure (showing data dimensionality and schematic representation of steps) and formula describing the model as well as references to the exact method are necessary.

2. Semantic vectors/embeddings when derived from collocational matrixes or with neural networks to some degree reflect the frequency of the words they encode. Vectors of more frequent words tend to be more similar to each other than vectors of less frequent words. In this dataset, were the adjectives and nouns matched on word form / lemma frequency? Could better decodability of nouns across contexts and times be simply due to their higher frequency and greater vector similarity (to each other), when comped to adjectives?

3. Since decoding happened within participants (not on pooled data) and the accuracies were simply averaged, this means that decoding variability between participants was not “accounted for” and observed effects cannot be generalised outside of this sample – the models trained on one participant cannot be used to predict data in another participant/s – and this needs to be acknowledged explicitly.

Minor

1. Introduction p.3 para.2 - “complex effects of ambiguity” is a bit ambiguous. Do you mean in sentences or in narratives? In complex and naturalistic listening conditions. Please clarify.

2. Methods section p.5 “spoken and signed language as regards to the neural correlates ..” did you mean “with regards to”?

3. Did any artefact rejection or blink / heart beat removal with ICA take place? I appreciate that when using decoding artefact rejection is not strictly compulsivity as with ERPs but it is good to state this explicitly and give some references.

4. The Statistical significance section on p. 8 is somewhat difficult to read, especially the second paragraph. Was ANOVA done first across all time points and then 0.05 cutoff applied (seems that way since F values were summed for cluster mass estimate)? The word ‘label’ is ambiguous, please clarify.

5. p.9 para 2 “… similar patterns will be picked out by regression.. ” not clear, not sure ‘patterns’ is the right word. Do you mean to say regression weights learned in one window will also generate above-chance performance in another window.

6. Results p.10 was responce accuracy above 90% for all participants?

7. Figure 2 – most of the figure 1 caption text belongs in the text of the results section, not in the figure caption. Also please explain exactly why you think the demand to attend to the background colour improved decodability for adjectives in lists.

8. p.11 “we found an interaction effect...” please explain

9. Figure 3 and 4 – does the black contouring indicate statistical significance? Again, please consider putting text that describes results out of the caption and into main text.

10. It would be useful throughout the manuscript to refer to not simply ‘representations’ but semantic or lexicon-semantics representation, since this is what model was trained to decode (as opposed to say phonological representations).

While revising your submission, please upload your figure files to the Preflight Analysis and Conversion Engine (PACE) digital diagnostic tool, https://pacev2.apexcovantage.com/. PACE helps ensure that figures meet PLOS requirements. To use PACE, you must first register as a user. Registration is free. Then, login and navigate to the UPLOAD tab, where you will find detailed instructions on how to use the tool. If you encounter any issues or have any questions when using PACE, please email PLOS at figures@plos.org. Please note that Supporting Information files do not need this step.

PLoS One. 2021 Mar 4;16(3):e0242754. doi: 10.1371/journal.pone.0242754.r002

Author response to Decision Letter 0


9 Feb 2021

We thank our reviewers for their thoughtful comments on our work. We have addressed them in the new revision, and discuss each comment below. Our response is also provided as a color-coded pdf, which is likely easier to read than what follows here.

Reviewer #1:

Variability of noun adjective representations in comprehension and production tasks.

*Invariability of noun representation across time in this production task is somewhat unexpected. In the phrasal context the combination of noun and adjective would elicit a different representation of the noun (a lamp that is red, not any lamp), and the combined representation would have to be broken down into their constituent representations (red, lamp) to produce the correct articulation. Considering that the original article shows a composition activity, noun representation when modified by the adjective should correspond to a different representation than the noun in isolation. In this sense brain representation of noun and adjective in the list context would be expected to be more consistent across time than noun in phrases.

We agree that under composition meaning changes. However, color adjectives are largely intersectional and in this case, would not be expected to have a large effect on the semantics of the noun. In addition, our analyses search for the part of the representation that stays the same. There could be additional brain activity associated with the adjective, and so long as it doesn't change or otherwise hinder the representation of the noun we would expect to see no difference to the decoding accuracy of the noun. We have updated the discussion (section “Role of Phrasal Composition in the Temporal and Contextual Generalizability of Noun Representations”) to reflect this.

Interestingly there is little variability for nouns in lists (was Figure 3A, now Fig 4A, Noun list TGM). We believe the absence of off-diagonal above-chance accuracy stems from the fact that the noun need not be maintained in the mental workspace to be composed. Rather, it can be stored away until articulation is imminent.

Moreover, how does this result on adjective variability and noun robustness of representation in the production task relate to the somewhat symmetrical result in the comprehension task studied in Fyshe et al., 2019.

We have added a discussion of this to the section “Role of Phrasal Composition in the Temporal and Contextual Generalizability of Noun Representations”

Finally, did the authors explore how are the decoding accuracies for nouns and adjectives when training and testing the models across the context conditions? This could provide important information on the effect of composition on noun and adjective representation.

In the paper we included what we felt were the best cross-condition tests of how the meaning of isolated words change as they are used in phrases. Figure 5 (was Figure 4) shows training in isolation and testing in either list or phrase, and shows that the isolated representations are quite similar to the phrasal representation, but not to the list representations.

Although this was not our primary question of interest, we had run list/phrase ANOVAs, which indicated a consistent effect of category (nouns more decodable than adjectives mid-epoch, and adjectives more decodable than nouns late epoch) with no interaction when comparing list/phrase to phrase/phrase, but an interaction when comparing list/list to list/phrase, supporting our within-context word category effects and our conclusion that phrasal context enhances decodability of nouns.

Accuracy before speech onset

The authors mention that the increased decoding accuracy for the adjective representation prior to word articulation is motor related. This would mean that the early representation of adjectives is also partly motor related? The hypothesis suggested by the authors would benefit from some discussion on word multimodal representation at early stages of word processing.

We do not think the early adjective signal is motor related, because there is no off-diagonal accuracy in any graph that connects a late (pre-utterance) period to an early period.

In relation to the methods sections some points should be made more clear and detailed so reproduction is possible:

M1. The preprocessing parameters seem to be different from the original paper (i.e: epochs length, filters). If this is the case, the authors should specify all of the information concerned with the preprocessing (i.e: trial rejection, all filter specifications). If the authors did not start their study from raw data I would suggest the authors to refer the readers to the original paper for the preprocessing details. Although note that in the original paper filter specifications that would allow replicating the preprocessing are missing (filter type, cutoff frequency, filter order, roll-off or transition bandwidth, direction of computation).

We have updated the text to more clearly state the differences between processing:

“Trials were epoched at 100ms before to 700ms after stimuli onset to avoid contamination via motion artifact coinciding with overt speech and noise was reduced using Continuously Adjusted Least-Squares Method (Adachi et al., 2001). Epochs were baseline corrected using the average of a 100ms interval prior to the stimulus onset. Unlike Blanco-Elorrieta et al., we rejected only those trials with erroneous responses. MEG signals were band-passed using a Butterworth filter of order 20 between 0.1Hz and 40Hz. We used no ICA artefact rejection or blink / heart beat removal, as such artefacts are less problematic in decoding studies.”

M2. Authors should be more explicit on the details of the permutation cluster analysis. The null distribution was constructed taking the cluster with the maximum statistic sum? or all clusters F sum were included?

If the statistic F corresponds to the interaction obtained by a 3x2 ANOVA for the time points 100-190ms, how were the main effect cluster times determined?, this should be more thoroughly explained and justified.

The nonparametric clustering permutation based on 2by3 ANOVAs is done as follows. No model retraining is involved. The procedure is as follows:

1. run a 2by3 ANOVA for every time point and obtain "f-ratio "and "p_value" for each effect ( word category (2), condition (3) and interaction)

2. find clusters that span longer than 15ms (3 consecutive time points) where p-value of each point is less than 0.05 , record sum(f-ratios) for each cluster

3. assign "cluster corrected p-value" based on sum(f-ratios) from the null distribution formed as below:

repeat 10*1000{

1. shuffle the 2by3 matrix (full shuffle not row-wise or column-wise)

2. calculate 2by3 ANOVA as before

3. for the above clusters, record sum(f-ratios)

}

4. report the largest cluster in 0-400 and 400-650 to account for early and late effects.

We have incorporated this description into the “Statistical Significance” section.

M3. I would suggest the authors to incorporate the statistical results for the 2x2 ANOVAs contrasting the isolated word stimuli to the phrases and list trials

We have incorporated the 2 x 2 into an appendix to the paper. Here are the results from those tests:

2 x 2 ANOVA comparing isolation/isolation against phrase/phrase:

Main effect of context at 245-295ms and 510-650ms;

Main effect of word-category at 120-355ms;

no interaction effect.

2 x 2 ANOVA comparing list/list against phrase/phrase:

main effect of context at 70-100ms and 440-470ms;

main effect of word-category at 250-350ms and 605-650ms;

interaction effect at 95-195ms

2 x 2 ANOVA comparing isolation/isolation against list/list:

main effect of context at 60-175ms and 405-650ms;

main effect of word-category at 290-360ms;

interaction effect at 145-190ms

Minor issues are detailed below separated by sections

METHODS

M4. Please specify the order of presentation for the different condition blocks.

The order of blocks was randomized across participants with the only constraint that two blocks of the same condition never appeared consecutively. We have clarified this in the paper.

M5. “with a sampling frequency of 1000Hz” This refers to the recording parameters not the filter specifications?

fixed

M6. Note 40HZ in all caps

fixed

M7. Misphrased: total of N choose 2 pairs in p.7

This is terminology typically used for binomial coefficients. For clarity, we included the mathematical notation with the exact number. We also mentioned the excluded test pairs in the manuscript. Epoch averaging yielded 4 averaged epochs per noun (5 nouns) and 20 averaged epochs in total. Due to repeated nouns among epochs, we have excluded test pairs with the same nouns in our analysis. There are 30 of such pairs. Same procedure followed for adjectives.

M8. Authors should cite software and statistical packages used to carry the analyses, providing information on versions, etc.

Decoding analysis was done in scikit-learn (Pedregosa et al. 2011). Statistical significance test was done in MNE-Python (Gramfort et al. 2013) and FieldTrip (Oostenveld et al. 2011). The code for all analysis and respective software versions is available at https://github.com/mahon94/compositionInBrain.

Citations have been added to the draft.

RESULTS

R1. This phrase is not clear: “A main effect of word-category was significant at 240-395ms, with higher decoding accuracies for nouns than adjectives and at 580-650, with higher decoding accuracies for adjectives than nouns (Figure 2, gray shading).”

Rephrased

R2. As separate models for each participant were trained authors should report the variability across subjects on top of the average model performance.

SEM appears as the shaded area surrounding the lines in our line charts (Figure 3, was Figure 2).

R3. Misphrased: “Isolated noun representations HAD DID not generalize well to list contexts, and did not show generalization across time”

Fixed

DISCUSSION

D1. Replace an for a in “of an adjectives” in p.15

Fixed

D2. Misphrased “which no sustained activations” in p.15

Fixed

Reviewer #2: The paper by Honari-Jahromi and colleagues presents an important and under-researched question in cognitive neuroscience of language – how stable are the neural representations of the semantic properties of adjectives and nouns across both temporal, contextual and combinatorial dimensions. To answer this question they use a MEG data and a combination of decoding techniques. Results they present are interesting and compelling. Overall I am very happy with both the quality of the paper, novelty of the question and the analysis used to answer it, however there are several areas that would require improvement and clarifications.

Major

1. The methods section describing the analysis used (p. 7) requires a lot of clarification. The stages of the decoding analysis are not clear from the text and at present it would be difficult to replicate it. For instance, what is the dimensionality of the data and predictor matrixes? What distance d was used for vector comparison (Cosine? Mahalanobis?)? In each regression (n=300) the predicted single value (the nth dimension of the semantic vector) is derived/estimated from channel x time matrix? Was the data averaged over these 100ms? Was here some dimensionality reduction performed on the sensors? Or was the stimulus value estimated by integration over all 100 time points and channels (multiplied by the learned decoder matrix)? I think for the benefit of the reader an explanatory figure (showing data dimensionality and schematic representation of steps) and formula describing the model as well as references to the exact method are necessary.

Details of this analysis are added to the paper. We train d independent ridge regression models to predict each dimension of a semantic space Y∈RN×dfrom the MEG dataset X∈RN×p where d=300 is dimensions of the Skip-gram vectors,N=20 is the total number of averaged epochs which are reshaped to vector of length p=c×t, for c=208 MEG gradiometer sensors with t=100 time samples. In each regression (n=300) the predicted single value (the nth dimension of the semantic vector) is estimated from the reshaped average epoch vector. The averaging of epochs is discussed in the preprocessing section. While we did not use any dimensionality reduction method, we use a well-known kernel method based on singular value decomposition to speed up training and regularized regression to prevent overfitting. As mentioned in the draft, we used cosine distance for 2-vs-2 tests. We have added a new figure (Figure 2) to help clarify these points.

2. Semantic vectors/embeddings when derived from collocational matrixes or with neural networks to some degree reflect the frequency of the words they encode. Vectors of more frequent words tend to be more similar to each other than vectors of less frequent words. In this dataset, were the adjectives and nouns matched on word form / lemma frequency? Could better decodability of nouns across contexts and times be simply due to their higher frequency and greater vector similarity (to each other), when comped to adjectives?

We controlled for frequency across word types. English frequency was extracted from Balota et al., 2007. Noun freq. mean = 14396 (sd = 8242); Adj freq. Mean = ; 14976 (t= -0.47, p = .640742).

Balota, D. A., Yap, M. J., Hutchison, K. A., Cortese, M. J., Kessler, B., Loftis, B., ... & Treiman, R. (2007). The English lexicon project. Behavior research methods, 39(3), 445-459.

We have added the clarification and citation to the paper.

3. Since decoding happened within participants (not on pooled data) and the accuracies were simply averaged, this means that decoding variability between participants was not “accounted for” and observed effects cannot be generalised outside of this sample – the models trained on one participant cannot be used to predict data in another participant/s – and this needs to be acknowledged explicitly.

Thank you for raising this point. We added this to the methodology section:

It should be noted that every model was trained separately for each participant. Thus, the patterns underlying the decoding accuracies we observed may not be stable across people, and our analysis did not test for such stability. Rather, we tested for the presence of a pattern, and if the patterns generalize across time and condition within a participant’s data. Our methodology then tests if the average decoding accuracy (a function of the participant-specific patterns) shows stable patterns across participants.

Minor

1. Introduction p.3 para.2 - “complex effects of ambiguity” is a bit ambiguous. Do you mean in sentences or in narratives? In complex and naturalistic listening conditions. Please clarify.

The intention was to refer back to the types of cases mentioned in the first paragraph. This has now been clarified.

Revision: “thus did not investigate ambiguous cases such as those just mentioned.”

2. Methods section p.5 “spoken and signed language as regards to the neural correlates ..” did you mean “with regards to”?

Fixed

3. Did any artefact rejection or blink / heart beat removal with ICA take place? I appreciate that when using decoding artefact rejection is not strictly compulsivity as with ERPs but it is good to state this explicitly and give some references.

We did not apply ICA, and updated the methodology to make this explicit

4. The Statistical significance section on p. 8 is somewhat difficult to read, especially the second paragraph. Was ANOVA done first across all time points and then 0.05 cutoff applied (seems that way since F values were summed for cluster mass estimate)? The word ‘label’ is ambiguous, please clarify.

This section was unclear. We’ve reworked it, and hope it is clearer now. We have two completely separate statistical tests:

1. permutation tests to determine if accuracy is above chance

2. anova test to find condition/word-category effects: we did not retrain any models for this. We permuted 2-vs-2 accuracies by randomly assigning word categories or condition labels.

5. p.9 para 2 “… similar patterns will be picked out by regression.. ” not clear, not sure ‘patterns’ is the right word. Do you mean to say regression weights learned in one window will also generate above-chance performance in another window.

Reworded

6. Results p.10 was responce accuracy above 90% for all participants?

Yes.

7. Figure 2 – most of the figure 1 caption text belongs in the text of the results section, not in the figure caption. Also please explain exactly why you think the demand to attend to the background colour improved decodability for adjectives in lists.

Thank you for pointing this out. Much of this text has been moved to the results section.

Naming the background color intuitively feels more effortful and less natural than naming the object-color. We find it plausible that this may increase neural demands in some way that could also increase decodability. It is a speculation of course, but we want to offer it as a possible way to think about this adjective effect. This note has now been added on p. 11.

8. p.11 “we found an interaction effect...” please explain

Clarified to which ANOVA the interaction belongs.

9. Figure 3 and 4 – does the black contouring indicate statistical significance? Again, please consider putting text that describes results out of the caption and into main text.

We have incorporated mention of the black outlines into the main text.

10. It would be useful throughout the manuscript to refer to not simply ‘representations’ but semantic or lexicon-semantics representation, since this is what model was trained to decode (as opposed to say phonological representations).

Changes made to the introduction, and throughout where needed.

Attachment

Submitted filename: PLoS ONE reviews.pdf

Decision Letter 1

Nicola Molinaro

12 Feb 2021

Neural representation of words within phrases: Temporal evolution of color-adjectives and object-nouns during simple composition

PONE-D-20-34763R1

Dear Dr. Fyshe,

We’re pleased to inform you that your manuscript has been judged scientifically suitable for publication and will be formally accepted for publication once it meets all outstanding technical requirements.

Within one week, you’ll receive an e-mail detailing the required amendments. When these have been addressed, you’ll receive a formal acceptance letter and your manuscript will be scheduled for publication.

An invoice for payment will follow shortly after the formal acceptance. To ensure an efficient process, please log into Editorial Manager at http://www.editorialmanager.com/pone/, click the 'Update My Information' link at the top of the page, and double check that your user information is up-to-date. If you have any billing related questions, please contact our Author Billing department directly at authorbilling@plos.org.

If your institution or institutions have a press office, please notify them about your upcoming paper to help maximize its impact. If they’ll be preparing press materials, please inform our press team as soon as possible -- no later than 48 hours after receiving the formal acceptance. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information, please contact onepress@plos.org.

Kind regards,

Nicola Molinaro, Ph.D.

Academic Editor

PLOS ONE

Acceptance letter

Nicola Molinaro

22 Feb 2021

PONE-D-20-34763R1

Neural representation of words within phrases:Temporal evolution of color-adjectives and object-nouns during simple composition

Dear Dr. Fyshe:

I'm pleased to inform you that your manuscript has been deemed suitable for publication in PLOS ONE. Congratulations! Your manuscript is now with our production department.

If your institution or institutions have a press office, please let them know about your upcoming paper now to help maximize its impact. If they'll be preparing press materials, please inform our press team within the next 48 hours. Your manuscript will remain under strict press embargo until 2 pm Eastern Time on the date of publication. For more information please contact onepress@plos.org.

If we can help with anything else, please email us at plosone@plos.org.

Thank you for submitting your work to PLOS ONE and supporting open access.

Kind regards,

PLOS ONE Editorial Office Staff

on behalf of

Dr. Nicola Molinaro

Academic Editor

PLOS ONE

Associated Data

    This section collects any data citations, data availability statements, or supplementary materials included in this article.

    Supplementary Materials

    S1 Appendix

    (DOCX)

    Attachment

    Submitted filename: PLoS ONE reviews.pdf

    Data Availability Statement

    The data underlying this study is available on OSF (https://osf.io/p7gc6/) and the code is available on GitHub (https://github.com/fyshelab/NeuralPhraseComposition).


    Articles from PLoS ONE are provided here courtesy of PLOS

    RESOURCES