Skip to main content
UKPMC Funders Author Manuscripts logoLink to UKPMC Funders Author Manuscripts
. Author manuscript; available in PMC: 2026 Jun 12.
Published in final edited form as: J Cogn Neurosci. 2024 Jul 1;36(8):1760–1769. doi: 10.1162/jocn_a_02182

Spatiotemporal Properties of Common Semantic Categories for Words and Pictures

Yulia Bezsudnova 1,, Andrew J Quinn 1, Syanah C Wynn 1,2, Ole Jensen 1
PMCID: PMC7619145  EMSID: EMS214049  PMID: 38739567

Abstract

◼The timing of semantic processing during object recognition in the brain is a topic of ongoing discussion. One way of addressing this question is by applying multivariate pattern analysis to human electrophysiological responses to object images of different semantic categories. However, although multivariate pattern analysis can reveal whether neuronal activity patterns are distinct for different stimulus categories, concerns remain on whether low-level visual features also contribute to the classification results. To circumvent this issue, we applied a cross-decoding approach to magnetoencephalography data from stimuli from two different modalities: images and their corresponding written words. We employed items from three categories and presented them in a randomized order. We show that if the classifier is trained on words, pictures are classified between 150 and 430 msec after stimulus onset, and when training on pictures, words are classified between 225 and 430 msec. The topographical map, identified using a searchlight approach for cross-modal activation in both directions, showed left lateralization, confirming the involvement of linguistic representations. These results point to semantic activation of pictorial stimuli occurring at ~150 msec, whereas for words, the semantic activation occurs at ~230 msec. ◼

Introduction

Humans are capable of recognizing and inferring the semantic category of a presented object regardless of the modality in which it is presented, whether through visual, auditory, or textual means.

A powerful way of studying how the human brain encodes the semantics of objects is to apply multivariate pattern analysis (MVPA) to human electrophysiological responses to stimuli of different conceptual categories (Groen, Dekker, Knapen, & Silson, 2022; Liuzzi, Aglinskas, & Fairhall, 2020; Cichy, Pantazis, & Oliva, 2014; Carlson, Tovar, Alink, & Kriegeskorte, 2013). The timing of the transition from visual representations of individual objects to more abstract semantic-type concepts is determined by identifying the time points at which the category of the object can be accurately predicted from the multivariate representation in the magnetoencephalography (MEG) or EEG data (Wardle & Baker, 2020; Wang, Kuperberg, & Jensen, 2018; Contini, Wardle, & Carlson, 2017; Kumar, Federmeier, Fei-Fei, & Beck, 2017; Proklova, Kaiser, & Peelen, 2016). However, drawing meaningful conclusions from the above-chance classification using stimuli from a single modality may be limited, as there are often multiple dimensions in which the two conditions differ (Frisby, Halai, Cox, Ralph, & Rogers, 2023; Peelen & Downing, 2023). For example, one can argue that perceptual information beyond semantic content influences classification outcomes. Furthermore, the temporal dynamic of object recognition may also be influenced by the specific experimental task employed in the study, such as a one-back task, category judgment, or detection task. As such, the precise timing of semantic category activation remains an open question (Dirani & Pylkkänen, 2020, 2023; Giari, Leonardelli, Tao, Machado, & Fairhall, 2020; Kaiser, Azzalini, & Peelen, 2016; Miozzo, Pulvermüller, & Hauk, 2015; MacGregor, Pulvermüller, Van Casteren, & Shtyrov, 2012; Simanova, Van Gerven, Oostenveld, & Hagoort, 2010).

Studying brain responses to stimuli from different modalities (e.g., words vs. images) mitigates the criticism related to the influence of perceptual features on classification, as these stimuli do not share common low-level features (Hauk, Davis, Ford, Pulvermüller, & Marslen-Wilson, 2006). In addition, demonstrating the existence of cross-modal generalization (training a classifier on one modality and classifying another) at specific time points will provide insights into the ongoing debate about how and when semantic information is encoded (Dirani & Pylkkänen, 2023; Federmeier, 2022; Hauk, Coutout, Holden, & Chen, 2012; Indefrey, 2011).

In studies involving image recognition, findings indicate that low-level perceptual features are activated within the first 100 msec, followed by the emergence of representations related to semantic categories (e.g., animal) by 150 msec (Miozzo et al., 2015; Cichy et al., 2014; Carlson et al., 2013; Clarke, Taylor, Devereux, Randall, & Tyler, 2013). However, as mentioned earlier, it is debated to which extent perceptual features influence the time course (Proklova, Kaiser, & Peelen, 2019). The timing of semantic activation in the textual modality has been investigated, and it is evident that low-level features of written words do not contain information about the semantic category (Dirani & Pylkkänen, 2023; Federmeier, 2022; Hauk et al., 2012; Indefrey, 2011). Some studies assert that semantic processing occurs at approximately 300– 500 msec after stimulus onset, as indicated by the N400 response (Grainger & Holcomb, 2009; Pylkkänen & Marantz, 2003). Others argue that information becomes available in a graded manner following stimulus onset, and the N400 serves as an indicator of postcomprehension processes associated with the semantic integration of stimuli into the context (Federmeier, 2022; DeLong & Kutas, 2020; Giari et al., 2020; Pulvermüller, Shtyrov, & Hauk, 2009). These arguments are based on studies pointing to components of semantic processing occurring around as early as 150–200 msec (Giari et al., 2020; Amsel, Urbach, & Kutas, 2013; Hauk et al., 2012; Amsel, 2011; del Prado Martín, Hauk, & Pulvermüller, 2006; Thorpe, Fize, & Marlot, 1996). Finding evidence for semantic representation shared across modalities would contribute to an improved understanding of the timeline of semantic activation but would also be aligned with the hub-and-spoke theory of semantic representations, where modality-specific representations interact via an amodal hub (Ralph, Jefferies, Patterson, & Rogers, 2017; Humphreys, Hoffman, Visser, Binney, & Lambon Ralph, 2015; Pobric, Jefferies, & Ralph, 2007). There is limited research done investigating the time of the common cross-modal generalization using M/EEG, and the reported temporal dynamic of cross-modal representation varies (Dirani & Pylkkänen, 2023; Iamshchinina, Karapetian, Kaiser, & Cichy, 2022; Giari et al., 2020; Leonardelli, Fait, & Fairhall, 2019; Simanova et al., 2010).

In one MEG study (Leonardelli et al., 2019), the categorization between famous places (Big Ben) and famous people (Brad Pitt) is studied in the picture and written word modalities. After each stimulus, participants were asked to perform either a shallow categorization task (Place or Person?) or a deeper semantic task (Italian or foreign?). Within the modality (modality-specific), category information robustly appears around 100 msec for pictures and 230 msec for words. The cross-modal (across modality) classification was studied only from words to pictures and revealed three significant clusters separated in time. Specifically, the authors suggest that cross-modal generalization unfolds through a three-stage process and the first shared representations are accessed at 200 msec using words and at 110 msec using pictures. However, brain activity from concepts such as famous places and people used in their paradigm might not generalize to more common objects. Furthermore, any findings of cross-modal representations across modalities could be driven by a shared category-judgment process because of a category-naming task rather than an automatic activation of semantic representations (Hebart, Bankson, Harel, Baker, & Cichy, 2018).

Another recent MEG study (Dirani & Pylkkänen, 2023) examining generalization between written words and pictures, using picture naming and word reading tasks, showed different results. Significant decoding (animal vs. tool) activates surprisingly early around 75 msec for pictures, and 95 msec for words. Cross-generalization occurs simultaneously for both modalities around 150 msec. The different time courses of the shared semantic representation between two studies (Dirani & Pylkkänen, 2023; Leonardelli et al., 2019) could be attributed to variations in participant tasks and experimental paradigms. Importantly, in the latter study (Dirani & Pylkkänen, 2023), a block presentation of categories was utilized to demonstrate cross-modal categorization. This might result in anticipatory effects making the semantic categorization occur earlier. On the basis of these considerations, the precise timing of the activation of the shared representation should be reevaluated in experiments where the category order is randomized while explicit category naming is not required.

In this work, we investigated the time course of semantic activation within and across pictural and textual stimulation presented in a randomized order. We analyzed MEG data using MVPA. To further eliminate any anticipatory biases, such as the possibility of stronger activation of general motor commands for the tools category, we chose to include multiple categories. Finally, we applied a search-light approach to identify the brain areas involved in the semantic representations.

Methods

Participants

Thirty-eight healthy adult participants took part in the study. Five participants had to be excluded because of extensive noise in the data, and two participants were excluded because of low accuracy during the task (accuracy less than 85%). Therefore, the final sample consisted of 31 participants (mean age = 21.77 years, SD = 3.31 years; 21 female participants). The number of participants was selected to match the average number of participants used in previous studies that employed MVPA (Dirani & Pylkkänen, 2023; Singer, Cichy, & Hebart, 2023; Leonardelli et al., 2019). The study was conducted at the Center for Human Brain Health in Birmingham, United Kingdom. All participants were native English speakers with normal or corrected-to-normal vision. The University of Birmingham Ethics Committee approved the study. The participants were provided written informed consent and received either £15 per hour or course credits as compensation for their participation.

Experimental Paradigm

The experimental design used here was an adapted version of the experimental paradigm used in a previous study (Iamshchinina et al., 2022). However, some stimuli were replaced and the task for the participant was modified. The stimulus set was composed of 48 objects. Each stimulus was presented as a picture and a written word. The objects were organized according to three dimensions of categories, each separated into two categorical divisions: size (big or small), movement (mobile or still), and nature (natural or man-made). Each object in the study belonged to each of the dimensions (e.g., a book is small, still, and man-made). The stimulus set was balanced such that each categorical division included one half of the stimulus set (24 objects). Hence, each set of categories (e.g., big, still, natural) included six objects. The selection of categorical divisions in this study was based on previous findings demonstrating reliable neural representations of semantic dimensions across categories (Iamshchinina et al., 2022; Konkle & Caramazza, 2013). Each object was presented in two modalities: written words and pictures. The size of the images was 400 × 400 pixels presented on a gray screen at a visual angle of 6°. In the stimuli set for textual modality (font: Arial, bold, 75), we did not control for the lexical properties of the words, which resulted in a mixture of high-frequency and low-frequency words. In addition, the length of the words was not controlled. Importantly, we did not find any significant differences (two-tailed t test, p > .3) in word frequency or word length obtained via The English Lexicon Project (Balota et al., 2007) across any pair of categorical divisions, namely, size,” “movement, and nature.

Each object (word or picture) was presented for 600 msec. Before the object presentation, the fixation cross was shown for 500–700 msec. The experiment was divided into nine blocks where either words (five blocks) or images (four blocks) were successively presented. The experimental design followed a consistent order, with word blocks always presented first, followed by image blocks, and so on (Figure 1B). After each block, a participant had a break. We included an extra word block based on previous reports that indicated a lower signal-to-noise ratio for classification in the textual modality (Dirani & Pylkkänen, 2023; Simanova et al., 2010). Each block consisted of 240 trials, with 48 of them being question trials, and lasted for around 6 min. Each stimulus was present 4 times in a block. During a probe question, two stimuli were displayed on the screen with one of them corresponding to the previously presented stimuli. Participants were asked to press either the left or right button to identify it. For image blocks, probe questions were presented in word modality and vice versa (Figure 1B). As such, a participant had to identify the correct picture during the written words and vice versa. This design aimed to keep participants engaged and stimulate deep cross-modal perception without explicitly addressing the category distinction between stimuli. Note that in each block the probe question was applied for every object once and alternative choices were unique as well.

Figure 1. Experimental paradigm.

Figure 1

(A) The stimulus set is composed of 48 objects that can be divided according to three dimensions (size: big/small; movement: still/mobile; nature: natural/man-made). (B) In both image and word blocks, participants were presented with stimuli in random order. In the picture block, participants viewed images of the objects, whereas in the word block, they read the corresponding words. The task (question trial) was presented randomly every fifth trial on average. Participants were required to identify whether the picture or word corresponded to a previously seen stimulus and press the appropriate button accordingly. When pictures were presented, the probe question was presented as a word and vice versa.

The stimuli presentation was implemented in MATLAB (The MathWorks) using the Psychophysics Toolbox (Brainard, 1997).

MEG Acquisition

MEG data were recorded using a 306-sensor TRIUX MEGIN Elekta Neuromag system, consisting of 204 orthogonal gradiometers and 102 magnetometers, with online band-pass filtering from 0.1 to 330 Hz and a sampling rate of 1000 Hz. Before collecting data, the positions of three anatomical fiducial points (nasion, left preauricular, and right preauricular points) were recorded using a Polhemus Fastrack electromagnetic digitizer system. Furthermore, we recorded the positions of four head-position indicator coils: two placed on the left and right mastoid bone, and two on the forehead, with a minimum separation of 3 cm between each coil. After completing these initial steps, participants were seated in an upright position with a 60° angled backrest within the MEG gantry. Electrodes were affixed about 2.5 cm from the outer canthus of each eye to capture horizontal EOG signals. To record the VEOG, a pair of electrodes was placed above and below the right eye in line with the pupil. The electrocardiography signals were captured using a set of electrodes positioned on both the left and right collarbones.

Data Preprocessing

The data are analyzed using the open-source toolbox MNE Python v1.4.2 (Gramfort et al., 2013) following the standards defined in the FLUX Pipeline (Ferrante et al., 2022).

First, blinks and muscle artifacts are annotated using EOG channels and magnetometers recordings, respectively. A semi-automatic detection algorithm was utilized to mark sensors with excessive artifacts (on average 4.5 channels per participant). Then, the data were low-pass filtered at 100 Hz to reduce head-position indicator coils artifacts. We did not use signal space separation (SSS) or Maxwell filtering (Taulu & Simola, 2006) because our previous research demonstrated that they negatively impact classification results (in preparation; Bezsudnova & Jensen, 2023). We attribute this decline in performance to the increase of the white noise in the data following SSS filtering. For a more detailed description of how the SSS filter amplifies the white noise in the data, please refer to Taulu, Simola, and Kajola (2005). Signal Space Projection (eight system-provided projectors) and then independent component analysis (ICA) algorithms were applied (Hyvärinen & Oja, 2000; Uusitalo & Ilmoniemi, 1997) to remove external interference and components associated with cardiac artifacts and eyeblinks. The identification process involved analyzing the time courses and topographies of the ICA components. On average, three components corresponding to two cardiac-related and one blink-related artifact were identified and removed for each subject. After ICA, the data were segmented into trials. Trials corresponding to the same stimulus were averaged together to construct super-trials. This step further enhances the signal-to-noise ratio of the data (Ashton et al., 2022; Guggenmos, Sterzer, & Cichy, 2018). We down-sampled the data to 500 Hz. The epochs were time-locked to the onset of the stimuli and were cropped to a time window of 100 msec before the stimuli and 700 msec after the stimuli onset. Each time point was represented as a 306-dimensional vector (channel data). We expanded this vector by including past 25 msec (12 time points) and future 25 msec (12 time points) of data, resulting in a 306 × 25 dimensional feature vector representing each time point. This procedure (termed delayed embedding) resulted in a more information-rich representation of the neural activity associated with a given stimulus by incorporating more time points (Cheng, 2021; Grootswagers, Wardle, & Carlson, 2017; Tyler, Cheung, Devereux, & Clarke, 2013; Chan, Halgren, Marinkovic, & Cash, 2011). Therefore, the feature vector is constructed from super-trials with embedded 50 msec of data. As the last step before applying the classifier to the data, the feature vectors were standardized by removing the mean and scaling to unit variance per sensor. Following this procedure, the data from gradiometers and magnetometers can be combined.

Classification Analysis

MVPA was applied to the preprocessed MEG data to classify the categories of the presented objects (Cichy et al., 2014; Carlson et al., 2013; Haynes & Rees, 2006). The feature vectors associated with different categories (e.g., mobile vs. still) and one modality (e.g., words) are labeled accordingly and used as input for the classifier. We employed a support vector machine (Cortes & Vapnik, 1995) from the Python module Scikit-learn to classify the data over time. The classification procedure relied on a fivefold cross-validation approach, and the performance was quantified using the area under the curve metric, which measures the classifier discriminative ability. This procedure was done separately for each subject.

First, we investigated the classification accuracy for each category dimension (movement,” “size, and nature) averaged across participants for each modality separately. The dimensions where both modalities showed a classification accuracy significantly above the chance level were selected for further analysis. We averaged the classification curves of the chosen categorical dimensions to obtain results that were less contaminated by low-level features (Iamshchinina et al., 2022). The same categorical dimensions were selected for cross-modal analysis. For cross-modal analysis, we also used time-generalized MVPA (King & Dehaene, 2014). We used the classification approach mentioned above with the exception that the classifier was training in one modality at one time point and testing it on another at a different time point. This procedure was done twice, once with the classifier trained on the words data and tested on the pictures data, and once where it was trained on the pictures data and tested on the words.

To explore the spatial distribution of representations across the MEG sensors over time, we employed the searchlight approach on sensor-level data (Leonardelli et al., 2019; Kriegeskorte, Mur, & Bandettini, 2008). In this approach, classification was performed on patches that are defined by all sensors within a 4-cm radius from a specific sensor (e.g., 4 cm from MEG 1423). This typically resulted in 15 sensors (consisting of gradiometers and magnetometers). Therefore, the feature vector had dimensions N × 25, where N is the number of sensors in the patch and 25 are time points from the delayed-embedding procedure. Patches were created for all possible sensor locations (rim sensors had fewer sensors in the patch). For each channel location, the classification accuracy from the relevant patch was averaged over a chosen time interval and this value was plotted on the topographical sensor map.

Statistical Analysis

To find significant time points for the modality-specific classification curves and simultaneous cross-generalization, we used a nonparametric, one-sampled permutation t test (one-tailed) against 50% chance level, controlled for multiple comparisons (Sassenhagen & Draschkow, 2019; Maris & Oostenveld, 2007) implemented in the GLMTools Python package (https://pypi.org/project/glmtools/). We set a cluster-forming threshold of 1.7, corresponding to an alpha threshold of .05. Clusters of t values that exceeded the cluster-forming threshold were formed based on direct adjacency in time (the minimum number of vertices in terms of time points in a cluster was set to 2) and summarized using the sum of t values within the cluster. The largest cluster from each of the N sign-flip permutations were computed to form a null distribution. A cluster in the observed t values was considered significant if its cluster stat (sum of t value with cluster) lay at or above the 95 th percentile of the null distribution. This corresponds to an alpha threshold of p = .05.

The same analysis was done for cross-modal time generalization results. Note that in this case, the clusters are based on direct adjacency in both axes.

Results

Within-modality Decoding

When considering pictures, we find robust decoding for every categorical dimension (size,” “movement, and nature) separately. However, for words, we found robust classification only for dimensions size (big/small) and movement (still/mobile), but not nature (natural/ man-made). In a previous study (Cichy et al., 2014), this category dimension also exhibited the lowest decoding accuracy when examining pictures. Other modalities such as written or spoken words have generally shown lower classification accuracy compared with pictures (Dirani & Pylkkänen, 2023; Iamshchinina et al., 2022), suggesting that the nature-dimension in textual modality would be very difficult to classify. We also speculate that the nature category we applied was too broad. We included any nature objects such as mountain,” “rainbow, and forest rather than, for example, more specific and standard objects from the animal world. All the results presented were therefore derived by averaging the classification results obtained from just two categorical dimensions (size and movement).

The classification accuracies for the picture modality and textual modality are shown in Figure 2A. As expected, a classifier that uses brain activity elicited by words showed lower performance compared with picture stimuli (Ghazaryan et al., 2023). For words, decoding is most pronounced around 240–350 msec after stimuli are presented. For pictures, decoding is most pronounced around 155–510 msec. These results indicate that semantic activation for words occurs approximately 100 msec later than for pictures. Note that the time estimation might be blurred because of the feature vector including ±25 msec of information; however, the relative relationships between modalities are preserved.

Figure 2.

Figure 2

(A) The time course of category decoding averaged over two categorical dimensions: size (big/small), and movement (still/ mobile). The blue line marks the classification curve for word modality; the black line marks the classification curve for picture modality. Significant clusters (p < .05; controlled for multiple comparisons over time) are shown as highlighted areas accordingly. (B) The topographical map of category decoding averaged over 300 ± 10 msec created using searchlight MVPA decoding categories of pictures. (C) Decoding categories of words. The colorcode indicates the classification accuracy (area under the curve) averaged over the time interval 300 ± 10 msec. The classification accuracy is averaged over two category dimensions: size (big/small) and movement (still/mobile).

For the time points (300 ± 10 msec) when the decoding accuracy is strongest, we show topographical maps of the classification accuracy using a searchlight approach. The sensors that contributed strongest to the overall accuracy of the picture category are located over bilateral temporal and parietal areas (Figure 2B). For the decoding word category, the informative sensors are located over the left temporal and left frontal parts of the brain (Figure 2C). This localization is in line with prior results using MEG in object decoding (Cichy et al., 2014) and word reading studies (Hultén et al., 2021; Kumar et al., 2017; Santi, Friederici, Makuuchi, & Grodzinsky, 2015).

Cross-modal Decoding

Time-generalized, cross-decoding results is shown in Figure 3. When training on pictures and testing on words, shared representation occurs between 225 and 430 msec (Figure 3A); when training on words and testing on pictures, shared representation occurs between 150 and 430 msec (Figure 3B). Interestingly, the most pronounced semantic decoding across modalities for both words and pictures becomes significant around the same time as the categories are classified within each modality (Figure 2A).

Figure 3.

Figure 3

Time generalization results for category information, where classifier was (A) trained on pictures and tested on words, and (B) trained on words and tested on pictures. The classification accuracy is averaged over two categorical dimensions: size (big/small) and movement (still/ mobile). Significant clusters are highlighted (one-sided permutation test, p < .05, corrected for multiple comparisons).

In addition, we showed the simultaneous (trained and tested on the same time points) cross-modal classification between textual and pictural modality in Figure 4A. When training on words and testing on pictures, the significant cluster emerged at 280–430 msec. When reversely training on pictures and testing on words, the cluster emerged at 330–430 msec.

Figure 4.

Figure 4

(A) The time course of simultaneous (trained and tested on the same time points) cross-modal category decoding averaged over two categorical dimensions: size (big/small) and movement (still/ mobile). The blue line marks the classification curve when classifiers were trained on pictures and tested on words; the black line marks the classification curve when trained on words and tested on pictures. Significant clusters (p < .05; controlled for multiple comparisons over time) are shown as highlighted areas accordingly. (B) Topographical map of cross-modal categorical decoding averaged over 400 ± 10 msec created using searchlight MVPA: when training on words, testing on pictures modality; (C) when training on pictures, testing on words modality. The colorcode indicates the classification accuracy averaged over the time interval 400 ± 10 msec. The classification accuracy is averaged over two categorical dimensions: size (big/small) and movement (still/mobile).

Next, we examined the topographical maps of the decoding accuracy at 400 ± 10 msec to check where the simultaneous cross-modal classification is the most pronounced (Figure 4B and C). We selected the time interval of 400 ± 10 msec because it corresponds to the peak decoding accuracy in Figure 4A. These maps for simultaneous cross-modal activation in both directions show more pronounced left lateralization (Figure 4B and C) compared with topographies from within-modality decoding (Figure 2B and C). Activation in the inferior parietal cortex is in line with Liu, Wang, Zhou, Ding, and Luo (2017).

In summary, when considering the within- and cross-modal results, this point to semantic activation of pictures and words occurs at ~150 msec and ~230 msec, respectively.

Discussion

In this study, we investigated the temporal dynamics of semantic processing when objects were presented as pictures and text. We assessed the qualitative similarities between categories elicited for each modality using MVPA on MEG data. The classification shows the most pronounced decoding activation for the pictures starting at 155 msec over the bilateral posterior part of the brain and for words after 240 msec over the left temporal and frontal parts of the brain (Figure 2). We show successful cross-modal classification using MEG data (Figure 3). If the classifier is trained on words, pictures are classified between 150 and 430 msec, and when training on pictures, words are classified between 225 and 430 msec. The topographical map for simultaneous cross-modal activation in both directions (from words to pictures and from pictures to words) reveals strong left lateralization (Figure 4B and C).

Semantic activation for pictures starts around 150 msec, consistent with findings reported in a previous study (Iamshchinina et al., 2022) that employed a similar set of stimuli. The temporal dynamics of representations elicited by words demonstrate a later activation around 240 msec, which is in line with Giari and colleagues (2020) and Leonardelli and colleagues (2019) and studies using the N400 paradigm where the semantic response starts to build up at 250 msec (Federmeier, 2022; Giari et al., 2020; Amsel et al., 2013; Hauk et al., 2012). The different timing of picture and word categorization can be explained by the fact that low-level visual features in written words are less indicative of semantic categories compared with pictures (Hauk et al., 2006; Job & Tenconi, 2002). Therefore, words may activate categorical information through different more time-consuming processes from those elicited by pictures (Dirani & Pylkkänen, 2023; Indefrey, 2011; Job & Tenconi, 2002). Note that decoding onsets of 75 msec for pictures and 95 msec for words acquired in the study (Dirani & Pylkkänen, 2023) is almost 70 msec earlier than what we have demonstrated here. This disparity could likely be attributed to the categorical presentation in blocks as in the study by Dirani and colleagues (Dirani & Pylkkänen, 2023), which facilitates faster category extraction.

The time course of cross-modal decoding (Figure 3) with off-diagonal time points structured as a rectangle suggests the shared semantic representation exhibits some sustained features common to both modalities. For word modality, for example, when training on pictures and testing on words (Figure 3A), sustained shared semantic representation emerges in the 225- to 430-msec interval, whereas for pictures, for example, when training on words and testing on pictures (Figure 3B), the shared representation emerges in 150- to 430-msec interval. The delayed semantic activation prompted by words compared with pictures is also evident in the within-modality classification results shown in Figure 2. We attribute this to words having abstract low-level orthographic features that are not informative about category assignment (Hauk et al., 2006). Therefore, words require orthographical processing before their semantics can be accessed (Dirani & Pylkkänen, 2023; Giari et al., 2020; Hauk et al., 2012; Pulvermüller et al., 2009) and the categorical information is activated slower than for pictures.

In summary, both within-modality classification results and cross-modal classification results show the difference in the time course of categorical classification between pictures and words, indicating the semantics extracted through different processes for pictures compared with words (Dirani & Pylkkänen, 2023; Federmeier, 2022; Indefrey, 2011).

As mentioned in the introduction of this article, the perception and representation of the stimuli might be influenced by the experimental paradigm. Tasks relying on 1-back comparisons may not elicit strong semantic processing of the stimuli, possibly explaining why a significant cross-modal classification was not found in a previous study (Iamshchinina et al., 2022). In a study (Giari et al., 2020) where no cross-modal decoding was found, the task forced subjects to focus on the similarity of the objects by rating them according to a broader category. This might not have elicited sufficiently deep conceptual processing of the images. In our experiment, participants had to identify the word that corresponds to the previously seen picture and vice versa; therefore, the association between the two modalities was encouraged. We acknowledge that this procedure serves to promote cross-model activation. The relationship between association or mental imagery and semantic activation remains an open question for future research (Xie, Kaiser, & Cichy, 2020; Shatek, Grootswagers, Robinson, & Carlson, 2019). We attribute comparable timing of modality-specific and cross-modal semantic decoding to linguistic activation when viewing both pictures and words because of the task design. As a next step, it would be interesting to explore paradigms using unique stimuli instead of repeated images of the objects to control for the pairing accumulating between words and corresponding pictures.

Linguistic activation when looking at both pictures and words also explains why the topographical maps from cross-modal decoding for both directions (Figure 4B and C) are located over the left regions of the brain, resembling the topographical map of word categorization. Words elicit abstract shared category representations because of their nonrepresentative, low-level feature, whereas categorization within pictorial modality is contaminated by perceptual features. Furthermore, the location of most informative sensors may include the left anterior temporal lobe (ATL) as shown with fMRI to be involved in semantic processing across various tasks and including cross-modal generalization between written words and corresponding images (Branzi, Humphreys, Hoffman, & Ralph, 2020; Humphreys et al., 2015). These studies support the hub- and-spoke theory (Ralph et al., 2017) in which the ATL supports shared representations. Although our MEG study is inconsistent with the hub-and-spoke theory for semantic representations, our topographical plots do point to an extended network supporting the semantic encoding going beyond the ATL. In the future, it would be interesting to investigate the role of each frequency band in the development of semantic representation (Xie et al., 2020; Bastos et al., 2015).

Conclusion

Our results demonstrate the time course of modality-independent semantic representations isolated from perceptual confounds. Specifically, we examined the time course of activation of semantic representations common for words and pictures. We found that the semantic activation of words occurs at ~230 msec, and for pictures, it occurs at ~150 msec. In the future, we will conduct the same experiment with an optically pumped magnetometer–MEG system to check our hypothesis that the new system can offer significant advantages in experiments designed for MVPA.

Acknowledgments

The computations described in this article were performed using the University of Birminghams BEAR Cloud service, which provides flexible resources for intensive computational work to the universitys research community. See https://www.birmingham.ac.uk/bear for more details. We thank Jonathan Winter for his support of MEG data acquisition. We also thank Dr. Yali Pan for the fruitful discussions.

Funding Information

This work was supported by the Biotechnology and Biological Sciences Research Council (https://dx.doi.org/10.13039/501100000268), grant number: BB/R018723/1; an Engineering and Physical Sciences Research Council award, grant number: EP/ T001046/1; Wellcome Trust Investigator Award in Science, grant number: 207550; a Wellcome Trust Discovery Award, grant number: 227420; and by the NIHR Oxford Health Biomedical Research Centre, grant number: NIHR203316.

Footnotes

Author Contributions

Yulia Bezsudnova: Conceptualization; Data curation; Formal analysis; Investigation; Methodology; Project administration; Software; Visualization; WritingOriginal draft; WritingReview & editing. Andrew Quinn: Formal analysis; Methodology; Software; WritingReview & editing. Syanah Wynn: Methodology; Software. Ole Jensen: Conceptualization; Funding acquisition; Methodology; Project administration; Resources; Supervision; WritingReview & editing.

Diversity in Citation Practices

Retrospective analysis of the citations in every article published in this journal from 2010 to 2021 reveals a persistent pattern of gender imbalance: Although the proportions of authorship teams (categorized by estimated gender identification of first author/last author) publishing in the Journal of Cognitive Neuroscience (JoCN) during this period were M(an)/ M = .407, W(oman)/ M = .32, M/ W = .115, and W/ W = .159, the comparable proportions for the articles that these authorship teams cited were M/ M = .549, W/ M = .257, M/ W = .109, and W/ W = .085 (Postle and Fulvio, JoCN, 34:1, pp. 1–3). Consequently, JoCN encourages all authors to consider gender balance explicitly when selecting which articles to cite and gives them the opportunity to report their articles gender citation balance.

Data Availability Statement

Raw data are available upon request, and code can be accessed on https://github.com/Y-Bezs/cross-modal-project.

References

  1. Amsel BD. Tracking real-time neural activation of conceptual knowledge using single-trial event-related potentials. Neuropsychologia. 2011;49:970–983. doi: 10.1016/j.neuropsychologia.2011.01.003. [DOI] [PubMed] [Google Scholar]
  2. Amsel BD, Urbach TP, Kutas M. Alive and grasping: Stable and rapid semantic access to an object category but not object graspability. Neuroimage. 2013;77:1–13. doi: 10.1016/j.neuroimage.2013.03.058. [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Ashton K, Zinszer BD, Cichy RM, Nelson CA, III, Aslin RN, Bayet L. Time-resolved multivariate pattern analysis of infant EEG data: A practical tutorial. Developmental Cognitive Neuroscience. 2022;54:101094. doi: 10.1016/j.dcn.2022.101094. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Balota DA, Yap MJ, Cortese MJ, Hutchison KA, Kessler B, Loftis B, et al. The English lexicon project. Behavior Research Methods. 2007;39:445–459. doi: 10.3758/bf03193014. [DOI] [PubMed] [Google Scholar]
  5. Bastos AM, Vezoli J, Bosman CA, Schoffelen J-M, Oostenveld R, Dowdall JR, et al. Visual areas exert feedforward and feedback influences through distinct frequency channels. Neuron. 2015;85:390–401. doi: 10.1016/j.neuron.2014.12.018. [DOI] [PubMed] [Google Scholar]
  6. Bezsudnova Y, Jensen O. Optimizing magnetometers arrays and pre-processing pipelines for multivariate pattern analysis. bioRxiv. 2023 doi: 10.1016/j.jneumeth.2024.110279. [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Brainard DH. The psychophysics toolbox. Spatial Vision. 1997;10:433–436. [PubMed] [Google Scholar]
  8. Branzi FM, Humphreys GF, Hoffman P, Ralph MAL. Revealing the neural networks that extract conceptual gestalts from continuously evolving or changing semantic contexts. Neuroimage. 2020;220:116802. doi: 10.1016/j.neuroimage.2020.116802. [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Carlson T, Tovar DA, Alink A, Kriegeskorte N. Representational dynamics of object vision: The first 1000 ms. Journal of Vision. 2013;13:1. doi: 10.1167/13.10.1. [DOI] [PubMed] [Google Scholar]
  10. Chan AM, Halgren E, Marinkovic K, Cash SS. Decoding word and category-specific spatiotemporal representations from MEG and EEG. Neuroimage. 2011;54:3028–3039. doi: 10.1016/j.neuroimage.2010.10.073. [DOI] [PMC free article] [PubMed] [Google Scholar]
  11. Cheng F. Using single-trial representational similarity analysis with EEG to track semantic similarity in emotional word processing. arXiv. 2021 doi: 10.48550/arXiv.2110.03529. [DOI] [Google Scholar]
  12. Cichy RM, Pantazis D, Oliva A. Resolving human object recognition in space and time. Nature Neuroscience. 2014;17:455–462. doi: 10.1038/nn.3635. [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Clarke A, Taylor KI, Devereux B, Randall B, Tyler LK. From perception to conception: How meaningful objects are processed over time. Cerebral Cortex. 2013;23:187–197. doi: 10.1093/cercor/bhs002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Contini EW, Wardle SG, Carlson TA. Decoding the time-course of object recognition in the human brain: From visual features to categorical decisions. Neuropsychologia. 2017;105:165–176. doi: 10.1016/j.neuropsychologia.2017.02.013. [DOI] [PubMed] [Google Scholar]
  15. Cortes C, Vapnik V. Support-vector networks. Machine Learning. 1995;20:273–297. doi: 10.1023/A:1022627411411. [DOI] [Google Scholar]
  16. DeLong KA, Kutas M. Comprehending surprising sentences: Sensitivity of post-N400 positivities to contextual congruity and semantic relatedness. Language, Cognition and Neuroscience. 2020;35:1044–1063. doi: 10.1080/23273798.2019.1708960. [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. del Prado Martín FM, Hauk O, Pulvermüller F. Category specificity in the processing of color-related and form-related words: An ERP study. Neuroimage. 2006;29:29–37. doi: 10.1016/j.neuroimage.2005.07.055. [DOI] [PubMed] [Google Scholar]
  18. Dirani J, Pylkkanen L. Lexical access in naming and reading: Spatiotemporal localization of semantic facilitation and interference using MEG. Neurobiology of Language. 2020;1:185–207. doi: 10.1162/nol_a_00008. [DOI] [Google Scholar]
  19. Dirani J, Pylkkanen L. The time course of cross-modal representations of conceptual categories. Neuroimage. 2023;277:120254. doi: 10.1016/j.neuroimage.2023.120254. [DOI] [PubMed] [Google Scholar]
  20. Federmeier KD. Connecting and considering: Electrophysiology provides insights into comprehension. Psychophysiology. 2022;59:e13940. doi: 10.1111/psyp.13940. [DOI] [PMC free article] [PubMed] [Google Scholar]
  21. Ferrante O, Liu L, Minarik T, Gorska U, Ghafari T, Luo H, et al. FLUX: A pipeline for MEG analysis. Neuroimage. 2022;253:119047. doi: 10.1016/j.neuroimage.2022.119047. [DOI] [PMC free article] [PubMed] [Google Scholar]
  22. Frisby SL, Halai AD, Cox CR, Ralph MAL, Rogers TT. Decoding semantic representations in mind and brain. Trends in Cognitive Sciences. 2023;27:258–281. doi: 10.1016/j.tics.2022.12.006. [DOI] [PubMed] [Google Scholar]
  23. Ghazaryan G, van Vliet M, Saranpää A, Lammi L, Lindh-Knuutila T, Hultén A, et al. Trials and tribulations when attempting to decode semantic representations from MEG responses to written text. Language, Cognition and Neuroscience. 2023:1–12. doi: 10.1080/23273798.2023.2219353. [DOI] [Google Scholar]
  24. Giari G, Leonardelli E, Tao Y, Machado M, Fairhall SL. Spatiotemporal properties of the neural representation of conceptual content for words and pictures—An MEG study. Neuroimage. 2020;219:116913. doi: 10.1016/j.neuroimage.2020.116913. [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Grainger J, Holcomb PJ. Watching the word go by: On the time-course of component processes in visual word recognition. Language and Linguistics Compass. 2009;3:128–156. doi: 10.1111/j.1749-818X.2008.00121.x. [DOI] [PMC free article] [PubMed] [Google Scholar]
  26. Gramfort A, Luessi M, Larson E, Engemann DA, Strohmeier D, Brodbeck C, et al. MEG and EEG data analysis with MNE-Python. Frontiers in Neuroscience. 2013;7:267. doi: 10.3389/fnins.2013.00267. [DOI] [PMC free article] [PubMed] [Google Scholar]
  27. Groen IIA, Dekker TM, Knapen T, Silson EH. Visuospatial coding as ubiquitous scaffolding for human cognition. Trends in Cognitive Sciences. 2022;26:81–96. doi: 10.1016/j.tics.2021.10.011. [DOI] [PubMed] [Google Scholar]
  28. Grootswagers T, Wardle SG, Carlson TA. Decoding dynamic brain patterns from evoked responses: A tutorial on multivariate pattern analysis applied to time series neuroimaging data. Journal of Cognitive Neuroscience. 2017;29:677–697. doi: 10.1162/jocn_a_01068. [DOI] [PubMed] [Google Scholar]
  29. Guggenmos M, Sterzer P, Cichy RM. Multivariate pattern analysis for MEG: A comparison of dissimilarity measures. Neuroimage. 2018;173:434–447. doi: 10.1016/j.neuroimage.2018.02.044. [DOI] [PubMed] [Google Scholar]
  30. Hauk O, Coutout C, Holden A, Chen Y. The time-course of single-word reading: Evidence from fast behavioral and brain responses. Neuroimage. 2012;60:1462–1477. doi: 10.1016/j.neuroimage.2012.01.061. [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Hauk O, Davis MH, Ford M, Pulvermüller F, Marslen-Wilson WD. The time course of visual word recognition as revealed by linear regression analysis of ERP data. Neuroimage. 2006;30:1383–1400. doi: 10.1016/j.neuroimage.2005.11.048. [DOI] [PubMed] [Google Scholar]
  32. Haynes J-D, Rees G. Decoding mental states from brain activity in humans. Nature Reviews Neuroscience. 2006;7:523–534. doi: 10.1038/nrn1931. [DOI] [PubMed] [Google Scholar]
  33. Hebart MN, Bankson BB, Harel A, Baker CI, Cichy RM. The representational dynamics of task and object processing in humans. eLife. 2018;7:e32816. doi: 10.7554/eLife.32816. [DOI] [PMC free article] [PubMed] [Google Scholar]
  34. Hultén A, van Vliet M, Kivisaari S, Lammi L, Lindh-Knuutila T, Faisal A, et al. The neural representation of abstract words may arise through grounding word meaning in language itself. Human Brain Mapping. 2021;42:4973–4984. doi: 10.1002/hbm.25593. [DOI] [PMC free article] [PubMed] [Google Scholar]
  35. Humphreys GF, Hoffman P, Visser M, Binney RJ, Lambon Ralph MA. Establishing task- and modality-dependent dissociations between the semantic and default mode networks. Proceedings of the National Academy of Sciences, USA. 2015;112:7857–7862. doi: 10.1073/pnas.1422760112. [DOI] [PMC free article] [PubMed] [Google Scholar]
  36. Hyvärinen A, Oja E. Independent component analysis: Algorithms and applications. Neural Networks. 2000;13:411–430. doi: 10.1016/s0893-6080(00)00026-5. [DOI] [PubMed] [Google Scholar]
  37. Iamshchinina P, Karapetian A, Kaiser D, Cichy RM. Resolving the time course of visual and auditory object categorization. Journal of Neurophysiology. 2022;127:1622–1628. doi: 10.1152/jn.00515.2021. [DOI] [PMC free article] [PubMed] [Google Scholar]
  38. Indefrey P. The spatial and temporal signatures of word production components: A critical update. Frontiers in Psychology. 2011;2:255. doi: 10.3389/fpsyg.2011.00255. [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Job R, Tenconi E. Naming pictures at no cost: Asymmetries in picture and word conditional naming. Psychonomic Bulletin & Review. 2002;9:790–794. doi: 10.3758/bf03196336. [DOI] [PubMed] [Google Scholar]
  40. Kaiser D, Azzalini DC, Peelen MV. Shape-independent object category responses revealed by MEG and fMRI decoding. Journal of Neurophysiology. 2016;115:2246–2250. doi: 10.1152/jn.01074.2015. [DOI] [PMC free article] [PubMed] [Google Scholar]
  41. King J-R, Dehaene S. Characterizing the dynamics of mental representations: The temporal generalization method. Trends in Cognitive Sciences. 2014;18:203–210. doi: 10.1016/j.tics.2014.01.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
  42. Konkle T, Caramazza A. Tripartite organization of the ventral stream by animacy and object size. Journal of Neuroscience. 2013;33:10235–10242. doi: 10.1523/JNEUROSCI.0983-13.2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
  43. Kriegeskorte N, Mur M, Bandettini PA. Representational similarity analysis—Connecting the branches of systems neuroscience. Frontiers in Systems Neuroscience. 2008;2:4. doi: 10.3389/neuro.06.004.2008. [DOI] [PMC free article] [PubMed] [Google Scholar]
  44. Kumar M, Federmeier KD, Fei-Fei L, Beck DM. Evidence for similar patterns of neural activity elicited by picture- and word-based representations of natural scenes. Neuroimage. 2017;155:422–436. doi: 10.1016/j.neuroimage.2017.03.037. [DOI] [PubMed] [Google Scholar]
  45. Leonardelli E, Fait E, Fairhall SL. Temporal dynamics of access to amodal representations of category-level conceptual information. Scientific Reports. 2019;9:239. doi: 10.1038/s41598-018-37429-2. [DOI] [PMC free article] [PubMed] [Google Scholar]
  46. Liu L, Wang F, Zhou K, Ding N, Luo H. Perceptual integration rapidly activates dorsal visual pathway to guide local processing in early visual areas. PLoS Biology. 2017;15:e2003646. doi: 10.1371/journal.pbio.2003646. [DOI] [PMC free article] [PubMed] [Google Scholar]
  47. Liuzzi AG, Aglinskas A, Fairhall SL. General and feature-based semantic representations in the semantic network. Scientific Reports. 2020;10:8931. doi: 10.1038/s41598-020-65906-0. [DOI] [PMC free article] [PubMed] [Google Scholar]
  48. MacGregor LJ, Pulvermüller F, Van Casteren M, Shtyrov Y. Ultra-rapid access to words in the brain. Nature Communications. 2012;3:711. doi: 10.1038/ncomms1715. [DOI] [PMC free article] [PubMed] [Google Scholar]
  49. Maris E, Oostenveld R. Nonparametric statistical testing of EEG- and MEG-data. Journal of Neuroscience Methods. 2007;164:177–190. doi: 10.1016/j.jneumeth.2007.03.024. [DOI] [PubMed] [Google Scholar]
  50. Miozzo M, Pulvermüller F, Hauk O. Early parallel activation of semantics and phonology in picture naming: Evidence from a multiple linear regression MEG study. Cerebral Cortex. 2015;25:3343–3355. doi: 10.1093/cercor/bhu137. [DOI] [PMC free article] [PubMed] [Google Scholar]
  51. Peelen MV, Downing PE. Testing cognitive theories with multivariate pattern analysis of neuroimaging data. Nature Human Behaviour. 2023;7:1430–1441. doi: 10.1038/s41562-023-01680-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
  52. Pobric G, Jefferies E, Ralph MAL. Anterior temporal lobes mediate semantic representation: Mimicking semantic dementia by using rTMS in normal participants. Proceedings of the National Academy of Sciences, USA. 2007;104:20137–20141. doi: 10.1073/pnas.0707383104. [DOI] [PMC free article] [PubMed] [Google Scholar]
  53. Proklova D, Kaiser D, Peelen MV. Disentangling representations of object shape and object category in human visual cortex: The animate–inanimate distinction. Journal of Cognitive Neuroscience. 2016;28:680–692. doi: 10.1162/jocn_a_00924. [DOI] [PubMed] [Google Scholar]
  54. Proklova D, Kaiser D, Peelen MV. MEG sensor patterns reflect perceptual but not categorical similarity of animate and inanimate objects. Neuroimage. 2019;193:167–177. doi: 10.1016/j.neuroimage.2019.03.028. [DOI] [PubMed] [Google Scholar]
  55. Pulvermüller F, Shtyrov Y, Hauk O. Understanding in an instant: Neurophysiological evidence for mechanistic language circuits in the brain. Brain and Language. 2009;110:81–94. doi: 10.1016/j.bandl.2008.12.001. [DOI] [PMC free article] [PubMed] [Google Scholar]
  56. Pylkkänen L, Marantz A. Tracking the time course of word recognition with MEG. Trends in Cognitive Sciences. 2003;7:187–189. doi: 10.1016/s1364-6613(03)00092-5. [DOI] [PubMed] [Google Scholar]
  57. Ralph MAL, Jefferies E, Patterson K, Rogers TT. The neural and computational bases of semantic cognition. Nature Reviews Neuroscience. 2017;18:42–55. doi: 10.1038/nrn.2016.150. [DOI] [PubMed] [Google Scholar]
  58. Santi A, Friederici AD, Makuuchi M, Grodzinsky Y. An fMRI study dissociating distance measures computed by Broca’s area in movement processing: Clause boundary vs. identity. Frontiers in Psychology. 2015;6:654. doi: 10.3389/fpsyg.2015.00654. [DOI] [PMC free article] [PubMed] [Google Scholar]
  59. Sassenhagen J, Draschkow D. Cluster-based permutation tests of MEG/EEG data do not establish significance of effect latency or location. Psychophysiology. 2019;56:e13335. doi: 10.1111/psyp.13335. [DOI] [PubMed] [Google Scholar]
  60. Shatek SM, Grootswagers T, Robinson AK, Carlson TA. Decoding images in the mind’s eye: The temporal dynamics of visual imagery. Vision. 2019;3:53. doi: 10.3390/vision3040053. [DOI] [PMC free article] [PubMed] [Google Scholar]
  61. Simanova I, van Gerven M, Oostenveld R, Hagoort P. Identifying object categories from event-related EEG: Toward decoding of conceptual representations. PLoS One. 2010;5:e14465. doi: 10.1371/journal.pone.0014465. [DOI] [PMC free article] [PubMed] [Google Scholar]
  62. Singer JJD, Cichy RM, Hebart MN. The spatiotemporal neural dynamics of object recognition for natural images and line drawings. Journal of Neuroscience. 2023;43:484–500. doi: 10.1523/JNEUROSCI.1546-22.2022. [DOI] [PMC free article] [PubMed] [Google Scholar]
  63. Taulu S, Simola J. Spatiotemporal signal space separation method for rejecting nearby interference in MEG measurements. Physics in Medicine & Biology. 2006;51:1759–1768. doi: 10.1088/0031-9155/51/7/008. [DOI] [PubMed] [Google Scholar]
  64. Taulu S, Simola J, Kajola M. Applications of the signal space separation method. IEEE Transactions on Signal Processing. 2005;53:3359–3372. doi: 10.1109/TSP.2005.853302. [DOI] [Google Scholar]
  65. Thorpe S, Fize D, Marlot C. Speed of processing in the human visual system. Nature. 1996;381:520–522. doi: 10.1038/381520a0. [DOI] [PubMed] [Google Scholar]
  66. Tyler LK, Cheung TPL, Devereux BJ, Clarke A. Syntactic computations in the language network: Characterizing dynamic network properties using representational similarity analysis. Frontiers in Psychology. 2013;4:271. doi: 10.3389/fpsyg.2013.00271. [DOI] [PMC free article] [PubMed] [Google Scholar]
  67. Uusitalo MA, Ilmoniemi RJ. Signal-space projection method for separating MEG or EEG into components. Medical and Biological Engineering and Computing. 1997;35:135–140. doi: 10.1007/BF02534144. [DOI] [PubMed] [Google Scholar]
  68. Wang L, Kuperberg G, Jensen O. Specific lexico-semantic predictions are associated with unique spatial and temporal patterns of neural activity. eLife. 2018;7:e39061. doi: 10.7554/eLife.39061. [DOI] [PMC free article] [PubMed] [Google Scholar]
  69. Wardle SG, Baker C. Recent advances in understanding object recognition in the human brain: Deep neural networks, temporal dynamics, and context. F1000Research. 2020;9 doi: 10.12688/f1000research.22296.1. F1000. [DOI] [PMC free article] [PubMed] [Google Scholar]
  70. Xie S, Kaiser D, Cichy RM. Visual imagery and perception share neural representations in the alpha frequency band. Current Biology. 2020;30:2621–2627. doi: 10.1016/j.cub.2020.07.023. [DOI] [PMC free article] [PubMed] [Google Scholar]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

Raw data are available upon request, and code can be accessed on https://github.com/Y-Bezs/cross-modal-project.

RESOURCES