Abstract
Neural responses in visual cortex to a central stimulus are modulated by its surrounding context. Although this contextual modulation is well established for simple stimuli, how it manifests during dynamic naturalistic scene perception remains unclear. Using fMRI, we examined how motion congruence and categorical similarity between central and surrounding scenes shape contextual modulation across the visual hierarchy. Central and surrounding scenes systematically varied in categorical similarity (identical exemplar, different exemplar of the same basic-level category, different basic-level category, or different superordinate category) and in motion direction (same or opposite direction). Neural responses were measured in primary visual cortex (V1), motion-selective cortex (hMT+), and scene-selective occipital and parahippocampal place areas (OPA and PPA). hMT+ showed robust motion-dependent contextual modulation, with stronger suppression for same-direction motion. V1 responses were sensitive to low-level similarity between center and surround, with reduced responses when center and surround were physically identical and moved in the same direction. In contrast, scene-selective regions (OPA and PPA) did not show univariate modulation by motion congruence or categorical similarity, but multivoxel activity patterns reflected the categorical relationship between center and surround. Multivariate analyses further showed that decoding accuracy in OPA and PPA increased as categorical similarity decreased. Together, these results demonstrate that contextual modulation in natural vision differs across visual regions, reflecting sensitivity to motion relationships in hMT+, low-level similarity in V1, and higher-level categorical structure in scene-selective cortex. These findings highlight how contextual influences operate at multiple levels of visual representation to support the processing of dynamic natural scenes.
Keywords: contextual modulation, fMRI, motion perception, multivariate pattern analysis, scene perception, surround suppression
Introduction
Sensitivity to visual stimuli is strongly influenced by the surrounding context. A classic example of such contextual influences is surround modulation, in which the response of a neuron to a stimulus presented in the receptive field (RF) center is affected by stimuli presented in the surround. This modulation can manifest as suppression (surround suppression) or facilitation (surround facilitation) [1–11]. These effects have traditionally been attributed to antagonistic center–surround interactions within the visual cortex, particularly in motion-selective regions such as the middle temporal area (MT) [e.g., MT hypothesis; see 12–18].
Center-surround interactions have been observed across multiple stages of the visual hierarchy. Surround modulation has been demonstrated in retinotopic areas of early visual cortex (V1-V4) as well as in motion-selective regions such as MT and MST, suggesting that it reflects a canonical computation within the visual system [19–23]. Importantly, the strength and form of these interactions depend on stimulus properties. Both neural [6, 15, 16, 18, 24] and behavioral studies [13–16, 25] show that low-level features such as size, contrast, and motion direction systematically influence the surround modulation. Typically, feature similarity between center and surround increases suppression, whereas dissimilarity reduces it and can even reverse the effect into facilitation. For example, neurons responding to a central grating are strongly suppressed by iso-oriented surrounds but can be facilitated when the surround is orthogonal [5, 26, 27]. A comparable pattern is observed for motion, where suppression is typically stronger when the center and surround drift in the same direction than when they move in opposite directions [5, 18, 23, 28–30].
Contextual influences extend beyond classical surround modulation paradigms. Functional MRI studies show that responses in early visual cortex are modulated by context during contour integration, with activity increasing when aligned contours must be detected within cluttered backgrounds [31, 32]. Beyond early visual cortex, similar effects have also been observed in higher-level object-selective regions. For example, activity in the lateral occipital complex (LOC) increases when spatially distributed visual elements are grouped into coherent contours or shapes [32, 33]. Together, these findings indicate that contextual modulation operates across multiple stages of visual processing and supports the integration of increasingly complex visual structure.
Importantly, the nature of contextual modulation varies across different regions of the visual system. In early visual areas, such effects are often linked to interactions within and beyond the classical receptive field [19, 20]. In contrast, contextual influences in higher-level visual regions are frequently associated with relationships between objects and scene elements within the broader visual environment [34–36]. These differences likely reflect the larger receptive fields in higher-level areas, which integrate information across broader regions of the visual field [37–39]. However, it remains an open question how these contextual interactions generalize to complex, dynamic scenes.
Our recent behavioral work using dynamic natural scenes provides evidence that such contextual interactions can depend on categorical relationships between center and surround scenes. Specifically, we found that surround suppression increased as the categorical similarity between the center and surround scenes increased [40]. These results are consistent with findings from scene categorization studies showing that categorically congruent contexts can facilitate recognition, whereas incongruent contexts impair performance [41, 42]. Consistent with prior work using simple stimuli [25, 30, 43], we also observed reduced suppression when the center and surround scenes moved in opposite directions. Together, these findings suggest that contextual modulation in natural vision may depend on both low-level visual features (e.g., motion direction) and higher-level relationships between scenes (e.g., categorical similarity).
While our behavioral findings demonstrate suppressive contextual interactions in dynamic natural scenes, neurophysiological studies using natural stimuli have reported both suppressive and facilitatory interactions under different experimental conditions. For example, in the macaque primary visual cortex (V1), suppression was stronger when the center and surround contained homogeneous natural images [44], whereas in the cat visual cortex, surrounds that are coherent with the central scene elicited facilitation [45]. Consistent with this, a human fMRI study shows that V1 responses to natural image fragments are enhanced when the surrounding context is coherent with the dominant scene [46]. Together, these findings indicate that contextual modulation is sensitive to the statistical properties, spatial organization, and coherence of natural stimuli. However, while prior work has demonstrated contextual modulation with static natural stimuli in V1, it remains unclear whether similar contextual influences are observed across different regions of the human visual system during the perception of dynamic natural scenes, and whether they are shaped by higher-level relationships such as categorical similarity between the center and surround scene.
To address this gap, we used functional magnetic resonance imaging (fMRI) to investigate how contextual relationships between central and surrounding scenes influence neural responses to dynamic natural scenes. We measured responses across multiple stages of visual processing, including primary visual cortex (V1), motion-selective cortex (hMT+), and the scene-selective occipital place area (OPA) and parahippocampal place area (PPA). This allowed us to directly compare how contextual modulation varies across regions associated with different aspects of visual scene processing, from low-level feature processing to higher-level scene representation.
Participants viewed central scenes presented together with surrounding scenes that varied in categorical similarity across four levels: identical exemplar, different exemplar from the same basic-level category, different basic-level categories within the same superordinate category, and different superordinate categories. We also manipulated motion congruence, comparing conditions in which the center and surround moved in the same versus opposite directions, to determine whether motion-dependent contextual effects observed with simple stimuli generalize to naturalistic scenes. By independently manipulating categorical similarity and motion congruence, our design systematically tests how contextual relationships at multiple levels of visual representation shape neural responses to dynamic natural scenes.
Methods
Participants
Twenty healthy volunteers (11 female, mean age = 26.7 years) participated in the study. One participant was excluded due to excessive head motion during scanning (see Preprocessing for motion-quality criteria), leaving a final sample of nineteen participants. The sample size (N = 19) is comparable to that used in previous studies investigating low-level surround modulation [14, 18, 27]. All reported normal or corrected-to-normal vision. Written informed consent was obtained before participation, and participants received monetary compensation. The study was approved by the Ethics Committee of Justus Liebig University Giessen. All experimental protocols were in accordance with the Declaration of Helsinki.
Stimuli
Stimuli were panoramic videos created by moving static scene images behind a circular occluder (908 × 699 pixels) at a speed of 6°/s. The central image was presented through a 1.9°-diameter aperture, surrounded by an annulus ranging from 2.5° to 10.4°. To separate the center and surround, the area between them (1.9°-2.5°) remained un-stimulated. A cosine envelope was applied at the aperture boundaries to minimize sharp transitions at the edges.
Scene images were selected from two superordinate categories (indoor: restaurants, museums; outdoor: parks, residential areas), each featuring two basic-level categories, with two exemplars per basic-level category, resulting in a total of eight scene images. The images shown in Figure 1 illustrate the stimulus set. The center scenes were always presented with a surrounding scene, except in center-only trials where the central stimulus appeared alone (see below). The resulting center–surround stimuli varied in their categorical relationship between center and surround, with four conditions: identical exemplar condition, where the center and surround images were identical; different exemplar condition, where the center and surround images belonged to the same basic-level category (e.g., museums) but were not identical; different basic-level category condition, where the center and surround images belonged to the same superordinate category (e.g., indoor) but were from different basic-level categories; and different superordinate category condition, where the center and surround images belonged to different superordinate categories (see Figure 1). Motion congruency was manipulated with two conditions: in half of the trials, the center and surround scenes moved in the same direction, and in the other half, they moved in opposite directions.
Figure 1. Stimuli, paradigm, and region of interest (ROI) definition.
A) The natural scene images used to create the stimuli were drawn from two superordinate categories (indoor and outdoor) and two basic-level categories within each (restaurants, museums, parks, and residential areas), with two exemplars per basic-level category. B) Example stimuli for each category condition, shown from left to right: identical exemplar, different exemplar, different basic-level category, different superordinate category. C) Experimental design. Each trial began with a fixation point (500 ms), which briefly turned red to signal the upcoming stimulus, followed by presentation of a scene video (360 ms). Intertrial intervals varied randomly between 3.7 and 5.6 s. On 10% of trials, a high-contrast Gabor patch appeared at a random location within either the center or surround region, and participants reported its detection with a keypress. D) Functional localization of ROIs. ROIs were defined individually using independent functional localizers. hMT+ was identified with the dynamic > static contrast (puncorrected < 0.001), and V1 with the horizontal > vertical contrast (puncorrected < 0.001). Sub-ROIs in hMT+ and V1 were further defined using a center–surround localizer (center > surround ; puncorrected < 0.0001). Scene-selective regions (OPA, PPA) were identified using a scene > object contrast (p < 0.05). Because no reliable center–surround effects were observed, full ROIs were used for these regions. The figure shows a representative cortical surface from a single subject’s right hemisphere, illustrating ROI locations.
Experimental Paradigm
We used a mixed event-related design in which center-only and center+surround trials were presented within the same run (Figure 1). In total, nine trial types (four category conditions × two motion directions, plus the center-only condition) were presented in randomized order. Center-only trials served as a baseline for quantifying contextual modulation by the surround.
Each trial began with a fixation point (500 ms), followed by a scene video (360 ms) on a mid-gray background. Participants were instructed to maintain fixation at the center of the display. Intertrial intervals varied randomly between 3.7 and 5.6 seconds, sampled from a uniform distribution. To ensure that participants attended to both center and surround, on 10% of trials a small Gabor patch (spatial frequency = 1 cycles/°, diameter = 0.7°, contrast = 98%) appeared at a random location within either the center or surround region. Participants reported its detection with a keypress, with a mean accuracy of 92.2% (SD = 7.6%). Each condition was presented eight times per run, and participants completed five runs. Each run lasted approximately 10 minutes, and short breaks were provided between runs.
Data Acquisition
MRI data were acquired on a 3T Siemens Magnetom Prisma scanner (Siemens Healthineers, Erlangen, Germany) equipped with a 64-channel head coil at the Bender Institute of Neuroimaging (BION), Justus Liebig University Giessen. High-resolution anatomical images were collected using a T1-weighted 3D sagittal MP-RAGE sequence (voxel size = 1 mm3 isotropic, 176 slices). Functional images were acquired with a T2*-weighted EPI sequence (TR = 1850 ms, TE = 30 ms, voxel size = 2.2 mm3, 58 slices). Visual stimuli were presented on a MR-compatible LED monitor (1920 × 1080 pixels, 120 Hz) positioned at the rear of the bore and viewed via a mirror mounted on the head coil at a distance of 140 cm. Stimuli were generated and presented using MATLAB (MathWorks, Natick, MA) and the Psychophysics Toolbox [47]. Participant responses were recorded using an MR-compatible fiber-optic response box. Each session began with an anatomical scan, followed by four localizer runs and five experimental runs, for a total duration of approximately 90 minutes.
Localizer Runs
hMT+
The hMT+ area of each participant was defined using established methods [48]. Stimuli consisted of 200 randomly positioned white dots presented within a 12° aperture on a black background. In dynamic blocks, the dots moved along one of four trajectories: radial (expansion–contraction), angular (clockwise–counterclockwise rotation), horizontal (left–right), or vertical (up–down). Motion direction changed every 1.85 s to prevent adaptation. In static blocks, the dots were identical but remained fixed in position. Each block lasted 14.8 s, and dynamic and static blocks alternated eight times per run. Participants maintained central fixation and performed a color-change detection task in which they reported changes in the color of the fixation point.
V1
The V1 area of each participant was defined using a flickering checkerboard-patterned wedge paradigm, similar to established retinotopic mapping methods [49, 50]. Because we were specifically interested in the V1/V2 boundary, we used alternating horizontal and vertical wedges instead of rotating or expanding stimuli [51, 52]. The run consisted of 14.8-s blocks of horizontal and vertical wedges presented in alternation, repeated ten times in the run. As in the hMT+ localizer, participants maintained fixation and performed a color-change detection task.
OPA and PPA
Scene-selective areas were defined using 3-s movie clips from four categories: faces, scenes, objects, and scrambled objects [53, 54]. Each category included 60 clips. Face clips featured seven children filmed against a black background in close-up, showing only their faces as they danced or interacted with toys or with adults who remained out of frame. Scene clips were recorded in 15 different locations, primarily pastoral settings, and filmed from a slowly moving car. Object clips depicted 15 distinct inanimate items (e.g., mobiles, wind-up toys, toy planes, tractors, and rolling balls). Scrambled-object clips were generated by dividing each object movie into a 15 × 15 grid and spatially shuffling the resulting segments within each frame. Participants completed eight blocks per category, each consisting of six videos randomly sampled from the full set of 60 exemplars in that category. Although the same actor, scene, or object could appear more than once, the large stimulus pool made such repetitions unlikely. As in the previous localizer runs, participants maintained fixation and reported changes in the color of the fixation point.
hMT+ and V1 Sub-ROI Localizer Run
Using an independent localizer run, we identified the voxels corresponding to the spatial location and size of the center stimuli as sub-ROIs within hMT+ and V1. This approach is commonly used to localize responses to low-level grating stimuli [14, 15, 18], but here we adapted it to ensure that the same spatially defined sub-ROIs could be applied in the analysis of naturalistic stimuli. In the localizer, participants viewed drifting high-contrast (98%) center and surround (i.e., annulus) gratings, matched in size and location to those used in the functional runs. The run consist of 14.8-s center blocks alternating with 14.8-s surround blocks, each separated by a 14.8-s blank period, repeated six times. As in the other localizers, participants maintained fixation and reported changes in the color of the fixation point.
Data Analysis
Preprocessing
Localizer Runs
Localizer data were preprocessed and analyzed using the FM-RIB Software Library (FSL) (www.fmrib.ox.ac.uk/fsl) and Freesurfer [55–57]. High-resolution anatomical images were skull-stripped with BET. Preprocessing steps for functional images included motion correction with MCFLIRT, high-pass temporal filtering (100s), and BET brain extraction. Spatial smoothing (4 mm FWHM) was applied only for the scene localizer used to define scene-selective region of interests (ROIs); no smoothing was applied for V1, hMT+, or sub-ROI localizers. Each participant’s functional images were aligned to their own high-resolution anatomical image and registered to the standard Montreal Neurological Institute (MNI) 2-mm brain using FLIRT. The 3D cortical surface was constructed from anatomical images for each participant using FreeSurfer’s recon-all command for visualizing statistical maps, anatomical delineation, and identifying ROIs.
Experimental Runs
Experimental data were preprocessed and analyzed in SPM12 (www.fil.ion.ucl.ac.uk/spm/). Preprocessing steps included geometric distortion correction with the SPM FieldMap toolbox, motion correction and coregistration of functional volumes to each participant’s T1-weighted structural image. Structural images were segmented and normalized to MNI 2-mm standard space, and the same transformation was applied to the functional images. For ROI-based univariate analyses, the unsmoothed functional data were used to preserve spatial specificity of voxel responses. For multivariate pattern analyses (MVPA), functional images were spatially smoothed with a 6-mm FWHM Gaussian kernel to improve signal-to-noise ratio while minimizing spatial blurring [58, 59].
To assess data quality, head motion was quantified for each volume using framewise displacement (FD). Volumes with FD greater than 0.5 mm were flagged as motion outliers [60]. Runs with more than 30% flagged volumes were discarded [61]. Participants with two or more runs discarded were excluded [62]. One participant met this exclusion criterion and was removed from all analyses. Importantly, none of the remaining participants had any runs excluded; all results reported below are based on the remaining N = 19 participants.
ROI Construction
For all ROI constructions, a general linear model (GLM) was applied using FSL’s FMRI Expert Analysis Tool (FEAT). The predicted fMRI response in each trial was computed assuming a double-gamma hemodynamic response function (HRF). Nuisance regressors for linear motion (derived from MCFLIRT) were also included in the model. For removing temporal autocorrelations, FILM prewhitening was applied [57]. Importantly, ROI definitions were based entirely on independent localizer runs and were therefore statistically independent from the experimental data used in subsequent analyses.
hMT+
The hMT+ ROI was defined individually for each participant using the motion localizer data. Statistical maps of the dynamic > static contrast were computed using FSL and thresholded at p < 0.001 (uncorrected) for each participant. The resulting activation maps were registered to each participant’s anatomical space and visualized on the cortical surface using FreeSurfer. Guided by the MT anatomical label from FreeSurfer, voxels showing stronger responses to dynamic than static dots near the ascending limb of the inferior temporal sulcus, consistent with the anatomical location of hMT+, were selected to define the hMT+ ROI and used as a mask for subsequent sub-ROI localization.
V1
The V1 ROI was defined individually for each participant using the retinotopic localizer data. Statistical maps for the horizontal > vertical meridian contrast were computed using FSL and thresholded at p < 0.001 (uncorrected) for each participant. Activation maps were registered to each participant’s anatomical space and visualized on the cortical surface using FreeSurfer. Guided by the V1 anatomical label from FreeSurfer, voxels located along the calcarine sulcus and consistent with the expected anatomical location of V1 were selected to define the ROI, which was used for subsequent analyses.
hMT+ and V1 sub-ROIs
To identify voxels corresponding specifically to the spatial location of the center stimulus, we analyzed data from the independent center–surround localizer run. Within the previously defined V1 and hMT+ ROIs, voxels showing significantly stronger responses to the center than to the surround grating (center > surround) were identified using a voxelwise threshold of p < 0.0001 (uncorrected). The resulting voxels were defined as V1 and hMT+ sub-ROIs and were subsequently used as masks for the analysis of the experimental runs.
OPA and PPA
For the scene-selective ROIs, we identified the voxels with the greatest scene versus face and object contrast using FSL. Clusters were identified using a voxelwise threshold of Z > 2.3 and a cluster-corrected significance threshold of p < 0.05, based on Gaussian random field theory [57]. The occipital place area (OPA) was localized as the scene-selective cluster on the lateral surface near the transverse occipital sulcus, and the parahippocampal place area (PPA) as the scene-selective cluster on the ventral visual cortex. The identified regions were used as OPA and PPA masks. Because the center vs. surround contrast did not yield reliable voxel-level responses in OPA and PPA, we used the full ROIs rather than sub-ROIs for the analysis of the experimental runs.
Analysis of Experimental Runs
Univariate Analysis
The experimental runs were analyzed in SPM12 using a general linear model (GLM). For each run, separate regressors were specified for each of the nine conditions, and six motion parameters from realignment were included as nuisance regressors. Predicted BOLD responses were modeled by convolving event onsets with a double-gamma hemodynamic response function (HRF), and statistical parametric maps (SPMs) were generated to assess the main effect of each condition across runs. To quantify fMRI responses, percent signal change (PSC) values were calculated for each condition and used as the primary measure of fMRI response in all regions of interest.
Percent signal change (PSC) was estimated from condition-specific beta values extracted within each ROI. For each regressor, beta estimates were scaled by the corresponding regressor amplitude and normalized by the session-specific ROI baseline, with resulting PSC estimates averaged across runs within condition.
To determine whether the center-only (i.e., baseline) stimulus evoked reliable activity, PSC values in the center-only condition were first compared against zero using one-sample t -tests. In ROIs that either showed a reliable response to the center-only condition or were defined using center-selective sub-ROI selection, contextual modulation was quantified using the suppression index (SI) to characterize the direction of modulation, defined as
| (1) |
where B. denotes percent signal change (PSC) for the given condition. Negative SI values indicate facilitation, positive values indicate suppression, and an SI of 0 indicates no modulation. The use of SI allows us to determine the direction of contextual modulation, distinguishing between suppressive and facilitatory effects of the surround relative to the center-only response.
Contextual modulation was assessed by testing SI values against zero using one-sample, two-tailed Student’s t -tests with FDR correction for multiple comparisons. We then performed a two-way Analysis of Variance (ANOVA) on SI values with categorical similarity (four levels: identical exemplar, different exemplar, different basic-level category, different superordinate category) and motion-direction congruence (two levels: same-direction, opposite-direction) as factors for each sub-ROI using SPSS Version 25 (IBM Corp., Armonk, NY). Finally, post hoc paired-sample t -tests were used to examine differences between conditions following significant ANOVA effects.
In ROIs that either did not show a reliable response to the center-only condition or were defined without center-selective sub-ROI selection, analyses focused on direct comparisons among the center+surround conditions rather than subtraction-based metrics relative to the center-only baseline. In these regions, contextual effects were assessed using a two-way ANOVA on PSC values rather than SI values.
Multivariate Analysis
To investigate distributed patterns of activity, we performed ROI-based decoding using CoSMoMVPA [63] on condition-wise GLM beta estimates. For each participant and ROI (hMT+, V1, OPA, and PPA), beta images from each run and condition were assembled into voxel pattern vectors. Classification was implemented using linear discriminant analysis (LDA). Datasets were z-scored within training folds and class-balanced within chunks (runs). Cross-validation used a leave-one-run-out partitioning scheme, and mean accuracy across folds was taken as the decoding score.
To assess whether the presence of a surround stimulus alters distributed activity patterns, we first performed two-way decoding between the center-only and center+surround conditions. Two-way decoding quantifies the separability of multivoxel response patterns between center-only and center+surround conditions. Higher decoding accuracy reflects greater differences in these patterns, due to changes in overall response magnitude and/or voxelwise pattern configuration. Decoding accuracy was therefore used as an index of the strength of contextual effects, although it does not provide information about their direction (i.e., suppression vs facilitation). For each ROI, decoding was performed separately for all combinations of scene categories and motion directions. To assess categorical effects, data were merged across motion directions and decoding accuracies were compared across the four category conditions using a oneway repeated-measures ANOVA. To assess motion effects, data were merged across categories and decoding accuracy was compared between same- and opposite-direction conditions using paired-samples t -tests.
Because OPA and PPA were defined as full functional ROIs rather than center-selective sub-ROIs, decoding center-only versus center+surround trials in these regions could be confounded by differences in the spatial extent of stimulation. To address this, we performed an additional analysis restricted to the center+surround conditions, thereby holding stimulation extent constant across conditions. Specifically, we conducted four-way classification to distinguish among the four category conditions (identical exemplar, different exemplar, different basic-level category, and different superordinate category). Data were collapsed across motion directions to increase statistical power. This analysis tests whether multivoxel activity patterns encode information about the categorical relationship between center and surround scenes independently of differences in peripheral stimulation.
Statistical significance relative to chance was assessed using permutation testing. For each participant, condition labels were randomly shuffled within runs, and the decoding analysis was repeated 1,000 times to generate a null distribution of accuracies. For group-level inference, shuffled accuracies were averaged across participants for each permutation to obtain a null distribution of group mean accuracies. The observed group mean accuracy was then compared with this distribution, and the empirical p-value was computed as the proportion of permutations in which the shuffled mean exceeded the observed value (p < 0.05).
Visual Feature Analysis
To assess whether the observed effects could be explained by low-level image properties, we conducted a visual feature analysis quantifying low-level visual similarity between center and surround stimuli. Following prior work using hierarchical models of early visual processing [64], we used an HMAX-based representation to characterize low-level visual structure. For each image, responses were summarized into a fixed-length feature vector, and pairwise cosine similarity between L2-normalized C1 feature vectors was computed for all center–surround image pairs. Similarity values were then averaged within each categorical condition (identical exemplar, different exemplar, different basic-level category, different superordinate category).
Results
Univariate analyses reveal distinct patterns of contextual modulation across visual regions
We first tested whether the center-only condition evoked reliable responses in each ROI (see PSC values in Figure 2). The hMT+ sub-ROI showed a reliable center-only response (t(18)=4.11, p < 0.001). Contextual modulation was quantified using the suppression index (SI; see Methods).
Figure 2.
Univariate fMRI responses. Percent signal change (PSC) responses are shown across four category conditions (identical exemplar, different exemplar, different basic-level category, and different superordinate category), two motion-direction conditions (same and opposite), and the center-only condition. Results are displayed for hMT+, V1, OPA, and PPA. Circles represent individual participant values. Error bars represent the standard error of the mean (SEM).
Figure 3 shows SI values for the hMT+ sub-ROI across the four categorical similarity conditions and two motion-direction conditions. SI values were significantly greater than zero only in the identical-exemplar condition for same-direction trials (t(18)=2.55, pFDR = 0.04), indicating the presence of suppression by the contextual surround.
Figure 3.
Univariate surround modulation effects. Suppression index (SI) values are shown across four category conditions: identical exemplar, different exemplar, different basic-level category, and different superordinate category, and two motion-direction conditions (same and opposite). Results are displayed for hMT+ and V1. Positive SI values indicate surround suppression; negative values reflect facilitation. Circles represent individual participant values. Error bars represent the standard error of the mean (SEM).
A two-way ANOVA revealed a significant main effect of motion direction (F(1,18)= 16.12, p = 0.001, η2 = 0.47), but no significant main effect of category (F(3,54)= 1.70, p = 0.18, η2 = 0.086), or interaction between category and direction (F(3,54)= 0.13, p = 0.94, η2 = 0.007). These results suggest that contextual suppression in hMT+ depends primarily on the relative motion direction between center and surround.
As expected, the categorical similarity of the center and surround had no effect on the suppression strength, consistent with hMT+ being primarily motion selective rather than sensitive to scene content. Because the categorical similarity neither modulated SI nor interacted with motion direction, SI values were averaged across the four category conditions (Figure 4). A paired-sample t -test on the averaged data confirmed that suppression was significantly stronger when the center and surround moved in the same direction than when they moved in opposite directions (t(18) = 4.02, p = 0.001). Together, these results demonstrate that hMT+ exhibits motion-dependent suppression under dynamic, natural scene conditions.
Figure 4.
Averaged univariate contextual modulation effects (N=19) in hMT+ and V1. hMT+ SI values averaged across the four category conditions for same- and opposite-direction trials. V1 SI values for the identical exemplar condition versus the mean of the three other (non-identical) conditions (i.e., different exemplar, different basic-level category, and different superordinate category conditions), shown separately for same- and opposite-direction trials. Circles represent individual participant values. Error bars represent the standard error of the mean (SEM).
In V1 sub-ROI, responses to the center-only condition were not significantly different from zero (t(18)=-1.37, p=0.19). Contextual modulation was nevertheless quantified using the suppression index (SI) within these independently defined center-selective sub-ROIs. Figure 3 shows the SI values for the V1 sub-ROI across four scene-category conditions and two motion-direction conditions. SI values were significantly lower than zero in all conditions except the identical-exemplar condition for same-direction trials (t(18)= −0.09, pFDR = 0.20).
Importantly, although center-only responses were not reliable in V1, responses varied systematically as a function of the relationship between center and surround stimuli. A two-way ANOVA revealed a significant interaction between motion direction and categorical similarity (F(3,54)=8.91, p< 0.001, η2 = 0.33), but no significant main effects of motion direction (F(1,18)=0.01, p=0.92, η2 = 0.001) or category (F(3,54)=0.1, p=0.96, η2 = 0.005).
Because the three non-identical category conditions (different exemplar, different basic-level category, and different superordinate category) did not differ from each other and showed no effect of motion direction (all ps>0.14), they were collapsed into a single “non-identical” condition, analyzed separately for same- and opposite-direction trials (Figure 4). Post-hoc comparisons confirmed that category × motion direction interaction was driven primarily by the identical-exemplar condition. Motion direction significantly affected responses only when the center and surround were identical, with reduced facilitation when the center and surround moved in the same direction compared to opposite directions (t(18)=3.49, pFDR = 0.003). Moreover, for same-direction trials, SI values were significantly lower for identical exemplars than for the collapsed non-identical conditions (t(18)=3.31, pFDR = 0.004). This pattern indicates that V1 responses depend jointly on motion direction and stimulus similarity, consistent with sensitivity to contextual structure.
Because center-only responses did not differ significantly from zero in OPA (t(18) = −0.90, p = 0.38) or PPA (t(18) = 0.73, p = 0.47), contextual effects in these regions were assessed using PSC comparisons among the center+surround conditions. In both regions, responses were comparable across conditions, and a two-way ANOVA revealed no significant main effects of motion direction (OPA: F(1,18)=1.47, p=0.24; PPA: F(1,18)=0.53, p=0.48), category (OPA: F(3,54)=0.13 p=0.94; PPA: F(3,54)=1.5, p=0.23), or their interaction (OPA: F(3,54)=1.28, p=0.29; PPA: F(3,54)=1.34, p=0.27).
These results indicate that responses in scene-selective regions were not reliably modulated by either motion congruence or categorical similarity at the univariate level under the present stimulus conditions.
Multivariate analyses reveal motion-dependent contextual modulation in hMT+
To assess how motion direction influences surround modulation, data were collapsed across the four category conditions prior to decoding to isolate motion-related effects. We then decoded the center-only versus center+surround conditions in V1, hMT+, OPA, and PPA under two motion-direction conditions: same-direction and opposite-direction.
As shown in Figure 5, decoding accuracy in hMT+ was significantly above chance only for the same-direction condition trials (p = 0.02). A paired-samples t -test revealed significantly higher decoding accuracy for same- than opposite-direction trials (t(18) = 3.07, p = 0.007), indicating that motion congruence enhances contextual modulation in hMT+ multivoxel responses. This multivariate pattern closely mirrors the univariate results, which showed stronger contextual suppression when the center and surround moved in the same direction. Together, these analyses demonstrate that hMT+ activity reflects motion-dependent contextual modulation both in overall response amplitude and in distributed activation patterns, consistent with classical motion-dependent modulation observed with simple stimuli [25, 30, 43].
Figure 5.
Multivoxel decoding accuracy (N=19) for center-only and center+surround classification for the two motion-direction conditions: same-direction and opposite-direction. Mean decoding accuracy is shown for hMT+, V1, OPA, and PPA. Circles represent individual participant values. Error bars indicate the standard error of the mean (SEM). The dashed line indicates chance-level (0.5).
Decoding accuracy was above chance in V1 and OPA for both same-direction and opposite-direction trials (all p < 0.05). However, paired-samples t -tests revealed no significant differences in decoding accuracy between same- and opposite-direction trials (all p > 0.20). As the univariate analysis in V1 revealed a motion-congruence effect only in the identical-exemplar condition, collapsing decoding analyses across category conditions may have reduced sensitivity to detect motion-dependent differences in multivoxel response patterns. Although OPA showed reliable above-chance decoding, the absence of direction-related differences in these regions suggests that their surround modulation reflects general contextual integration rather than motion congruence.
Multivariate analyses reveal category-dependent contextual representations in OPA and PPA
To assess how category similarity influences surround modulation, data were collapsed across the two motion-direction conditions prior to decoding. We then decoded the center-only versus center+surround conditions in V1, hMT+, OPA, and PPA across four category conditions: identical exemplar, different exemplar, different basic-level category, and different superordinate category.
As shown in Figures 6, classification accuracy was significantly above chance for all category conditions in V1 (all p < 0.05). In hMT+, accuracy was also above chance for all category conditions (all p < 0.05) except the different basic-level category condition (p = 0.19). Despite robust decoding performance, one-way repeated-measures ANOVAs revealed no effect of category in either region (V1: F(3, 54) = 1.93, p = 0.13, η2 = 0.10; hMT+: F(3, 54) = 0.56, p = 0.64, η2 = 0.03). These results indicate strong surround modulation in both areas, but this modulation did not depend on categorical similarity between center and surround, consistent with the roles of V1 and hMT+ in processing low-level visual features such as edges, orientation, contrast, and motion rather than higher-level, category-selective information.
Figure 6.
Multivoxel decoding accuracy (N=19) for the center-only versus center+surround classification across the four category conditions: identical exemplar, different exemplar, different basic-level category, and different superordinate category. Mean decoding accuracy is shown for hMT+, V1, OPA, and PPA. Circles represent individual participant values. Error bars represent the standard error of the mean (SEM). The dashed line indicates chance-level (0.5).
As shown in Figure 6, the classification accuracy was significantly above chance in OPA for all category conditions (all p < 0.05), except the identical-exemplar condition (p = 0.11). A one-way repeated-measures ANOVA revealed a main effect of category, F(3, 54) = 4.12, p = 0.01, η2 = 0.19. Post-hoc comparisons showed higher decoding for the different superordinate-category condition than for both the identical- and different-exemplar conditions (all t(18) > 3.01, p < 0.02), and marginally higher decoding for the different basic-level category condition than for the identical-exemplar condition (t(18) = 2.04, p = 0.06). Decoding did not differ between the identical- and different-exemplar conditions, t(18) = 0.70, p = 0.51.
These results suggest that surround modulation in OPA increases with categorical distance. Multivoxel responses were more strongly influenced by the surround when the center and surround scenes belonged to different categories, indicating that OPA is sensitive not only to the presence of contextual surround input but also to its semantic relationship with the center. This categorical sensitivity aligns with our behavioral findings [40], in which surround modulation increased as categorical similarity decreased.
PPA showed a similar pattern, with decoding accuracy increasing as the categorical dissimilarity between center and surround increased. Decoding was above chance for the different basic-level and superordinate category conditions (both p < 0.05), marginal for the different exemplar condition (p = 0.08), and not significant for the identical-exemplar condition (p = 0.33). The main effect of category was significant, F(3, 54) = 2.93, p = 0.04, η2 = 0.14. Post-hoc comparisons revealed higher decoding accuracy for the different superordinate-category condition than for the identical-exemplar condition (t(18) = 2.59, p = 0.02), and higher decoding accuracy for the different basic-level category condition than for the identical-exemplar condition (t(18) = 2.18, p = 0.04). Compared with OPA, surround modulation in PPA exhibited a more gradual and weaker increase in decoding accuracy with categorical distance. Taken together, these results extend surround modulation effects to the scene-selective cortex, demonstrating that surround modulation in these regions is shaped by higher-level categorical relationships.
To assess whether scene-selective regions reflect contextual relationships between center and surround scenes independent of differences in retinotopic extent, we performed a four-way classification analysis restricted to center+surround trials. This ensured that peripheral stimulation was held constant across conditions. We decoded the category relationship between center and surround scenes (identical exemplar, different exemplar, different basic-level category, and different superordinate category), collapsing across motion-direction conditions to isolate category-related effects.
As shown in Figure 7, decoding accuracy in V1 and hMT+ did not differ significantly from chance (both p > 0.47), indicating that these regions did not reliably represent categorical relationships between center and surround scenes. In contrast, both scene-selective regions showed sensitivity to categorical relationships between center and surround scenes. Decoding accuracy was significantly above chance in both OPA and PPA (both p < 0.05). Together, these results suggest that activity patterns in scene-selective cortex reflect contextual relationships between central and surrounding scenes, demonstrating that contextual modulation in these regions is shaped by higher-level categorical structure.
Figure 7.
Multivoxel decoding accuracy (N=19) for four-way classification of center–surround category relationships (identical exemplar, different exemplar, different basic-level category, and different superordinate category). Circles represent individual participant values. Error bars represent the standard error of the mean (SEM). The dashed line indicates chance-level (0.25).
Visual Feature Analysis
To assess whether the observed effects could be explained by low-level image properties, we conducted a visual feature analysis quantifying low-level visual similarity between center and surround stimuli. Following prior work using hierarchical models of early visual processing [64], we used an HMAX-based representation to characterize low-level visual structure.
Mean low-level similarity between center and surround stimuli was comparable across category conditions (identical exemplar: M = 1; different exemplar: M = 0.93; different basic-level: M = 0.91; different superordinate: M = 0.90). A permutation test revealed no significant effect of categorical condition on similarity (p = 0.12). These results indicate that the categorical differences observed in the neural data are unlikely to be explained by systematic differences in low-level visual similarity between center and surround stimuli.
Discussion
The present study investigated how motion congruence and categorical similarity between center and surround scenes shape neural responses to dynamic natural scenes. Using fMRI, we measured neural responses while systematically varying the contextual relationship between center and surround across four levels of categorical similarity (identical exemplar, different exemplar, different basic-level category, and different superordinate category) and two levels of motion congruence (same vs. opposite direction). Univariate analyses revealed distinct patterns of contextual modulation across visual regions. In hMT+, responses showed robust motion-dependent suppression, with stronger suppression when the center and surround moved in the same direction. In V1, responses varied with the low-level similarity between center and surround stimuli, showing reduced responses (e.g., weaker facilitation) when the two were identical and moved in the same direction. In contrast, scene-selective regions (OPA and PPA) did not show univariate modulation by either motion congruence or categorical similarity. Multivariate analyses complemented these findings. In hMT+, decoding between center-only and center+surround conditions was stronger for same-than opposite-direction motion, indicating enhanced sensitivity to motion congruence at the level of distributed activity patterns. In scene-selective regions, multivoxel patterns distinguished among category conditions, with decoding accuracy increasing as categorical similarity between center and surround decreased, indicating sensitivity to categorical relationships.
These results extend previous work on contextual influences to dynamic natural scenes. While contextual modulation has been extensively characterized using simple stimuli such as gratings [5, 12, 15], its neural basis under naturalistic conditions has remained largely unexplored. The present findings demonstrate that neural responses to natural scenes are shaped by contextual relationships between central and surrounding visual input. Importantly, these contextual effects varied across visual regions, suggesting that contextual influences operate differently depending on the type of information represented in each region. In V1 and hMT+, contextual effects were primarily driven by low-level feature relationships, such as stimulus similarity and motion congruence. In contrast, scene-selective regions showed sensitivity to higher-level contextual structure, as reflected in the categorical relationship between center and surround scenes.
Motion-dependent contextual modulation in hMT+
hMT+ exhibited stronger suppression when the center and surround moved in the same direction than when they moved in opposite directions. Decoding analyses further confirmed that this contextual modulation was driven specifically by the motion direction, independent of categorical similarity. This pattern aligns with classical neu-rophysiological findings showing that MT neurons show stronger suppression for same-direction motion and weaker or no suppression when the center and surround move in opposite directions [5, 23, 28, 29, 65]. Such motion-dependent modulation likely reflects local inhibitory interactions among direction-selective neurons in MT, consistent with the MT hypothesis of surround modulation [12]. By extending these low-level findings to naturalistic conditions, our results provide neural evidence that motion-based center–surround interactions are recruited during complex scene processing in hMT+. Furthermore, the close correspondence between our neural and behavioral findings using identical stimuli [40] suggests that perceptual suppression during natural vision may reflect these motion-sensitive inhibitory mechanisms in hMT+.
Contextual interaction in V1 reflects physical similarity
In V1, responses varied systematically with the relationship between center and surround stimuli. Specifically, responses were reduced for identical exemplars relative to other category conditions in same-direction trials, indicating that greater low-level similarity between center and surround is associated with weaker responses. Moreover, motion direction affected responses only in the identical-exemplar condition, with lower responses when the center and surround moved in the same direction compared to opposite directions.
This pattern is consistent with classical findings showing stronger suppression for highly similar or homogeneous stimuli, such as iso-oriented gratings [1, 2, 4, 26, 27, 66]. The reduced responses for identical center–surround pairs therefore suggest that V1 is sensitive to low-level similarity between center and surround scenes. The motion-dependent effects also resemble modulation patterns reported in motion-selective cortex. These effects may reflect horizontal interactions within V1 or feedback signals from motion-selective areas such as hMT+, consistent with evidence that motion information can influence early visual processing [19, 67, 68].
Findings from animal studies using naturalistic stimuli provide converging evidence for this interpretation. For example, Coen-Cagli, Kohn, and Schwartz [44] reported that macaque V1 exhibits stronger suppression when the center and surround consist of homogeneous natural image regions. This aligns broadly with our findings, which also reveal contextual modulation driven by physical similarity, although it manifested as reduced facilitation rather than increased suppression. Consistent with this, multivariate analyses further showed robust surround modulation in V1, but this effect did not depend on categorical similarity or motion congruence. Given that the univariate motion effect was observed only in the identical-exemplar condition, collapsing across category conditions in the decoding analysis may have reduced sensitivity to detect motion-dependent differences in multivoxel response patterns.
On the other hand, Onat, Jancke, and König [45] used voltage-sensitive dye imaging in the cat visual cortex to show that V1 responses to natural scene clips were facilitated when the surrounding context was drawn from the same scene. Similarly, a human fMRI study has shown that V1 responses are modulated by the consistency of surrounding scene context, with enhanced responses for scene-coherent compared to incoherent configurations [46]. While this effect is not directly equivalent to our manipulation, it likewise indicates that V1 responses are sensitive to relationships between center and surround in naturalistic stimuli. However, in our study, responses were reduced in the identical (i.e., coherent) condition compared to non-identical conditions. This difference may reflect differences in the contextual manipulations, stimulus properties, and task demands across studies.
The nature of contextual effects in V1 has been shown to vary across studies, with both suppressive and facilitatory influences reported depending on stimulus configuration, task demands, and attentional state [7, 9, 66, 69, 70]. In the present study, responses to the center-only condition were not reliably different from zero, limiting the interpretation of the observed facilitation effects relative to the baseline. One possible explanation for the absence of a reliable center-only response is the nature of the stimuli and task. The central stimulus consisted of natural scene images rather than high-contrast, simple stimuli such as gratings, and was relatively small and briefly presented, which may have produced weaker responses in V1 and a reduced signal-to-noise ratio.
Importantly, the absence of a reliable center-only response does not imply a lack of contextual effects. Univariate responses varied systematically with center–surround similarity, and multivariate analyses further revealed robust surround modulation in V1. Together, these findings indicate that V1 responses were sensitive to center–surround relationships even when the center stimulus alone did not elicit a strong univariate response. Overall, the pattern of results suggests that contextual effects in V1 primarily reflect sensitivity to low-level physical similarity rather than higher-level categorical relationships
Category-dependent contextual effects in scene-selective cortex
While univariate analyses revealed no category or motion effects, the multivariate analyses showed clear category-dependent contextual effects in scene-selective regions. Both OPA and PPA were sensitive to categorical relationships between center and surround scenes, with multivoxel responses more strongly influenced by the surround when the center and surround belonged to different categories. This pattern was most pronounced in OPA. PPA showed a similar but weaker graded effect with increasing categorical distance, possibly reflecting the lower signal-to-noise ratio typically observed in ventral visual regions [e.g., 71, 72]. Importantly, decoding revealed no motion-related effects in either region, suggesting that response differences in scene-selective cortex are primarily associated with categorical rather than motion-related factors.
The visual feature analysis showed that similarity between center and surround did not differ reliably across non-identical categorical conditions, making it unlikely that these effects are driven by low-level visual features. One potential concern is that the observed effects could instead reflect the recruitment of voxels with retinotopic selectivity for peripheral stimulation. To address this, we conducted an additional analysis restricted to the center+surround conditions, thereby holding the spatial extent of stimulation constant across conditions. This analysis revealed significant decoding across category conditions in both OPA and PPA, indicating that distributed response patterns encode information about center–surround relationships beyond differences in peripheral stimulation. In addition, each surround image appeared equally often across category conditions, making it unlikely that decoding differences reflect stimulus-specific preferences. Together, these findings suggest that the observed effects reflect sensitivity to contextual relationships between central and surrounding scene information rather than differences in peripheral stimulation or low-level visual differences.
Our categorical similarity manipulation aligns with established frameworks of visual categorization that distinguish between superordinate and basic-level representations. Superordinate categories (e.g., indoor vs. outdoor) capture broad semantic distinctions, whereas basic-level categories correspond to more specific and perceptually meaningful scene classes [73, 74]. Converging evidence indicates that such categorical structure is reflected in visual cortex, with coarse distinctions associated with large-scale response differences and finer distinctions relying on more subtle representational variations [74]. Importantly, the three category levels used here establish a clear ordinal structure, such that categorical distance increases monotonically from different exemplars within a basic-level category, to different basic-level categories, and finally to superordinate categories. This graded organization provides a principled and convenient way to parametrize categorical distance [75]. Consistent with this framework, our recent behavioral results, obtained with the same stimulus set and category manipulation [40], revealed a graded increase in contrast thresholds from identical exemplars to different exemplars within the same basic-level category, and further across basic- and superordinate-level categories. This graded pattern is consistent with differences in categorical distance and provides converging support for the validity of our category manipulation. Importantly, these behavioral results are consistent with the present neural findings, indicating that contextual influences vary systematically with categorical dissimilarity.
Similar sensitivity to categorical context has been demonstrated in prior neuroimaging work on scene and object perception. Contextual incongruence between scene elements has been shown to alter neural responses in parahippocampal and occipitotemporal regions [42, 76, 77]. For example, Peyrin et al. [42] reported that categorically congruent (e.g., both man-made) but physically dissimilar (e.g., buildings vs. streets) peripheral scenes disrupted central scene categorization and increased activation in inferior frontal and occipitotemporal cortices. Consistent with this, we observed stronger context-dependent differences even when the center and surround belonged to the same superordinate category (e.g., two indoor scenes) but differed at the basic level (e.g., restaurant vs. museum).
Contextual influences across visual cortex may support efficient scene representation
Real-world scenes show systematic spatial and semantic regularities (e.g., a bathroom typically contains a sink, mirror, and shower arranged within a characteristic layout). The visual system exploits these regularities to integrate information efficiently across multiple representational levels, supporting rapid understanding of complex environments [34, 36, 78]. Computational and neurophysiological models suggest that such contextual integration relies on recurrent and inhibitory interactions that can modulate sensory input according to its relevance [68, 79, 80]. Consistent with this view, behavioral and neuroimaging work demonstrates that coherent scene structure facilitates both scene and object perception [36, 81–86]. Together, these findings suggest that contextual mechanisms may enhance informative signals while attenuating redundant input, thereby promoting stable and efficient visual processing in natural environments.
Our results extend this framework by showing that categorically incongruent surrounds are associated with stronger contextual modulation in scene-selective cortex than congruent surrounds. Whereas studies using simple stimuli have attributed contextual effects to the suppression of statistical redundancies in low-level features [19, 20, 87–89], the present findings suggest that, at higher representational levels, responses vary systematically with categorical relationships between center and surround. In natural scenes, categorically incongruent information may enhance neural responses because it violates contextual expectations and conveys higher informational value, whereas with simple stimuli, identical information may be suppressed because it is predictable and redundant. In contrast, V1 showed a pattern more consistent with low-level surround modulation, where facilitation decreased when the center and surround were physically identical, indicating a greater influence of visual redundancy than categorical structure. Both forms of modulation may therefore reflect a shared computational principle: optimizing neural representations by emphasizing informative differences while attenuating predictable input. Both forms of modulation may therefore reflect a shared computational principle: optimizing neural representations by emphasizing informative differences while attenuating predictable input.
Together, these results are consistent with a functional differentiation of contextual influences across visual cortex, likely shaped by the functional specialization of distinct visual pathways. Early visual areas (V1) are primarily governed by physical similarity, mid-level motion areas (hMT+) are sensitive to motion congruence, and higher-level scene-selective regions (OPA and PPA) integrate categorical structure. This pattern is broadly consistent with predictive and efficient coding accounts [19, 79, 80], whereby different regions of the visual system may selectively suppress predictable or redundant input and enhance informative discrepancies.
Limitations and future directions
Although the present study provides neural evidence for category- and motion-dependent contextual effects in natural vision, several questions remain open for future investigation. First, although the visual feature analysis indicated that low-level visual similarity did not differ reliably across categorical conditions, we did not explicitly manipulate low-level image properties such as spatial frequency or amplitude spectrum. It therefore remains possible that more specific or localized visual features, not captured by the present analysis, contribute to surround modulation [42]. Future work could examine these factors more directly by varying physical properties and categorical coherence independently to better disentangle their respective influences.
Second, while we used sub-ROI localizers in V1 and hMT+ to identify voxels corresponding to the center stimulus, this approach could not be applied to OPA and PPA due to their larger receptive fields. Consequently, the ROIs in these regions may have included voxels responsive to both center and surround, so that the effects cannot be directly interpreted as modulatory effects on the representation of the central stimulus. Moreover, scene-selective areas are thought to contain fine-grained functional subfields with heterogeneous tuning to spatial scale, visual field position, and category information [e.g., 38, 90]. The absence of univariate category effects in OPA and PPA may therefore reflect this spatial and functional heterogeneity, where the averaging across voxels with different category or positional preferences could mask category-specific responses. Category-specific modulatory effects may instead depend on multivariate response patterns in scene-selective cortex, necessitating pattern-based analysis approaches for characterizing this fine-scale organization. Category-specific modulatory effects may instead depend on multivariate response patterns in scene-selective cortex, necessitating pattern-based analysis approaches for characterizing this fine-scale organization. As decoding performance alone does not specify which stimulus or response dimensions drive classification, the present results should be interpreted as consistent with categorical similarity effects rather than as definitive evidence for a specific representational mechanism.
Moreover, although our stimuli consisted of natural scenes, the applied drifting motion does not fully reflect the complexity of motion in everyday visual experience. While coherent motion can occur during self-motion or eye movements, natural environments typically involve more complex and heterogeneous motion patterns. Thus, our stimuli provide a controlled approximation of naturalistic motion, and future work should extend this approach using more complex, dynamic videos to better capture these dynamics.
Finally, the relatively small central aperture (1.9°) may have limited the amount of scene information available and potentially constrained explicit recognition of detailed scene content. Although similar stimulus sizes are commonly used in studies of center–surround interactions [12, 15], and prior work demonstrates that meaningful scene information can be extracted from comparable visual extents [41, 91], the present study did not include a direct measure of perceptual scene categorization. However, the stimuli and design closely matched those used in our recent behavioral study [40], in which participants reliably categorized central scenes and exhibited robust contextual effects under comparable conditions, providing indirect evidence that the central stimuli conveyed perceptually relevant information. Future work combining neuroimaging with concurrent behavioral measures will be important to more directly establish the relationship between neural responses and perceptual experience.
Conclusions
Using dynamic natural scenes, we show that contextual influences on neural responses depend on both motion-direction congruence and categorical relationships between center and surround, but that these effects differ systematically across functionally distinct visual regions. In hMT+, univariate and multivariate analyses revealed motion-dependent contextual modulation, with stronger effects for same-direction than opposite-direction motion, consistent with classical center–surround interactions in motion-selective cortex. In V1, responses were sensitive to low-level similarity between center and surround, with reduced responses when the two were physically identical and moved in the same direction. In scene-selective cortex (OPA and PPA), multivoxel activity patterns encoded the categorical relationship between center and surround scenes, with decoding accuracy increasing as categorical similarity between center and surround decreased. Together, these findings indicate that contextual modulation on natural scene processing differs across cortical regions, shifting from sensitivity to low-level feature relationships in early and motion-selective cortex to sensitivity to higher-level categorical structure in scene-selective regions. This reflects the distinct representational roles of these regions and suggests that contextual influences operate at multiple levels of visual representation. More broadly, this pattern is consistent with predictive and efficient coding accounts, whereby different regions of the visual system attenuate predictable input and emphasize informative discrepancies to support perceptual stability and efficient scene understanding.
New & Noteworthy.
Visual perception depends on contextual interactions across the visual field. Using fMRI and dynamic natural scenes, we show that contextual modulation differs across the visual system: it is motion-dependent in hMT+, reflects low-level similarity in V1, and is shaped by categorical relationships between center and surround in scene-selective cortex. Together, these findings reveal differences across functionally distinct visual regions in how contextual information is represented during natural vision.
Acknowledgements
This work was supported by the Deutsche Forschungsgemeinschaft (DFG), grants SFB/TRR 135 (project no. 222641018); KA4683/5-1 (project no. 518483074); and under Germany’s Excellence Strategy (EXC 3066/1, “The Adaptive Mind”, project no. 533717223). It was further supported by an European Research Council (ERC) Starting Grant (PEP, ERC-2022-STG 101076057). Views and opinions expressed are those of the authors only and do not necessarily reflect those of the European Union or the European Research Council. Neither the European Union nor the granting authority can be held responsible for them.
MR imaging for this study was performed at the Bender Institute of Neuroimaging (BION) at Justus Liebig University Giessen, Germany.
We thank Tugce Dalmis for assistance with data collection.
Footnotes
Author Contributions
M.K. and D.K. conceived and designed the research; M.K. performed the experiments and analyzed the data; M.K., D.P., and D.K. interpreted the results of the experiments; M.K. prepared the figures and drafted the manuscript; M.K., D.P., and D.K. edited and revised the manuscript; M.K., D.P., and D.K. approved the final version of the manuscript.
Competing interests
The authors declare no competing interests.
Data Availability
The data that support the findings of this study are available at the following DOI: (https://doi.org/10.5281/zenodo.17774893).
References
- 1.DeAngelis GC, Freeman RD, Ohzawa I. Length and width tuning of neurons in the cat’s primary visual cortex. Journal of Neurophysiology. 1994;71(1):347–374. doi: 10.1152/jn.1994.71.1.347. [DOI] [PubMed] [Google Scholar]
- 2.Walker Gary A, Ohzawa Izumi, Freeman Ralph D. Asymmetric Suppression Outside the Classical Receptive Field of the Visual Cortex. Journal of Neuroscience. 1999;19(23):10536–10553. doi: 10.1523/JNEUROSCI.19-23-10536.1999. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Xing Jing, Heeger David J. Measurement and modeling of center-surround suppression and enhancement. Vision Research. 2001;41:571–583. doi: 10.1016/s0042-6989(00)00270-4. [DOI] [PubMed] [Google Scholar]
- 4.Cavanaugh James R, Bair Wyeth, Anthony Movshon J. Nature and interaction of signals from the receptive field center and surround in macaque V1 neurons. Journal of Neurophysiology. 2002;88(4):2530–2546. doi: 10.1152/jn.00692.2001. [DOI] [PubMed] [Google Scholar]
- 5.Cavanaugh James R, Bair Wyeth, Movshon J Anthony. Selectivity and Spatial Distribution of Signals From the Receptive Field Surround in Macaque V1 Neurons. Journal of Neurophysiology. 2002;88:2547–2556. doi: 10.1152/jn.00693.2001. [DOI] [PubMed] [Google Scholar]
- 6.Zenger-Landolt Barbara, Heeger David J. Response Suppression in V1 Agrees with Psychophysics of Surround Masking. Journal of Neuroscience. 2003;23(17):6884–6893. doi: 10.1523/JNEUROSCI.23-17-06884.2003. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Williams Adrian L, Singh Krishna D, Smith Andrew T. Surround Modulation Measured With Functional MRI in the Human Visual Cortex. Journal of Neurophysiology. 2003;89(1):525–533. doi: 10.1152/jn.00048.2002. [DOI] [PubMed] [Google Scholar]
- 8.Petrov Yury, Carandini Matteo, McKee Suzanne. Two Distinct Mechanisms of Suppression in Human Vision. Journal of Neuroscience. 2005;25(38):8704–8707. doi: 10.1523/JNEUROSCI.2871-05.2005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Ichida Jennifer M, et al. Response Facilitation from the “Suppressive” Receptive Field Surround of Macaque V1 Neurons. Journal of Neurophysiology. 2007;98(4):2168–2181. doi: 10.1152/jn.00298.2007. [DOI] [PubMed] [Google Scholar]
- 10.Pihlaja M, et al. Quantitative multifocal fMRI shows active suppression in human V1. Human Brain Mapping. 2008;29(9):1001–1014. doi: 10.1002/hbm.20442. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Shushruth S, et al. Strong recurrent networks compute the orientation tuning of surround modulation in the primate primary visual cortex. The Journal of Neuroscience. 2012;32:308–321. doi: 10.1523/JNEUROSCI.3789-11.2012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Tadin Duje, et al. Perceptual consequences of centre-surround antagonism in visual motion processing. Nature. 2003;424:312–315. doi: 10.1038/nature01800. [DOI] [PubMed] [Google Scholar]
- 13.Tadin Duje, et al. Improved Motion Perception and Impaired Spatial Suppression following Disruption of Cortical Area MT/V5. Journal of Neuroscience. 2011;31(4):1279–1283. doi: 10.1523/JNEUROSCI.4121-10.2011. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Er Gorkem, Pamir Zahide, Boyaci Huseyin. Distinct patterns of surround modulation in V1 and hMT+ NeuroImage. 2020;220:117084. doi: 10.1016/j.neuroimage.2020.117084. [DOI] [PubMed] [Google Scholar]
- 15.Schallmo Michael-Paul, et al. Suppression and facilitation of human neural responses. eLife. 2018;7:1–23. doi: 10.7554/eLife.30334. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Turkozer Halide B, Pamir Zahide, Boyaci Huseyin. Contrast Affects fMRI Activity in Middle Temporal Cortex Related to Center–Surround Interaction in Motion Perception. Frontiers in Psychology. 2016;7:1–8. doi: 10.3389/fpsyg.2016.00454. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Pack Christopher C, Hunter J Nicholas, Born Richard T. Contrast Dependence of Suppressive Influences in Cortical Area MT of Alert Macaque. Journal of Neurophysiology. 2005;93(3):1809–1815. doi: 10.1152/jn.00629.2004. [DOI] [PubMed] [Google Scholar]
- 18.Kiniklioglu Mert, Boyaci Huseyin. hMT+ activity predicts the effect of spatial attention on surround suppression. Journal of Vision. 2025;25(4):12. doi: 10.1167/jov.25.4.12. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Angelucci Alessandra, et al. Circuits and Mechanisms for Surround Modulation in Visual Cortex. Annual Review of Neuroscience. 2017;40:425–451. doi: 10.1146/annurev-neuro-072116-031418. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Nurminen Lauri, Angelucci Alessandra. Multiple Components of Surround Modulation in Primary Visual Cortex: Multiple Neural Circuits with Multiple Functions? Vision Research. 2014;104:47–56. doi: 10.1016/j.visres.2014.08.018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Hallum Lauren E, Movshon J Anthony. Surround suppression supports second-order feature encoding by macaque V1 and V2 neurons. Vision Research. 2014;104:24–35. doi: 10.1016/j.visres.2014.10.004. ISSN: 0042-6989. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.Sundberg Kristy A, Mitchell Jude F, Reynolds John H. Spatial attention modulates center-surround interactions in macaque visual area V4. Neuron. 2009;61(6):952–963. doi: 10.1016/j.neuron.2009.02.023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Allman J, Miezin F, McGuinness E. Direction- and velocity-specific responses from beyond the classical receptive field in the middle temporal visual area (MT) Perception. 1985;14:105–126. doi: 10.1068/p140105. [DOI] [PubMed] [Google Scholar]
- 24.Shushruth S, et al. Different orientation tuning of near- and far-surround suppression in macaque primary visual cortex mirrors their tuning in human perception. The Journal of Neuroscience. 2013;33(1):106–119. doi: 10.1523/JNEUROSCI.2518-12.2013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Kiniklioglu Merve, Boyaci Huseyin. Increasing the spatial extent of attention strengthens surround suppression. Vision Research. 2022;199:108074. doi: 10.1016/j.visres.2022.108074. [DOI] [PubMed] [Google Scholar]
- 26.Serrano-Pedraza Ignacio, Grady John P, Read Jenny CA. Spatial frequency bandwidth of surround suppression tuning curves. Journal of Vision. 2012;12(6):24. doi: 10.1167/12.6.24. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Schallmo Michael-Paul, et al. The effects of orientation and attention during surround suppression of small image features: A 7 Tesla fMRI study. Journal of Vision. 2016;16(10):19. doi: 10.1167/16.10.19. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 28.Born RT, Tootell RBH. Segregation of global and local motion processing in primate middle temporal visual area. Nature. 1992;357:497–499. doi: 10.1038/357497a0. [DOI] [PubMed] [Google Scholar]
- 29.Lamme VAF. The neurophysiology of figure-ground segregation in primary visual cortex. Journal of Neuroscience. 1995;15:1605–1616. doi: 10.1523/JNEUROSCI.15-02-01605.1995. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Paffen Chris LE, et al. Center-surround inhibition and facilitation as a function of size and contrast at multiple levels of visual motion processing. Journal of Vision. 2005;5:571–578. doi: 10.1167/5.6.8. [DOI] [PubMed] [Google Scholar]
- 31.Qiu Cheng, et al. Responses in early visual areas to contour integration are context dependent. Journal of Vision. 2016;16(8):19. doi: 10.1167/16.8.19. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Altmann Christian F, Heinrich HBülthoff, Kourtzi Zoe. Perceptual organization of local elements into global shapes in the human visual cortex. Current Biology. 2003;13(4):342–349. doi: 10.1016/s0960-9822(03)00052-6. [DOI] [PubMed] [Google Scholar]
- 33.Murray Scott O, et al. Shape perception reduces activity in human primary visual cortex. Proceedings of the National Academy of Sciences of the United States of America. 2002;99(23):15164–15169. doi: 10.1073/pnas.192579399. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Bar Moshe. Visual objects in context. Nature Reviews Neuroscience. 2004;5(8):617–629. doi: 10.1038/nrn1476. [DOI] [PubMed] [Google Scholar]
- 35.Oliva Aude, Torralba Antonio. The role of context in object recognition. Trends in Cognitive Sciences. 2007;11(12):520–527. doi: 10.1016/j.tics.2007.09.009. [DOI] [PubMed] [Google Scholar]
- 36.Kaiser Daniel, et al. Object Vision in a Structured World. Trends in Cognitive Sciences. 2019;23(8):672–685. doi: 10.1016/j.tics.2019.04.013. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Levy Ifat, et al. Center-Periphery Organization of Human Object Areas. Nature Neuroscience. 2001;4(5):533–539. doi: 10.1038/87490. ISSN: 1097-6256. [DOI] [PubMed] [Google Scholar]
- 38.Silson Edward H, et al. A Retinotopic Basis for the Division of High-Level Scene Processing between Lateral and Ventral Human Occipitotemporal Cortex. The Journal of Neuroscience. 2015;35(34):11921–11935. doi: 10.1523/JNEUROSCI.0137-15.2015. ISSN: 0270-6474. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Nasr Shahin, et al. Scene-Selective Cortical Regions in Human and Nonhuman Primates. The Journal of Neuroscience. 2011;31(39):13771–13785. doi: 10.1523/JNEUROSCI.2792-11.2011. ISSN: 0270-6474. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Kiniklioglu Mert, Kaiser Daniel. Characterizing surround suppression with dynamic natural scenes. Journal of Vision. 2026;26(3):9. doi: 10.1167/jov.26.3.9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Faurite Chloé, et al. Interaction between central and peripheral vision: Influence of distance and spatial frequencies. Journal of Vision. 2024;24(1):3. doi: 10.1167/jov.24.1.3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Peyrin Carole, et al. Semantic and physical properties of peripheral vision are used for scene categorization in central vision. Journal of Cognitive Neuroscience. 2021;33(5):799–813. doi: 10.1162/jocn_a_01689. [DOI] [PubMed] [Google Scholar]
- 43.Paffen Chris LE, Alais David, Verstraten Frans AJ. Center–surround inhibition deepens binocular rivalry suppression. Vision Research. 2005;45:2642–2649. doi: 10.1016/j.visres.2005.04.018. [DOI] [PubMed] [Google Scholar]
- 44.Coen-Cagli Ruben, Kohn Adam, Schwartz Odelia. Flexible gating of contextual influences in natural vision. Nature Neuroscience. 2015;18(11):1648–1655. doi: 10.1038/nn.4128. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Onat Selim, Jancke Dirk, Peter König. Cortical long-range interactions embed statistical knowledge of natural sensory input: a voltage-sensitive dye imaging study. F1000Research. 2013;2:51. doi: 10.12688/f1000research.2-51.v2. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Mannion Damian J, Kersten Daniel J, Olman Cheryl A. Scene coherence can affect the local response to natural images in human V1. European Journal of Neuroscience. 2015;42(11):2895–2903. doi: 10.1111/ejn.13082. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Brainard David H. The Psychophysics Toolbox. Spatial Vision. 1997;10:433–436. [PubMed] [Google Scholar]
- 48.Huk AC, Dougherty RF, Heeger DJ. Retinotopy and functional sub-division of human areas mt and mst. Journal of Neuroscience. 2002;22(5):7195–7205. doi: 10.1523/JNEUROSCI.22-16-07195.2002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Engel SA, Glover GH, Wandell BA. Retinotopic organization in human visual cortex and the spatial precision of functional mri. Cerebral Cortex. 1997;7(2):181–192. doi: 10.1093/cercor/7.2.181. [DOI] [PubMed] [Google Scholar]
- 50.Sereno MI, et al. Borders of multiple visual areas in humans revealed by functional magnetic resonance imaging. Science. 1995;268(5212):889–893. doi: 10.1126/science.7754376. [DOI] [PubMed] [Google Scholar]
- 51.Greenberg Adam S, et al. Visuotopic cortical connectivity underlying attention revealed with white-matter tractography. Journal of Neuroscience. 2012;32(8):2773–2782. doi: 10.1523/JNEUROSCI.5419-11.2012. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Slotnick Scott D, Yantis Steven. Efficient acquisition of human retinotopic maps. Human Brain Mapping. 2003;18(1):22–29. doi: 10.1002/hbm.10077. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Pitcher David, et al. Differential selectivity for dynamic versus static information in face-selective cortical regions. NeuroImage. 2011;56(4):2356–2363. doi: 10.1016/j.neuroimage.2011.03.067. [DOI] [PubMed] [Google Scholar]
- 54.Emel Küçük, et al. Moving and Static Faces, Bodies, Objects, and Scenes Are Differentially Represented across the Three Visual Pathways. Journal of Cognitive Neuroscience. 2024;36(12):2639–2651. doi: 10.1162/jocn_a_02139. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Dale AM, Fischl B, Sereno MI. Cortical Surface-Based Analysis: I. Segmentation and Surface Reconstruction. NeuroImage. 1999;9(2):179–194. doi: 10.1006/nimg.1998.0395. [DOI] [PubMed] [Google Scholar]
- 56.Fischl B, Sereno MI, Dale AM. Cortical Surface-Based Analysis: II: Inflation, Flattening, and a Surface-Based Coordinate System. NeuroImage. 1999;9(2):195–207. doi: 10.1006/nimg.1998.0396. [DOI] [PubMed] [Google Scholar]
- 57.Woolrich MW, et al. Temporal Autocorrelation in Univariate Linear Modeling of FMRI Data. NeuroImage. 2001;14(6):1370–1386. doi: 10.1006/nimg.2001.0931. [DOI] [PubMed] [Google Scholar]
- 58.Kriegeskorte Nikolaus, Goebel Rainer, Bandettini Peter. Information-based functional brain mapping. Proceedings of the National Academy of Sciences of the United States of America. 2006;103(10):3863–3868. doi: 10.1073/pnas.0600244103. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Op de Beeck Hans P. Against hyperacuity in brain reading: spatial smoothing does not hurt multivariate fMRI analyses? NeuroImage. 2010;49(3):1943–1948. doi: 10.1016/j.neuroimage.2009.02.047. [DOI] [PubMed] [Google Scholar]
- 60.Power Jonathan D, et al. Spurious but systematic correlations in functional connectivity MRI networks arise from subject motion. NeuroImage. 2012;59(3):2142–2154. doi: 10.1016/j.neuroimage.2011.10.018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Parkes Lara, et al. An evaluation of the efficacy, reliability, and sensitivity of motion correction strategies for resting-state functional MRI. NeuroImage. 2018;171:415–436. doi: 10.1016/j.neuroimage.2017.12.073. [DOI] [PubMed] [Google Scholar]
- 62.Ciric Rastko, et al. Benchmarking of participant-level confound regression strategies for the control of motion artifact in studies of functional connectivity. NeuroImage. 2017;154:174–187. doi: 10.1016/j.neuroimage.2017.03.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Oosterhof Nikolaas N, Connolly Andrew C, Haxby James V. CoSMoMVPA: Multi-Modal Multivariate Pattern Analysis of Neuroimaging Data in Matlab/GNU Octave. Frontiers in Neuroinformatics. 2016;10:27. doi: 10.3389/fninf.2016.00027. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Riesenhuber Maximilian, Poggio Tomaso. Hierarchical Models of Object Recognition in Cortex. Nature Neuroscience. 1999;2(11):1019–1025. doi: 10.1038/14819. [DOI] [PubMed] [Google Scholar]
- 65.Kastner S, Nothdurft HC, Pigarev IN. Neuronal correlates of pop-out in cat striate cortex. Vision Research. 1995;37:371–376. doi: 10.1016/s0042-6989(96)00184-8. [DOI] [PubMed] [Google Scholar]
- 66.Flevaris Anastasia V, Murray Scott O. Attention Determines Contextual Enhancement versus Suppression in Human Primary Visual Cortex. The Journal of Neuroscience. 2015;35(35):12273–12280. doi: 10.1523/JNEUROSCI.1409-15.2015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Jean-Michel Hupé, et al. Feedback Connections Act on the Early Part of the Responses in Monkey Visual Cortex. Journal of Neurophysiology. 2001;85(1):134–145. doi: 10.1152/jn.2001.85.1.134. ISSN: 0022-3077. [DOI] [PubMed] [Google Scholar]
- 68.Angelucci Alessandra, Bressloff Paul C. The contribution of feedforward, lateral and feedback connections to the classical receptive field center and extra-classical receptive field surround of primate V1 neurons. Progress in Brain Research. 2006;154:93–120. doi: 10.1016/S0079-6123(06)54005-1. [DOI] [PubMed] [Google Scholar]
- 69.Nurminen Lauri, et al. Area Summation in Human Visual System: Psychophysics, fMRI, and Modeling. Journal of Neurophysiology. 2009;102(5):2900–2909. doi: 10.1152/jn.00201.2009. [DOI] [PubMed] [Google Scholar]
- 70.Nurminen Lauri, Kilpeläinen Markku, Vanni Simo. Fovea-Periphery Axis Symmetry of Surround Modulation in the Human Visual System. PLOS ONE. 2013;8(2):1–11. doi: 10.1371/journal.pone.0057906. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Rua Catarina, et al. Improving fMRI in Signal Drop-Out Regions at 7 T by Using Tailored Radio-Frequency Pulses: Application to the Ventral Occipito-Temporal Cortex. MAGMA (New York, NY) 2018;31(2):257–267. doi: 10.1007/s10334-017-0652-x. ISSN: 1352-8661. [DOI] [PubMed] [Google Scholar]
- 72.Winawer Jonathan, et al. Mapping hV4 and Ventral Occipital Cortex: The Venous Eclipse. Journal of Vision. 2010;10(5):1–1. doi: 10.1167/10.5.1. ISSN: 1534-7362. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Rosch Eleanor, et al. Basic objects in natural categories. Cognitive Psychology. 1976;8(3):382–439. [Google Scholar]
- 74.Grill-Spector Kalanit, Weiner Kevin S. The functional architecture of the ventral temporal cortex and its role in categorization. Nature Reviews Neuroscience. 2014;15(8):536–548. doi: 10.1038/nrn3747. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Chen L, Cichy Radoslaw M, Kaiser Daniel. Coherent categorical information triggers integration-related alpha dynamics. Journal of Neurophysiology. 2024;131(4):619–625. doi: 10.1152/jn.00450.2023. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 76.Rémy Frédéric, et al. Incongruent Object/Context Relationships in Visual Scenes: Where Are They Processed in the Brain? Brain and Cognition. 2014;84(1):34–43. doi: 10.1016/j.bandc.2013.10.008. ISSN: 0278-2626. [DOI] [PubMed] [Google Scholar]
- 77.Faivre Nathan, et al. Imaging Object-Scene Relations Processing in Visible and Invisible Natural Scenes. Scientific Reports. 2019;9(1):4567. doi: 10.1038/s41598-019-38654-z. ISSN: 2045-2322. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 78.Võ Melissa L-H. The meaning and structure of scenes. Vision Research. 2021;181:10–20. doi: 10.1016/j.visres.2020.11.003. [DOI] [PubMed] [Google Scholar]
- 79.Rao RPN, Ballard Dana H. Predictive Coding in the Visual Cortex: A Functional Interpretation of Some Extra-Classical Receptive-Field Effects. Nature Neuroscience. 1999;2:79–87. doi: 10.1038/4580. [DOI] [PubMed] [Google Scholar]
- 80.Gilbert Charles D, Li Wu. Top-down influences on visual processing. Nature Reviews Neuroscience. 2013;14(5):350–363. doi: 10.1038/nrn3476. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 81.Davenport Jessica L, Potter Mary C. Scene consistency in object and background perception. Psychological Science. 2004;15(8):559–564. doi: 10.1111/j.0956-7976.2004.00719.x. [DOI] [PubMed] [Google Scholar]
- 82.Võ Melissa L-H, Wolfe Jeremy M. Differential electrophysiological signatures of semantic and syntactic scene processing. Psychological Science. 2013;24(9):1816–1823. doi: 10.1177/0956797613476955. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 83.Kaiser Daniel, Peelen Marius V. Transformation from independent to integrative coding of multi-object arrangements in human visual cortex. NeuroImage. 2018;169:334–341. doi: 10.1016/j.neuroimage.2017.12.065. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 84.Kaiser Daniel, Häberle Greta, Cichy Radoslaw M. Cortical Sensitivity to Natural Scene Structure. Human Brain Mapping. 2020;41(5):1286–1295. doi: 10.1002/hbm.24875. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 85.Kaiser Daniel, Häberle Greta, Cichy Radoslaw M. Coherent natural scene structure facilitates the extraction of task-relevant object information in visual cortex. NeuroImage. 2021;240:118365. doi: 10.1016/j.neuroimage.2021.118365. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 86.Chen Liuqing, Cichy Radoslaw M, Kaiser Daniel. Semantic Scene-Object Consistency Modulates N300/400 EEG Components, but Does Not Automatically Facilitate Object Representations. Cerebral Cortex. 2022;32(16):3553–3567. doi: 10.1093/cercor/bhab433. [DOI] [PubMed] [Google Scholar]
- 87.Coen-Cagli Ruben, Dayan Peter, Schwartz Odelia. Cortical Surround Interactions and Perceptual Salience via Natural Scene Statistics. PLOS Computational Biology. 2012;8(3):e1002405. doi: 10.1371/journal.pcbi.1002405. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 88.Vinje William E, Gallant Jack L. Sparse Coding and Decorrelation in Primary Visual Cortex during Natural Vision. Science. 2000;287(5456):1273–1276. doi: 10.1126/science.287.5456.1273. [DOI] [PubMed] [Google Scholar]
- 89.Schwartz Odelia, Simoncelli Eero P. Natural Signal Statistics and Sensory Gain Control. Nature Neuroscience. 2001;4(8):819–825. doi: 10.1038/90526. [DOI] [PubMed] [Google Scholar]
- 90.Baldassano Christopher, Beck Diane M, Fei-Fei Li. Differential Connectivity Within the Parahippocampal Place Area. NeuroImage. 2013;75:228–237. doi: 10.1016/j.neuroimage.2013.02.073. ISSN: 1053-8119. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 91.Trouilloud Audrey, et al. Influence of physical features from peripheral vision on scene categorization in central vision. Visual Cognition. 2022;30(6):425–442. doi: 10.1080/13506285.2022.2087814. [DOI] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
The data that support the findings of this study are available at the following DOI: (https://doi.org/10.5281/zenodo.17774893).








