Skip to main content
Proceedings of the National Academy of Sciences of the United States of America logoLink to Proceedings of the National Academy of Sciences of the United States of America
. 2002 Jul 9;99(15):9813–9818. doi: 10.1073/pnas.152318799

Modeling of automatic capture and focusing of visual attention

Teuvo Kohonen 1,*
PMCID: PMC125026  PMID: 12107284

Abstract

An explanation, based on simple analysis of the spatiotemporal variations of the visual environment, is given to the automatic capture and focusing of visual attention. It is assumed that the transmittance for the sensory signals is modulated by separate control circuits that sample input from the same area of the visual field but at a lower resolution. When these circuits detect significant spatial and/or temporal variations, they “open gates” for the more accurate information arising from the same area. If the variations are related to the spatial resolution, which varies within wide limits over the retina, the visual field is “opened” up to a radius where it captures the most salient structures of the image. If the temporal variations of the signals are further emphasized, the high spatial frequencies begin to dominate. If then the gaze is moved by a small amount, the transmittance of the foveal signal paths is activated strongest.


The term attention has been used for many different psychological, behavioral, and physiological acts and states, such as focused awareness and certain activated conditions of the nervous system. In this context, it may be expedient to define attention as a set of those neural functions by which sensory information is selected and emphasized for perception.

Many phenomena related to attention have been described already by the ancient and medieval philosophers such as Aristotle, Lucretius, and Descartes. In 1740, the German psychologist C. Wolff noted that the greater the visual attention, the smaller the part of the visual field to which it extends. For a review, cf. ref. 1.

The objective of the work in presentation has been to point out that a very simple automatic selection mechanism, supposed to be at work already in the ascending visual signal paths, may underlie several familiar phenomena that are usually attributed to visual attention. By means of simple modeling approaches, an explicit explanation is given in this paper to the following phenomena: first, automatic activation of a subset of visual signal paths, equivalent to an “attentional window,” such that the width of the window is defined by the local variances of the visual signals; second, narrowing of the “attentional window” when small saccadic eye movements, voluntary or involuntary, are made (this effect can be shown to ensue from the same model, when the primary signals are further high-pass filtered); and third, shifting of the “attentional window” when strong or novel stimuli (distractors) occur eccentrically in the visual field.

General Modeling Hypotheses

To carry out the simulations presented in this paper, we have first to construct some general modeling hypotheses. They are of three kinds: anatomical and physiological facts, which will be simplified and paraphrased for theoretical approaches; assumptions that are biologically motivated; and consequences of the former.

Facts.

(i) Visual signals (in particular in primates) are propagated from the retina to the visual cortex via the thalamus along signal paths that mediate various visual features. These features are related to the different locations of the retina, and the topological order of the signal paths is preserved all of the way through (2, 3).

(ii) The sampling areas of the retinal ganglia are known to vary within wide limits over the retina. The signal paths that start in the center of the retina (fovea) are able to carry signals that represent denser variations in the spatial domain, whereas the peripheral paths are better fit to smoother variations that have a larger extension in the spatial domain.

(iii) The retina is mapped onto the thalamus in such a way that the fovea projects onto the greater portion, and relatively few cells are devoted to the peripheral retina (4). A similar magnification factor is associated with the projection of the retina onto the striate cortex, which is generally known to be such that the radial coordinates of the retina are roughly transformed into rectangular coordinates on the cortex. (For a detailed mapping in the cat brain, cf., e.g., ref. 5; for a recent theoretical account, cf., e.g., ref. 6.)

Assumptions.

(iv) The most essential theoretical hypothesis made in this paper is that with the main signal paths there must be associated other circuits that modulate the transmittance of the former for the signals. Such neural circuits exist at least in the thalamus (lateral geniculate body; ref. 7). We assume that each control circuit monitors and controls its own subset of signal paths up to a certain radius in the lateral direction. Further, we assume that the control circuits somehow analyze spatial and temporal variations of the signals in their neighborhood, at a resolution that is proportional to the corresponding resolution at the retina, and open an “attentional window” according to this analysis. The control circuits shall further compete mutually in the sense that preference for the transfer of signals is given to those paths, the associated control circuits of which detect the greater variations in the signals in its neighborhood. The rest of the neighboring paths are inhibited or somehow neglected. Competitive functions exist in many places in the nervous systems, but here the competition can also be more indirect, defined, e.g., in the “higher” parts of the brain.

Consequences.

(v) According to iii, if the control circuits reside in the thalamus or the visual cortex, their control effect, when related to the central part of the retina, is confined to a narrower area, whereas the effect is much wider in the peripheral parts.

(vi) If the gaze is directed at a small object or densely located details concentrated on a small area, and there are no salient visual structures in its surroundings, the control circuits associated with the narrow central areas of the retina are activated more strongly than the peripheral control circuits. Accordingly, an attentional window around the fovea is then opened, whereas the rest of the signal paths are inhibited or otherwise suppressed. Conversely, the broader or smoother the structures of the visual scene around the gaze, the broader the attentional window.

(vii) If the image is translated by a small amount on the retina and the control circuits are assumed sensitive to transient changes in the signals, the high-frequency components in the differential image are emphasized. Accordingly, the relative changes caused by this translation are more prominent in the control circuits that modulate signals in the central area of the retina, where the spatial resolution is higher, than in the peripheral areas with lower resolution, respectively.

(viii) Thus, if the gaze is shifted by a small amount, activation of the foveal signal paths in relation to the peripheral ones is emphasized, which is experienced subjectively as narrowing of the visual field and concentration of attention.

The Channel Concept

When we inspect what neural organizations might exist for the selection and emphasis of signal paths, we may take advantage of the fact that the visual system (and other sensory systems as well) is known to make use of various channels. A channel, by definition, is a theoretical subsystem for the transfer of a restricted type of information. Then we may imagine that the signal paths starting at the retina and ending up on the visual cortex are organized in spatially ordered, functionally separate channels. A channel is here identified with a set of signal paths, the transmittance of which is controlled by a common control circuit, as, e.g., delineated in Fig. 1. It is not necessary that the channels are separated: a signal path may belong to different channels and be subordinated to several control circuits.

Figure 1.

Figure 1

Rendering of a neural “channel.” P, Principal neuron. There is a common control circuit for this subset of P cells. The dotted arrows represent eventual interaction between the control circuits.

Comparison of Variations in Different Image Areas

Variable Sampling Density.

Hypothesis iv stated that the control circuits should monitor local variations in the signal paths. However, if we want to compare the contents of different parts of the visual field, we have to take into account how the spatial resolution varies over the retina. A small pattern projected on the fovea may elicit a response roughly similar to what its magnified version in the periphery would do, and, theoretically at least, the equivalent magnification can then be taken as proportional to the distance of the pattern from the center of the retina. We do not yet consider details such as different color or motion sensitivities over the retina. Also the true resolution, when measured in terms of contrast sensitivities, may be somewhat different (8).

Spatial variations in signals over subareas of images may be measured in different ways, e.g., by spatial frequency spectra, components of the wavelet transforms (9), or various entropies or other measures of information. We have obtained rather good results when comparing the variances of the samples taken from the images. As pointed out later, variance-sensitive neural circuits can be very simple. However, to relate any of these measures to the biological vision, the samples should always be taken with a spatial density that corresponds to the assumed resolution in the respective subfield of vision.

Sampling of Photographic Images.

We used in simulations photographic images where the pixels were defined in an orthogonal grid. No feature extraction was thereupon yet applied: only the pixel values were modulated by the control circuits. The variable sampling density and resolution, however, had to be taken into account in some way in the control circuit. Consider the grid shown in Fig. 2, which shall correspond to a subarea of the image, and which we call the sampling grid. This grid shall correspond to a control circuit to which a set of local signal paths (delineated by the dashed circle in Fig. 2) is subordinated.

Figure 2.

Figure 2

One example of the sampling grids used for the gating of signal transfer in simulations. The small dots correspond to pixels. Over each of the seven square areas, the average avi, i = 1… 7 of the pixels is computed, whereafter the variance of the avi is evaluated. The dashed circle, which corresponds to the circles of Fig. 5, delineates the set of signal paths modulated by this sampling grid.

Because the density of the pixels cannot be changed, the diameter of this sampling grid shall be selected to correspond to the desired spatial resolution in a particular image area: around the assumed direction of the gaze, the sampling grid shall be smaller and have fewer pixels, whereas the diameter of the grid shall be selected wider and more pixels must be covered with increasing distance from the direction of the gaze (cf. Fig. 5). The same number of subsets of pixels (seven) over each sampling grid was defined, and the averages avi, i = 1… 7 of the pixels over these subsets were computed. The variance for each sampling grid was evaluated on the basis of the avi. Averaging corresponds to low-pass filtering.†

Figure 5.

Figure 5

Placement of the channels over a hypothetical model retina. A control circuit with corresponding (effective) diameter is associated with each circle.

Interaction of Sampling with Signal Paths.

So far, we have not yet considered how the sampling grid and the main signal paths should be combined. These systems were assumed to be organizationally separate. The attentional windows do not have sharp borders, and the sampling grid shall have only a diffuse effect on the main signal paths, the average range of which was delineated by the dashed circle in Fig. 2. When applied to photographs, we have taken for the local gating function, controlled by the sampling grid, the softer Gaussian form.

Let Gatek mean the controllable vectorial gating function corresponding to the kth sampling grid. Let Image mean the vector of pixels in the whole image. The gating function shall multiply the image pixelwise. We are looking for the resulting gated image as some kind of nonlinear superposition of selected parts of the primary input image. If the variance computed from the pixels by the kth sampling grid is denoted Variancek, then the gated image, regarded as the vector Result, may be expressed as

graphic file with name M1.gif 1

where f1 and f2 are still undefined, eventually nonlinear vector functions, and ⊗ means the Hadamard product, or the componentwise product of the corresponding vectors.

At this point, it may be of some interest to note that if we take f2 to be a function that forms the pth power of each vector component, and if f1 forms the corresponding pth root of each component, respectively, then, when we denote

graphic file with name M2.gif 2

and, when p → ∞, Eq. 1 becomes

graphic file with name M3.gif 3

i.e., only the signals over the best-matching channel are transferred.

This work is based on the idea that the best-matching channel can always be determined in some way. The nonlinear “thresholding” by Eqs. 2 and 3, however, is only one possibility and not a very stable one; a more powerful method would be if the channels can be made to compete through some neural mechanisms, as discussed later.

Optimization Approaches

Optimal Width of Attentional Window.

Our first demonstration of the channel operations is not based on real anatomical neural structures. It is a more abstract, theoretical study, the purpose of which is to show that the spatial variations in the input image can mathematically define a proper width of the attentional window. To that end, we first consider a number of alternative channels of different width, located concentrically around the center of the “retina.” In other words, we shall first discuss the theoretical problem of whether the “optimal” width of the attentional window can be found as a solution to an optimization problem, where the solution for the width of the control grid maximizes the variance of the local averages avi of the pixel values, denoted Variance. Let us again call the image data vector Image. Let Grid(w) mean the choice for the grid with width w; then the “optimal” width is defined as

graphic file with name M4.gif 4

A robust optimization of wo in Eq. 4 was carried out over a discrete set of five sampling grids, with their widths varying from 10 to 80 pixels, respectively.

In the first series of simulations illustrated in Fig. 3, we demonstrate the “optimal” width of the attentional window, when the gaze was directed at various objects of different widths: the house, the tower, the palm, the window, and the telephone pole, respectively.

Figure 3.

Figure 3

Demonstration of the opening of attentional windows, the widths of which were automatically determined by the structures present in the area around the gaze. First picture: Original image. The rest of the pictures show attentional windows, when the gaze was directed to the front house, the tower of the back house, the palm, the paned window, and the telephone pole, respectively.

Narrowing of Attentional Window.

The next phenomenon that is explainable by the optimization approach is the narrowing of the attentional window when the gaze is moved, voluntarily or involuntarily, by a small amount.

Let us assume that every sampling grid, to some extent, has also high-pass filter properties, i.e., it enhances transient (phasic) values of the signals it samples. Let these temporal variations of the signals ensue from the shifts of the gaze, i.e., translations of the input image over the sampling grids.

To give first a theoretical explanation to the phenomenon we study, let us consider, without much loss of generality, a one-dimensional grid and evaluate what effect a small translation of the “image” on it will have. Let us assume for simplicity that the amount of translation is a multiple of the grid spacing and at any rate smaller than the width of any of the subareas over which the averages avi are computed.

Consider, for instance, the one-dimensional “image” consisting of the numerical pixel values [3, 1, 4, 1, 5, 9, 2, 6]. Let the sampled part of it be [4, 1, 5, 9]. Let the image be translated by one pixel, and let the new set of samples be [1, 5, 9, 2]. The difference of the averages of the new and the old samples, respectively, is (1 + 5 + 9 + 2 − 4 − 1 − 5 − 9)/4. Thus, a number of the middle pixels cancel each other, and only the edge pixels contribute to the difference. It may then be obvious that the larger the sampling grid, the smaller the relative contribution of the edge pixels to the difference of the averages in general. For other configurations of the sampling grid and for general image data, it is more cumbersome to carry out a similar analysis, but one may safely conjecture that, if the original pixels are bounded and have certain very general statistical properties, the expected averages avi, computed from the difference of the translated images, when the translation is small, decrease along with the increasing width w of the sampling grid.

A similar result is obtained if one considers the spatial frequencies of the images: if the translation is small, the absolute value of the difference is approximately proportional to the Euclidean norm of its gradient, in which high spatial frequencies are enhanced in proportion to the frequency.

In the evaluation of the optimal width wo from Eq. 4, the variances computed from the avi for the difference image thus decrease with the width of the grid, too, and the optimal width wo is decreased.

When only a fraction of the previous image is subtracted from the new image, a similar shift of wo toward smaller values, although a weaker one, can be seen. This effect is then reflected as a narrowing of the attentional window. In Fig. 4, a sequence of images is shown, where the subtracted fraction was 50% of each previous image.

Figure 4.

Figure 4

Automatic narrowing of the attentional window, when the variances were computed on the basis of images from which 50% of the previously sampled translated image was subtracted. First picture: The original image. The three other pictures form a sequence, in which the gaze was shifted in steps, the size of which became successively smaller.

Neural Implementation

Neural Detection of Variance.

The neural implementation of the variance detector is not a particularly demanding task. Consider that the variance of a set of variables {xi | i = 1… n} is expressible as the sum of squares of the xi, divided by n, minus the square of the average of the xi. For this situation, we need two types of neurons: excitatory ones that add the signals quadratically, and inhibitory ones where the addition is linear.

To activate the signal paths according to local variations in the signals, one could also analyze local spatial frequencies instead of variations in local averages. The spatial frequency analysis might seem to be directly amenable to neural implementation, because various kinds of frequency filters are known to exist in the sensory systems. For instance, the simple cells of the visual cortex are known to be selective to lines with different orientation, but they can also be regarded as bandpass filters for various spatial frequencies (10, 11). They seem to be modelable as Gabor filters (12, 13). An adaptive model, in which neural detectors very much resembling the Gabor filters are learned from natural image data, is the ASSOM system suggested by this author (14, 15). All these filter functions, however, are very computation intensive in simulations.

Competing Channels.

The possibilities for the handling of theoretical optimization tasks may be rather limited in the neural realms. Especially if the optimization must take place very fast, during a typical period of 100 ms, a realistic neural method would be competition among a set of discrete, alternative, parametrically different subsystems that are performing the same task in parallel. In this method, when the subsystem called the winner first performs the task, it blocks out its competitors from that task (16, 17).

It is natural to think that the channels are distributed spatially over the retina according to its spatial resolution. They may also overlap partially. The simplified configuration used in the next simulations is shown in Fig. 5.

As mentioned in hypotheses iii and v, it will be necessary to point out that, if the principal cells of Fig. 1 and their control circuits reside physically on higher levels of the visual system, they need not have very different physical sizes to comply with the retinal subareas of different size delineated in Fig. 5. On the contrary, because the central areas of the retina have a larger representation (higher magnification factor) on the thalamus and the primary visual cortex than the periphery of the retina, the signal paths starting at very unequal areas on the retina may converge to bundles of principal cells of roughly equal diameter on the higher levels.

If the channels, as depicted in Fig. 5, are made to compete, the control circuits should somehow be able to communicate mutually to compare their activity states, either directly, as indicated by the dotted arrows in Fig. 1, or maybe under a central control. In the latter case, if the comparison is made on a higher level (eventually in higher areas of the cerebral cortex), its result must be fed back to the gating circuits.

Attentional Window as an Activated Subset of Channels.

Above, we defined the optimal width of the attentional window as the width of the sampling grid for which the variance of the avi was maximized. Next, we shall consider a more concrete “biological” case in which the set of channels is fixed and their sizes and positions are exemplified in Fig. 5. For each channel, a sampling grid of corresponding diameter is associated.

Instead of looking for a single optimal channel as before, we now determine a combination of k activated channels over which the variance of the avi is highest. In the neural-network theory, this scheme would be called “k winners.” In this way, whereas most of the channels are located eccentrically with respect to the direction of the gaze, the combination of the activated channels defines a more or less symmetric (usually noncircular) attentional window.

In the simulations presented in Fig. 6, we thus use the 33-channel “retina” of Fig. 5 and let five highest-variance channels define the attentional window. As can be seen, the five channels together tend to emphasize a part of the visual field where some meaningful patterns are present.

Figure 6.

Figure 6

Examples of attentional windows spanned by a combination of five activated channels. The black cross indicates the direction of the gaze. Notice how the prominent patterns act as distractors, attenuating other parts of the visual field.

It is also discernible that, if the variance in the central part of the visual field is low, prominent eccentric patterns tend to attenuate weaker parts of the visual field, which can then be interpreted as the distraction of attention by the prominent eccentric objects.

Discussion

In the simulations carried out in this work, the monitoring and gating function of the control circuit was still studied on a rather abstract level, without trying to specify where it could reside in the brain. The monitoring function was also assumed as very simple, namely, one that is able to analyze variances in subsets of neighboring pixels, and enhance the transients in these variances. If we had aimed at the demonstration of more complex effects such as the emphasis (“pop out”) of deviating features or patterns in the visual field, the control circuits should have been made sensitive to such features and patterns. This is a possible continuation to our studies, and, for instance, variances in color signals can be computed readily. However, for the detectors of more complicated features such as lines and motion, more complicated neural components, and orders of magnitude more computing power in simulations would be needed.

In trying to simulate the combined effect of many different functions, the main problem at least in the primate visual cortex is that the various features are analyzed in different subareas. However, the divergence to these subareas takes place after the primary visual cortex, area V1, whereas the proposed gating mechanism could occur anywhere along the retina-thalamus-V1 pathway. The results of all these analyses should then be made available to the same relaying stations. Although many kinds of feedback are known to exist in the visual system, their structures are not known with sufficient accuracy to be amenable to concrete computer simulations.

The above problem has been discussed in the prevailing theories of selective attention that are based on synchronized oscillations (18–22). The main emphasis in the latter has been an attempt to explain the binding problem, i.e., how the features of very different kinds, while being detected in different places, can be combined into a unified perceived item. Although the synchronization of neural oscillations has been demonstrated in simple simulation models, it is not yet clear how these oscillations in the first place could be excited selectively, in relation to the sensory signals. Maybe the present article could give a hint to that.

The binding effects had to be left outside the scope of this article. It might also turn out that only very simple features are used in the ascending signal paths where the reactions must be fast (23), whereas if the more complex features are analyzed on higher levels and the results are fed back to the gating circuits, their effect on attention might be delayed. This possibility could be tested experimentally.

Above, the theory of automatic capture and focusing of visual attention was discussed only as a computational-neuroscience problem. Nonetheless, it might have an even more direct impact on active computer vision and the retinal implants. In the present plans to design prostheses for vision, a video image is projected onto the visual cortex (24) or eventually directly onto the optic nerve. However, it will be very hard for the blind person to accept all of the artificially defined optic information present in the visual scene. If the visual information would be reduced and preprocessed by the attention principle as described in this paper, the person would perceive only the most important components of the visual scene. It is then plausible that he/she would learn to “see” easier and faster.

Acknowledgments

Special thanks are due to Professor Riitta Hari for the comments on the manuscript. This work was done under the auspices of Helsinki University of Technology.

Footnotes

†

In the demonstrations, the variance of the avi was actually divided by the mean of the pixels over the sampling grid.

References

  • 1.Hatfield G. In: Visual Attention. Wright R D, editor. New York: Oxford Univ. Press; 1998. pp. 3–25. [Google Scholar]
  • 2.Hubel D H, Wiesel T N. J Physiol. 1961;155:385–398. doi: 10.1113/jphysiol.1961.sp006635. [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 3.Hubel D H, Wiesel T N. J Comp Neurol. 1972;146:421–450. doi: 10.1002/cne.901460402. [DOI] [PubMed] [Google Scholar]
  • 4.Nicholls J G, Martin A R, Wallace B G. From Neuron to Brain. 3rd Ed. Sunderland, MA: Sinauer; 1992. pp. 596–597. [Google Scholar]
  • 5.Tusa R J, Palmer L A, Rosenquist A C. J Comp Neurol. 1978;177:213–235. doi: 10.1002/cne.901770204. [DOI] [PubMed] [Google Scholar]
  • 6.Bresslof P C, Cowan J D, Golubitsky M, Thomas P J, Wiener M C. Neural Comput. 2002;14:473–491. doi: 10.1162/089976602317250861. [DOI] [PubMed] [Google Scholar]
  • 7.Shepherd G M, editor. The Synaptic Organization of the Brain. New York: Oxford Univ. Press; 1990. [Google Scholar]
  • 8.Virsu V, Rovamo J. Exp Brain Res. 1979;37:475–494. doi: 10.1007/BF00236818. [DOI] [PubMed] [Google Scholar]
  • 9.Daubechies I. IEEE Trans Inf Theory. 1990;36:961–1005. [Google Scholar]
  • 10.Daugman J. Vision Res. 1980;20:847–856. doi: 10.1016/0042-6989(80)90065-6. [DOI] [PubMed] [Google Scholar]
  • 11.Daugman J. J Opt Soc Am A. 1985;2:1160–1169. doi: 10.1364/josaa.2.001160. [DOI] [PubMed] [Google Scholar]
  • 12.Marc̆elja S. J Opt Soc Am. 1980;70:1297–1300. doi: 10.1364/josa.70.001297. [DOI] [PubMed] [Google Scholar]
  • 13.Jones J P, Palmer L A. J Neurophysiol. 1987;58:1233–1258. doi: 10.1152/jn.1987.58.6.1233. [DOI] [PubMed] [Google Scholar]
  • 14.Kohonen T. Biol Cybern. 1996;75:281–291. [Google Scholar]
  • 15.Kohonen T, Kaski S, Lappalainen H. Neural Comput. 1997;9:1321–1344. [Google Scholar]
  • 16.Yuille A, Geiger D. In: The Handbook of Brain Theory and Neural Networks. Arbib M, editor. Cambridge, MA: MIT Press; 1995. pp. 1056–1060. [Google Scholar]
  • 17.Kaski S, Kohonen T. Neural Networks. 1994;7:973–984. [Google Scholar]
  • 18.von der Malsburg C, Schneider W. Biol Cybern. 1986;54:29–40. doi: 10.1007/BF00337113. [DOI] [PubMed] [Google Scholar]
  • 19.Eckhorn R, Bauer R, Jordan W, Brosch M, Kruse W, Munk M, Reitboeck H J. Biol Cybern. 1988;60:121–130. doi: 10.1007/BF00202899. [DOI] [PubMed] [Google Scholar]
  • 20.Chawanya T, Aoyagi T, Nishikawa I, Okuda K, Kuramoto Y. Biol Cybern. 1993;68:483–490. doi: 10.1007/BF00200807. [DOI] [PubMed] [Google Scholar]
  • 21.Niebuhr E, Koch C, Rosin C. Vision Res. 1993;33:2789–2802. doi: 10.1016/0042-6989(93)90236-p. [DOI] [PubMed] [Google Scholar]
  • 22.Corchs S, Deco G. Neural Networks. 2001;14:981–990. doi: 10.1016/s0893-6080(01)00055-7. [DOI] [PubMed] [Google Scholar]
  • 23.Steinman B A, Steinman S B, Lehmkuhle S. Vision Res. 1997;37:17–23. doi: 10.1016/s0042-6989(96)00151-4. [DOI] [PubMed] [Google Scholar]
  • 24.Loeb G E. In: Handbook of Brain Theory and Neural Networks. Arbib M A, editor. Cambridge, MA: MIT Press; 1995. pp. 768–772. [Google Scholar]

Articles from Proceedings of the National Academy of Sciences of the United States of America are provided here courtesy of National Academy of Sciences

RESOURCES